Method for operating electronic device supporting ai-assistance function and electronic device supporting same
The electronic device integrates multiple AI assistance functions by mapping scene objects to personas, enhancing the versatility and effectiveness of voice user interfaces in interacting with users across diverse domains.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- SAMSUNG ELECTRONICS CO LTD
- Filing Date
- 2025-11-18
- Publication Date
- 2026-06-04
AI Technical Summary
Existing voice user interfaces lack the ability to effectively integrate multiple AI assistance functions for diverse domains and scenarios, limiting their versatility and effectiveness in interacting with users.
An electronic device equipped with a camera, microphone, and processor is capable of detecting scene objects mapped to different personas, checking their relationship, and integrating these personas to generate responses across multiple information databases, enabling seamless interaction and response generation based on user queries.
Enhances the capability of voice user interfaces to provide integrated AI assistance across various domains by mapping and integrating personas, improving user interaction and response relevance.
Smart Images

Figure KR2025019076_04062026_PF_FP_ABST
Abstract
Description
Method of operating an electronic device supporting AI assistance functions and an electronic device supporting the same
[0001] The embodiments disclosed in this document relate to methods for utilizing AI assistance functions.
[0002] A voice user interface can interact with an intelligent personal assistant (IPA) or virtual assistant (VA) (or AI assistant function) running on a voice command device. A voice command device is a device that can be controlled by a voice user interface. The voice user interface can support interaction between humans and electronic devices by using speech recognition to understand voice commands (e.g., spoken commands) and questions, and by outputting responses using text-to-speech. With advancements in automatic speech recognition (ASR) and natural language understanding (NLU), voice user interfaces are becoming increasingly popular in automobiles, mobile devices (e.g., smartphones, tablets, watches, etc.), home appliances (e.g., washing machines, dryers, etc.), and entertainment devices (e.g., smart TVs, smart speakers).
[0003] The information described above may be provided as related art for the purpose of aiding understanding of the present disclosure. No claim or determination is made as to whether any of the foregoing may be applied as prior art in relation to the present disclosure.
[0004] An electronic device according to one embodiment disclosed in this document includes a microphone (182), a camera (170), a memory (130), and at least one processor (120) operatively connected to the microphone, the camera, and the memory. At least one instruction stored in the memory and configured to be executed by the at least one processor to perform the operation of the electronic device may be configured to detect a first scene object mapped to a first persona and a second scene object mapped to a second persona in a scene acquired through the camera, check whether the relationship between the first scene object and the second scene object satisfies a set condition, and if the set condition is satisfied, map an integrated persona that combines the first persona and the second persona to at least one of the first scene object and the second scene object. Here, the first persona may be mapped to a first AI (artificial intelligence) assist function that generates a first response to a user query through a first domain corresponding to a first information database, the second persona may be mapped to a second AI assist function that generates a second response to the user query through a second domain corresponding to a second information database different from the first information database, and the integrated persona may be mapped to a third AI assist function that generates a third response to the user query through a domain corresponding to the first information database and the second information database.
[0005] A method of operating an electronic device according to an embodiment disclosed in this document may include, in a scene acquired through a camera, detecting a first scene object mapped to a first persona and a second scene object mapped to a second persona; checking whether the relationship between the first scene object and the second scene object satisfies a set condition; and if the set condition is satisfied, mapping an integrated persona, which is the integration of the first persona and the second persona, to at least one of the first scene object and the second scene object.
[0006] A computer-readable recording medium disclosed in this document may store at least one instruction such that, when executed by an electronic device, the electronic device performs the operation of detecting a first scene object mapped to a first persona and a second scene object mapped to a second persona in a scene acquired through a camera, the operation of checking whether the relationship between the first scene object and the second scene object satisfies a set condition, and if the set condition is satisfied, the operation of mapping an integrated persona, which integrates the first persona and the second persona, to at least one of the first scene object and the second scene object.
[0007] An electronic device according to one embodiment disclosed in this document includes a microphone (182), a camera (170), a memory (130), and at least one processor (120) operatively connected to the microphone, the camera, and the memory. At least one instruction stored in the memory and configured to be executed by the at least one processor to perform the operation of the electronic device is configured to detect a first scene object in a scene collected through the camera, receive a mapping query related to persona mapping to the first scene object, determine the type of persona to be mapped to the first scene object based on keywords included in the mapping query, and map the determined persona to the first scene object. The persona may include an artificial intelligence (AI) assistance function that generates a response to a user query based on an information database related to a domain set based on the keywords.
[0008] In relation to the description of the drawings, the same or similar reference numerals may be used for identical or similar components.
[0009] Figure 1 illustrates an example of augmented reality.
[0010] FIG. 2 illustrates a block diagram of an electronic device according to one embodiment.
[0011] FIG. 3a shows a perspective view of an electronic device according to one embodiment.
[0012] FIG. 3b illustrates an example of internal hardware of an electronic device according to one embodiment.
[0013] FIG. 4a shows a rear perspective view of an electronic device according to one embodiment.
[0014] FIG. 4b shows a front perspective view of an electronic device according to one embodiment.
[0015] FIG. 5a shows a front perspective view of a detachable electronic device according to one embodiment.
[0016] FIG. 5b shows a rear perspective view of a detachable electronic device according to one embodiment.
[0017] FIG. 6 illustrates examples of electronic devices according to an embodiment.
[0018] FIG. 7 is a diagram showing an example of a processor configuration of an electronic device according to one embodiment.
[0019] FIG. 8 is a diagram showing an example of persona operation according to one embodiment.
[0020] FIG. 9 is a diagram showing an example of a scene object and persona mapping method according to one embodiment.
[0021] FIG. 10 is a diagram showing an example of a method of using a persona according to one embodiment.
[0022] FIG. 11 is a diagram showing an example of a persona change method according to one embodiment.
[0023] FIG. 12 is a diagram showing an example of a situation related to persona mapping according to one embodiment.
[0024] FIG. 13 is a diagram illustrating an example of a situation related to persona mapping verification according to one embodiment.
[0025] FIG. 14 is a diagram showing an example of a first situation related to a persona activation setting according to one embodiment.
[0026] FIG. 15 is a diagram showing an example of a second situation related to a persona activation setting according to one embodiment.
[0027] FIG. 16 is a diagram showing an example of a first situation related to the use of a persona according to one embodiment.
[0028] FIG. 17 is a diagram showing an example of a second situation related to the use of a persona according to one embodiment.
[0029] FIG. 18 is a diagram showing an example of a second situation related to persona mapping according to one embodiment.
[0030] FIG. 19 is a diagram showing an example of a fourth situation related to persona mapping according to one embodiment.
[0031] FIG. 20 is a diagram showing an example of a fifth situation related to persona mapping according to one embodiment.
[0032] FIG. 21 is a diagram illustrating an example of a situation related to persona grouping and operation according to one embodiment.
[0033] FIG. 22 is a diagram showing an example of a second situation related to persona grouping and operation according to one embodiment.
[0034] FIG. 23 is a diagram illustrating an example of a situation related to persona inheritance according to one embodiment.
[0035] FIG. 24 is a diagram illustrating an example of a situation related to gesture-based persona mapping according to one embodiment.
[0036] FIG. 25 is a diagram showing an example of an electronic device operation method related to persona mapping according to one embodiment.
[0037] FIG. 26 is a diagram showing another example of an electronic device operation method related to persona mapping according to one embodiment.
[0038] FIG. 27 is a diagram showing an example of an electronic device operation method related to persona integration according to one embodiment.
[0039] Hereinafter, embodiments of the present invention are described with reference to the accompanying drawings. However, this is not intended to limit the present invention to specific embodiments and should be understood to include various modifications, equivalents, and / or alternatives of the embodiments of the present invention.
[0040] The embodiments of the present invention described below may be provided based on electronic devices related to augmented reality so as to provide user services in an augmented reality (e.g., virtual reality, augmented reality, or mixed reality) situation. For example, the embodiments of the present invention may provide AI assistance functions by classifying them by specific fields through electronic devices, and may support mapping at least some of the multiple AI assistance functions specialized for specific fields to each of the scene objects in augmented reality. Additionally or generally, the electronic devices of the present invention may support integrated AI assistance functions by integrating multiple AI assistance functions specialized for specific fields based on at least one of the placement relationship or semantic relationship of multiple scene objects in a real-world situation (or augmented reality situation). In the following description, various types of electronic devices and environments to which the embodiments of the present invention may be applied are mentioned based on FIGS. 1 to 6, various embodiments related to the operation of the electronic devices of the present invention are described in FIGS. 7 to 24, and examples related to the method of operating the electronic devices are described in FIGS. 25 to 27. The embodiments described below with reference to each drawing may be combined or mixed with embodiments described in other drawings.
[0041] Figure 1 illustrates an example of extended reality.
[0042] Referring to FIG. 1, according to one embodiment, an electronic device (10) may be configured to provide extended reality. Extended reality may include augmented reality, mixed reality, and / or virtual reality. For example, the electronic device (10) may provide augmented reality and / or mixed reality by outputting a virtual object mapped to a real space (or real world) where the user (1) is located. For example, the electronic device (10) may provide augmented reality by outputting a virtual object within any virtual space. For example, the electronic device (10) may provide extended reality by configuring at least some of the surrounding spaces of the user (1) as a virtual space and the remaining parts as a real space.
[0043] In the present disclosure, “actual space” may include physically existing space and physical objects within the physical space. In the present disclosure, “virtual space” may be referred to as a virtual space provided by an electronic device (10). The virtual space and the actual space may be distinguished based on a visual representation provided by the electronic device (10). For example, if the electronic device (10) captures and provides the actual space, the provided space may be referred to as the actual space. For example, if the electronic device (10) visually reconstructs the actual space, the provided space may be referred to as the virtual space.
[0044] In the present disclosure, a “virtual object” may be referred to as a graphic object that does not exist in real space but is output by an electronic device (10). The electronic device (10) may output the virtual object in real space and / or virtual space.
[0045] In the example of FIG. 1, the electronic device (10) can output a first virtual object (8a) and / or a second virtual object (8b) in real space. For example, the electronic device (10) can output the first virtual object (8a) and / or the second virtual object (8b) mapped to a real object (e.g., a laptop (3) on a desk (2)). In the present disclosure, “mapping output” may include outputting a virtual object by mapping it to a real object or a real location. For example, the electronic device (10) can output the first virtual object (8a) and / or the second virtual object (8b) by mapping it to the laptop (3). The electronic device (10) can output the first virtual object (8a) and / or the second virtual object (8b) to a location adjacent to the laptop (3) when viewed by the user (1). The electronic device (10) can output the first virtual object (8a) and / or the second virtual object (8b) so that the first virtual object (8a) and / or the second virtual object (8b) have a depth similar to that of the laptop (3) (or are observed by the user (1) at a similar depth).
[0046] According to one embodiment, the electronic device (10) may support multiple modalities. For example, the electronic device (10) may be configured to receive and process various types of inputs. For example, 'modality' may refer to a channel for interaction between the electronic device (10) and the user (1). In this case, voice input and text input via an interface (e.g., a virtual keyboard) may be considered different modalities. For example, 'modality' may refer to the format of data input into an artificial intelligence model (e.g., a generative artificial intelligence model). For example, the voice input of the user (1) may be converted into text data via STT (speech to text) and NLU (natural language understanding) and input into the artificial intelligence model. From the perspective of the artificial intelligence model, voice input and text input via an interface (e.g., a virtual keyboard) may be considered the same modality. In the following, 'modality' may refer to a channel (or input channel) between a user (1) and an electronic device (10) and / or a format of input data for an artificial intelligence model. In the present disclosure, 'modality' may be referred to as 'input type'.
[0047] In one example, the electronic device (10) may be a wearable device. The electronic device (10) may acquire voice input from a user (1) using a microphone. The electronic device (10) may acquire image input using a camera. The electronic device (10) may acquire user input using an interface (e.g., a physical interface including buttons and / or a touchpad). The electronic device (10) may acquire user input from an external device (not shown) that is communicably connected. The electronic device (10) may be configured to acquire voice input, image input and / or user input, and to process the acquired input.
[0048] According to one embodiment, the electronic device (10) can provide an image having depth (hereinafter referred to as a depth image). For example, the electronic device (10) can adjust the depth of an image (e.g., an image of real space) obtained through a camera. For example, the electronic device (10) can detect (or obtain) depth information of real space and, based thereon, provide an image that reflects the depth of real space.
[0049] In the following, an electronic device (10) that provides a depth image having depth can be described with reference to various drawings.
[0050] FIG. 2 illustrates a block diagram of an electronic device according to one embodiment.
[0051] Referring to FIG. 2, according to one embodiment, an electronic device (10) may include at least one processor (120) (hereinafter referred to as processor (120)), at least one memory (130) (hereinafter referred to as memory (130)), at least one sensor circuit (140) (hereinafter referred to as sensor circuit (140)), at least one display (160) (hereinafter referred to as display (160)), at least one camera (170) (hereinafter referred to as camera (170)), at least one interface (180) (hereinafter referred to as interface (180)), and / or at least one communication circuit (190) (hereinafter referred to as communication circuit (190)).
[0052] The components of the electronic device (10) described above are merely one example, and depending on the example, the electronic device (10) may be implemented to have more components than the components shown in FIG. 2 or fewer components.
[0053] For example, at least some components of the electronic device (10) that are not shown in FIG. 2 but can be composed of components of the electronic device (10) (e.g., battery, acoustic output module and / or antenna module) may be included in the composition of the electronic device (10). Additionally or optionally, at least one of the components of the electronic device (10) described above may be integrated with other components.
[0054] According to one embodiment, the processor (120) may be connected communically, electrically, operatively, or functionally to a memory (130), a sensor circuit (140), a display (160), a camera (170), an interface (180), and / or a communication circuit (190). In various embodiments of the present disclosure, when one component is connected “operatively” to another component, it may mean that the component is connected to enable the other component to operate. For example, the component may enable the other component by transmitting a control signal to the other component directly or through another component. In various embodiments of the present disclosure, when one component is connected “functionally” to another component, it may mean that the component is connected to enable the function of the other component. For example, the component may enable the function of the other component by transmitting a control signal to the other component directly or through another component.
[0055] The processor (120) may include at least one processing circuit. For example, the processor (120) may include an application processor (AP), a central processing unit (CPU), an image signal processor (ISP), a graphics processing unit (GPU), a neural processing unit (NPU), a tensor processing unit (TPU), and / or a communication processor (CP). The processor (120) may include at least one chip or one chipset. In the present disclosure, the processor (120) may be referred to as a hardware component having an architecture by at least one processing circuit. For example, the processor (120) may be placed on a substrate (e.g., a printed circuit board) located inside the electronic device (10) and may communicate with other components of the electronic device (10) through at least one conductive path formed in the substrate, a flexible printed circuit board (FPCB), a cable, and / or any communication channel (e.g., a wireless communication channel).
[0056] The memory (130) can store instructions. When the instructions are executed by the processor (120), they can cause the electronic device (10) to perform various operations. For example, the instructions can cause the electronic device (10) to perform various operations by being executed individually or collectively by at least one processor. In various embodiments of the present disclosure, the operation of the electronic device (10) may be referred to as an operation performed by the processor (120) by executing instructions stored in the memory (130). The memory (130) may be referred to as a hardware component for storing data.
[0057] The sensor circuit (140) may be configured to detect information associated with the electronic device (10), e.g., context information. The context information may include optical information of the surroundings of the electronic device (10), movement information of the electronic device (10), attachment status information of the electronic device (10), location information of the electronic device (10), and / or information of adjacent objects. For example, the sensor circuit (140) may include an image sensor (141), an inertial sensor (143), a position sensor (145), a proximity sensor (147), and / or a depth sensor (149). The configuration of the sensor circuit (140) shown in FIG. 2 is exemplary and the embodiments of the present disclosure are not limited thereto.
[0058] The image sensor (141) may be configured to detect optical signals. For example, the image sensor (141) may be configured to detect light signals, infrared signals, and / or ultraviolet signals.
[0059] The inertial sensor (143) may be configured to detect movement of the electronic device (10). For example, the inertial sensor (143) may include an inertial measurement unit (IMU) configured to detect the orientation, magnetic force, angular rate, and / or specific force of the electronic device (10). The inertial sensor (143) may include an accelerometer, a gyroscope, and / or a magnetometer.
[0060] The location sensor (145) may include any sensor configured to acquire the geographical location of the electronic device (10). For example, the location sensor (145) may acquire geographical location information based on a global navigation satellite system (GNSS). For example, the location sensor (145) may acquire geographical location information based on the reception strength and / or angle of arrival of an external signal. In one example, the electronic device (10) may acquire geographical location information based on reception information from a base station.
[0061] The proximity sensor (147) may be configured to detect an external object in close proximity to the electronic device (10). The electronic device (10) may use the proximity sensor (147) to detect the blockage status of the electronic device (10). For example, the proximity sensor (147) may include a capacitive sensor, a photoelectric sensor, a touch switch, an illuminance sensor, and / or an inductive sensor.
[0062] The depth sensor (149) can obtain depth information for one or more objects contained in (or existing in) the actual space. For example, one or more objects may include fixed structures such as trees, banners, or buildings, as well as non-fixed objects such as animals or people.
[0063] According to one embodiment, depth information may correspond to the distance from the depth sensor (149) to a specific object. In this regard, the depth sensor (149) may include a light emitting part that emits light having a predetermined wavelength and a sensor part (640) that detects reflected light that is reflected back from the specific object. For example, the depth value may be larger the further the distance from the depth sensor (149) to the specific object.
[0064] According to one embodiment, the depth sensor (149) can acquire depth information for the actual space corresponding to the field of view (FOV) of the user (1) or the camera (170). For example, the depth sensor (149) can acquire depth information using at least one of the time of flight (ToF) method, the structured light method, and the stereo image method. According to an embodiment, the depth sensor (149) may be integrated with an image sensor (141).
[0065] The display (160) may include at least one pixel configured to output an image. In one example, the display (160) may include a plurality of displays. For example, the display (160) may include a left-eye display and a right-eye display. The display (160) may include a front display and / or a rear display. The display (160) may include at least one of a see-through display, a flexible display, a rollable display, a foldable display and / or a rigid display. In one example, the display (160) may include at least one projector for projecting an image.
[0066] The camera (170) can capture still images and / or video. According to one embodiment, if the electronic device (10) includes a plurality of cameras, each of the plurality of cameras may have at least one of the facing direction, magnification, or field of view.
[0067] According to one embodiment, at least one of the plurality of cameras (e.g., a shooting camera) may be configured to be provided at a first location of the electronic device (10) to acquire an image of a first direction. For example, the image of the first direction may include actual space corresponding to the user's (1) field of vision. Additionally, at least another of the plurality of cameras (e.g., an eye-tracking camera) may be configured to be provided at a second location of the electronic device (10) to acquire an image of a second direction. For example, the image of the second direction may include the user's eyes. Additionally, at least yet another of the plurality of cameras (e.g., a motion recognition camera) may be provided at a third location of the electronic device (10) to acquire an image of a third direction. For example, the image of the third direction may include the body of the user (1) performing a specific gesture.
[0068] In one example, the electronic device (10) can select a camera for acquiring an image from a plurality of cameras. The electronic device (10) can select a camera for acquiring an image based on the user's context information (e.g., speech and / or direction of movement). For example, the electronic device (10) can identify an image containing an object corresponding to the user's speech among the plurality of cameras through image recognition. The electronic device (10) can acquire an image corresponding to the voice input using the camera that acquired the identified image. For example, the electronic device (10) can select a camera among the plurality of cameras that faces the direction of movement of the electronic device (10). In one example, the electronic device (10) can select one camera among the plurality of cameras according to the user input.
[0069] The interface (180) may include at least one device configured to receive input. For example, the interface (180) may include a touch circuit (181) (e.g., a touch screen display) configured to receive touch input. The interface (180) may include at least one microphone (e.g., a first microphone (182) and / or a second microphone (183)) configured to receive voice input. In one example, the electronic device (10) can identify the location of the speaker of the voice input (e.g., the relative direction of the speaker to the electronic device (10)) by performing beamforming using the first microphone (182) and the second microphone (183). The interface (180) may include at least one device for output. For example, the interface (180) may include a haptic module for tactile output, at least one speaker for sound output (e.g., a speaker (184)), and / or an indicator. The interface (180) may include any human interface device (HID). For example, the interface (180) may include a button (185). According to one embodiment, the processor (120) may be configured to receive input using the interface (180) and to process the received input.
[0070] The communication circuit (190) may be configured to perform short-range wireless communication and / or long-range wireless communication. The communication circuit (190) may include a network interface card (NIC). The processor (120) may communicate with other external electronic devices based on wireless communication and / or wired communication, for example, using the communication circuit (190). The processor (120) may communicate with external devices via an internet protocol (IP) network, for example, using the communication circuit (190).
[0071] FIG. 3a shows a perspective view of an electronic device according to one embodiment.
[0072] Referring to FIG. 3a, an electronic device (300) (e.g., electronic device (10)) according to one embodiment may include at least one display (350) (e.g., display (160)) and a frame (or housing) supporting at least one display (350). For example, the electronic device (300) may be worn on a part of a user's body. For example, the electronic device (300) may include a glasses-type device configured to provide extended reality.
[0073] The frame may be formed as a physical structure that allows the electronic device (300) to be worn on the user's body. The frame may be configured so that when the user wears the electronic device (300), the first display (350-1) and the second display (350-2) can be positioned to correspond to the user's left and right eyes.
[0074] The frame may include a first rim (301) covering at least a portion of a first display (350-1), a second rim (302) covering at least a portion of a second display (350-2), a bridge (303) positioned between the first rim (301) and the second rim (302), a first pad (311) positioned along a portion of the edge of the first rim (301) from one end of the bridge (303), a second pad (312) positioned along a portion of the edge of the second rim (302) from the other end of the bridge (303), a first temple (304) extending from the first rim (301) and fixed to a portion of the wearer's ear, and a second temple (305) extending from the second rim (302) and fixed to a portion of the ear opposite to the ear.
[0075] The first pad (311) and the second pad (312) may come into contact with a part of the user's nose, and the first temple (304) and the second temple (305) may come into contact with a part of the user's face and a part of the ear.
[0076] The temples (304, 305) can be rotatably connected to the rim through the hinge units (306, 307) of FIG. 3B. The first temple (304) can be rotatably connected to the first rim (301) through a first hinge unit (306) positioned between the first rim (301) and the first temple (304). The second temple (305) can be rotatably connected to the second rim (302) through a second hinge unit (307) positioned between the second rim (302) and the second temple (305).
[0077] The frame may include a contact area (320) in which at least a portion comes into contact with a part of the user's body when the user wears the electronic device (300). For example, the contact area (320) may include a nose pad (310), a first temple (304), and a second temple (305).
[0078] The electronic device (300) can provide visual information to a user through at least one display (350). For example, at least one display (350) may include a transparent or translucent lens. At least one display (350) may include a first display (350-1) and / or a second display (350-2) spaced apart from the first display (350-1). For example, the first display (350-1) and the second display (350-2) may be positioned at locations corresponding to the user's left and right eyes, respectively.
[0079] FIG. 3b illustrates an example of internal hardware of an electronic device according to one embodiment.
[0080] Referring to FIGS. 3a and 3b, according to one embodiment, the electronic device (300) may include hardware that performs various functions (e.g., hardware described through FIG. 2). For example, the hardware may include a battery module (370), an antenna module (375), optical devices (382, 384), speakers (392-1, 392-2), microphones (394-1, 394-2, 394-3), a light-emitting module (not shown), and / or a printed circuit board (390). The aforementioned hardware may be placed within a frame.
[0081] At least one display (350) may form a display area on the lens to provide a user wearing the electronic device (300) with visual information that is distinct from the visual information, along with the visual information contained in the external light passing through the lens. According to an embodiment, the lens may be formed based on at least one of a Fresnel lens, a pancake lens, or a multi-channel lens.
[0082] A display area formed by at least one display (350) may be formed on the second surface (332) among the first surface (331) and the second surface (332) of the lens. When a user wears the electronic device (300), external light may be transmitted to the user by being incident on the first surface (331) and transmitted through the second surface (332). As another example, at least one display (350) may display a virtual reality image to be combined with a real-world image transmitted through external light.
[0083] The virtual reality image output from at least one display (350) can be transmitted to the user's eye through one or more hardware included in the electronic device (300) (e.g., optical devices (382, 384) and / or at least one waveguide (333, 334)).
[0084] According to one embodiment, the electronic device (300) may include waveguides (333, 334) that diffract light transmitted from at least one display (350) and relayed by optical devices (382, 384) and transmit it to a user.
[0085] Waveguides (333, 334) may be formed based on at least one of glass, plastic, or polymer. A nano pattern may be formed on the exterior or at least a portion of the interior of the waveguides (333, 334). The nano pattern may be formed based on a polygonal and / or curved grating structure. Light incident on one end of the waveguides (333, 334) may be propagated to the other end of the waveguides (333, 334) by the nano pattern.
[0086] Waveguides (333, 334) may include at least one diffractive element (e.g., DOE (diffractive optical element), HOE (holographic optical element)) and at least one reflective element (e.g., a reflective mirror). For example, waveguides (333, 334) may be placed within an electronic device (300) to guide a screen displayed by at least one display (350) to the user's eye. For example, the screen may be transmitted to the user's eye based on total internal reflection (TIR) occurring within the waveguides (333, 334).
[0087] According to one embodiment, microphones (394-1, 394-2, 394-3) (e.g., a first microphone (182) and / or a second microphone (183)) of an electronic device (300) are positioned on at least a portion of a frame to acquire a sound signal. A first microphone (394-1) positioned on a nose pad (310), a second microphone (394-2) positioned on a second rim (302), and a third microphone (394-3) positioned on a first rim (301) are shown in FIG. 3b, but the number and placement of the microphones (394) are not limited to the embodiment of FIG. 3b. If there are two or more microphones (394) included in the electronic device (300), the electronic device (300) can identify the direction of the sound signal using a plurality of microphones positioned on different portions of the frame.
[0088] In one embodiment, the camera (340) (e.g., camera (170)) may include an eye tracking camera (ET CAM) (340-1), a motion recognition camera (340-2), and / or a shooting camera (340-3). The shooting camera (340-3), the eye tracking camera (340-1), and the motion recognition camera (340-2) may be positioned at different locations on the frame and may perform different functions.
[0089] The eye tracking camera (340-1) can output data indicating the gaze of a user wearing the electronic device (300). For example, the electronic device (300) can detect the gaze from an image containing the user's pupils obtained through the eye tracking camera (340-1). An example in which the eye tracking camera (340-1) is positioned toward the user's right eye is shown in FIG. 3b, but the embodiment is not limited thereto, and the eye tracking camera (340-1) may be positioned solely toward the user's left eye or toward both eyes.
[0090] A shooting camera (340-3) can capture a real image or background to be matched with a virtual image in order to implement augmented reality or mixed reality content. The shooting camera can capture an image of a specific object located at the position where the user is looking and provide the image to at least one display (350). The at least one display (350) can display a single image in which information regarding a real image or background including the image of the specific object obtained using the shooting camera and a virtual image provided through optical devices (382, 384) are superimposed. In one embodiment, the shooting camera may be placed on a bridge (303) positioned between the first rim (301) and the second rim (302).
[0091] The eye tracking camera (340-1) can achieve more realistic augmented reality by tracking the gaze of a user wearing the electronic device (300), thereby matching the user's gaze with visual information provided to at least one display (350). For example, when the user looks straight ahead, the electronic device (300) can naturally display environmental information related to the user's front on at least one display (350) at the location where the user is situated. The eye tracking camera (340-1) may be configured to capture an image of the user's pupil to determine the user's gaze. For example, the eye tracking camera (340-1) may receive a gaze detection light reflected from the user's pupil and track the user's gaze based on the position and movement of the received gaze detection light. In one embodiment, the eye tracking camera (340-1) may be positioned at locations corresponding to the user's left and right eyes. For example, the eye-tracking camera (340-1) may be positioned within the first rim (301) and / or the second rim (302) to face the direction in which the user wearing the electronic device (300) is located.
[0092] The motion recognition camera (340-2) can provide a specific event on a screen provided on at least one display (350) by recognizing the movement of the user's entire body or part thereof, such as the user's torso, hands, or face. The motion recognition camera (340-2) can recognize the user's motion (e.g., gesture recognition), acquire a signal corresponding to said motion, and provide a display corresponding to said signal on at least one display (350). Accordingly, the electronic device (300) can perform a designated function based on identifying the signal corresponding to said motion. In one embodiment, the motion recognition camera (340-2) may be placed on the first rim (301) and / or the second rim (302).
[0093] According to one embodiment, an electronic device (300) can analyze an object included in a real-world image collected through a camera (340-1), combine a virtual object corresponding to an object among the analyzed objects that is the target of augmented reality provision, and display it on at least one display (350). The virtual object may include at least one of text and an image regarding various information related to the object included in the real-world image. The electronic device (300) can analyze the object based on a multi-camera such as a stereo camera. For the object analysis, the electronic device (300) can execute ToF and / or SLAM (simultaneous localization and mapping) supported by the multi-camera.
[0094] The camera (340) included in the electronic device (300) is not limited to the eye-tracking camera (340-1) and / or motion recognition camera (340-2) described above. For example, the electronic device (300) can identify external objects included within the field of view by using a shooting camera (340-3) positioned toward the user's field of view. The identification of external objects by the electronic device (300) can be performed based on a sensor for identifying the distance between the electronic device (300) and the external object, such as a depth sensor (e.g., depth sensor (149)), and / or a ToF sensor. The camera (340) positioned toward the field of view may support an autofocus function and / or an optical image stabilization (OIS) function. For example, the electronic device (300) may include a camera (340) (e.g., a face tracking camera) positioned toward the face to acquire an image including the face of a user wearing the electronic device (300).
[0095] The battery module (370) can supply power to the electronic components of the electronic device (300). In one embodiment, the battery module (370) may be placed within the first temple (304) and / or the second temple (305). For example, the battery module (370) may be a plurality of battery modules (370). The plurality of battery modules (370) may each be placed in the first temple (304) and the second temple (305). In one embodiment, the battery module (370) may be placed at the end of the first temple (304) and / or the second temple (305).
[0096] The antenna module (375) can transmit a signal or power to the outside of the electronic device (300) or receive a signal or power from the outside. The antenna module (375) can be electrically and / or operatively connected to a communication circuit (e.g., communication circuit (190)) within the electronic device (300). In one embodiment, the antenna module (375) may be placed within the first temple (304) and / or the second temple (305). For example, the antenna module (375) may be placed close to one side of the first temple (304) and / or the second temple (305).
[0097] Speakers (392-1, 392-2) (e.g., speaker (184)) can output an acoustic signal to the outside of the electronic device (300). In one embodiment, the speakers (392-1, 392-2) may be placed within a first temple (304) and / or a second temple (305) to be positioned adjacent to the ears of a user wearing the electronic device (300). For example, the electronic device (300) may include a second speaker (392-2) positioned adjacent to the user's left ear within the first temple (304), and a first speaker (392-1) positioned adjacent to the user's right ear within the second temple (305).
[0098] The electronic device (300) may include a printed circuit board (390) on a PCB. The PCB (390) may be included in at least one of a first temple (304) or a second temple (305). The PCB (390) may include an interposer disposed between at least two sub-PCBs. One or more hardware components included in the electronic device (300) may be disposed on the PCB (390). The electronic device (300) may include a flexible PCB (FPCB) for interconnecting the hardware components.
[0099] The electronic device (300) may further include configurations not illustrated in relation to FIGS. 3a and 3b. For example, the electronic device (300) may further include a light source (e.g., LED) that emits light toward a subject (e.g., user's eyes, face, and / or an object outside the field of view) being photographed using the camera (340). The light source may include an LED of infrared wavelength. The light source may be placed in at least one of the frame and hinge units (306, 307).
[0100] For example, the electronic device (300) can identify an external object touching the frame (e.g., a user's fingertip) and / or a gesture performed by said external object by using a touch sensor, a grip sensor, and / or a proximity sensor formed on at least a portion of the surface of the frame.
[0101] For example, the electronic device (300) may include at least one of a gyroscope sensor, a gravity sensor, and / or an acceleration sensor (e.g., sensor circuit (140)) for detecting the posture of the electronic device (300) and / or the posture of a body part (e.g., head) of a user wearing the electronic device (300).
[0102] FIG. 4a illustrates a rear perspective view of an electronic device according to one embodiment. FIG. 4b illustrates a front perspective view of an electronic device according to one embodiment. With reference to FIG. 4a and 4b, the electronic device (400) may correspond to an example of the electronic device (10) of FIG. 1.
[0103] Referring to FIG. 4a, according to one embodiment, a first surface (410) of an electronic device (400) (e.g., electronic device (10)) may have a wearable form on a part of a user's body (e.g., the face of the user). Although not illustrated, the electronic device (400) may further include a strap for securing to a part of a user's body and / or one or more temples (e.g., the first temple (304) and / or the second temple (305) of FIG. 3a and FIG. 3b). A first display (350-1) for outputting an image to the user's left eye and a second display (350-2) for outputting an image to the user's right eye may be disposed on the first surface (410). The electronic device (400) may further include a rubber or silicone packing formed on the first surface (410) to prevent interference by light different from light emitted from the first display (350-1) and the second display (350-2) (e.g., ambient light).
[0104] According to one embodiment, the electronic device (400) may include cameras (440-1, 440-2) (e.g., camera (170)) for photographing and / or tracking both eyes of an adjacent user. The cameras (440-1, 440-2) may be referred to as ET cameras. According to one embodiment, the electronic device (400) may include cameras (440-3, 440-4) (e.g., camera (170)) for photographing and / or recognizing the face of a user. The cameras (440-3, 440-4) may be referred to as FT cameras.
[0105] Referring to FIG. 4b, on a second surface (420) opposite to the first surface (410) of FIG. 4a, a camera (e.g., cameras (440-5, 440-6, 440-7, 440-8, 440-9, 440-10)) (e.g., camera (170)) and / or a sensor (e.g., depth sensor (430)) may be placed to obtain information related to the external environment of the electronic device (400). For example, cameras (440-5, 440-6, 440-7, 440-8, 440-9, 440-10) may be placed on the second surface (420) to recognize external objects different from the electronic device (400). For example, by using cameras (440-9, 440-10), the electronic device (400) can acquire images and / or media to be transmitted to each of the user's two eyes. Camera (440-9) may be placed on the second surface (420) of the electronic device (400) to acquire an image to be displayed through a second display (350-2) corresponding to the right eye among the two eyes. Camera (440-10) may be placed on the second surface (420) of the electronic device (400) to acquire an image to be displayed through a first display (350-1) corresponding to the left eye among the two eyes.
[0106] According to one embodiment, the electronic device (400) may include a depth sensor (430) (e.g., sensor circuit (140)) disposed on a second surface (420) to identify the distance between the electronic device (400) and an external object. By using the depth sensor (430), the electronic device (400) can obtain spatial information (e.g., depth map) for at least a portion of the field of vision of a user wearing the electronic device (400).
[0107] Although not shown, a microphone (e.g., a first microphone (182) and / or a second microphone (183)) for acquiring sound output from an external object may be placed on the second surface (420) of the electronic device (400). Depending on the embodiment, the number of microphones may be one or more.
[0108] As described above, an electronic device (400) according to one embodiment may have a form factor for being worn on a user's head. The electronic device (400) may provide a user experience based on augmented reality, virtual reality, and / or mixed reality while being worn on the head. Using a first display (350-1) and a second display (350-2) positioned toward each of the user's two eyes, the electronic device (400) may display a screen containing an external object. Within the screen, the electronic device (400) may register the external object or display the result of executing a function assigned to the external object.
[0109] FIG. 5a shows a front perspective view of a detachable electronic device according to one embodiment. FIG. 5b shows a rear perspective view of a detachable electronic device according to one embodiment.
[0110] Referring to FIGS. 5a and 5b, an electronic device (500) (e.g., electronic device (10)) according to one embodiment may be configured to be detachably attached to a housing (501). For example, the electronic device (500) may be inserted into a slot (510) of the housing (501) or may be physically coupled to the housing (501) by any structure of the housing (501).
[0111] When combined, the display (550) (e.g., display (160)) of the electronic device (500) can be viewed through a left eye lens (551-1) and a right eye lens (551-2). The electronic device (500) can output a left eye image to an area of the display (550) corresponding to the left eye lens (551-1) and output a right eye image to an area of the display (550) corresponding to the right eye lens (551-2).
[0112] Although not illustrated, the housing (501) may further include a strap for securing to a part of the user's body and / or one or more temples (e.g., a first temple (304) and / or a second temple (305) in FIG. 3a and 3b). The housing (501) may further include a rubber or silicone packing to prevent interference by ambient light.
[0113] According to one embodiment, the electronic device (500) may include a camera (540) (e.g., camera (170)) and / or a sensor (e.g., depth sensor (149)) for acquiring information related to the external environment of the electronic device (500). The housing (501) may include a camera hole (541) for the camera (540) to be exposed to the outside when combined with the electronic device (500). For example, the electronic device (500) may acquire spatial information (e.g., depth map) for at least a portion of the field of view of a user wearing the electronic device (500) using a plurality of cameras.
[0114] Although not shown, the electronic device (500) may include a microphone (e.g., a first microphone (182) and / or a second microphone (183)) for acquiring sound output from an external object.
[0115] The housing (501) may provide a structure for physical coupling with the electronic device (500) and for user wearing. In one example, the housing (501) may be electrically coupled with the electronic device (500). In this case, the housing (501) may include at least some of the components of the electronic device (10) described above in relation to FIG. 2. The electronic device (500) may receive information obtained by the components of the housing (501) through communication with the housing (501) or output information using the components of the housing (501).
[0116] FIG. 6 illustrates examples of electronic devices according to an embodiment.
[0117] With reference to FIG. 6, the functions of the electronic device (10) described above in relation to FIG. 2 may be implemented through a plurality of electronic devices. In the examples of FIG. 3a through FIG. 5b, the electronic device (300, 400, 500) is described as a head-mounted device (HMD), but the embodiments of the present disclosure are not limited thereto. In the example of FIG. 6, a user (601) may use a plurality of electronic devices connected to communicate with each other.
[0118] For example, visual information may be acquired by a camera of a wearable device (610) (e.g., artificial intelligence (AI)), a mobile device (620), and / or smart glasses (600). For example, auditory information may be acquired by a microphone of a wearable device smart glasses (600), a wearable device (610), a mobile device (620), and / or an audio device (630). For example, visual information may be provided through a display of a device having a display (e.g., smart glasses (600) and / or a mobile device (620)). For example, auditory information may be provided through a device having a speaker (e.g., smart glasses (600), a wearable device (610), a mobile device (620), and / or an audio device (630)).
[0119] The “electronic device” of the present disclosure may include a device capable of providing an extended reality to a user (1), matching a voice object indicated by the user (1)’s speech with a scene object provided through or by the electronic device, and supporting an AI assistance function based on the matching result. Additionally, or generally, if there are multiple scene objects configured to support an AI assistance function, the electronic device may integrate an AI assistance function mapped to each scene object in correspondence with the relationship between the scene objects (e.g., at least one relationship among a placement relationship or a semantic relationship) and support the integrated AI assistance function.
[0120] The examples described above in connection with FIGS. 3a through 6 are examples of forms of electronic devices and are for illustrative purposes only and are not intended to limit the embodiments of the present disclosure described below. A person skilled in the art will understand that, in addition to the electronic devices in the examples described above in connection with FIGS. 3a through 5b, various forms of electronic devices may perform at least one of the embodiments of the present disclosure described below. Furthermore, a person skilled in the art will understand that, as described above in connection with FIG. 6, a plurality of electronic devices may perform the embodiments of the present disclosure described below through communication.
[0121] Hereinafter, with reference to various drawings, an electronic device (10) and a method of operation (or management) thereof that provide at least one embodiment related to the utilization of a composite AI assistance function (or assistant) will be described. In the following description, the electronic device 10 is described as the basis, but the electronic device 10 may be replaced with or operated in combination with at least one of the electronic devices 300, 400, 500 and 501, 600, 610, 620 and 630 described in other drawings to support the embodiment described below.
[0122] FIG. 7 is a diagram showing an example of a processor configuration of an electronic device according to one embodiment. The processor (120) configuration of the electronic device (10) described in FIG. 7 may be configured so that the electronic device performs an operation according to at least one of the various embodiments of the present invention by executing at least one program module (e.g., a software module, an application, an API (application programming interface), or a plurality of instructions) stored in memory (130).
[0123] Referring to FIGS. 1 to 7, a processor (120) of an electronic device (10) according to one embodiment may include a recognition unit (121) and an analysis unit (122). The recognition unit (121) and the analysis unit (122) may be composed of at least one software-type program module as described above, but the present description is not limited thereto. For example, at least some of the recognition unit (121) and / or the analysis unit (122) may be composed of a hardware processor.
[0124] According to one embodiment, the recognition unit (121) may include a configuration for recognizing the surrounding conditions of the electronic device (10). The recognition unit (121) may be composed of at least one module. For example, the recognition unit (121) may include a voice recognition module (121a), an object recognition module (121b), a user input processing module (121c), a gaze processing module (121d), and / or a gesture recognition module (121e). Alternatively, at least one of the modules included in the recognition unit (121) may be implemented as a single integrated module. For example, the voice recognition module (121a) and the object recognition module (121b) may be implemented as a single integrated recognition module, and the user input processing module (121c) may be implemented as a single integrated processing module including the gaze processing module (121d) and the gesture recognition module (121e). In this way, the configuration of the recognition unit (121) can be implemented by flexibly differentiating into various modules or by implementing an integrated module depending on the electronic device (10).
[0125] According to one embodiment, the voice recognition module (121a) can collect audio signals generated around the electronic device (10) and detect the speech portion of the user (1) in the collected audio signals. For example, within a detectable range of the microphone of the electronic device (10) (e.g., at least one of the first microphone (182) and / or the second microphone (183) of FIG. 2), the electronic device (10) can detect the human voice portion and perform voice recognition on the detected speech portion of the user (1) to convert it into text. For example, the voice recognition module (121a) can be activated when the electronic device (10) is worn at a designated location of the user (1). The voice recognition module (121a) can activate the microphone in relation to the collection of speech of the user (1) and perform voice recognition on the collected speech portion of the user (1) using the microphone. The voice recognition module (121a) can transmit the voice recognition result to the persona management module (122c). As another example, the voice recognition module (121a) may be activated when a designated voice command is collected. A database related to voice recognition may be stored in memory (130) or stored in an external server device and referenced through the communication circuit (190) of the electronic device (10). Additionally or generally, the voice recognition module (121a) may collect ambient sounds occurring around the electronic device (10) (e.g., within a certain distance from the electronic device (10)) and recognize the surrounding situation by recognizing the collected ambient sounds. For example, if the ambient situation information (or surrounding environment information) recognized by the voice recognition module (121a) corresponds to set information (e.g., a clapping sound, a sound corresponding to a set beat, or a set music sound), the ambient situation information may be transmitted to the persona management module (122c) as a user input.
[0126] According to one embodiment, the object recognition module (121b) can recognize scene objects in a scene (e.g., a captured image, or a camera image, a specific video or still image) obtained through the electronic device (10). For example, the object recognition module (121b) can distinguish and recognize a background object and at least one scene object placed on the background object for a scene (or video) obtained through the electronic device (10). For example, the object recognition module (121b) can extract at least one scene object based on the boundary line of at least one scene object placed on the background object, and detect attributes of the scene object (e.g., feature points and / or type of scene object) based on at least one of the shape, size, pattern, and color of the extracted at least one scene object. In this regard, the object recognition module (121b) can acquire a captured image (or a corresponding scene) for the front (or the direction in which the camera of the electronic device (10) is positioned to take an image) using a camera (e.g., camera (170) of FIG. 2) included in the electronic device (10), and can detect at least one scene object by performing an analysis on the acquired captured image (or a corresponding scene).
[0127] According to one embodiment, the user input processing module (121c) can process explicit user input. For example, the user input processing module (121c) can perform processing on user input received through an interface of the electronic device (10) (e.g., the interface (180) of FIG. 2). User input processed by the user input processing module (121c) can be transmitted to the persona management module (122c).
[0128] According to one embodiment, the gaze processing module (121d) can track and detect the gaze of the user (1). In this regard, the gaze processing module (121d) may operate at least one camera capable of capturing the user's (1) eyes (or pupils). For example, when the user (1) is wearing the electronic device (10), the gaze processing module (121d) can activate a camera capable of capturing the user's (1) eyes and determine which direction the user's (1) gaze is directed based on image analysis obtained through the camera. Additionally, or generally, the gaze processing module (121d) may detect a scene object (e.g., an object that overlaps at least partially or is directed at at least partially around the object) that corresponds to the user's (1) gaze and the scene obtained by the camera facing forward of the electronic device (10). For example, the gaze processing module (121d) can detect a scene object that overlaps at least a portion of the user's (1) gaze direction within the scene acquired by the camera, or a scene object located within a certain angle range based on the user's (1) gaze direction. Alternatively, the gaze processing module (121d) can detect a scene object located within a certain distance based on the user's (1) gaze direction within the scene acquired by the camera.
[0129] In relation to gaze processing and / or detection of scene objects in the line of sight, the electronic device (10) may include at least two cameras (e.g., a camera that captures the eyes for user gaze detection, and a camera that captures the direction in which scene objects are located).
[0130] According to one embodiment, the gesture recognition module (121e) can track and detect the gestures of the user (1). In this regard, the gesture recognition module (121e) can analyze images acquired by a camera (e.g., camera (170) of FIG. 2) capable of capturing at least a part of the user's (1) body (or a specific object that can be captured by the electronic device (10) while the electronic device (10) is being worn and can be recognized as a gesture). As an example, the gesture recognition module (121e) can recognize gestures based on the user's (1) hand or foot, or other parts of the user's (1) body (or a corresponding model or object). In this regard, the gesture recognition module (121e) can store a database necessary for gesture recognition in memory (130) and check whether a body part related to the user's (1) gesture is detected in the analysis of the images acquired by the camera. The gesture recognition module (121e) can recognize the occurrence of a gesture based on image comparison or feature point comparison of objects when the user's (1) body takes a motion corresponding to a specific gesture stored in the database.
[0131] According to one embodiment, the analysis unit (122) can determine the surrounding conditions of a user (1) wearing an electronic device (10) by performing an analysis on various information collected by the recognition unit (121). The analysis unit (122) may be composed of at least one module. For example, the analysis unit (122) may include an object ID recognition module (122a), an object grouping module (122b), a persona management module (122c), and / or a similar object management module (122d). Here, at least some of the modules constituting the analysis unit (122) may be integrated with other modules to be implemented as a single integrated module. For example, the object ID recognition module (122a) and the object grouping module (122b) may be implemented as a single integrated object processing module, and the persona management module (122c) and the similar object management module (122d) may be implemented as a single integrated management module. The configuration of the above analysis unit (122) can be flexibly changed depending on the type or hardware performance of the electronic device (10).
[0132] According to one embodiment, the object ID recognition module (122a) can recognize object IDs for scene objects recognized by the object recognition module (121b). In this regard, the object recognition module (121b) can acquire a captured image (or scene, scene image, camera image) for one direction of the electronic device (10), and analyze the scene corresponding to the acquired captured image to determine whether a set scene object is present in the current scene (or the image currently being captured). In relation to checking whether a set scene object exists in the current scene, the object ID recognition module (122a) can search a scene object database stored in memory (130). The scene object database may include attributes of at least one scene object to which an AI assistance function is mapped (e.g., feature points and / or type information), and information about AI assistance functions mapped to at least one scene object. The object ID recognition module (122a) can identify attributes (e.g., feature points) for a scene object detected in the current scene, search for a previously stored scene object that matches the scene object database using the identified attributes, and determine a stored scene object that is identical or similar to a set threshold value or higher as a scene object included in the current scene. The object ID recognition module (122a) can identify ID information mapped to the determined scene object by referring to the scene object database stored in memory (130) and determine which AI assistance function is mapped.
[0133] According to one embodiment, the object grouping module (122b) can perform object grouping when there are multiple scenes with at least one scene object detected by the object ID recognition module (122a). For example, the object grouping module (122b) can receive information regarding scene objects included in the current scene from the object ID recognition module (122a). When multiple scene objects are included in the current scene, the object grouping module (122b) can check whether there are scene objects among the multiple scene objects that satisfy a specified condition (or a set condition). For example, the object grouping module (122b) can group multiple scene objects when the separation distance between multiple scene objects is within a set distance. Alternatively, the object grouping module (122b) can group multiple scene objects when multiple scene objects are placed together within another set scene object. Alternatively, the object grouping module (122b) can group multiple scene objects included in a set pattern when multiple scene objects are placed in a set pattern. Alternatively, the object grouping module (122b) may group multiple scene objects placed within a scene if their meaning (or interrelationship) is greater than or equal to a set threshold value. According to one embodiment, the object grouping module (122b) may group multiple scene objects included in a single angle of view of a camera based on the user's (1) gaze information, multiple scene objects located within a certain distance based on the user's (1) gaze information, multiple scene objects included within a certain angle range based on the user's (1) gaze information, and / or multiple scene objects located in a specific space distinct from the surroundings.According to one embodiment, the object grouping module (122b) can group multiple scene objects when the current user (1) is located at a designated location (e.g., the user (1)'s house, office, or a designated specific location) in addition to at least one of the various conditions described above (e.g., multiple scene objects are within a certain distance, within a single angle of view, within a certain range relative to the user's line of sight, or multiple scene objects are located in a specific space).
[0134] For example, the object grouping module (122b) can group scene objects corresponding to a book and a notebook, scene objects corresponding to a notebook and a writing tool, scene objects corresponding to a flowerpot and a watering can, and scene objects corresponding to a board and a marker when the correlation is greater than or equal to a set threshold value. The score related to the correlation may be assigned a higher score the closer the object is to a specific field or topic based on the specific field or topic associated with each scene object.
[0135] For example, if there are multiple scene objects, the object grouping module (122b) can group each of the multiple scene objects selected by the user (1). The object grouping module (122b) can support the selection of scene objects within a certain distance in real space. Alternatively, the object grouping module (122b) can group multiple scene objects selected by the user (1) regardless of distance. The object grouping module (122b) can group at least some of the multiple scene objects according to the user's preference stored in memory. For example, the object grouping module (122b) can check the characteristics of the grouped scene objects selected by the previous user, automatically group objects having the same or similar characteristics, or recommend grouping and group the recommended scene objects based on user confirmation. The grouped objects can be activated when one or more of the grouped scene objects are activated, or they can be activated according to the activation method (or method, input for activation) of the grouped scene objects set by the user (1).
[0136] According to one embodiment, the persona management module (122c) may perform at least one of the following operations: mapping a persona and an AI assistance function to at least one of the scene objects; unmapping the persona and AI assistance function mapped to the scene objects; changing the persona and AI assistance function of the scene objects; unintegrating the grouped scene objects; inheritance of the persona mapped to the scene object or the integrated persona mapped to the scene objects; and activating the AI assistance function in response to a user utterance activating a persona set on a specific scene object.
[0137] According to one embodiment, the persona may be mapped to a scene object included in the scene in response to user input. The persona may be mapped (or connected) to an AI assistant function that generates a response to a query. Accordingly, the state in which the persona is activated may include a state in which the AI assistant function mapped to the persona is prepared to generate a response. The type of the AI assistant function may be defined (or determined, described) by the persona. The AI assistant function may generate a response based on an information database of a specific field. The information database of the specific field may be stored in a specific memory device or server device, and the AI assistant function may access the information database through a domain that defines the access path to the specific memory device or server device, and generate a response corresponding to the query based on the information database of the specific field.
[0138] According to one embodiment, as described above, a scene object is mapped to a persona, and the persona is mapped (or connected) to an AI assistance function, and the AI assistance function may include a structure configured to allow a specific AI module to use a specific information database. The specific field may be determined in correspondence with the definition of the persona. For example, the persona may function as a command capable of calling an AI assistance function configured to provide a response corresponding to a query by a user (1) based on at least some of the various information databases, such as academic fields (e.g., economics, society, ethics, politics, history, physical education, …), hobby fields, special skills fields, people, and places.
[0139] According to one embodiment, the persona may be defined by terms or names that the user (1) can intuitively understand or recognize as a specific information database field. Alternatively, the persona may be constructed based on at least some keywords included in the user's (1) mapping query (e.g., a query uttered by the user to map AI assistance functions of a specific field to a scene object) during the process of mapping to a scene object. The name of such a persona may be changed in response to a change request from the user (1). Alternatively, at least some of the name defining the persona may be changed in response to a change in the user's (1) settings, and the specific field may be changed in response to the change in the name defining the persona.
[0140] For example, if a persona is defined by the term “plant,” the persona can be mapped (or connected) to an AI assistant function that generates responses based on an information database related to the plant field. If a persona is defined by the term “botanist,” a combination of “plant” and “scholar,” the persona can be mapped to an AI assistant function that generates responses based on an information database containing detailed information about the plant field at the scholar level. Alternatively, the persona can be a term (or command, function executor) that activates a specific knowledge field domain to provide a response to a query based on a database (or DB server device) of that specific knowledge field domain (e.g., the plant domain). As described above, when the AI assistant function generates a response to a query, the accessible information databases may vary depending on the change in the term defining the persona.
[0141] The persona management module (122c) may map the detected scene object to a designated persona and a designated AI assistant function, and perform numbering (e.g., assigning a number) or naming (e.g., assigning a name) for the scene object when the scene object detected by the object recognition module (121b) and the user input transmitted by the user input processing module (121c) satisfy the set conditions. The user input may include at least one of input signals, such as voice commands, gaze, gestures, and / or other input means. The persona management module (122c) may release the mapping (or assignment) when a user input satisfying the set conditions occurs for a scene object to which a specific persona and AI assistant function are mapped (or assigned). Alternatively, the persona management module (122c) may control the change of the type of AI assistant function or the change of the type of scene object when a user input satisfying the set conditions occurs for a scene object to which a specific persona and AI assistant function are mapped.
[0142] According to one embodiment, the persona management module (122c) may map (or assign) an integrated persona to scene objects grouped by the object grouping module (122b) and store and / or manage mapping information. According to one embodiment, the persona management module (122c) may perform unmapping processing on the integrated persona mapped to the scene objects when a setting condition related to ungrouping is satisfied for the grouped scene objects (e.g., when at least one scene object among the scene objects is moved and the distance between the scene objects exceeds a reference value). At least one of the following operations will be described through the following embodiments: the mapping operation of the scene object, the persona, and the AI assistance function; the unmapping operation of the persona and the AI assistance function mapped to the scene object; the modification operation of the persona and the AI assistance function for the scene object; the integration of the personas and / or the unintegration operation of the integrated persona; and the inheritance operation of the persona.
[0143] The similar object management module (122d) can detect attributes (e.g., feature points and / or types) of a scene object to which a persona and AI assistance function managed by the persona management module (122c) is mapped, and can support management of copying or inheriting the persona and AI assistance function mapped to the scene object to other scene objects that have the same attributes (feature points and / or types) as the scene object or have similarity within a set range. For example, when the object recognition module (121b) detects a scene object, the similar object management module (122d) can search the memory (130) (or a designated external server device) to see if there is a stored scene object with attributes (e.g., feature points or the same type) that are the same or similar to the detected scene object. When a stored scene object of the same or similar type is found, the similar object management module (122d) can control the application of the persona and AI assistance function mapped to the previously stored scene object to the currently detected scene object. In this process, the similar object management module (122d) can check with the user (1) whether to apply the same persona and AI assistance function mapped to the previously saved scene object, and control the application of the persona and AI assistance function to the currently detected scene object through the user's (1) confirmation.
[0144] Additionally or generally, the electronic device (10) may further include an LLM (123) (large language module). Alternatively, the electronic device (10) may be configured to be connected via a communication circuit (190) to an external server device including the LLM (123). Alternatively, the electronic device (10) may further include other artificial intelligence (or artificial intelligence model) in addition to the LLM (123). The other artificial intelligence (or artificial intelligence model) may include at least one artificial intelligence (or artificial intelligence model) that can be used in the process of video / image analysis, object recognition, or gesture recognition. For example, the other artificial intelligence may utilize an artificial intelligence model (e.g., at least one of a CNN (Convolutional Neural Network), GAN (Generative Adversarial Network), or VAE (Variational Auto Encoder)) capable of analyzing an image (or an image captured through a camera) in real time and detecting, recognizing, separating, or generating specific objects within the image in real time.
[0145] The above LLM (123) can understand various queries and support corresponding command processing in relation to at least one of the mapping of personas according to the embodiment of the present description, utilization of AI assistance functions through personas, integration of personas (or integration of AI assistance functions mapped to personas), and change or deletion of personas and AI assistance functions.
[0146] As described above, the electronic device (10) according to the embodiment of the present invention detects at least one scene object in a scene collected by a camera and can map a persona and an AI assistance function to the at least one scene object in response to a set user input. In relation to the mapping of the persona and AI assistance function, the electronic device (10) may include a learned AI module capable of supporting the AI assistance function. Alternatively, the electronic device (10) may form a communication channel with an external server device in which the learned AI module is stored in relation to the support of the AI assistance function, and may support the AI assistance function through the external server device. Furthermore, the electronic device (10) may further support at least one of a mapping release operation and a mapping change operation for the persona and AI assistance function, and may specifically provide more useful information to the user by integrating multiple personas or multiple AI assistance functions in response to the relationships of the scene objects.
[0147] FIG. 8 is a diagram showing an example of persona operation according to one embodiment.
[0148] Referring to FIGS. 1 through 8, according to one embodiment, an electronic device (10) can check whether a persona designation condition is satisfied as in operation 801. Here, the persona may be a term that calls (or activates) a connected AI assistant function. The AI assistant function may include an AI module capable of generating a response to a query based on an information database of a specific field defined by the persona.
[0149] When the persona designation condition is satisfied, the electronic device (10) in operation 803 may designate a persona to a designated scene object. Alternatively, the electronic device (10) may map an AI assistance function related to user input as a persona to a scene object (or designated scene object) designated by the user (1). The persona includes a command that can call an AI assistance function, or a term that allows the user to activate a function set on a specific scene object, and the function set on the scene object may include a function that can activate at least one of a specific knowledge field domain (or a DB server device defined as a domain) and an AI assistance function that provides a response to the user's query through the specific knowledge field domain.
[0150] In operation 805, the electronic device (10) may store information about the state in which a persona is mapped to a designated scene object in memory (e.g., memory (130) of FIG. 2). At this time, the electronic device (10) may store at least some of the location information of the scene object (or location and / or orientation information of the scene object in the scene, location information of the scene object acquired), and the time information of the persona mapping, along with the storage of information about the designated scene object and the persona mapping state (or the state in which a persona is mapped to a scene object).
[0151] In operation 807, the electronic device (10) can perform a persona calling method designation. In this regard, the electronic device (10) can define a user input (e.g., voice input, touch or virtual touch, or instruction via finger or gaze) that can direct or select a scene object to which the persona is mapped. The operation of defining the type of user input may be performed during the process of the user (1) mapping the persona to the scene object or may be changed through a setting.
[0152] According to one embodiment, when a user (1) possessing or wearing an electronic device (10) generates user input for activating a scene object to which a persona is mapped, the electronic device (10) can activate the persona in response to the user input, as in operation 809. For example, the electronic device (10) can activate an AI module providing an AI assistant function stored in memory (130) in relation to the activation of an AI assistant function connected to the persona mapped to the scene object and the persona. Alternatively, the electronic device (10) can form a communication channel with an external server device in which an AI module providing an AI assistant function is stored and request the external server device to activate the AI assistant function. According to various embodiments, the electronic device (10) can recognize at least one of the user's gaze information, the user's hand movement information, and / or the physical arrangement of the scene object, and activate an AI assistant function (AI assistant) according to the recognized situation.
[0153] In operation 811, the electronic device (10) can perform persona grouping (or grouping of scene objects, or grouping of AI assistance functions connected to personas, or grouping of information databases accessed by AI assistance functions) if there are multiple scene objects included in the current scene and the multiple scene objects are in a relationship that satisfies a set condition (e.g., a relationship where the distance between multiple scene objects (e.g., physical distance between scene objects in reality or distance between scene objects in the scene) is placed within a specified distance). With respect to distance detection, the electronic device (10) can perform distance detection based on at least one of the distance from the scene shooting location to the scene objects (e.g., depth using a depth camera), the focal length of the camera that captured the scene objects (e.g., focal length of an RGB camera), and / or the distance between scene objects obtained in the scene.
[0154] According to one embodiment, the process related to persona grouping (or grouping of scene objects) may be recommended by the electronic device (10) (e.g., outputting text or an image related to the recommendation to the scene) and approved by a user (1) confirmation action (e.g., approved by user input for recommendation approval). When an integrated persona is created in response to persona grouping, the type of user input for calling the integrated persona (or grouped persona, hereinafter described as integrated persona) may be specified through an action that specifies the method of calling the integrated persona (or grouped persona, hereinafter described as integrated persona) in action 807. In this regard, the electronic device (10) may provide the user (1) with a screen interface (or an augmented reality-based scene) capable of defining the type of user input related to calling the integrated persona, and may determine the type of user input according to the user's selection. Alternatively, as another example, at least one of the calling methods of the personas prior to integration may be inherited as the integrated persona calling method. In this regard, the electronic device (10) can output information regarding the calling method of the personas prior to integration to the scene so that the user (1) can select it. The integrated persona can manage the AI assistance functions of the grouped personas in an integrated manner. For example, an integrated persona can be created by integrating a first scene object mapped to a first persona with a first AI assistance function set to generate a response by referencing a first information database, and a second scene object mapped to a second persona with a second AI assistance function set to generate a response by referencing a second information database. The integrated persona may be mapped to a third AI assistance function that manages the first domain related to the first information database and the second domain related to the second information database in an integrated manner, and generates a response to a user query based on at least one of the first domain and the second domain (or using at least one of the first information database and the second information database).
[0155] According to one embodiment, when a persona is activated (or when an AI assistant function is prepared to be used with the persona connected to the persona), the user (1) inputs a query (or user query, user utterance, user text input, user image input) related to the use of the AI assistant function connected to the persona, and the electronic device (10) can collect the query in an 813 operation. In relation to the input of the query, the electronic device (10) can activate a microphone, collect user utterance, and perform voice recognition on the utterance to collect the content of the query.
[0156] FIG. 9 is a diagram showing an example of a scene object and persona mapping method according to one embodiment.
[0157] Referring to FIGS. 1 through 9, an electronic device (10) according to one embodiment is activated in response to user (1) operation or in response to set scheduling information and can collect various surrounding information. For example, the electronic device (10) can collect a scene (900a) (or a captured image, camera image, image signal, a still image or video stored in memory (130), or a scene image), and can collect user input (900b) (user input signal) based on at least some of various input means (e.g., input devices such as a keyboard, mouse, touchscreen, jog shuttle, jog stick, or physical button). Additionally, the electronic device (10) can collect surrounding audio signals (900c) (audio signal) using a microphone.
[0158] In operation 901, the electronic device (10) can perform object recognition on the collected scene (900a) to detect at least one scene object. The electronic device (10) can generate an object list for at least one scene object identified as a result of the object recognition.
[0159] In operation 902, the electronic device (10) can perform gesture recognition from the scene (900a). In this regard, the electronic device (10) can check whether the scene (900a) contains a scene object set (e.g., a scene object related to a gesture). For example, the electronic device (10) can check whether the scene (900a) contains a hand, finger, foot, or corresponding object corresponding to the set gesture. If the scene (900a) contains a hand, the electronic device (10) can generate a hand tracking result.
[0160] In operation 903, the electronic device (10) can perform gaze processing. In this regard, the electronic device (10) can check whether there is a scene object containing a user's eye related to gaze processing in the scene (900a). If the scene (900a) contains a scene object containing an eye, the electronic device (10) can generate a gaze tracking result by tracking the gaze based on the direction indicated by the pupil in the scene object corresponding to the eye.
[0161] In operation 904, when an audio signal (900c) is collected, the electronic device (10) can perform voice recognition on the collected audio signal (900c). In this regard, the electronic device (10) can activate a voice recognition module (121a) for voice recognition on the audio signal (900c) and generate a voice recognition result (e.g., text query).
[0162] In operation 905, the electronic device (10) can perform processing of user input. For example, the electronic device (10) can collect at least one of gesture information through operation 902, user gaze information through operation 903, text information corresponding to user utterance through operation 904, and user input through at least some of various input means. Based on the collected information or user input, the electronic device (10) can generate a user's pointing location on an image coordinate system acquired by a camera (e.g., camera (170) in FIG. 2) and a user query (e.g., a mapping query requesting persona assignment for a scene object).
[0163] In operation 906, the electronic device (10) can define an object ID and / or category for a scene object referred to by the user based on an object list of scene objects obtained through operation 901 and a user's point location and / or user query collected through operation 905, and can map a persona based on the user query. According to one embodiment, in the process of operation 906, the electronic device (10) can finally determine which scene object the user is interacting with by combining the location information of objects recognized in the object recognition operation (operation 901) and the user input processed in the user input processing operation (operation 905).
[0164] In operation 907, the persona management module (122c) of the electronic device (10) can map a specific persona to a scene object based on at least some of the user point location, object ID (or category), and user query. The object ID is intended to distinguish a scene object recognized during the object recognition process from other scene objects and may be determined by user input or randomly generated by the electronic device (10). The category (or category of a scene object, or type of a scene object) is an attribute of a scene object (e.g., feature point or type) and can be used for persona inheritance (or inheritance of AI assistance functions). According to one embodiment, regarding the use of a scene object category, the electronic device (10) can check the category of a scene object designated by the user and check if there is a previously mapped persona based on that category. If there is a persona mapped to the category of the designated scene object, the electronic device (10) can inherit (or assign, assign, set) the persona mapped to the category as the persona of the scene object designated by the user. The above user query can be used to determine the type of persona. For example, the electronic device (10) can determine the name of the persona (or calling term, calling command) and the type of AI assistance function connected to the persona through the above user query (or at least some of the keywords included in the user query). The persona management module (122c) of the electronic device (10) can store and manage the persona in memory (130) when the persona is mapped to a specific scene object. According to one embodiment, when the electronic device (10) obtains a user query requesting the designation of a persona for a scene object, it can identify the scene object designated by the user through an object ID recognition operation (906 operation), and assign the persona requested in the user query to the designated scene object and store it in memory (130).As an example, the electronic device (10) can store at least part of the user query, scene objects, and personas in memory (130).
[0165] FIG. 10 is a diagram showing an example of a method of using a persona according to one embodiment.
[0166] Referring to FIGS. 1 to 10, in relation to a persona usage method according to one embodiment, an electronic device (10) can activate a camera in response to a request by a user (1) or in response to an operation by a user (1) and collect a scene regarding the shooting direction. The scene may include a still image or a video (or a preview image).
[0167] When a scene is acquired, the electronic device (10) (e.g., the object recognition module (121b) of FIG. 7) can perform object recognition in operation 1001. The electronic device (10) can perform image recognition (or image recognition) of the scene through artificial intelligence to recognize, track, and / or distinguish at least one of scene objects included in the scene. For example, regarding object recognition, the electronic device (10) can distinguish the background and at least one scene object placed in the background for the scene, and classify the scene objects based on the boundary line of at least one scene object. Additionally or generally, the electronic device (10) can extract feature points of the scene objects and detect object information matching the extracted feature points from an object information database (or an external server device holding the object information database) already stored in memory (130) to determine the type of scene object (or type of scene object, name of scene object). When object recognition for a scene is completed, the electronic device (10) can store an object list (e.g., an object list of multiple scene objects) in memory (130).
[0168] In operation 1003, the electronic device (10) (e.g., the object ID recognition module (122a) of FIG. 7) can recognize the IDs of scene objects obtained during the object recognition process. In this regard, the electronic device (10) can determine whether at least one scene object mapped to a pre-registered persona matches a scene object detected in the current scene. In this regard, the electronic device (10) can check the ID of the scene object detected in the scene and determine whether there is a scene object that matches the object ID of the scene objects mapped to the persona stored in memory (130). The ID of the scene object detected in the current scene can be assigned based on the feature points and / or type (or attribute) of the scene object. In this regard, if the feature points and / or types (or attributes) of a scene object detected in the current scene are the same (or similar) as the feature points and / or types (or attributes) of a scene object to which a persona is mapped in another scene, the electronic device (10) may assign an ID to the scene object detected in the current scene that is the same as or at least partially similar to the ID assigned to the scene object to which a persona is mapped in another scene.
[0169] In operation 1005, an electronic device (10) (e.g., the object grouping module (122b) of FIG. 7) can perform persona grouping when there are multiple scene objects in a scene and the multiple scene objects satisfy a set relationship. In this regard, the electronic device (10) checks whether there are multiple scene objects to which personas are mapped, and if the scene objects to which each persona is mapped satisfy a set relationship, it can group each persona and integrate the grouped personas. During the persona grouping process, the electronic device (10) can be controlled to integrate AI assistance functions. For example, the electronic device (10) can be configured to integrate the information DB of the AI assistance functions to which the personas are connected and to generate a response through the integrated information DB.
[0170] In action 1007, the electronic device (10) (e.g., the persona management module (122c) of FIG. 7) can determine whether to activate a persona. For example, if the electronic device (10) has a grouped integrated persona, in action 1009, it can activate the integrated persona (or integrated AI assistant function) corresponding to the grouped persona. Or, if there is no grouped integrated persona, in action 1009, the electronic device (10) can generate a response based on the AI assistant function connected to the persona corresponding to the selected scene object or directed by user input. According to one embodiment, an electronic device (10) (e.g., the persona management module (122c) of FIG. 7) may store information regarding persona changes in memory (e.g., the memory (130) of FIG. 2). For example, when a persona is integrated and an integrated persona is created, the electronic device (10) may store which personas the created integrated persona was integrated with, when and where, and by which activation command it was defined. Additionally, when an integrated persona is released, the electronic device (10) may separate the released integrated personas into their respective personas and store at least one of the time and location of the release of the integrated persona.
[0171] When one or more scene objects mapped to a persona are detected in the image collected by the camera, the electronic device (10) can determine whether they can be locally grouped by inferring the physical distance between adjacent designated scene objects mapped to other personas. The electronic device (10) can group the scene objects mapped to the grouped persona (or integrated persona) so that they are simultaneously activated or deactivated. In this regard, the electronic device (10) may include a persona-designated scene object activation determination module. The electronic device (10) can identify the scene object that the user (1) is interacting with by using a user-designated object recognition module. For example, the user-designated object recognition module can determine which scene object the user (1) intends to interact with by synthesizing user inputs such as the user's (1) gaze movements, hand movements, and user (1) voice input. A scene object mapped to a persona can be activated by the user's (1) gaze (or user designation). Alternatively, the scene object mapped to the persona may be activated by a different type of interaction other than gazing, depending on the user's (1) settings.
[0172] FIG. 11 is a diagram showing an example of a persona change method according to one embodiment.
[0173] Referring to FIGS. 1 through 11, in operation 1101, the electronic device (10) (e.g., the user input processing module (121c) of FIG. 7) performs processing for user input, and in operation 1103, the electronic device (10) (e.g., the persona management module (122c) of FIG. 7) may perform persona change processing. For example, the electronic device (10) may receive user input related to persona change in a user interface (or augmented reality) context related to persona change. For example, in operations 1101 and 1103, the electronic device (10) may receive user input requesting a change to a persona mapped to a specific scene object (e.g., user input for scene object selection and / or user query (or utterance) requesting a persona change).
[0174] In operation 1105, the electronic device (10) (e.g., the persona management module (122c) of FIG. 7) can reconfigure a prompt to be delivered to an AI assistant function related to the persona when a persona change for a scene object is requested. For example, the electronic device (10) can configure a prompt for a persona of a different characteristic related to the scene object when a change request is made to map a persona of a different characteristic to the scene object. The electronic device (10) can provide the reconfigured prompt in response to the persona change request to an agent managing the AI assistant function.
[0175] An agent managing AI assistance functions may change AI assistance functions based on a change in the persona mapped to the scene object according to the prompt reconfiguration in operation 1107. For example, when a first persona supporting a first AI assistance function is mapped to a first scene object, and a request is made to change the first persona to a second persona supporting a second AI assistance function, the electronic device (10) (e.g., the persona management module (122c) of FIG. 7) may process the prompt reconfiguration and / or mapping of the second AI assistance function related to the use of the second AI assistance function while mapping the second persona supporting the second AI assistance function to the first scene object.
[0176] When the management (or change operation) of the AI assistance function is completed, the electronic device (10) (e.g., the persona management module (122c) of FIG. 7) can activate the AI assistance function in operation 1109. As another example, the electronic device (10) may include a first additional processor related to the activation of the AI assistance function and a second additional processor (e.g., a persona activation support module) that supports the activated AI assistance function. Subsequently, when the user (1) generates a specific user query while specifying a scene object, the electronic device (10) transmits the user query to the changed AI assistance function based on the scene object to which the changed persona is mapped, and the changed AI assistance function can generate a response to the user query and output it in operation 1111. The output operation for the above response may be output via a screen on a display included in the electronic device (10) (e.g., the display (160) of FIG. 2) or as an audio signal through a speaker of the electronic device (10) (e.g., the speaker (184) of FIG. 2).
[0177] FIG. 12 is a diagram showing an example of a situation related to persona mapping according to one embodiment.
[0178] Referring to FIGS. 1 through 12, a user (1) may be carrying or wearing an electronic device (10) capable of providing augmented reality. The electronic device (10) (or the processor (120) of FIG. 2) is activated in response to user (1) operation and can output content corresponding to augmented reality so that the user (1) can view it. As previously described, the augmented reality may include a situation in which at least some of the objects located in real space can be viewed as they are or through a camera via the electronic device (10), and / or a situation in which virtual content that is viewable by the user (1) but does not have a physical form in reality is displayed. The objects (or things) located in real space may be provided to the user (1) via the electronic device (10) as a scene (1201) in which first to fourth scene objects (1211, 1221, 1231, 1241) are placed. Accordingly, at least some of the first to fourth scene objects (1211, 1221, 1231, 1241) placed in the scene (1201) that the user (1) views through the electronic device (10) may exist as actual objects in real space. At least some of the other first to fourth scene objects (1211, 1221, 1231, 1241) placed in the scene (1201) may not exist as actual objects in real space. Additionally, all objects and backgrounds located in real space may be included in the scene (1201) that the user (1) views. Depending on the configuration, only some of all objects and backgrounds located in real space may be included in the scene (1201) that the user (1) views.
[0179] According to one embodiment, regarding persona mapping, a user (1) may perform an utterance corresponding to a first mapping query (1_sc1). The first mapping query (1_sc1) may include a first request information requesting to map a first persona (e.g., a botanist) associated with a first AI assistance function to the first scene object (1211) while indicating or selecting specific scene objects included in a scene (1201) (e.g., a first scene object (1211) corresponding to a vase and / or a second scene object (1221) corresponding to a clock) and / or a second request information requesting to map a second persona (e.g., a schedule notification, or a schedule informant) associated with a second AI assistance function to the second scene object (1221). According to one embodiment, an electronic device (10) may receive input from the user (1) for mapping a persona to a scene object (or an object in a captured image of a real object). For example, an electronic device (1) (e.g., the voice recognition module (121a) of FIG. 7) can acquire a user's voice command and convert the voice signal for the voice command into text based on the voice recognition function. According to one embodiment, the electronic device (1) can identify the intent (user intent) and entity (key word) included in the converted text and, based on this, identify a request to assign a persona to a specific scene object or a category of scene objects (or real objects, things) of a specific category (or type). As described above, when the electronic device (10) first receives input from the user (1) to map a specific persona to a specific real object, it is possible to perform the corresponding task through the user's voice command.For example, when a user (1) wears an electronic device (1) corresponding to AR glasses and speaks, “Attach an Assistant that can ask about the plant in that vase,” the electronic device (10) can convert the voice signal into text, “Attach an Assistant that can ask about the plant in that vase,” through a voice recognition function. The electronic device (10) can determine the user intent contained in the converted text, and the speech can be interpreted as a request to assign a persona to a specific object (ID) or a specific type of object (Category).
[0180] In relation to the persona mapping above, the electronic device (10) can utilize keywords included in the first request information (e.g., transparent vase is now a botanist) of the first mapping query (1_sc1) to search for a first domain (or a first information DB, or a memory device or server device where the first information DB is stored, or a path and authority for accessing the memory device or server device) containing information related to said keywords, and can connect a first AI assistance function that generates a response based on said first domain (e.g., first information DB) to said first persona. Similarly, the electronic device (10) may utilize keywords included in the second request information of the first mapping query (1_sc1) (e.g., clock schedule notification from now until tomorrow) to search for a second domain (e.g., second information DB, or a memory device or server device where the second information DB is stored, or a path and authority for accessing the memory device or server device) containing information related to said keywords, and may connect a second AI assistance function that generates a response based on said second domain (e.g., second information DB) to said second persona. Regarding the generation of a persona name, the electronic device (10) may generate a persona name using at least some of the keywords included in the first request information or the second request information.
[0181] In the above domain search process, the electronic device (10) checks the characteristics of the keywords included in the first mapping query (1_sc1), and if the characteristics of the keywords require the utilization of a database stored in an external server device, it can connect an AI assistance function capable of generating a response based on the external server device to the first persona. Alternatively, if it is determined that the characteristics of the keywords included in the first mapping query (1_sc1) require the utilization of information stored in the memory (130) of the electronic device (10), the electronic device (10) can connect an AI assistance function capable of generating a response based on the information stored in the memory (130) to the second persona. Additionally or generally, the electronic device (10) checks the characteristics of the keywords in the first mapping query (1_sc1), and if it is determined that the keywords require the utilization of the memory (130) of the electronic device (10) and the database of the external server device, it can map a persona capable of generating a response by utilizing the information stored in the memory (130) and the database of the external server device to the corresponding scene object.
[0182] In relation to performing the above-described operation, the electronic device (10) can extract feature points for the scene objects (1211, 1221, 1231, 1241) during the recognition process of the scene objects (1211, 1221, 1231, 1241), and detect object names corresponding to the scene objects (1211, 1221, 1231, 1241) in the object recognition DB based on the feature points. For example, the electronic device (10) can collect information regarding the characteristics of the scene objects (1211, 1221, 1231, 1241) and names that can identify the scene objects (1211, 1221, 1231, 1241) during the object recognition process. Based on this, the electronic device (10) can detect a first scene object (1211) corresponding to a transparent vase included in the scene (1201) based on information regarding “transparent” and “vase.” Similarly, the electronic device (10) can recognize a second scene object (1221) as a watch based on object analysis.
[0183] According to one embodiment, the electronic device (10) may operate at least one sensor during the process of mapping a persona to a scene object. For example, the electronic device (10) may determine the object (1) referred to by the user (1) through various sensors such as a motion detection sensor, an eye tracking (ET) sensor, and an RGB camera sensor. For example, based on sensor data obtained from a motion detection sensor, the electronic device (10) may determine whether the user (1) is designating a specific object using a gesture. Based on the ET sensor data, the electronic device (10) may more accurately determine the object referred to by the user by using the user's (1) gaze information (or gaze information, gaze information).
[0184] The electronic device (10) can perform an analysis of characteristic elements of a specific scene object to process a request from a user (1). For example, once the identification of a scene object designated by the user is completed through various sensors, the electronic device (10) can acquire characteristic points of the scene object using at least one of an RGB camera and a depth sensor (ToF sensor; Time of Flight). In this process, the electronic device (10) can acquire data such as the type, color, included text, shape, and / or texture of the scene object, and perform object classification that maps a persona based on this.
[0185] The electronic device (10) may provide detailed information that allows searching for a list of personas that can be mapped to each scene object (1211, 1221, 1231, 1241) of a scene (1201) and / or descriptions of AI assistance functions corresponding to the personas, in relation to persona mapping, before receiving a first mapping query (1_sc1), during the process of receiving the first mapping query (1_sc1), or during the process of mapping a first persona (e.g., botanist) and a second persona (e.g., schedule notification) to a first scene object (1211) and a second scene object (1221), respectively, through the first mapping query (1_sc1).
[0186] As another example, when scene objects (1211, 1221, 1231, 1241) are recognized while the user (1) is viewing a scene (1201), the electronic device (10) can identify scene objects that are not mapped to a persona and provide a list of personas that can be mapped to the scene objects through augmented reality. Additionally or generally, the electronic device (10) can provide a history of the personas used by the user (1) to date and / or a list of scene objects mapped to personas through augmented reality.
[0187] According to one embodiment, the user (1) may perform persona mapping based on the attributes (or types, categories) of specific scene objects. For example, when the user (1) generates a mapping query using the utterance (or keyword, voice command) “Vases correspond to botanists,” the electronic device (10) may not only map a persona corresponding to a botanist to the scene object corresponding to a vase included in the current scene (1201), but also map a persona corresponding to a botanist to all vase types. For example, if there are multiple vases in the scene (1201), the electronic device (10) may map a persona corresponding to a botanist to all of the multiple vases. Alternatively, if a vase that was not included in the scene (1201) is newly included in the scene (1201) or if the shooting range of the electronic device (10) is changed so that a new scene including the new vase is collected, the electronic device (10) can automatically map a persona corresponding to a botanist to the vase included in the new scene even without a mapping request from the user (1).
[0188] FIG. 13 is a diagram illustrating an example of a situation related to persona mapping verification according to one embodiment.
[0189] Referring to FIGS. 1 through 13, a user (1) may wear (or carry) an electronic device (10) and supply power. The electronic device (10) may activate a camera in response to the user's (1) operation and capture a set direction to acquire a scene (1201) (or camera image, scene image). The electronic device (10) may output the acquired scene (1201) so that the user can view it. The scene (1201) may include at least a portion of an image captured by the camera of the electronic device (10). The scene (1201) may include at least one scene object (e.g., a first scene object (1211) and / or a second scene object (1221)). Additionally, or generally, as described above in FIG. 12, the scene (1201) may include other scene objects in addition to the first scene object (1211) and / or the second scene object (1221).
[0190] According to one embodiment, a user (1) may generate at least one of a first confirmation query (1_Q1) or a second confirmation query (1_Q2) while watching a scene (1201). The first confirmation query (1_Q1) may include a speech (or voice signal, audio signal) specifying a specific scene object (e.g., a second scene object (1221)). Alternatively, the first confirmation query (1_Q1) may include content specifying a specific scene object (e.g., a clock) and content regarding a specific query (e.g., a query regarding the type of persona). In relation to the analysis of the first confirmation query (1_Q1), the electronic device (10) may activate a speech recognition module (121a), perform text conversion on the speech of the user (1), and then extract content selecting the type of scene object (e.g., a clock) and query content (e.g., what the persona is) from the text.
[0191] According to one embodiment, the second confirmation query (1_Q2) may include a motion query in which the user (1) points to (e.g., using a finger, hand, controller, ring, and / or specific device) or gazes at a second scene object (1221) while wearing the electronic device (10). In this regard, the electronic device (10) may include at least one camera among a first camera capable of capturing a set first direction (e.g., a forward direction in which the scene (1201) is captured) and a second camera positioned to capture the user's (1) eye (or pupil). When an object corresponding to the user's (1) finger is detected among the images acquired through the first camera, the electronic device (10) may detect a scene object (e.g., a second scene object (1221)) that overlaps within a certain range with the direction pointed by the user's (1) finger among the scene objects included in the scene (1201). Alternatively, the electronic device (10) may detect the gaze of the user (1) and detect a second scene object (1221) located in the direction in which the gaze is directed. If there are multiple scene objects in the direction of the user's gaze or the direction pointed by the finger, the electronic device (10) may select one scene object (e.g., a second scene object (1221)) among the multiple scene objects by referring to the content of selecting a scene object (e.g., a watch) in the first confirmation query (1_Q1).
[0192] According to one embodiment, when at least one of a first verification query (1_Q1) or a second verification query (1_Q2) is collected, the electronic device (10) can determine a scene object selected by the user (1) (e.g., a second scene object (1221)) based on the collected at least one verification query, and generate and output a first response (1_Rsp1) (e.g., a response related to the selected second scene object (1221)) according to the query content. In this process, the electronic device (10) can verify the user query using an LLM (123) and generate a first response (1_Rsp1) corresponding thereto. As an example, the electronic device (10) can check persona-related information set in relation to the current scene (1201) in memory (130) and generate a first response (1_Rsp1) based on the persona-related information. The memory (130) can store and manage persona management information for at least one scene object included in the scene (1201).
[0193] According to one embodiment, the electronic device (10) may output a first response (1_Rsp1) through a speaker (e.g., speaker (184) in FIG. 2). Additionally or generally, the electronic device (10) may output object description information (1222) (or keywords, object indicator keywords, persona call keywords, persona description keywords) in relation to (or linked to) a second scene object (1221) in a scene (1201). The object description information (1222) may be displayed adjacent to the second scene object (1221) within a certain distance on the scene viewed by the user (1), or may be displayed at a position spaced apart from the second scene object (1221) by using an indicator line pointing to the second scene object (1221). Alternatively, the object description information (1222) may be displayed such that at least a portion overlaps with the second scene object (1221) at the viewing angle of the user (1). The object description information (1222) may be information describing the persona of the corresponding scene object. Alternatively, the object description information (1222) may be information indicating the persona mapped to the scene object.
[0194] FIG. 14 is a diagram showing an example of a first situation related to a persona activation setting according to one embodiment.
[0195] Referring to FIGS. 1 through 14, regarding the persona activation setting, a user (1) may possess an electronic device (10) or wear it at a designated location (e.g., on the head). While wearing the electronic device (10), the user (1) may perform a first activation trigger action (1_hv1) regarding a specific scene object (e.g., a first scene object (1211)) included in the scene. For example, as illustrated, the user (1) may grasp the first scene object (1211) placed within the scene using their hand. The electronic device (10) may capture the first activation trigger action (1_hv1) using a camera. The first activation trigger action (1_hv1) may include, for example, a set gesture action.
[0196] According to one embodiment, the electronic device (10) may collect a first activation query (1_at1) or a second activation query (1_at2) using a microphone. The first activation query (1_at1) or the second activation query (1_at2) may be obtained from the utterance of a user (1). The electronic device (10) may perform speech recognition on the utterance of the user (1) collected through the microphone, perform text conversion, and collect the content included in the first activation query (1_at1) or the second activation query (1_at2) (e.g., an activation trigger for a persona mapped to a first scene object (1211). The first activation query (1_at1) or the second activation query (1_at2) may define a trigger capable of activating a persona mapped to a specific scene object. For example, the first activation query (1_at1) may include content (or keywords, or voice commands) describing the first activation trigger action (1_hv1). The second activation query (1_at2) may include content (or keywords, or voice commands) specifying the location of a specific scene object in relation to persona activation.
[0197] According to one embodiment, the electronic device (10) may generate a persona activation trigger by integrating the first activation trigger action (1_hv1) and the activation query (e.g., the first activation query (1_at1) or the second activation query (1_at2)). For example, based on the first activation trigger action (1_hv1) and the first activation query (1_at1), the electronic device (10) may subsequently automatically activate the persona mapped to the first scene object (1211) (e.g., activate the AI assistance function mapped to the persona) while the user (1) holds the first scene object (1211). Alternatively, the electronic device (10) may, based on the first activation trigger action (1_hv1) and the second activation query (1_at2), automatically activate the persona mapped to the first scene object (1211) (e.g., activate the AI assistance function mapped to the persona) when the first scene object (1211) maintains a specified position while the user (1) subsequently views a scene containing the first scene object (1211).
[0198] According to one embodiment, the electronic device (10) may stop automatic persona activation when the first scene object (1211) moves out of a designated location. Alternatively, the electronic device (10) may provide guidance that the persona cannot be activated due to a change in the designated location when the first scene object (1211) moves out of a designated location. Additionally, or generally, the electronic device (10) may be controlled to output guidance information requesting the first scene object (1211) to be relocated to a designated location for persona activation related to the first scene object (1211) (e.g., including guide information guiding the location where the first scene object (1211) should be for persona activation).
[0199] FIG. 15 is a diagram showing an example of a second situation related to a persona activation setting according to one embodiment.
[0200] Referring to FIGS. 1 through 15, in relation to persona activation settings, a user (1) may carry an electronic device (10) or wear it in a designated location (e.g., head) and perform a second activation trigger action (1_hv2) regarding a specific scene object (e.g., a first scene object (1211)) included in a scene (1201). For example, as illustrated, the user (1) may gaze at the first scene object (1211) placed within the scene (1201) for a certain period of time using their gaze. The electronic device (10) may capture the second activation trigger action (1_hv2) using a camera (e.g., a camera positioned to capture the user's (1) eyes (or pupils)).
[0201] According to one embodiment, the electronic device (10) may output an activation confirmation query (1_at_Q1) when the second activation trigger operation (1_hv2) is maintained for a specified time. The activation confirmation query (1_at_Q1) may be output, for example, through a speaker or on a scene (1201) that the user (1) is viewing. For example, the activation confirmation query (1_at_Q1) may include content for checking the persona activation settings of a first scene object (1211). The activation confirmation query (1_at_Q1) may be displayed within a certain distance of the first scene object (1211), indicate the first scene object (1211), or be displayed so as to overlap at least a portion with the first scene object (1211).
[0202] According to one embodiment, when a user (1) performs a confirmation (e.g., confirmation via user voice input, or confirmation via touch or other user input) regarding an activation confirmation query (1_at_Q1), the electronic device (10) may output guidance information regarding a second activation trigger action (1_hv2) related to the activation of a persona (e.g., a configured AI assistant function) of a first scene object (1211). For example, the electronic device (10) may output information at a certain location in the scene (1201) instructing the user (1) to gaze at the first scene object (1211) for a set amount of time or longer in order to activate the persona of the first scene object (1211) at a certain interval after the user (1) wears the electronic device (10). Alternatively, the electronic device (10) may output information guiding the persona mapping status of the first scene object (1211) when the user (1) gazes at the first scene object (1211), and information guiding a second activation trigger action (1_hv2) to activate the persona mapped to the first scene object (1211) (e.g., an audio signal or displayed at a designated location in the scene (1201)).
[0203] FIG. 16 is a diagram showing an example of a first situation related to the use of a persona according to one embodiment.
[0204] Referring to FIGS. 1 through 16, in relation to the use of a persona, a user (1) can activate an electronic device (10) and view a scene (1201) including a first scene object (1211) through the electronic device (10). The scene (1201) may include a plurality of scene objects as previously described. The electronic device (10) may provide an extended reality including the scene (1201) by combining at least a portion of an image collected by a camera with at least a portion of reality (a reality visible through see-through or a reality visible without undergoing separate image processing by the device). Alternatively, the electronic device (10) may provide the scene (1201) collected by the camera so that the user can view only that scene.
[0205] According to one embodiment, the user (1) may utter a first user query (1_user_Q) related to a first scene object (1211) included in a scene (1201). The electronic device (10) may keep a microphone capable of recognizing user utterances active, and when the user query (1_user_Q) is collected, may perform voice recognition for the user query (1_user_Q) using a voice recognition module (121a). The electronic device (10) may identify the content of the question regarding the user query (1_user_Q) using an LLM (123) and generate and / or output a second response (1_Rsp2) using a persona mapped to the scene object. For example, if the user query (1_user_Q) contains content regarding “direction,” the electronic device (10) may automatically determine the direction based on the user's (1) current location and the user's (1) gaze. Additionally, the electronic device (10) can detect a scene object corresponding to a plant among the scene objects included in the scene (1201) when the term “plant” included in the user query (1_user_Q) is detected. The electronic device (10) can detect a specific scene object (e.g., a first scene object (1211)) included in the scene (1201) by combining the content (or keyword) regarding the “direction” and the content (or keyword) regarding the “plant”. When multiple scene objects included in the user query (1_user_Q) are detected, the electronic device (10) can output information guiding the selection of one of the multiple scene objects or generate a response for the multiple scene objects.
[0206] According to one embodiment, in relation to the generation of a second response (1_Rsp2), the electronic device (10) can activate a first persona (e.g., an AI assistant function corresponding to a botanist) mapped to a first scene object (1211) and generate a second response (1_Rsp2) regarding the “plant” based on the first persona. When the AI assistant function mapped to the first persona is configured to generate a response based on a database of an external server device, the electronic device (10) forms a communication channel with the external server device, and the AI assistant function mapped to the first persona can generate and output the second response (1_Rsp2) based on information stored in the external server device that can provide information related to the “plant” (e.g., displaying text corresponding to the second response (1_Rsp2) at a designated location in the scene (1201) or outputting an audio signal corresponding to the second response (1_Rsp2).
[0207] FIG. 17 is a diagram showing an example of a second situation related to the use of a persona according to one embodiment.
[0208] Referring to FIGS. 1 through 17, regarding the use of personas, a user (1) can activate an electronic device (10) and view a scene (1201) including a first scene object (1211) and a second scene object (1221) through the electronic device (10). The first scene object (1211) and the second scene object (1221) included in the scene (1201) can maintain a set relationship (e.g., a relationship where the distance between the scene objects is within a first distance). When the first scene object (1211) to which the first persona is mapped and the second scene object (1221) to which the second persona is mapped maintain a set relationship, the electronic device (10) can group the first scene object (1211) and the second scene object (1221) and create an integrated persona (1_hv3) (or a third persona) by grouping the first persona and the second persona. The above integrated persona (1_hv3) may be mapped to a third AI assistant function that integrates a first AI assistant function mapped to the first persona and a second AI assistant function mapped to the second persona. The name of the integrated persona (1_hv3) may be configured to include at least a part of the name of the first persona and the name of the second persona. The above third AI assistant function (or integrated AI assistant function) may include an AI module that generates an integrated response based on an integrated domain (e.g., an integrated domain that integrates a first domain referenced by the AI assistant function mapped to the first persona for response generation and a second domain referenced by the second AI assistant function mapped to the second persona for response generation). Alternatively, the third AI assistance function may include an AI module that generates at least one of the integrated response, a response generated based on a first domain referenced by the AI assistance function mapped to the first persona for response generation, and a response generated based on a second domain referenced by the second AI assistance function mapped to the second persona for response generation.
[0209] In a situation where personas are integrated, the user (1) may utter a first user query (1_user_Q) related to a first scene object (1211) included in the scene (1201). The electronic device (10) keeps a microphone capable of recognizing user utterances in an active state, and when the first user query (1_user_Q) is collected, it may perform voice recognition for the first user query (1_user_Q) using a voice recognition module (121a). The electronic device (10) may use an LLM (123) to identify the content of the question regarding the first user query (1_user_Q) and check the persona mapping status related to the first scene object (1211). If the integrated persona is in a mapped state, the electronic device (10) may generate and output a third response (1_Rsp3) based on the integrated persona. In this process, the electronic device (10) (e.g., LLM (123) of FIG. 7) can analyze the first user query (1_user_Q) to detect the target (e.g., plant) of the third response (1_Rsp3).
[0210] According to one embodiment, in relation to the generation of a third response (1_Rsp3), the electronic device (10) identifies a second scene object (1221) grouped with a first scene object (1211) and can generate a third response (1_Rsp3) based on an integrated persona (1_hv3) (or a third AI assistance function mapped to the integrated persona (1_hv3)) corresponding to the grouped first scene object (1211) and second scene object (1221). The third AI assistance function mapped to the integrated persona (1_hv3) can generate a third response (1_Rsp3) based on a first domain (e.g., a plant-related database of an external server device) that contributes to the generation of a response for the first persona and a second domain (e.g., user schedule information stored in memory (130)) that contributes to the generation of a response for the second persona. For example, the third response (1_Rsp3) may include user schedule information regarding the “plant” (or plant keyword) indicated by the user query (1_user_Q), along with information regarding the “plant”. According to one embodiment, the third AI assist function may additionally utilize external domain information in connection with the generation of the third response (1_Rsp3). For example, the external domain information is information that can access an external server capable of providing information associated with the first object and / or the second object, and the third AI assist function may obtain specified information (e.g., date information, and / or other external environment information) from the external server.
[0211] According to one embodiment, the integrated persona may be updated. For example, a new scene object (e.g., a third scene object) may be moved into the scene (1201) that the user (1) is viewing and appear within the scene (1201). When the third scene object is moved and a set condition is satisfied that allows the first scene object (1211), the second scene object (1221), and the third scene object to be grouped, the electronic device (10) may map a new integrated persona, which integrates the first persona, the second persona, and the third persona mapped to the third scene object, to the first scene object (1211), the second scene object (1221), and the third scene object.
[0212] As an example, the first persona may be mapped to a first AI (artificial intelligence) auxiliary function (or first AI module) that generates a first response to a user query through a first domain corresponding to a first information database, and the second persona may be mapped to a second AI auxiliary function (or second AI module) that generates a second response to the user query through a second domain corresponding to a second information database. The integrated persona, formed by integrating the first persona and the second persona, may be mapped to an integrated AI auxiliary function (or integrated AI module) that generates a third response to the user query through a domain corresponding to the first information database and the second information database. The third persona may be mapped to a third AI auxiliary function (or third AI module) that generates a fourth response to a user query through a third domain corresponding to a third information database. The above new integrated persona may be mapped to a new integrated AI assistant function (or new integrated AI module) that generates and provides the fifth response for the specific user query based on a database integrating the first information database, the second database, and the third information database.
[0213] The integrated persona may be configured to generate and output a response based on the integrated database for the specific user query upon receiving the specific user query. Alternatively, the integrated persona may be configured to generate and output a first response based on the first persona, a second response based on the first persona, and a response based on the integrated persona all for the specific user query upon receiving the specific user query.
[0214] The new integrated persona may be configured to generate and output the response to the specific user query based on a database that integrates the first information database related to the first persona, the second database related to the second persona, and the third information database related to the third persona, when it receives a user query. As an example, the new integrated persona may be configured to generate and output the first response based on the first persona, the second response based on the second persona, the third response based on the third persona, and the response based on the three integrated information databases for the specific user query. Additionally or generally, the new integrated persona may generate and output at least one of the response based on the first and second personas, the response based on the second and third personas, and the response based on the first and third personas.
[0215] FIG. 18 is a diagram showing an example of a second situation related to persona mapping according to one embodiment.
[0216] Referring to FIGS. 1 through 18, a user (1) may wear (or possess) an electronic device (10) and supply power to the electronic device (10). Alternatively, the user (1) may wear (or possess) the electronic device (10) to which power has been supplied. The electronic device (10) may activate a camera in response to the user's (1) operation and capture a set direction to acquire a scene (1201) (or camera image, scene image). The scene (1201) may include, for example, a specific scene object (e.g., a second scene object (1221)) and object description information (1222) describing the specific scene object. The electronic device (10) may output the acquired scene (1201) so that the user (1) can view it. The electronic device (10) may provide a persona mapping mode or a persona mapping menu in relation to persona mapping. Support for the menu or mode may be applied in the same or similar way in other embodiments.
[0217] While watching a scene (1201), the user (1) can generate a first mapping trigger action (1_hv4) (e.g., pointing at the second scene object (1221) with a finger) that directs a specific scene object (e.g., a second scene object (1221)) and a third mapping query (1_sc3) for persona mapping (e.g., an utterance that maps a specific scene object to a persona). The electronic device (10) can detect the second scene object (1221) directed by the first mapping trigger action (1_hv4) through a camera and map a second persona to which a second AI assistance function (e.g., a schedule notification function) is connected to the second scene object (1221) based on voice recognition of a voice signal collected through a microphone. The electronic device (10) may output object detail information (1222) (e.g., text-schedule notification) within the scene (1201) or output it as audio information (or both) in relation to the second persona mapping guidance of the second scene object (1221).
[0218] According to one embodiment, while watching a scene (1201), the user (1) may generate a second mapping trigger action (1_hv5) (e.g., a second scene object (1221)) that directs a specific scene object (e.g., a second scene object (1221)) and a third mapping query (1_sc3) for persona mapping (e.g., an utterance that maps a specific scene object to a persona). The electronic device (10) detects the second scene object (1221) directed by the second mapping trigger action (1_hv5) through a camera and maps a second persona to which a second AI assistance function (e.g., a schedule notification function) is connected to the second scene object (1221) based on voice recognition of a voice signal collected through a microphone. The electronic device (10) can output object detail information (1222) (e.g., text-schedule notification) within the scene (1201) in relation to the second persona mapping guidance of the second scene object (1221).
[0219] FIG. 19 is a diagram showing an example of a fourth situation related to persona mapping according to one embodiment.
[0220] Referring to FIGS. 1 through 19, a user (1) possesses a second electronic device (12) (e.g., a portable electronic device or a smartphone), and the second electronic device (12) can output a first mapping screen (12_scr) related to persona mapping to a display unit. The first mapping screen (12_scr) may correspond to a scene (1201) provided by another electronic device (10) (or correspond to at least a part of the scene (1201)). Alternatively, the first mapping screen (12_scr) may correspond to a scene (or captured image) captured by a camera included in the second electronic device (12) in a set direction (or include at least a part of the captured image). As an example, the first mapping screen (12_scr) may include a second scene object image (1221_img) corresponding to the second scene object (1221) described above in FIG. 18.
[0221] The user (1) can generate a third mapping trigger action (1_hv6) (e.g., an action of touching the second scene object image (1221_img) displayed on the display of the second electronic device (12)) for selecting a specific scene object image (e.g., a second scene object image (1221_img)) through a first mapping screen (12_scr) displayed on the display of the second electronic device (12), and a third mapping query (1_sc3) for persona mapping (e.g., an utterance indicating that a specific scene object is mapped to a persona). The second electronic device (12) can receive the third mapping trigger action (1_hv6) through a display having a touch function (e.g., a touchscreen). Additionally, the second electronic device (12) can collect a query related to persona mapping related to the second scene object image (1221_img) based on voice recognition of a voice signal collected through a microphone. When the third mapping trigger operation (1_hv6) and the third mapping query (1_sc3) are collected, the second electronic device (12) can map a second persona connected to a second AI assistance function (e.g., a schedule notification function) to a second scene object image (1221_img). Additionally, the second electronic device (12) can output guidance information on the mapping of the second persona of the second scene object image (1221_img) on the first mapping screen (12_scr) of the display unit.
[0222] FIG. 20 is a diagram showing an example of a fifth situation related to persona mapping according to one embodiment.
[0223] Referring to FIGS. 1 through 20, a user (1) may use a third electronic device (13) (e.g., a desktop computer device or a stationary computer device). The third electronic device (13) may output a second mapping screen (13_scr) related to persona mapping on a display panel. The second mapping screen (13_scr) may correspond to a scene (1201) provided by another electronic device (10) (or correspond to at least a part of the scene (1201)). Alternatively, the second mapping screen (13_scr) may correspond to a scene (or captured image) captured by a camera functionally connected to the third electronic device (13) (or include at least a part of the captured image). As an example, the second mapping screen (13_scr) corresponds to at least a part of the scene (1201) described above in FIG. 18, and the second scene object image (1221_img) included in the second mapping screen (13_scr) may correspond to at least a part of the second scene object (1221) included in the scene (1201) of FIG. 18.
[0224] The user (1) can generate a fourth mapping trigger action (1_hv7) (e.g., an action of selecting the second scene object image (1221_img) using a mouse) and a third mapping query (1_sc3) for persona mapping (e.g., an utterance that maps a specific scene object to a persona) through a second mapping screen (13_scr) displayed on a display panel of the third electronic device (13). The third electronic device (13) can receive the fourth mapping trigger action (1_hv7) through a display panel having a touch function (e.g., a touch panel). Additionally, the third electronic device (13) can collect a query related to persona mapping related to the second scene object image (1221_img) based on voice recognition of a voice signal collected through a microphone. When the third electronic device (13) collects the fourth mapping trigger operation (1_hv7) and the third mapping query (1_sc3), it can map a second persona connected to a second AI assistance function (e.g., a schedule notification function) to the second scene object image (1221_img). Additionally, the third electronic device (13) can output guidance information on the mapping of the second persona of the second scene object image (1221_img) to the second mapping screen (13_scr) of the display panel.
[0225] FIG. 21 is a diagram illustrating an example of a situation related to persona grouping and / or operation according to one embodiment.
[0226] Referring to FIGS. 1 to 21, a user (1) may use an electronic device (10) (e.g., a head-worn electronic device). The electronic device (10) may construct and output a second scene (2201) based on a captured image collected by a camera in response to the direction being directed. As an example, the user (1) may view a second scene (2201) including a plurality of scene objects (2211, 2221, 2231) output by the electronic device (10). The second scene (2201) may include, for example, a plurality of scene objects (2211, 2221, 2231) and hand objects (2291, 2292) corresponding to the hands of the user (1) using at least some of the plurality of scene objects (2211, 2221, 2231). As an example, the plurality of scene objects (2211, 2221, 2231) may include, for instance, a fifth scene object (2211), a sixth scene object (2221), and a seventh scene object (2231). The first hand object (2291) may grasp the fifth scene object (2211), and the second hand object (2292) may grasp the sixth scene object (2221). The electronic device (10) can analyze and / or determine the relationship between the fifth scene object (2211) and the first hand object (2291) (e.g., the relationship in which the first hand object (2291) grasps the fifth scene object (2211)) and the relationship between the sixth scene object (2221) and the second hand object (2292) (e.g., the relationship in which the second hand object (2292) grasps the sixth scene object (2221)) based on the hand shape of the first hand object (2291) and the hand shape of the second hand object (2292).
[0227] Additionally or generally, the electronic device (10) may provide the second scene (2201) to the user (1) and provide object detail information (2212, 2222, 2232) mapped to a plurality of scene objects (2211, 2221, 2231). For example, the electronic device (10) can output a second scene (2201) including a fifth object detail (2212) related to a fifth scene object (2211) (e.g., information describing a fifth persona mapped to the fifth scene object (2211), such as a blue pen: mathematician), a sixth object detail (2222) related to a sixth scene object (2221) (e.g., information describing a sixth persona mapped to the sixth scene object (2221), such as a notebook: calculator), and a seventh object detail (2232) related to a seventh scene object (2231) (e.g., information describing a seventh persona mapped to the seventh scene object (2231), such as a cup: teacher).
[0228] The electronic device (10) can identify the relationship between the plurality of scene objects (2211, 2221, 2231) (e.g., a relationship placed within a certain distance) and integrate the plurality of scene objects (2211, 2221, 2231) to map an integrated persona. For example, the electronic device (10) can map an integrated persona, which is formed by integrating the fifth persona mapped to the fifth scene object (2211), the sixth persona mapped to the sixth scene object (2221), and the seventh persona mapped to the seventh scene object (2231), to at least one of the grouped plurality of scene objects (2211, 2221, 2231). The integrated persona can be mapped to an integrated AI assistant function configured to generate a response based on an integrated integrated domain. Here, the integrated domain may include at least some of the domain (or information DB) used by the 5th AI assistant function supporting the 5th scene object (2211) for generating a response, the domain used by the 6th AI assistant function supporting the 6th scene object (2221) for generating a response, and the domain used by the 7th AI assistant function supporting the 7th scene object (2231) for generating a response.
[0229] The electronic device (10) may collect a user query (User_Q) based on a sixth scene object (2221). When the user query (User_Q) is collected, the electronic device (10) may transmit the user query (User_Q) to an integrated AI assistant function mapped to an integrated persona (or transmit it as input to an AI module supporting the integrated AI assistant function). The AI module supporting the integrated AI assistant function may generate a fourth response (2202) to the user query (User_Q) using the integrated domain. The electronic device (10) may output the fourth response (2202) onto the second scene (2201). According to one embodiment, the user query (User_Q) may include at least some of the text inputs by user speech and / or user writing. Such diversity regarding the user query (User_Q) may be applied in the same or similarly in other embodiments of this description.
[0230] Meanwhile, in the description of FIG. 21 above, the numbering assigned to scene objects (e.g., 5th, 6th, 7th), the numbering assigned to personas (e.g., 5th, 6th, 7th), and the numbering assigned to AI assistance functions (e.g., 5th, 6th, 7th) are listed to distinguish the terms and may not have meaning in order or be connected to other embodiments.
[0231] FIG. 22 is a diagram illustrating an example of a second situation related to persona grouping and / or operation according to one embodiment.
[0232] Referring to FIGS. 1 to 22, a user (1) may use an electronic device (10) (e.g., a head-worn electronic device). The electronic device (10) may construct and output a third scene (2301) based on a captured image collected by a camera in response to the direction being directed. As an example, the user (1) may view a third scene (2301) including a plurality of scene objects (2311, 2321, 2331, 2341). The plurality of scene objects (2311, 2321, 2331, 2341) may include an eighth scene object (2311), a ninth scene object (2321), and a tenth scene object (2331), and an eleventh scene object (2341) in which a portion of the eighth to tenth scene objects (2311, 2321, 2331) is contained. As at least some of the 8th to 10th scene objects (2311, 2321, 2331) are housed in the 11th scene object (2341), the 8th to 10th scene objects (2311, 2321, 2331) can satisfy a grouping condition (e.g., a condition where scene objects are located within a certain distance).
[0233] According to one embodiment, among the scene objects stored in the 11th scene object (2341), the 8th scene object (2311) and the 9th scene object (2321) may be scene objects to which the 8th persona and the 9th persona are mapped, respectively. For example, the 8th scene object (2311) may include a scene object to which the 8th persona (e.g., a function capable of calling or activating an AI assistance function that provides a response to a user query based on an information database capable of providing information related to physics) is assigned, as described by the 8th object detailed information (2312) (e.g., black pen, indicating a state in which a physicist persona is assigned). When the user utters “black pen,” the 8th persona is activated to call the relevant AI assistance function, and the AI assistance function can generate and provide a response to a query entered after the command for activating the 8th persona (e.g., black pen) based on the physics-related information database. The above-mentioned ninth scene object (2321) may include a scene object to which a ninth persona is assigned (e.g., a function capable of calling or activating an AI assistance function that provides a response to a user query based on an information database capable of providing information related to a poet), as described by the ninth object detailed information (2322) (e.g., green pen, indicating a state in which a poet persona is assigned). When the user utters “green pen,” the ninth persona is activated to call the relevant AI assistance function, and the AI assistance function can generate and provide a response to a query entered after the command for activating the ninth persona (e.g., green pen) based on the poet-related information database.
[0234] As another example, the 10th scene object (2331) may be a scene object that does not have a separate persona mapped to it. The electronic device (10) can identify the 8th to 10th scene objects (2311, 2321, 2331) placed within the 11th scene object (2341), and identify the scene objects with personas mapped to them (e.g., the 8th scene object (2311) and the 9th scene object (2321)) based on information stored in the persona management module (122c) and memory (130). The electronic device (10) may provide the user (1) with an inquiry to determine whether to create an integrated persona only for the 8th scene object (2311) and the 9th scene object (2321) with personas mapped to them among the 8th to 10th scene objects (2311, 2321, 2331). When a user (1) generates a user input that determines the creation of an integrated persona, the electronic device (10) can map an integrated persona, which combines the eighth persona and the ninth persona mapped to the eighth scene object (2311) and the ninth scene object (2321) respectively, to at least one of the eighth scene object (2311) and the ninth scene object (2321).
[0235] Additionally or generally, when an integrated persona that combines the 8th scene object (2311) and the 9th scene object (2321) is mapped, the user (1) may map a new 10th persona to the 10th scene object (2331). The electronic device (10) checks the relationship between the 8th to 10th scene objects (2311, 2321, 2331) located within the 11th scene object (2341), and if the setting conditions for grouping are satisfied (e.g., when the scene objects are located within a certain distance), it may ask the user (1) whether to integrate the 10th persona mapped to the 10th scene object (2331) into the integrated persona that combines the 8th scene object (2311) and the 9th scene object (2321). When a user (1) directs integration, the electronic device (10) may map an integrated persona, which integrates the 8th to 10th personas, to at least one scene object among the 8th to 10th scene objects (2311, 2321, 2331). Accordingly, an integrated AI assistant function (or integrated AI module) that generates a response based on an integrated domain may be mapped to the integrated persona that integrates the 8th to 10th personas. Here, the integrated domain may include an 8th domain (e.g., an information DB, or a memory device or server device storing the information DB, or a path and authority to access the memory device or server device) where the 8th AI assistant function mapped to the 8th persona generates a response, a 9th domain where the 9th AI assistant function mapped to the 9th persona generates a response, and a 10th domain where the 10th AI assistant function mapped to the 10th persona generates a response.
[0236] For example, as the number and variety of information databases (or domains in various fields) increases, the content of the response to the query may become more specific and detailed, or the quantity of the response content (e.g., quantity of text or images) may increase or become more sophisticated. As the number of information databases decreases, the content of the response may become more simplified. The integrated response provided by the AI assistant function connected to the integrated persona in this description may have more content, be more sophisticated, specific, or detailed than the individual responses provided by the AI assistant functions connected to each persona prior to integration, and in some cases, may include more user-friendly content for using the electronic device (10). For example, the content of the integrated response may have at least one of quality and quantity higher or greater than the response content of each persona.
[0237] According to various embodiments, an integrated persona mapped to scene objects may be demapped as the scene objects move. For example, in a situation where an integrated persona is mapped to an 8th scene object (2311), a 9th scene object (2321), and a 10th scene object (2331), the 8th scene object (2311) may be moved and relocated to a distance greater than a reference value from the 9th scene object (2321) and the 10th scene object (2331). The electronic device (10) may demap the integrated persona mapped to the 8th scene object (2311), the 9th scene object (2321), and the 10th scene object (2331) as the 8th scene object (2311) moves. Additionally or generally, the electronic device (10) can determine whether the ninth scene object (2321) and the tenth scene object (2331) satisfy the grouping condition while the eighth scene object (2311) is repositioned to a distance greater than or equal to a reference value. If the ninth scene object (2321) and the tenth scene object (2331) satisfy the grouping condition, the electronic device (10) can map a new integrated persona, which integrates the ninth scene object (2321) and the tenth scene object (2331), to the ninth scene object (2321) and the tenth scene object (2331).
[0238] As described above, the electronic device (10) can unmap the integrated persona mapped to the scene objects when the grouping condition of the grouped scene objects is released. In this operation, when the electronic device (10) receives a user query, it can generate a response based on the persona mapped to each scene object and provide it to the user (1). For example, if a user query is received before the ungrouping of the 8th scene object (2311), the 9th scene object (2321), and the 10th scene object (2331), the electronic device (10) can generate and provide an integrated response to the user query based on the integrated persona corresponding to the grouping of the 8th scene object (2311), the 9th scene object (2321), and the 10th scene object (2331). Alternatively, the electronic device (10) may output at least one of the following responses together with the integrated response: a persona-based response mapped to the eighth scene object (2311), a persona-based response mapped to the ninth scene object (2321), a persona-based response mapped to the tenth scene object (2331), a persona-based response integrating the eighth scene object (2311) and the ninth scene object, a persona-based response integrating the eighth scene object (2311) and the tenth scene object, and a persona-based response integrating the ninth scene object (2321) and the tenth scene object.
[0239] When the 8th scene object (2311) is relocated to a distance greater than a reference value from the 9th scene object (2321) and the 10th scene object (2331), and a user query is received, the electronic device (10) can generate a response corresponding to the scene object specified by the user query and output the generated response (e.g., output via at least one of voice, image, or vibration). For example, if the user query specifies either the 9th scene object (2321) and / or the 10th scene object (2331), the electronic device (10) can generate and output a response based on a new integrated persona that integrates the personas mapped to the 9th scene object (2321) and the 10th scene object (2331). If the user query specifies the 8th scene object (2311), the electronic device (10) can generate and output a response based on the persona mapped (or assigned) to the 8th scene object (2311). When a user query is received without a separate scene object being specified, the electronic device (10) may output at least one of a persona-based response mapped to the 8th scene object (2311), a new integrated persona-based response that integrates the personas mapped to the 9th scene object (2321) and the 10th scene object (2331). Alternatively, when a user query is received without a separate scene object being specified, the electronic device (10) may output at least one of a persona-based response mapped to the 8th scene object (2311), a persona-based response mapped to the 9th scene object (2321), a persona-based response mapped to the 10th scene object (2331), and a new integrated persona-based response that integrates the personas mapped to the 9th scene object (2321) and the 10th scene object (2331).
[0240] FIG. 23 is a diagram illustrating an example of a situation related to persona inheritance according to one embodiment.
[0241] Referring to FIGS. 1 to 23, a user (1) may use an electronic device (10) (e.g., a head-worn electronic device). The electronic device (10) may construct and output a fourth scene (2401) based on a captured image collected by a camera in response to the direction being directed. As an example, the user (1) may view a fourth scene (2401) corresponding to a living space. The fourth scene (2401) may include, for example, a 12th scene object (2411), a 13th scene object (2421), and a 14th scene object (2431), and the fourth scene (2401) may include detailed information (2412) of the 12th scene object (2411), detailed information (2422) of the 13th scene object (2421), and detailed information (2432) of the 14th scene object (2431). The above 12th to 14th object details (2412, 2422, 2432) can be defined as cases where a persona is mapped to each of the 12th to 14th scene objects (2411, 2421, 2431). Accordingly, the user (1) can intuitively recognize whether the scene objects are persona mapped based on the existence of the 12th to 14th object details (2412, 2422, 2432).
[0242] The above 12th to 14th object details (2412, 2422, 2432) may include keywords indicating attributes (e.g., feature points or types) of scene objects (2411, 2421, 2431). Alternatively, the above 12th to 14th object details (2412, 2422, 2432) may include at least one keyword capable of inferring a persona mapped to the scene objects (2411, 2421, 2431). For example, the above 12th object details (2412) may include information capable of recognizing (or understanding, inferring, estimating) that the 12th persona mapped to the 12th scene object (2411) is connected to an AI assistance function that generates a response based on information regarding the person “Da Vinci”. The above 13th object detail information (2422) may include information that can recognize that the 13th persona mapped to the 13th scene object (2421) is connected to an AI assistant function that generates a response based on information regarding the person “Archimedes”. The above 14th object detail information (2432) may include information that can recognize that the 14th persona mapped to the 14th scene object (2431) is connected to an AI assistant function that generates a response based on information regarding the person “Einstein”. The electronic device (10) may store and manage information mapped to the attributes of the 12th to 14th scene objects (2411, 2421, 2431) (e.g., at least one of the feature points of the scene object, the type of the scene object, the name of the scene object, and the type of the scene object) for the above-described 12th to 14th object detail information (2412, 2422, 2432). According to one embodiment, the embodiment of the present invention supports placing objects such as figures (e.g., corresponding to scene objects when photographed with a camera) that suit the user's (1) taste in a personalized space (e.g., Workspace) and using them as personas.For example, when a persona assigned to a specific scene object is activated by making eye contact (or focusing gaze, selecting using gaze) with a specific scene object (e.g., at least one of the 12th to 14th scene objects (2411, 2421, 2431)), the electronic device (10) may support a service that allows the user (1) to ask questions and receive responses through conversation with a designated person (e.g., the "Da Vinci" person corresponding to the persona assigned to the 12th scene object (2411), the "Archimedes" person corresponding to the persona assigned to the 13th scene object (2421), and the "Einstein" person corresponding to the persona assigned to the 14th scene object (2431). As an example, the AI assistance functions activated by the personas assigned to each scene object (2411, 2421, 2431) can provide different responses to the same query from the user (1) based on the thoughts and philosophy of each person.
[0243] According to one embodiment, when the user (1) moves to another location, the electronic device (10) can support viewing another scene corresponding to the moved location. In the other scene, if a scene object corresponding to the 12th to 14th scene objects (2411, 2421, 2431) exists, the electronic device (10) can perform persona inheritance automatically or through user confirmation. As an example, if a doll identical to the 12th scene object (2411) or a doll judged (or presumed) to be similar within a certain range exists in the location where the user (1) moved, the electronic device (10) can map the 12th persona to the identical or similar doll in the other location automatically or in response to user approval.
[0244] According to one embodiment, when a persona is mapped to other scene objects included in the fourth scene (2401), such as a keyboard, a book, a flowerpot, and a picture frame, and after the user (1) moves to another location, if at least one of the keyboard, book, flowerpot, and picture frame is included in the scene obtained at that location, the electronic device (10) can automatically map the persona mapped in the fourth scene (2401) to a scene object having the same attribute (e.g., type of scene object, feature point of scene object, name of scene object, type of scene object), or proceed with persona mapping after the user confirms whether to map.
[0245] FIG. 24 is a diagram illustrating an example of a situation related to gesture-based persona mapping according to one embodiment.
[0246] Referring to FIGS. 1 through 24, a user (1) can map a persona to a temporary scene object that is movable and temporarily appears in the scene. For example, a user (1) wearing an electronic device (10) can position a first hand pose corresponding to a first gesture within the camera shooting range of the electronic device (10). The electronic device (10) captures the first hand pose and provides a scene containing a first temporary scene object (2511) to the user (1), and the user (1) can map a first persona to the first temporary scene object (2511). In response to the first persona mapping, the electronic device (10) can output first temporary object description information (2512) (e.g., British boy) related to the first temporary scene object (2511) by including it in the scene. For example, the first persona (e.g., British boy) mapped to the first temporary scene object (2511) can activate an AI assistance function that provides a response to the user's (1) query based on person information related to the British boy shared on the internet.
[0247] In a similar manner, the electronic device (10) captures a second hand pose different from the first hand pose and provides a scene containing a second temporary scene object (2521) to the user (1), and can map a second persona to the second temporary scene object (2521) through the user (1)'s persona mapping action (e.g., the persona mapping action described in the preceding drawings). The electronic device (10) can map second temporary object description information (2522) (e.g., Summary Man) related to the second temporary scene object (2521) to the second persona. The second temporary object description information (2522) (e.g., Summary Man) can be output while the second temporary scene object (2521) is provided as a scene or at the instruction of the user (1). For example, a second persona (e.g., Summary Man) mapped to a second temporary scene object (2521) can activate an AI assistance function that provides a response to a user's (1) query based on person information related to Summary Man shared on the internet (or information about Summary Man accessible via the internet).
[0248] According to one embodiment, the electronic device (10) may capture a third hand pose different from the first and second hand poses and provide a scene including a third temporary scene object (2531) to the user (1). Through the user's (1) persona mapping action (e.g., the persona mapping action described in the preceding drawings), the electronic device (10) may map a third persona to the third temporary scene object (2531) and map third temporary object description information (2532) (e.g., elementary school teacher) related to the third temporary scene object (2531) to the third persona. The third temporary object description information (2532) (e.g., elementary school teacher) may be output while the third temporary scene object (2531) is provided as a scene or at the direction of the user (1). For example, a third persona (e.g., an elementary school teacher) mapped to a third temporary scene object (2531) can activate an AI assistance function that provides a response to the user's (1) query based on average person information (or information of a person having a level of knowledge as an elementary school teacher) defined as a universal or general elementary school teacher, based on the city or country where the user (1) resides and / or knowledge shared via the internet.
[0249] According to one embodiment, the electronic device (10) may capture a fourth hand pose different from the first to third hand poses and provide a scene including a fourth temporary scene object (2541) to the user (1). Through the user (1)'s persona mapping operation (e.g., the persona mapping operation described in the preceding drawings), the electronic device (10) may map the fourth persona to the fourth temporary scene object (2541) and map the fourth temporary object description information (2542) (e.g., Coder) related to the fourth temporary scene object (2541) to the fourth persona. The fourth temporary object description information (2542) (e.g., Coder) may be output while the fourth temporary scene object (2541) is provided as a scene or at the direction of the user (1). For example, a fourth persona (e.g., Coder) mapped to a fourth temporary scene object (2541) can activate an AI assistance function that provides a response to the user's (1) query based on average person information (or information of a person having a level of knowledge as a programmer or code writer) defined as a universal or general programmer (or code writer) based on knowledge standards shared through the city or country where the user (1) resides and / or the internet.
[0250] According to one embodiment, the electronic device (10) may capture a fifth hand pose different from the first to fourth hand poses and provide a scene including a fifth temporary scene object (2551) to the user (1). Through the user (1)'s persona mapping action (e.g., the persona mapping action described in the preceding drawings), the electronic device (10) may map the fifth persona to the fifth temporary scene object (2551) and map the fifth temporary object description information (2552) (e.g., artist) related to the fifth temporary scene object (2551) to the fifth persona. The fifth temporary object description information (2552) (e.g., artist) may be output while the fifth temporary scene object (2551) is provided as a scene or at the direction of the user (1). For example, a fifth persona (e.g., artist) mapped to a fifth temporary scene object (2551) can activate an AI assistance function that provides a response to the user's (1) query based on the city or country where the user (1) resides, and / or knowledge standards shared on the internet, and average person information defined as a universal or general artist (or information of a person having a level of knowledge as an artist).
[0251] According to one embodiment, the electronic device (10) may capture a sixth hand pose different from the first to fifth hand poses and provide a scene including a sixth temporary scene object (2561) to the user (1). Through the user (1)'s persona mapping operation (e.g., the persona mapping operation described in the preceding drawings), the electronic device (10) may map the sixth persona to the sixth temporary scene object (2561) and map the sixth temporary object description information (2562) (e.g., Rock N Roll Star) related to the sixth temporary scene object (2561) to the sixth persona. The sixth temporary object description information (2562) (e.g., Rock N Roll Star) may be output while the sixth temporary scene object (2561) is provided as a scene or at the direction of the user (1). For example, the sixth persona (e.g., Rock N Roll Star) mapped to the sixth temporary scene object (2561) can activate an AI assistance function that provides a response to the user's (1) query based on specific person information defined as Rock N Roll Star in the city or country where the user (1) resides, the internet, or a specific broadcasting medium.
[0252] According to one embodiment, when a user (1) needs a specific persona mapped to the aforementioned hand posture while viewing a specific scene provided by the electronic device (10), the user (1) may position the hand in the specific posture within the camera shooting range of the electronic device (10) while assuming the specific hand posture. When the hand in the specific posture enters the shooting range, the electronic device (10) may output a temporary scene object corresponding to the hand in the specific posture and temporary object description information mapped to the temporary scene object. Additionally or generally, the electronic device (10) may check the relationship between the scene object included in the current scene and the temporary scene object, and if the set integration condition is satisfied, group the scene object and the temporary scene object and map a corresponding integration persona. Subsequently, when a user query is collected through user speech, the electronic device (10) may generate and output a response to the user query based on an AI assistance function based on the integration persona.
[0253] As described above, the electronic device (10) according to one embodiment supports utilizing a persona-based AI assistance function mapped to a specific temporary scene object by possessing a specific temporary scene object or creating it through a gesture.
[0254] FIG. 25 is a diagram showing an example of an electronic device operation method related to persona mapping according to one embodiment.
[0255] Referring to FIGS. 1 through 25, in relation to an electronic device operation method for persona mapping according to one embodiment, in operation 2571, a processor (120) of the electronic device (10) (e.g., voice recognition module (121a) of FIG. 7) can receive a voice command input from a user (1). For example, when a user intends to specify a specific scene object and an AI assistance function of a specific persona through a voice command, the processor (120) can collect and recognize the user's voice command.
[0256] In operation 2573, the processor (120) of the electronic device (10) (e.g., the voice recognition module (121a) of FIG. 7) can confirm (or determine) whether the user (1) has an intention to request a persona assignment based on the voice recognition result. If it is not a request for a persona assignment, a designation function can be performed in operation 2575. The designation function may include, for example, an AI assistance function using a persona.
[0257] In operation 2577, when it is determined that the user (1) intends to assign an AI assistance function of a specific persona to a specific scene object (or is a persona assignment request), the processor (120) of the electronic device (10) can identify the "referred object" and "assigned Persona" from the user's (1) query.
[0258] Once the designated object and designated persona are identified, in operation 2579, the processor (120) of the electronic device (10) (e.g., the object recognition module (121b) of FIG. 7) can collect multimodal data using at least one sensor (e.g., the sensor circuit (140) of FIG. 2) and perform scene object extraction through the multimodal data. For example, the processor (120) can extract scene objects using at least some of gaze information, gesture information, and / or image information.
[0259] In operation 2581, the processor (120) of the electronic device (10) (e.g., the similar object management module (122d) of FIG. 7) can perform feature point extraction and classification of the extracted scene object. For example, the processor (120) can perform feature point extraction and classification of the scene object based on RGB information and TOF information.
[0260] Next, in operation 2583, the processor (120) of the electronic device (10) (e.g., the object ID (identification) recognition module (122a) and / or similar object management module (122d) of FIG. 7) can determine whether the request is for designation to an object ID or a request for designation to an object Category.
[0261] When receiving user input designating a specific object ID, in operation 2585, the processor (120) of the electronic device (10) (e.g., the persona management module (122c) of FIG. 7) can assign a designated persona to the designated object (ID) and store it in a database (DB).
[0262] When receiving user input specifying an object Category, in operation 2587, the processor (120) of the electronic device (10) (e.g., the similar object management module (122d) of FIG. 7) can assign a specified persona to the designated object (Category) and store it in a database (DB).
[0263] As described above, the processor (120) of the electronic device (10) according to one embodiment may store data for each scene object depending on whether the user wants to map an AI assistance function to a specific object (ID) or to an object of the same type as the designated object (or designated object category), and may also store the type of AI assistance function mapped to the stored data. For example, in the case of a specific scene object, the processor (120) may store feature data that can distinguish the scene object, and in the case of a designated object, may store attribute data that the object possesses.
[0264] FIG. 26 is a diagram showing another example of an electronic device operation method related to persona mapping according to one embodiment.
[0265] Referring to FIGS. 1 through 26, in relation to an electronic device operation method for persona mapping according to one embodiment, in operation 2601, the processor (120) of the electronic device (10) can collect an audio signal. In this regard, the electronic device (10) can activate a microphone to maintain a state in which ambient audio signals can be collected. When the audio signal is collected, the processor (120) can check in operation 2603 whether the persona mapping setting state is present. In this regard, the electronic device (10) can support a menu and / or mode related to the persona mapping setting. Alternatively, the electronic device (10) can check whether a speech related to the persona mapping setting is collected among user speech. If the persona mapping setting state is not present, in operation 2605, the electronic device (10) can support the performance of a designated function. For example, the processor (120) of the electronic device (10) can execute a user function corresponding to the collected audio signal and output a screen accordingly. Alternatively, the processor (120) may perform speech recognition on the collected audio signal and transmit the speech recognition result as input to an AI assistance function designated as a user query.
[0266] When the persona mapping setting state is active, in operation 2607, the processor (120) of the electronic device (10) can collect mapping queries. For example, the processor (120) can extract mapping queries related to persona mapping from the audio signals collected in operation 2601. Alternatively, in operation 2607, the processor (120) can output guidance information requesting additional user utterances related to persona mapping, and if the user (1) proceeds with additional utterances, perform voice recognition for the additional utterances and collect mapping queries related to persona mapping based on the voice recognition results.
[0267] In operation 2609, the processor (120) of the electronic device (10) may detect a scene object to map to a persona according to a mapping query. In relation to scene object detection, the electronic device (10) may classify a background and at least one scene object included in a scene (or camera image) captured by a camera, and select a scene object indicated by the user (1) among the at least one scene object. For example, the processor (120) may select a specific scene object included in the scene as a scene object for persona mapping based on at least one of the selections made by the user (1) via gesture, touch, or voice input. Additionally, or generally, operations 2607 and / or operations 2609 may be performed in any order. For example, the operation 2609 may precede the operation 2607. Alternatively, the operations 2607 and / or operations 2609 may be performed simultaneously.
[0268] In operation 2611, the processor (120) of the electronic device (10) can perform an attribute (e.g., feature points and / or classification) of a selected scene object. In this regard, the processor (120) can perform image analysis on the scene object to extract the attribute of the scene object. The processor (120) can detect a scene object having the attribute through an object recognition DB with respect to the extracted attribute. Based on the information stored in the object recognition DB, the processor (120) can record and / or store at least one of the type, name, size, shape, pattern, and color of the scene object. Additionally, the processor (120) can assign a unique ID to the scene object having the attribute based on the object recognition DB.
[0269] In operation 2613, the processor (120) of the electronic device (10) may perform persona mapping and / or storage on a scene object. For example, regarding persona mapping, the processor (120) may perform keyword extraction from a mapping query obtained in operation 2607. Based on the extracted keywords, the processor (120) may connect an AI assistance function highly relevant to at least one extracted keyword to the persona of the scene object. Additionally, the processor (120) may generate object description information for the scene object based on the extracted keywords and map the object description information to the scene object. The object description information may be output to the scene viewed by the user (1) according to the settings or in response to a user call.
[0270] In operation 2615, the processor (120) of the electronic device (10) can check whether a termination event occurs to terminate the persona mapping. If no termination event occurs, the processor (120) can branch back to operation 2601 and re-perform the following operations. If a termination event to terminate the persona mapping occurs (e.g., user input instructing to end the mode, unwearing of the electronic device (10)), the electronic device (10) can perform a function termination.
[0271] In the process of performing the above-described operation, the electronic device (10) may, after recognizing and classifying scene objects included in the scene, construct an object list and generate and / or provide to the user (1) a persona list that can be mapped to the scene objects included in the object list. The persona list may include at least one persona preferred through a previous persona mapping history. Alternatively, the persona list may include recommended personas corresponding to the history of personas being mapped by attributes (e.g., feature points and / or types) of the scene objects. Alternatively, a persona list containing personas currently available to the electronic device (10) may be displayed on the user's (1) viewing screen (e.g., scene). Alternatively, personas used by a certain number of users may be displayed on the scene as a persona list.
[0272] According to one embodiment, when a user (1) attempts to specify an AI assistance function for a specific scene object and a specific persona through a voice command, the electronic device (10) can determine the user's (1) current intention based on the user's (1) voice command. If it is determined that the user (1) has an intention to specify an AI assistance function for a specific scene object and a specific persona, the electronic device (10) can identify the scene object referred to by the user (1) using various sensors. Once the scene object referred to by the user (1) is identified, the electronic device (10) extracts the feature points of the scene object and can then determine whether the user (1) wants to map an AI assistance function to a specific scene object (ID) or to an object of the same type as the referred object (category). The electronic device (10) can store data for each scene object and also store the type of AI assistance function mapped to the stored data. For example, the electronic device (10) can store feature data that can distinguish a specific scene object in the case of a specific scene object, and can store attribute data that the scene object has in the case of a designated object (or specific attribute, specific type, specific category).
[0273] At least one of the above-described operations 2601 to 2615 is executed by at least one instruction stored in memory (130), and the at least one instruction may be set so that the electronic device (10) performs a specific operation.
[0274] FIG. 27 is a diagram showing an example of an electronic device operation method related to persona integration according to one embodiment.
[0275] Referring to FIGS. 1 through 27, in a method for operating an electronic device related to persona integration, in operation 2701, the processor (120) of the electronic device (10) can receive and collect user input. In relation to receiving user input, the processor (120) can activate at least one of the input means, activate a microphone in relation to voice input, or activate a camera in relation to gesture input. Alternatively, the processor (120) can perform input signal, audio signal, and image collection input through at least one of the input means, microphone, and camera in an activated state.
[0276] In operation 2703, the processor (120) of the electronic device (10) can determine whether the received user input is a user input related to the activation of an AI assistance function. If the user input is not related to the activation of an AI assistance function, the processor (120) can process the performance of a function according to the type or characteristics of the user input in operation 2705. For example, the processor (120) can play specific augmented reality content according to the user input and provide it for the user (1) to watch. Alternatively, the processor (120) can perform an internet connection according to the user input, receive service-related data provided by a specific server device, and provide it for the user (1) to watch.
[0277] When user input regarding the activation of AI assistance functions is received, in operation 2707, the processor (120) of the electronic device (10) may perform scene recognition. In this regard, the processor (120) may use a camera to acquire a camera image for a set direction and construct a scene that the user (1) can view. The scene may include at least a portion of the camera image captured by the camera. Alternatively, the scene may include the camera image and / or additionally set virtual content. In relation to the scene recognition, the processor (120) may detect the background of the scene and / or at least one scene object included in the background.
[0278] In operation 2709, the processor (120) of the electronic device (10) can determine whether a scene object exists in the scene. In relation to scene object detection, the processor (120) can detect scene objects that satisfy set conditions, excluding the background, and perform classification of the detected scene objects. In relation to object classification, the processor (120) can determine whether there is a scene object mapped to the detected information based on an object recognition DB. If there is detected information that matches the information stored in the object recognition DB, the processor (120) can obtain at least one of the type, name, size, shape, color, and pattern of the scene object.
[0279] If no scene object exists, in operation 2711, the processor (120) of the electronic device (10) can process the performance of a function based on additional user input. For example, the processor (120) can play specific content depending on the type of additional user input. Alternatively, if the additional user input is a specific gesture corresponding to a temporary scene object, the processor (120) can recognize the gesture as a temporary scene object and proceed to step 2713.
[0280] If scene objects exist, in operation 2713, the processor (120) of the electronic device (10) can determine whether there are multiple scene objects. If there are multiple scene objects in the scene currently being viewed, in operation 2715, the processor (120) of the electronic device (10) can detect the distance between the scene objects. Alternatively, the processor (120) may determine whether there are other conditions set for grouping in addition to detecting the distance between the scene objects. Alternatively, regarding the grouping of scene objects, multiple conditions (e.g., conditions where the size of the scene objects is greater than or equal to a certain size and located within a certain distance) may be included. Additionally, or generally, each of the multiple scene objects included in the scene may include a state in which a persona is mapped through the persona mapping process described above with reference to FIG. 26.
[0281] After detecting the distance between scene objects, in operation 2717, the processor (120) of the electronic device (10) can check whether the detected distance is greater than or equal to a set reference value (Th1). If the detected distance is less than the reference value (Th1), in operation 2719, the processor (120) of the electronic device (10) can perform an integrated persona call. According to one embodiment, if the distance between a first scene object mapped to a first persona and a second scene object mapped to a second persona is within the reference value (Th1), the processor (120) can call an integrated persona that integrates the first persona and the second persona. In relation to the integrated persona call, the processor (120) can integrate a first domain supporting a first AI assistance function mapped to the first persona and a second domain supporting a second AI assistance function mapped to the second persona to integrate a database for generating a response. The processor (120) can control the third AI assistant function related to the integrated persona to generate a response based on an integrated database in which the domains are integrated during the process of generating a response. Alternatively, an AI module supporting the third AI assistant function may be configured to generate a response based on an integrated database.
[0282] According to one embodiment, when the detection distance between scene objects is greater than or equal to a reference value (Th1), in operation 2721, the processor (120) of the electronic device (10) may call a plurality of designated personas. For example, the processor (120) may call a first persona corresponding to a first scene object and a second persona corresponding to a second scene object, respectively. Subsequently, when a user query is input, the processor (120) may be controlled to generate and output a first response and / or a second response to the user query based on the first persona and / or the second persona, respectively, if no scene object is specified. When a specific scene object is specified as the user query is input (e.g., the first scene object is selected by user input), the processor may be controlled to generate and output a first response to the user query based on the first persona, respectively.
[0283] In operation 2713, if there are no multiple scene objects in the scene, for example, if there is only one scene object, in operation 2723, the processor (120) of the electronic device (10) may call a designated persona corresponding to the scene object. For example, if the scene contains only a fourth scene object to which a fourth persona is mapped, the processor (120) may activate an AI module that supports a fourth AI assistance function based on the fourth persona associated with the fourth scene object. Here, the fourth scene object included in the scene may include an object to which a fourth persona has been mapped through the persona mapping process described in FIG. 26. Alternatively, the electronic device (10) may automatically map a specific persona based on the attributes of the fourth scene object (e.g., at least one of feature points and / or types). The automatic persona mapping process may be performed by referring to the history of persona mapping performed by the user (1). For example, if a scene object having the same or similar attributes within a certain range as a fourth scene object included in a first scene is included in a second scene, the electronic device (10) can automatically map a persona identical to the persona mapped to the fourth scene object to the scene object included in the second scene. Alternatively, the electronic device (10) can check with the user (1) whether they agree to the same persona mapping, and through the user (1)'s confirmation, perform persona mapping to the scene object of the second scene.
[0284] Meanwhile, in the description related to some of the embodiments described above, the processor (120) activating an AI module is exemplified, wherein the operation of activating the AI module may include at least one of the operation of activating an AI module stored in the memory (130) of the electronic device (10) and the operation of forming a communication channel with an AI module stored in an external server device to provide a response service based on the AI module stored in the external server device. Accordingly, the operation of activating the AI module may be replaced or added with the operation of supporting a response service based on the AI module stored in the external server device.
[0285] In operation 2725, the processor (120) of the electronic device (10) can check whether a termination event occurs to terminate the use of the electronic device (10) according to the embodiment of the present description (e.g., scene recognition, scene object detection, persona calling for a single scene object, persona calling for multiple scene objects, or at least one of a combined persona calling for multiple scene objects having a relationship of setting conditions). If no termination event occurs, the processor (120) may branch back to before operation 2701 and re-perform the following operation. When a termination event occurs, the processor (120) may process the termination of the function.
[0286] At least one of the above-described operations 2701 to 2725 is executed by at least one instruction stored in memory (130), and the at least one instruction may be set so that the electronic device (10) performs a specific operation.
[0287] According to one embodiment, the electronic device (10) can check whether the AI assistance function is activated through the user's (1) gaze or voice call. Alternatively, the electronic device (10) can activate the AI assistance function through the user's (1) gaze or gaze information along with the user's (1) voice call. For example, if the user (1) gazes at a scene object mapped with the AI assistance function for a certain period of time or longer, or if the user gazes at the scene object mapped with the AI assistance function while simultaneously speaking a designated call word (e.g., Hi bixby), the electronic device (10) can recognize this and perform a process to activate the AI assistance function.
[0288] For example, the electronic device (10) can analyze sensor information to search for scene objects mapped with AI assistance functions. For example, the electronic device (10) can search for objects on the user's (1) field of view (FOV) using an RGB camera and a TOF sensor. Additionally, the electronic device (10) can analyze whether there are objects in a specific area using a gaze sensor. For example, the electronic device (10) can search for a pointable object, such as the user's (1) hand or a pen, through a camera, and check whether there are scene objects mapped with AI assistance functions in the area indicated by the pointable object.
[0289] According to one embodiment, if an AI assistance function is designated to a specific scene object (ID), the electronic device (10) can search for whether the scene object (ID) designated by the user (1) exists in the area where the scene object is designated. If the AI assistance function is mapped to a scene object (e.g., a cup) designated by the user (1), the electronic device (10) can check whether the scene object (e.g., a cup) designated by the user (1) exists in at least a part of the scene. If the user (1) sets the AI assistance function to an object type (category), the electronic device (10) can check whether an object (or scene object) corresponding to the object type designated by the user (1) exists in at least a part of the scene. For example, if the user (1) maps the AI assistance function to an object with the attribute of a cup, the electronic device (10) can check whether an object with the attribute of a cup exists in at least a part of the scene.
[0290] When a scene object to which the user (1) has mapped an AI assistance function exists within the scene, the electronic device (10) can determine whether the scene object is multiple or singular. When there are multiple objects to which the user (1) has mapped each AI assistance function, the electronic device (10) can determine the distance between the scene objects through a sensor. When multiple objects exist within a specific distance or less, the electronic device (10) can call each AI assistance function mapped to the multiple objects by integrating them into one.
[0291] For example, in a situation where a user (1) calls an AI assistant function through gaze information, if there is a cup A mapped to a plant Q&A domain and a cup B mapped to a math Q&A domain at the location where the user's (1) gaze is recognized, the electronic device (10) measures the distance between cup A and cup B. If the distance is less than a certain distance, an AI assistant function having a persona capable of handling both the plant Q&A and math Q&A domains (or mapped to a persona) is called, and if the distance is greater than a certain distance, an AI assistant function capable of handling plant Q&A and an AI assistant function capable of handling math Q&A can be called, respectively. In this process, the user (1) can specify a specific AI assistant function by gazing at a specific scene object through gaze input.
[0292] Meanwhile, although the above-described FIGS. 25 to 27 describes an electronic device operation method for persona mapping and / or grouping of personas described herein, the electronic device operation method described herein is not limited thereto. For example, the electronic device (10) described herein may support each electronic device operation method corresponding to various persona usage situations described above in FIGS. 12 to 24, and in this regard, the electronic device (10) may be configured to perform an operation regarding the electronic device operation method described in each embodiment based on a processor (120) that executes at least one instruction stored in memory (130).
[0293] Additionally or generally, although various embodiments related to the use of personas described in FIGS. 7 through 27 have been described above, the present description is not limited thereto, and the embodiments described in each figure may be understood as embodiments in which at least a part is combined or integrated with embodiments described in other figures. For example, an electronic device operation method supporting an AI assistance function according to an embodiment of the present description may support performing at least one of the following: changing the name of a persona or changing a mapped scene object, inheriting a mapped persona to another scene object, creating an integrated persona for multiple personas, and generating and / or outputting at least one response to a query, following a persona mapping operation. Furthermore, the electronic device operation method may perform at least one of releasing a persona mapped to a scene object and / or releasing an integrated persona after at least one of the various operations described above.
[0294] According to one embodiment among the various embodiments described above, the electronic device according to the embodiment of the present invention comprises a microphone (182), a camera (170), a memory (130), and at least one processor (120). At least one instruction stored in the memory and configured to be executed by the at least one processor to perform the operation of the electronic device may be configured to detect a first scene object mapped to a first persona and a second scene object mapped to a second persona in a scene acquired through the camera, check whether the relationship between the first scene object and the second scene object in the scene satisfies a set condition, and if the set condition is satisfied, map an integrated persona that integrates the first persona and the second persona to at least one of the first scene object and the second scene object. Here, the first persona may be mapped to a first AI (artificial intelligence) assist function that generates a first response to a user query through a first domain corresponding to a first information database, the second persona may be mapped to a second AI assist function that generates a second response to the user query through a second domain corresponding to a second information database, and the integrated persona may be mapped to a third AI assist function that generates a third response to the user query through a domain corresponding to the first information database and the second information database.
[0295] According to one embodiment, the at least one command may be configured to perform the integrated persona mapping when the distance between the first scene object and the second scene object is within a set distance.
[0296] According to one embodiment, the at least one command may be configured to generate and output only the third response for the specific user query when receiving a specific user query.
[0297] According to one embodiment, the at least one command may be configured to generate and output the first response, the second response, and the third response for the specific user query when it receives a specific user query.
[0298] According to one embodiment, the at least one command may be configured to disable the integrated persona mapping if the first scene object or the second scene object moves and the set condition is not satisfied.
[0299] According to one embodiment, the at least one command is configured such that a third scene object moves to check whether the relationship between the first scene object, the second scene object, and the third scene object satisfies the set condition, and if the relationship satisfies the set condition, a new integrated persona is created by integrating the first persona, the second persona, and the third persona mapped to the third scene object, and the new integrated persona is mapped to the first scene object, the second scene object, and the third scene object, and the third persona is characterized by being mapped to a third AI (artificial intelligence) assistance function that generates a fourth response to a user query through a third domain corresponding to a third information database.
[0300] According to one embodiment, the at least one command is characterized by being configured to generate and output the fifth response for the specific user query based on a database integrating the first information database, the second database, and the third information database, using the new integrated persona when a specific user query is received.
[0301] According to one embodiment, the at least one command is characterized by being configured to generate and output at least one of the first response, the second response, the third response, and the fourth response for the specific user query when a specific user query is received.
[0302] A method for operating an electronic device according to one embodiment of the various embodiments described above includes, in a scene acquired through a camera, detecting a first scene object to which a first persona is mapped and a second scene object to which a second persona is mapped; checking whether the relationship between the first scene object and the second scene object in the scene satisfies a set condition; and if the set condition is satisfied, mapping an integrated persona, which is formed by integrating the first persona and the second persona, to at least one of the first scene object and the second scene object. The first persona may be mapped to a first AI (artificial intelligence) auxiliary function that generates a first response to a user query through a first domain corresponding to a first information database, the second persona may be mapped to a second AI auxiliary function that generates a second response to the user query through a second domain corresponding to a second information database, and the integrated persona may be mapped to a third AI auxiliary function that generates a third response to the user query through a domain corresponding to the first information database and the second information database.
[0303] According to one embodiment, the mapping operation may include the integrated persona mapping operation when the distance between the first scene object and the second scene object is within a set distance.
[0304] According to one embodiment, the method may further include the operation of receiving a specific user query and the operation of generating and outputting the third response for the specific user query.
[0305] According to one embodiment, the method may further include the operation of receiving a specific user query and the operation of generating and outputting the first response, the second response, and the third response for the specific user query.
[0306] According to one embodiment, the method may further include the operation of releasing the integrated persona mapping if the first scene object or the second scene object moves and the set condition is not satisfied.
[0307] The above method further comprises the operation of moving a third scene object to check whether the relationship between the first scene object, the second scene object, and the third scene object satisfies the set condition; the operation of creating a new integrated persona by integrating the first persona, the second persona, and the third persona mapped to the third scene object when the relationship satisfies the set condition; and the operation of mapping the new integrated persona to the first scene object, the second scene object, and the third scene object, wherein the third persona is characterized by having a third AI (artificial intelligence) assistance function mapped to it that generates a fourth response to a user query through a third domain corresponding to a third information database.
[0308] According to one embodiment, the method is characterized by further including the operation of receiving a specific user query and the operation of generating and outputting the fifth response for the specific user query based on a database that integrates the first information database, the second database, and the third information database based on the new integrated persona.
[0309] According to one embodiment, the method is characterized by further including the operation of receiving a specific user query and the operation of generating and outputting at least one of the first response, the second response, the third response, and the fourth response for the specific user query.
[0310] A computer-readable recording medium according to an embodiment of the present invention stores at least one instruction such that, when executed by an electronic device, the electronic device performs the operation of detecting a first scene object mapped to a first persona and a second scene object mapped to a second persona in a scene acquired through a camera, the operation of checking whether the relationship between the first scene object and the second scene object in the scene satisfies a set condition, and if the set condition is satisfied, the operation of mapping an integrated persona, which integrates the first persona and the second persona, to at least one of the first scene object and the second scene object. The first persona is mapped to a first AI (artificial intelligence) auxiliary function that generates a first response to a user query through a first domain corresponding to a first information database, the second persona is mapped to a second AI auxiliary function that generates a second response to the user query through a second domain corresponding to a second information database, and the integrated persona is mapped to a user query through a domain corresponding to the first information database and the second information database A third AI assistance function that generates a third response can be mapped.
[0311] An electronic device according to an embodiment of the present invention comprises a microphone (182), a camera (170), a memory (130), and at least one processor (120). At least one instruction stored in the memory and configured to be executed by the at least one processor to perform the operation of the electronic device is configured to detect a first scene object in a scene collected through the camera, receive a mapping query related to persona mapping to the first scene object, determine the type of persona to be mapped to the first scene object based on keywords included in the mapping query, and map the determined persona to the first scene object. The persona may include an AI (artificial intelligence) assistance function that generates a response to a user query based on an information database related to a domain set based on the keywords.
[0312] According to one embodiment, the at least one command may be configured to collect another scene through the camera, detect a second scene object in the other scene, analyze the attributes of the second scene object, and if the second scene object has the same attributes as the first scene object, to inherit the first persona to the second scene object.
[0313] According to one embodiment, the at least one command may be configured to detect a mapping trigger operation indicating the first scene object and to map the persona to the first scene object in response to the mapping trigger operation.
[0314] According to one embodiment, the method according to the various embodiments disclosed herein may be provided by being included in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)) or an application store (e.g., Play Store). TM It can be distributed online (e.g., downloaded or uploaded) through ) or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily created on a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.
[0315] According to various embodiments, each component (e.g., module or program) of the components described above may include a singular or multiple entities. According to various embodiments, one or more of the components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Generally or additionally, multiple components (e.g., module or program) may be integrated into a single component. In this case, the integrated component may perform one or more functions of each of the components of the multiple components in the same or similar manner as those performed by the corresponding component among the multiple components prior to the integration. According to various embodiments, operations performed by the module, program, or other components may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.
Claims
1. In an electronic device, Mike (182); Camera (170); Memory (130); It includes at least one processor (120) operatively connected to the microphone, the camera, and the memory; and At least one instruction stored in the memory and executed by at least one processor to perform the operation of the electronic device is, In a scene acquired through the camera, detect a first scene object mapped to a first persona and a second scene object mapped to a second persona, and Check whether the relationship between the first scene object and the second scene object satisfies the established conditions, and When the above-mentioned set condition is satisfied, the integrated persona formed by integrating the first persona and the second persona is configured to be mapped to at least one of the first scene object and the second scene object, and The above-mentioned first persona is mapped to a first AI (artificial intelligence) assistance function that generates a first response to a user query through a first domain corresponding to a first information database, and The second persona above is mapped to a second AI assistance function that generates a second response to the user query through a second domain corresponding to a second information database different from the first information database, and The above integrated persona is an electronic device supporting an AI assistance function, characterized in that a third AI assistance function is mapped to a domain corresponding to a first information database and a second information database to generate a third response to the user query.
2. In Paragraph 1, The above at least one instruction is, An electronic device supporting an AI assistance function, characterized by being configured to perform the integrated persona mapping when the distance between the first scene object and the second scene object is within a set distance.
3. In Paragraph 1, The above at least one instruction is, When a specific user query is received, An electronic device supporting an AI assistance function, characterized by being configured to generate and output at least one of the first response, the second response, or the third response for the specific user query.
4. In Paragraph 1, The above at least one instruction is, When a specific user query is received, An electronic device supporting an AI assistance function, characterized by being configured to generate and output the third response for the specific user query mentioned above.
5. In Paragraph 1, The above at least one instruction is, An electronic device supporting an AI assistance function, characterized in that the integrated persona mapping is disabled when the first scene object or the second scene object moves and the set condition is not satisfied.
6. In Paragraph 1, The above at least one instruction is, The third scene object moves to check whether the relationship between the first scene object, the second scene object, and the third scene object satisfies the set condition, and If the above relationship satisfies the above-set condition, a new integrated persona is created by integrating the first persona, the second persona, and the third persona mapped to the third scene object, and The above new integrated persona is configured to map to the above first scene object, the above second scene object, and the above third scene object, and The above-mentioned third persona is an electronic device supporting an AI assistance function, characterized in that a third AI (artificial intelligence) assistance function that generates a fourth response to a user query through a third domain corresponding to a third information database is mapped thereto.
7. In Paragraph 6, The above at least one instruction is, An electronic device supporting an AI assistance function, characterized by being configured to generate and output the fifth response to the specific user query based on a database integrating the first information database, the second database, and the third information database, using the new integrated persona when a specific user query is received.
8. In Paragraph 6, An electronic device supporting an AI assistance function, characterized in that the above-mentioned at least one command is configured to generate and output at least one of the first response, the second response, the third response, and the fourth response for the specific user query when a specific user query is received.
9. An operation to detect a first scene object mapped to a first persona and a second scene object mapped to a second persona in a scene acquired through a camera; An operation to check whether the relationship between the first scene object and the second scene object satisfies the established conditions; When the above-mentioned set condition is satisfied, the operation of mapping the integrated persona, which is the integration of the first persona and the second persona, to at least one of the first scene object and the second scene object; is included. The above-mentioned first persona is mapped to a first AI (artificial intelligence) assistance function that generates a first response to a user query through a first domain corresponding to a first information database, and The second persona above is mapped to a second AI assistance function that generates a second response to the user query through a second domain corresponding to a second information database different from the first information database, and A method for operating an electronic device that supports an AI assistance function, characterized in that the above-mentioned integrated persona is mapped to a third AI assistance function that generates a third response to the user query through a domain corresponding to a first information database and a second information database.
10. In Paragraph 9, The above mapping operation is, A method for operating an electronic device that supports an AI assistance function, characterized by including the operation of mapping the integrated persona when the distance between the first scene object and the second scene object is within a set distance.
11. In Paragraph 9, When a specific user query is received, the operation of generating and outputting the third response for the specific user query; A method for operating an electronic device that supports an AI assistance function, characterized by further including at least one of the following operations: receiving a specific user query and generating and outputting the first response, the second response, and the third response for the specific user query.
12. In Paragraph 9, A method for operating an electronic device that supports an AI assistance function, further comprising: an action of releasing the integrated persona mapping when the first scene object or the second scene object moves and the set condition is not satisfied.
13. In Paragraph 9, An operation to check whether the relationship between the first scene object, the second scene object, and the third scene object satisfies the set condition by moving the third scene object; If the above relationship satisfies the above-set condition, the action of creating a new integrated persona by integrating the first persona, the second persona, and the third persona mapped to the third scene object; The operation of mapping the new integrated persona to the first scene object, the second scene object, and the third scene object is further included. A method for operating an electronic device that supports an AI assistance function, characterized in that the above-mentioned third persona is mapped to a third AI (artificial intelligence) assistance function that generates a fourth response to a user query through a third domain corresponding to a third information database.
14. In Paragraph 13, When a specific user query is received, an action of generating and outputting the fifth response for the specific user query based on a database integrating the first information database, the second database, and the third information database based on the new integrated persona; or A method for operating an electronic device that supports an AI assistance function, characterized by further including at least one operation of: receiving a specific user query and generating and outputting at least one of the first response, the second response, the third response, and the fourth response for the specific user query.
15. In a computer-readable recording medium, When the above computer-readable recording medium is executed by an electronic device, the electronic device, An operation to detect a first scene object mapped to a first persona and a second scene object mapped to a second persona in a scene acquired through a camera; An operation to check whether the relationship between the first scene object and the second scene object satisfies the established conditions; If the above-set condition is satisfied, at least one instruction is stored to perform the operation of mapping an integrated persona, which is the integration of the first persona and the second persona, to at least one of the first scene object and the second scene object. The above-mentioned first persona is mapped to a first AI (artificial intelligence) assistance function that generates a first response to a user query through a first domain corresponding to a first information database, and The second persona above is mapped to a second AI assistance function that generates a second response to the user query through a second domain corresponding to a second information database different from the first information database, and A computer-readable recording medium characterized in that the above-mentioned integrated persona is mapped to a third AI assistance function that generates a third response to the user query through a domain corresponding to the first information database and the second information database.