Directing a virtual agent based on user's eye behavior
By displaying a virtual agent on an electronic device and guiding its movements using an eye tracker, the problem of inconvenient user interaction is solved, resulting in a more intuitive and accurate user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- APPLE INC
- Filing Date
- 2022-07-19
- Publication Date
- 2026-06-16
AI Technical Summary
Interactions between users and virtual agents are often cumbersome, leading to user discomfort, inaccurate input, and a lack of intuitive interaction methods.
By displaying a virtual agent associated with the user on an electronic device and using eye trackers to obtain eye behavior data, the virtual agent can be guided to perform actions based on this data and contextual information.
It improves the intuitiveness and accuracy of user interaction, reduces user fatigue and discomfort, and enhances the user experience.
Smart Images

Figure CN122218945A_ABST
Abstract
Description
[0001] This application is a divisional application of invention patent application No. 202210854916.8, filed on July 19, 2022, entitled "Guiding Virtual Agents Based on User's Eye Behavior". Technical Field
[0002] This disclosure relates to the display of virtual agents, and more specifically to controlling virtual agents based on the user's eye behavior. Background Technology
[0003] In various scenarios, the device displays a virtual agent, and the user interacts with the virtual agent by providing user input to the device. However, user interaction is often cumbersome, such as when the user's hand moves across the environment toward the virtual agent to select or manipulate it. Therefore, user interaction can cause user discomfort, leading to inaccurate user input and a degraded overall user experience. Furthermore, the device lacks mechanisms to enable intuitive user interaction with the virtual agent. Summary of the Invention
[0004] According to some specific embodiments, a method is performed at an electronic device having one or more processors, non-transitory memory, and a display. The method includes displaying a virtual agent associated with a first viewing frustum on the display. The first viewing frustum includes a user avatar associated with a user, and the user avatar includes visual representations of one or more eyes. The method includes, while displaying the virtual agent associated with the first viewing frustum, obtaining eye-tracking data indicative of eye behavior associated with the user's eyes, updating the visual representations of one or more eyes based on the eye behavior, and guiding the virtual agent to perform actions based on the updated scene information associated with the electronic device.
[0005] According to some embodiments, an electronic device includes one or more processors, non-transitory memory, and a display. One or more programs are stored in the non-transitory memory and configured to be executed by one or more processors, and the one or more programs include instructions for performing or causing operations to be performed in any of the methods described herein. According to some embodiments, a non-transitory computer-readable storage medium stores instructions that, when executed by one or more processors of the electronic device, cause the device to perform or cause operations to be performed in any of the methods described herein. According to some embodiments, an electronic device includes means for performing or causing operations to be performed in any of the methods described herein. According to some embodiments, an information processing apparatus for use in an electronic device includes means for performing or causing operations to be performed in any of the methods described herein. Attached Figure Description
[0006] To better understand the various specific implementations described, reference should be made to the following detailed description in conjunction with the accompanying drawings, wherein similar reference numerals indicate corresponding parts in all the drawings.
[0007] Figure 1 It is a block diagram based on some specific implementation examples of portable multi-functional devices.
[0008] Figures 2A to 2U These are examples of how virtual agents are guided to perform various actions based on the user's corresponding eye behavior, according to specific implementations.
[0009] Figure 3 This is an example of a flowchart illustrating a method for guiding virtual agents to perform various actions based on the user's eye behavior, according to some specific implementations. Detailed Implementation
[0010] In various scenarios, the device displays a virtual agent, and the user interacts with the virtual agent by providing user input to the device. For example, the device includes a limb tracker that tracks the movement of the user's hand, and the device manipulates the display of the virtual agent based on this tracking. However, user interaction is often cumbersome, such as when the user's hand moves across the environment toward the virtual agent. Therefore, user interaction can cause physical or other discomfort to the user, such as hand fatigue. User discomfort often leads the user to provide inaccurate (e.g., unintended) input to the device, resulting in manipulation of the virtual agent that does not reflect the user's intent (or a lack thereof).
[0011] In contrast, the various specific embodiments disclosed herein include methods, systems, and electronic devices for guiding virtual agents to perform actions based on user eye behavior (e.g., eye movements) and scene information. For this purpose, the electronic device may include an eye tracker that acquires eye-tracking data indicative of user eye behavior. For example, eye behavior indicates focal position or eye patterns, such as rapid saccades or microsaccades. As an example, the electronic device determines that the user's focal position (e.g., gaze) is directed to a specific object, and in response, the electronic device guides the virtual agent to move toward that specific object or guides the virtual agent to move its own gaze toward that specific object. This action may include changing the appearance of one or more virtual eyes of the virtual agent, moving the virtual agent's body or head, emitting a sound from the virtual agent, and so on. For example, eye behavior includes a change from a first focal position to a second focal position. Thus, the electronic device guides the virtual agent's virtual eyes to change to a focal position less than a threshold distance from the second focal position. Examples of scene information include environment type (e.g., virtual reality (VR) environment, augmented reality (AR) environment, mixed reality (MR) environment), scene atmosphere (e.g., a dark and quiet room), information about objects within the scene, scene location (e.g., outdoors vs. indoors), and so on. Detailed Implementation
[0013] Reference will now be made in detail to specific embodiments, examples of which are illustrated in the accompanying drawings. Numerous specific details are shown in the following detailed description in order to provide a full understanding of the various described embodiments. However, it will be apparent to those skilled in the art that the various described embodiments can be practiced without these specific details. In other instances, well-known methods, processes, components, circuits, and networks are not described in detail so as not to unnecessarily obscure various aspects of the embodiments.
[0014] It will also be understood that, although in some cases the terms “first,” “second,” etc., are used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first contact may be named a second contact, and similarly, a second contact may be named a first contact, without departing from the scope of the various specific embodiments described. Both the first contact and the second contact are contacts, but they are not the same contact unless the context clearly indicates otherwise.
[0015] The terminology used in the description of the various embodiments herein is for the purpose of describing particular embodiments only and is not intended to be limiting. As used in the description of the various embodiments and in the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the term “and / or” as used herein refers to and covers any and all possible combinations of one or more of the associated listed items. It will also be understood that the terms “includes,” “including,” “comprises,” and / or “comprising” as used in this specification specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0016] As used herein, depending on the context, the term "if" is optionally interpreted as meaning "when," "at," "in response to determination," or "in response to detection." Similarly, depending on the context, the phrase "if it is determined..." or "if [the stated condition or event] is optionally interpreted as meaning "when it is determined..." or "in response to determination..." or "when [the stated condition or event] is detected," or "in response to the detection of [the stated condition or event]."
[0017] A physical environment refers to the physical world that people can sense and / or interact with without the aid of electronic devices. A physical environment can include physical features such as physical surfaces or physical objects. For example, a physical environment corresponds to a physical park that includes physical trees, physical buildings, and physical people. People can directly sense and / or interact with a physical environment through senses such as sight, touch, hearing, taste, and smell. Conversely, an extended reality (XR) environment refers to a fully or partially simulated environment that people sense and / or interact with via electronic devices. For example, an XR environment can include augmented reality (AR) content, mixed reality (MR) content, virtual reality (VR) content, etc. In the case of an XR system, a subset of a person's physical motion or a representation thereof is tracked, and in response, one or more features of one or more virtual objects simulated in the XR system are adjusted in a manner consistent with at least one physical law. For example, an XR system can detect head movement and, in response, adjust the graphical content and sound field presented to the person in a manner similar to how such views and sounds change in a physical environment. For example, an XR system can detect movement of electronic devices (e.g., mobile phones, tablets, laptops, etc.) that present the XR environment, and in response, adjust the graphical content and sound field presented to the user in a manner similar to how such views and sounds would change in a physical environment. In some cases (e.g., for accessibility reasons), an XR system can adjust the characteristics of the graphical content in the XR environment in response to representations of physical motion (e.g., voice commands).
[0018] Many different types of electronic systems enable people to sense and / or interact with a variety of XR environments. Examples include head-mounted systems, projection-based systems, head-up displays (HUDs), vehicle windshields with integrated display capabilities, windows with integrated display capabilities, displays shaped like lenses designed to be placed on a person's eyes (e.g., similar to contact lenses), headphones / earpieces, speaker arrays, input systems (e.g., wearable or handheld controllers with or without haptic feedback), smartphones, tablets, and desktop / laptop computers. Head-mounted systems may have an integrated opaque display and one or more speakers. Alternatively, head-mounted systems may be configured to receive external opaque displays (e.g., smartphones). Head-mounted systems may incorporate one or more imaging sensors for capturing images or video of the physical environment, and / or one or more microphones for capturing audio of the physical environment. Head-mounted systems may have transparent or semi-transparent displays instead of opaque displays. Transparent or semi-transparent displays may have a medium through which light representing the image is directed to the person's eyes. The display can utilize digital light projection, OLED, LED, uLED, liquid crystal on silicon, laser scanning light source, or any combination of these technologies. The medium can be an optical waveguide, holographic medium, optical combiner, optical reflector, or any combination thereof. In some implementations, transparent or translucent displays can be configured to selectively become opaque. Projection-based systems can employ retinal projection technology, which projects graphic images onto the human retina. Projection systems can also be configured to project virtual objects onto a physical environment, such as as holograms or on a physical surface.
[0019] Figure 1This is a block diagram of an example of a portable multi-functional device 100 (sometimes referred to herein as "electronic device 100" for brevity) according to some specific implementation. Electronic device 100 includes a memory 102 (e.g., one or more non-transitory computer-readable storage media), a memory controller 122, one or more processing units (CPUs) 120, a peripheral device interface 118, an input / output (I / O) subsystem 106, a display system 112, an inertial measurement unit (IMU) 130, an image sensor 143 (e.g., a camera), a contact strength sensor 165, an audio sensor 113 (e.g., a microphone), an eye-tracking sensor 164 (e.g., included within a head-mounted device (HMD)), a limb-tracking sensor 150, and other input or control devices 116. In some specific implementations, electronic device 100 corresponds to one of a mobile phone, tablet computer, laptop computer, wearable computing device, head-mounted device (HMD), head-mounted housing (e.g., electronic device 100 slides to or is otherwise attached to a head-mounted housing), etc. In some implementations, the head-mounted housing is shaped to form a receiver for receiving electronic equipment 100 with a display.
[0020] In some embodiments, the peripheral interface 118, one or more processing units 120, and memory controller 122 are optionally implemented on a single chip, such as chip 103. In other embodiments, they are optionally implemented on separate chips.
[0021] I / O subsystem 106 couples input / output peripherals on electronic device 100, such as display system 112 and other input or control devices 116, to peripheral interface 118. I / O subsystem 106 optionally includes display controller 156, image sensor controller 158, intensity sensor controller 159, audio controller 157, eye-tracking controller 160, one or more input controllers 152 for other input or control devices, IMU controller 132, limb tracking controller 180, and privacy subsystem 170. One or more input controllers 152 receive electrical signals from / send electrical signals to other input or control devices 116. Other input control devices 116 optionally include physical buttons (e.g., push-buttons, rocker buttons, etc.), dial pads, slide switches, joysticks, click wheels, etc. In some alternative embodiments, one or more input controllers 152 may optionally be coupled (or not coupled) to any of the following: a keyboard, an infrared port, a universal serial bus (USB) port, a stylus, a finger wearable device, and / or a pointing device such as a mouse. One or more buttons may optionally include push-buttons. In some embodiments, other input or control devices 116 include a positioning system (e.g., GPS) that obtains information about the position and / or orientation of the electronic device 100 relative to a particular object. In some embodiments, other input or control devices 116 include depth sensors and / or time-of-flight sensors that obtain depth information characterizing physical objects within the physical environment. In some embodiments, other input or control devices 116 include an ambient light sensor that senses ambient light from the physical environment and outputs corresponding ambient light data.
[0022] Display system 112 provides input and output interfaces between electronic device 100 and user. Display controller 156 receives electrical signals from display system 112 and / or sends electrical signals to display system 112. Display system 112 displays visual output to user. Visual output optionally includes graphics, text, icons, video, and any combination thereof (sometimes referred to herein as "computer-generated content"). In some embodiments, some or all of the visual output corresponds to user interface objects. As used herein, the term "visual representation" refers to a user-interactive graphical user interface object (e.g., a graphical user interface object configured to respond to input directed to the graphical user interface object). Examples of user-interactive graphical user interface objects include, but are not limited to, buttons, sliders, icons, selectable menu items, switches, hyperlinks, or other user interface controls.
[0023] Display system 112 may have a touch-sensitive surface, sensor, or sensor array to accept input from a user based on tactile and / or tactile contact. Display system 112 and display controller 156 (along with any associated modules and / or instruction set in memory 102) detect contact on display system 112 (and any movement or interruption of that contact) and translate the detected contact into interaction with user interface objects (e.g., one or more soft keys, icons, web pages, or images) displayed on display system 112. In an exemplary embodiment, the point of contact between display system 112 and the user corresponds to the user's finger or a finger-wearable device.
[0024] Display system 112 optionally employs LCD (Liquid Crystal Display) technology, LPD (Light Emitting Polymer Display) technology, or LED (Light Emitting Diode) technology, but other display technologies are used in other specific embodiments. Display system 112 and display controller 156 optionally employ any of a variety of touch sensing technologies now known or to be developed hereafter, as well as other proximity sensor arrays or other elements for determining one or more points of contact with display system 112, to detect contact and any movement or interruption thereof. These various touch sensing technologies include, but are not limited to, capacitive, resistive, infrared, and surface acoustic wave technologies.
[0025] Users may choose to use any suitable object or accessory, such as a stylus, a wearable finger device, or a finger, to interact with the display system 112. In some embodiments, the user interface is designed to work with finger-based contact and gestures, which may be less precise than stylus-based input due to the larger contact area of a finger on a touchscreen. In some embodiments, the electronic device 100 translates coarse finger-based input into precise pointer / cursor positions or commands to perform the actions desired by the user.
[0026] The audio circuitry also receives electrical signals converted from sound waves by the audio sensor 113 (e.g., a microphone). The audio circuitry converts the electrical signals into audio data and transmits the audio data to the peripheral interface 118 for processing. The audio data is optionally retrieved by the peripheral interface 118 from and / or transmitted to the memory 102 and / or the RF circuitry. In some specific implementations, the audio circuitry also includes a headset jack. This headset jack provides an interface between the audio circuitry and a removable audio input / output peripheral device, such as an output-only headset or a headset with both outputs (e.g., a mono or binaural headset) and inputs (e.g., a microphone).
[0027] The inertial measurement unit (IMU) 130 includes an accelerometer, a gyroscope, and / or a magnetometer to measure various force, angular rate, and / or magnetic field information relative to the electronic device 100. Therefore, depending on the specific implementation, the IMU 130 detects one or more positional change inputs of the electronic device 100, such as the electronic device 100 being rocked, rotated, or moved in a specific direction.
[0028] Image sensor 143 captures still images and / or video. In some embodiments, optical sensor 143 is located on the back of electronic device 100, opposite to the touchscreen on the front of electronic device 100, allowing the touchscreen to be used as a viewfinder for still image and / or video image acquisition. In some embodiments, another image sensor 143 is located on the front of electronic device 100, enabling the acquisition of images of the user (e.g., for selfies, for video conferencing while the user is viewing other video conferencing participants on the touchscreen, etc.). In some embodiments, the image sensor is integrated within the HMD. For example, image sensor 143 outputs image data representing physical objects (e.g., physical agents) within the physical environment.
[0029] A contact strength sensor 165 detects the strength of a contact on electronic device 100 (e.g., a touch input on a touch-sensitive surface of electronic device 100). The contact strength sensor 165 is coupled to a strength sensor controller 159 in I / O subsystem 106. The contact strength sensor 165 optionally includes one or more piezoresistive strain gauges, capacitive force sensors, electro-force sensors, piezoelectric sensors, optical force sensors, capacitive touch-sensitive surfaces, or other strength sensors (e.g., sensors for measuring the force (or pressure) of a contact on a touch-sensitive surface). The contact strength sensor 165 receives contact strength information (e.g., pressure information or a substitute for pressure information) from the physical environment. In some embodiments, at least one contact strength sensor 165 is arranged juxtaposed with or adjacent to the touch-sensitive surface of electronic device 100. In some embodiments, at least one contact strength sensor 165 is located on the side of electronic device 100.
[0030] Eye-tracking sensor 164 detects the eye gaze of a user of electronic device 100 and generates eye-tracking data indicating the user's gaze location. In various specific embodiments, the eye-tracking data includes data indicating a fixed point (e.g., a point of attention) of the user on a display panel, such as a display panel within a head-mounted device (HMD), a head-mounted housing, or a head-up display.
[0031] The limb tracking sensor 150 acquires limb tracking data indicating the position of a user's limbs. For example, in some embodiments, the limb tracking sensor 150 corresponds to a hand tracking sensor that acquires hand tracking data indicating the position of a user's hand or fingers within a specific object. In some embodiments, the limb tracking sensor 150 utilizes computer vision techniques to estimate limb pose based on camera images.
[0032] In various embodiments, electronic device 100 includes a privacy subsystem 170 that includes one or more privacy setting filters associated with user information, such as user information included in limb tracking data, eye gaze data, and / or body position data associated with the user. In some embodiments, privacy subsystem 170 selectively prevents and / or restricts electronic device 100 or parts thereof from acquiring and / or transmitting user information. To this end, privacy subsystem 170 receives user preferences and / or choices from the user in response to prompting the user to make user preferences and / or choices. In some embodiments, privacy subsystem 170 prevents electronic device 100 from acquiring and / or transmitting user information unless and until informed consent is obtained from the user. In some embodiments, privacy subsystem 170 anonymizes (e.g., scrambles or obfuscates) certain types of user information. For example, privacy subsystem 170 receives user input specifying which types of user information privacy subsystem 170 anonymizes. As another example, privacy subsystem 170 may include certain types of user information, including sensitive and / or identifying information, through user-specified (e.g., automatic) anonymization.
[0033] Figures 2A to 2U These are examples of how virtual agents are guided to perform various actions based on the user's corresponding eye movements, according to specific implementation details. For example... Figure 2A As shown, the physical environment 200 includes a first physical wall 201, a second physical wall 202, and a physical side cabinet 204. The long edge (length) of the physical side cabinet 204 is substantially parallel to the second physical wall 202, and the short edge (width) of the physical side cabinet 204 is substantially parallel to the first physical wall 201.
[0034] The physical environment 200 also includes a user 50 holding an electronic device 210. The electronic device 210 includes a display 212 associated with a visible area 214 of the physical environment 200. The visible area 214 includes a portion of a first physical wall 201, a portion of a second physical wall 202, and a physical side cabinet 204. In some implementations, the electronic device 210 corresponds to a mobile device, such as a smartphone, tablet, wearable device, etc. The user 50 includes eyes 52, and the user 50's other eye is... Figure 2A Not shown in the image.
[0035] In some embodiments, electronic device 210 corresponds to a head-mounted device (HMD) including an integrated display (e.g., a built-in display). In some embodiments, electronic device 210 includes a head-mounted housing. In various embodiments, the head-mounted housing includes an attachment area to which another device having a display can be attached. In various embodiments, the head-mounted housing is shaped to form a receiver for receiving another device (e.g., electronic device 210) including a display. For example, in some embodiments, electronic device 210 slides / snapes into or otherwise attaches to the head-mounted housing.
[0036] In some specific implementations, electronic device 210 includes an image sensor, such as a scene camera. The image sensor can capture image data characterizing the physical environment 200. The image data can correspond to images or image sequences (e.g., video streams). Electronic device 210 may include the ability to combine image data with computer-generated content (e.g., ... Figure 2D The compositing system shown includes a virtual baseball 222 and a virtual agent 224. In some specific embodiments, the electronic device 210 includes a rendering system (e.g., a graphics processing unit (GPU)) that renders objects to generate corresponding computer-generated content.
[0037] In some embodiments, electronic device 210 includes a perspective display. The perspective display allows ambient light from physical environment 200 to pass through it, and the representation of the physical environment is a function of the ambient light. In some embodiments, the perspective display is an additional display that allows optical visibility through a physical surface, such as an optical HMD (OHMD). For example, unlike pure synthesis using image data, a perspective display is able to reflect projected images from the display while allowing the user's vision to pass through it.
[0038] like Figure 2B As shown, in some specific embodiments, electronic device 210 includes eye tracker 214. Eye tracker 214 acquires eye-tracking data indicating eye behavior associated with the eyes 52 of user 50. For example, eye behavior indicates one or more of the following: gaze, focal (e.g., fixation) position, eye movement, etc. Figure 2B As shown, based on eye-tracking data, eye tracker 214 determines the first user's gaze 54a. (As...) Figure 2C As shown, the first user's line of sight 54a intersects with the first physical wall 201 at the first focal point position 56a. In other words, the user 50's eye 52 is focused on a point or part of the first physical wall 201.
[0039] like Figure 2DAs shown, in some embodiments, electronic device 210 operates according to an operating environment 220, such as the XR environment described above. For this purpose, in some embodiments, electronic device 210 acquires image data characterizing the physical environment 200 via an image sensor. The image sensor may have a field of view that substantially corresponds to the viewable area 214 of display 212. Therefore, the image data includes a corresponding representation of the physical features of the physical environment 200. Thus, the operating environment 220 includes corresponding representations of the first physical wall 201, the second physical wall 202, and the physical side cabinet 204. Furthermore, the operating environment 220 includes various computer-generated content, including a virtual baseball 222, a virtual dog 224 residing on the physical side cabinet 204, and a user avatar 230. In some embodiments, electronic device 210 synthesizes the image data with the computer-generated content to generate the operating environment 220. In some embodiments, electronic device 210 displays the corresponding representation of the physical features on display 212, and further displays the virtual baseball 222 and the virtual dog 224.
[0040] User avatar 230 is associated with user 50 (e.g., visually represented). Therefore, user avatar 230 includes a visual representation of eyes 232 that can represent the eyes 52 of user 50. In some specific implementations, electronic device 210 determines the first avatar gaze 234a associated with user avatar 230 based on a first user gaze 54a associated with user 50. For example, as... Figure 2D As shown, the first avatar's line of sight 234a intersects with the first point 236a of the first physical wall 201, which roughly corresponds to the first focal point position 56a associated with the user 50's eye 52.
[0041] In some implementations, electronic device 210 participates in a coexistence session with another electronic device, enabling electronic device 210 and the other electronic device to operate concurrently according to operating environment 220. Therefore, the other electronic device can display a user avatar 230, and electronic device 210 can display a user avatar representing the user of the other electronic device.
[0042] The virtual dog 224 includes a virtual eye 226 associated with a first-view frustum 228a. Note that the first-view frustum 228a includes a user avatar 230. In other words, the virtual dog 224 focuses on the area of the operating environment 220 including the user avatar 230, enabling the virtual dog 224 to respond to the eye behavior of the user avatar 230. The electronic device 210 determines the eye behavior of the user avatar 230 based on the corresponding tracked eye behavior of the user 50's eyes 52.
[0043] like Figure 2EAs shown, eye tracker 214 tracks user 50's eye 52 and determines the change from a first user gaze 54a to a second user gaze 54b. The second user gaze 54b intersects with a second focal position 56b of the physical environment 200, which corresponds to points above and to the right of the physical side cabinet 204.
[0044] Based on the change in the second focal position 56b, the electronic device 210 updates the visual representation of the user avatar 230's eyes 232, such as... Figure 2F As shown. That is, the visual representation of eye 232 changes from the first avatar's gaze 234a to the second avatar's gaze 234b, which roughly corresponds to the second user's gaze 54b. Note that the second avatar's gaze 234b intersects with the second point 236b on the virtual dog 224. In other words, the focus of the user avatar 230 is guided to the virtual dog 224.
[0045] In some implementations, the focus of the user avatar 230 is directed to the virtual dongle 224, and the electronic device 224 activates the virtual dongle 224 (e.g., enabling the virtual dongle 224 to perform actions). In some implementations, such as Figure 2F and Figure 2G As shown, electronic device 210 changes the appearance of the virtual dongle 224 from a solid line boundary to a dashed line boundary to indicate activation. Changing the appearance of the virtual dongle 224 on display 212 provides feedback to user 50 that the virtual dongle 224 has been activated, thereby reducing the likelihood of user 50 providing subsequent input attempting to activate the virtual dongle 224, and thus reducing the resource utilization of electronic device 210.
[0046] like Figure 2H As shown, eye tracker 214 tracks user 50's eye 52 and determines the change from second user gaze 54b to third user gaze 54c. The third user gaze 54c intersects with a third focal point 56c of the physical environment 200. The third focal point 56c corresponds to a point on the top surface of the physical side cabinet 204.
[0047] Based on the change in the third focal position 56c, the electronic device 210 updates the visual representation of the user avatar 230's eyes 232, such as... Figure 2I As shown. That is, the visual representation of eye 232 changes from the second avatar's gaze 234b to the third avatar's gaze 234c, which roughly corresponds to the third user's gaze 54c. Note that the third avatar's gaze 234c intersects with the third point 236c on the virtual baseball 222. In other words, the focus of user avatar 230 is directed to the virtual baseball 222.
[0048] like Figure 2JAs shown, eye tracker 214 tracks user 50's eyes 52 and determines the change from a third user gaze 54c to a second user gaze 54b associated with a second focal point 56b. For example, user 50's eye behavior corresponds to a rapid saccade originating from the position of virtual dog 224 (such as...). Figure 2E As shown), move it to the position of physical side cabinet 204 (as shown). Figure 2H As shown), and move back to the position of virtual dog 224 (as shown). Figure 2J (As shown). Based on the change in the second focal position 56b, the electronic device 210 updates the visual representation of the user avatar 230's eyes 232 back to the second avatar gaze 234b associated with the second point 236b corresponding to the virtual dog 224, as shown. Figure 2K As shown.
[0049] Depending on the specific implementation, based on the eye behavior of the user avatar 230, the electronic device 210 guides the virtual dog 224 to perform one or more actions. For example, based on the visual representation of the eyes 232 moving from the virtual dog 224 to the virtual baseball 222 and back to the virtual dog 224, the electronic device 210 guides the virtual agent 224 to change the appearance of the virtual eyes 226. As an example, such as Figure 2L As shown, the electronic device 210 guides the virtual dog 224 to change its virtual eye 226 from a first viewing frustum 228a to a second viewing frustum 228b. For this purpose, in some embodiments, the electronic device 210 selects the second viewing frustum 228b to include the virtual baseball 222 because the user avatar 230 was previously focused on the virtual baseball 222. In other words, the electronic device 210 guides the virtual dog 224 to change its gaze target to roughly match the user avatar 230's previous gaze.
[0050] like Figure 2M As shown, eye tracker 214 tracks user 50's eye 52 and determines the change from a second user gaze 54b to a first user gaze 54a associated with a first focal point 56a. Based on the change in the first focal point position 56a, electronic device 210 updates the visual representation of user avatar 230's eye 232 to the first avatar gaze 234a associated with the second point 236a, as... Figure 2N As shown. Furthermore, because the user 50's focus has moved away from the virtual baseball 224 (e.g., for at least a threshold amount of time), the electronic device 210 guides the virtual dog 224 from the second viewing frustum 228b to the first viewing frustum 228a, as... Figure 2N As shown. The first viewing frustum 228a includes a user avatar 230, and therefore the virtual dog 224 can view the user avatar 230 and wait for further instructions from the user avatar 230.
[0051] Depending on the specific implementation, electronic device 210 guides virtual dog 224 to perform one or more actions based on one or more corresponding duration thresholds associated with the focus position. Examples of utilizing duration thresholds include... Figures 20 to 2U As shown in the image. Figure 2O As shown, eye tracker 214 tracks user 50's eye 52 and determines changes in the user's gaze 54c from a first user gaze 54a to a third user gaze 54c associated with a third focal point 56c on the surface of physical side cabinet 204. Based on the change in the third focal point position 56c, electronic device 210 updates the visual representation of user avatar 230's eye 232 from the first avatar gaze 234a to the third avatar gaze 234c, as... Figure 2P As shown. The third avatar's line of sight 234c is associated with the third point 236c located on the virtual baseball 222.
[0052] like Figure 2Q As shown, based on the change in point 236c, the electronic device 210 guides the virtual dog 224 to update the virtual eye 226 to change from the first viewing frustum 228a to the second viewing frustum 228b, as referenced. Figure 2K and Figure 2L As stated above.
[0053] In some specific implementations, electronic device 210 determines that the user avatar 230's eyes 232 maintain focus on a third point 236c (on the virtual baseball 222) for at least a first threshold duration. Based on the satisfaction of the first threshold duration, electronic device 210 guides the virtual dog 224 to move toward the virtual baseball 222, as... Figure 2R The first moving line 240 in the middle indicates this. Figure 2S The completion of the movement from virtual dog 224 to virtual baseball 222 is shown.
[0054] Furthermore, in some specific implementations, the electronic device 210 determines that the user avatar 230's eyes 232 maintain focus on the third point 236c (on the virtual baseball 222) for at least a second threshold duration, which is greater than a first threshold duration. For example, the first threshold duration is two seconds from when the user avatar 230's eyes 232 initially focus on the virtual baseball 222, while the second threshold duration is four seconds from when the user avatar 230's eyes 232 initially focus on the virtual baseball 222. Based on the satisfaction of the second threshold duration, the electronic device 210 guides the virtual dog 224 to bring the virtual baseball 222 to the user avatar 230, as... Figure 2T The second moving line 242 is shown in the diagram. Figure 2UThe illustration shows the completion of the movement of the virtual dog 224 and the virtual baseball 222 toward the user avatar 230. Furthermore, the electronic device 210 guides the virtual dog 224 to change its virtual eye 226 from being associated with a second viewing frustum 228b to a third viewing frustum 228c. The third viewing frustum 228c includes the eye 232 of the user avatar 230, enabling the virtual dog 224 to receive additional directions from the user avatar 230 (eye 232).
[0055] Figure 3 This is an example of a flowchart illustrating a method 300 in which a guided virtual agent performs various actions based on a user's eye behavior, according to some specific implementation. In various implementations, method 300 or a portion thereof is performed by an electronic device (e.g., electronic device 210). In various implementations, method 300 or a portion thereof is performed by a mobile device, such as a smartphone, tablet, or wearable device. In various implementations, method 300 or a portion thereof is performed by a head-mounted display (HMD). In some implementations, method 300 is performed by processing logic components (including hardware, firmware, software, or combinations thereof). In some implementations, method 300 is performed by a processor executing code stored in a non-transitory computer-readable medium (e.g., memory).
[0056] As shown in box 302, method 300 includes displaying a virtual agent associated with a first-view frustum on a display. Examples of virtual agents include various computer-generated entities such as humans, animals, humanoids, robots, apes, anthropomorphic entities, etc. As shown in box 304, the first-view frustum includes a user avatar associated with a user. The user avatar can provide a graphical representation of the user. For example, refer to... Figure 2C and Figure 2D The virtual agent corresponds to a virtual dog 224, which has a first visual frustum 228a, the first visual frustum including a user avatar 230 associated with the user 50. The user avatar includes visual representations of one or more eyes, such as... Figure 2D The visual representation of the eyes 232 of the user avatar 230 shown. The visual representation of one or more eyes may correspond to computer-generated eyes (e.g., the general eyes of an individual) or to the actual eyes of the user overlaid on the user avatar. For example, an electronic device captures an image of the user's eyes, identifies the eyes in the image (e.g., via computer vision), and overlays the eyes onto the user avatar.
[0057] As shown in box 306, method 300 includes obtaining eye-tracking data that indicates eye behavior associated with a user's eyes. Eye behavior can indicate the current focus position, such as the location where the user is looking or focusing within the physical environment. For example, referencing Figure 2CThe eye tracker 214 determines, based on eye tracking data, that the user 50's eye 52 is focused on a first focal position 56a located on the first physical wall 201.
[0058] As another example, as shown in box 308, eye behavior includes the movement of the user's eye from a first focal position to a second focal position. For example, refer to... Figure 2C and Figure 2E The eye tracker determines that eye 52 moves from a first focal position 56a to a second focal position 56b. As shown in box 310, in some embodiments, this movement includes rapid saccades, such as reference... Figure 2E , Figure 2H and Figure 2J As described above, rapid salivation can move between a primary focal position and a secondary focal position, such as moving from the origin to an object of interest and back to the origin. For example, the eye gaze is initially directed to the ground, moves to a virtual dog, and then moves backward toward the ground. Other examples of eye movement include smooth tracking, convergence and divergence, and vestibular eye movement.
[0059] As shown in box 312, method 300 includes updating the visual representations of one or more eyes based on eye behavior. For example, refer to... Figure 2D and Figure 2F The electronic device 210 is based on the corresponding movement of the user 50's eye 52 (in) Figure 2C and Figure 2E (As shown in the figure) the visual representation of eye 232 is changed from being directed to the first point 236a to being directed to the second point 236b.
[0060] As shown in box 314, in some implementations, method 300 includes determining an eye behavior indication activation request based on eye-tracking data, and activating a virtual agent in response to determining the eye behavior indication activation request. Once activated, the virtual agent can be guided to perform an action. In some implementations, the activation request corresponds to an avatar focused on the virtual agent. For example, see reference... Figure 2F and Figure 2G In response to determining that the second focus position 236b is on the virtual dongle 224, the electronic device 210 activates the virtual dongle 224. In some specific implementations, the activation request corresponds to focusing on the virtual agent for at least a threshold amount of time.
[0061] As shown in box 316, method 300 includes guiding a virtual agent to perform a first action based on updated and associated scene information of the electronic device. For example, the first action includes a change in the virtual agent's head pose, such as the virtual agent moving its head toward the user avatar. As another example, the first action includes the virtual agent emitting an audible sound, such as a virtual dog bark. In some implementations, the electronic device performing method 300 includes one or more environmental sensors that output environmental data, and method 300 includes determining scene information based on the environmental data. Examples of one or more environmental sensors include image sensors, depth sensors, simultaneous localization and mapping (SLAM) sensors, visual inertial odometry (VIO) sensors, global positioning system (GPS) sensors, etc.
[0062] In some implementations, when scene information indicates a first environment type, the first action corresponds to the first action type, and when scene information indicates a second environment type different from the first environment type, the first action corresponds to the second action type. Examples of environment types include virtual reality (VR) environments, augmented reality (AR) environments, mixed reality (MR) environments, etc. Other examples of scene information include scene atmosphere (e.g., a dark and quiet room), information about objects within the scene, scene location (e.g., outdoors vs. indoors), and so on. Furthermore, scene information may indicate mapping (e.g., a grid) that indicates multiple physical objects and surfaces, such as those determined based on SLAM data, point cloud data, etc. As an example, when scene information indicates a quiet environment, the electronic device guides the virtual agent to limit the volume of any sounds produced (e.g., emitted) by the virtual agent. As another example, when scene information indicates that a physical object blocks the straight path between the virtual agent and a focal point associated with the user's eye, the electronic device guides the virtual agent to move along a path avoiding the physical object in order to reach the focal point.
[0063] As shown in box 318, in some embodiments, the first action includes changing the appearance of one or more virtual eyes. As shown in box 320, in some embodiments, changing the appearance of one or more virtual eyes includes changing one or more virtual eyes from a first viewing frustum to a second viewing frustum. For example, refer to... Figures 2H to 2LBased on the user 50's eye 52 changing the focus between the virtual dog 224 and the virtual baseball 222, the electronic device 210 instructs the virtual dog 224 to change the association of its virtual eye 226 from that of a first viewing frustum 228a to that of a second viewing frustum 228b (including the virtual baseball 222). As another example, changing the appearance of one or more virtual eyes includes changing the color of one or more virtual eyes, enlarging one or more virtual eyes, shrinking one or more virtual eyes, etc. For example, based on scene information indicating that the virtual dog is outdoors, and based on the user's eye movement toward the ball, the electronic device guides the virtual dog to enlarge its eyes to indicate that the virtual dog is excitedly playing with the ball outdoors.
[0064] As shown in box 322, in some implementations, the first action includes the movement of the virtual agent from a first position within the operating environment to a second position within the operating environment. In some implementations, the movement of the virtual agent is based on detecting that the user's eyes maintain focus at a specific focal position for at least a threshold duration. As an example, based on determining that the user avatar 230's eyes 232 maintain focus at a third point 236c (on the virtual baseball 222) for at least a first threshold duration, the electronic device 210 guides the virtual dog 224 to move toward the virtual baseball 222, as shown. Figure 2R As indicated by the first moving line 240. Continuing this example, the electronic device 210 determines that the user avatar 230's eyes 232 maintain focus on the third point 236c (on the virtual baseball 222) for at least a second threshold duration, which is longer than the first threshold duration. Based on the second threshold duration, the electronic device 210 guides the virtual dog 224 to bring the virtual baseball 222 to the user avatar 230, as indicated by the first moving line 240. Figure 2T The second moving line 242 is shown in the diagram. Therefore, based on maintaining focus on a specific point or area for different durations, the electronic device can guide the virtual agent to perform different corresponding actions.
[0065] This disclosure describes various features, none of which alone can achieve the benefits described herein. It should be understood that the various features described herein can be combined, modified, or omitted, as will be apparent to those skilled in the art. Other combinations and sub-combinations beyond those specifically described herein will be apparent to those skilled in the art and are intended to form part of this disclosure. Various methods are described herein in conjunction with various flowchart steps and / or stages. It should be understood that in many cases, certain steps and / or stages can be combined such that multiple steps and / or stages shown in the flowchart can be performed as a single step and / or stage. Additionally, certain steps and / or stages can be divided into additional sub-components to be performed independently. In some cases, the order of steps and / or stages can be rearranged, and certain steps and / or stages can be omitted entirely. Furthermore, the methods described herein should be understood to be broadly interpretable, such that additional steps and / or stages beyond those shown and described herein can also be performed.
[0066] Some or all of the methods and tasks described herein can be performed and fully automated by a computer system. In some cases, the computer system may include multiple different computers or computing devices (e.g., physical servers, workstations, storage arrays, etc.) that communicate and interoperate via a network to perform the functions described herein. Each such computing device typically includes a processor (or multiple processors) that executes program instructions or modules stored in memory or other non-transitory computer-readable storage media or devices. The various functions disclosed herein may be implemented in such program instructions, but alternatively, some or all of the disclosed functions may be implemented in the computer system's dedicated circuitry (e.g., ASIC, FPGA, or GP-GPU). In cases where the computer system includes multiple computing devices, these devices may be located in the same location or not. The results of the disclosed methods and tasks can be persistently stored by converting physical storage devices such as solid-state memory chips and / or disks into different states.
[0067] The various processes defined herein take into account options for obtaining and using users' personal information. For example, such personal information may be used to provide improved privacy screens on electronic devices. However, the extent to which such personal information is collected should be based on the user's informed consent. As described herein, users should understand and control the use of their personal information.
[0068] Personal information will be used by the appropriate parties only for lawful and reasonable purposes. Parties using such information will comply with privacy policies and practices that are at least in accordance with applicable laws and regulations. Furthermore, such policies should be comprehensive, user-accessible, and considered to meet or exceed government / industry standards. In addition, parties may not distribute, sell, or otherwise share such information except for any reasonable and lawful purpose.
[0069] However, users can limit the extent to which parties can access or otherwise obtain their personal information. For example, settings or other preferences can be adjusted so that users can decide whether their personal information can be accessed by various entities. Furthermore, while some of the characteristics defined herein are described in the context of the use of personal information, aspects of these characteristics can be implemented without the need for such information. For example, if user preferences, account names, and / or location history are collected, this information can be obfuscated or otherwise generalized so that it does not identify the corresponding user.
[0070] This disclosure is not intended to be limited to the specific embodiments shown herein. Various modifications to the specific embodiments described herein will be apparent to those skilled in the art, and the general principles defined herein can be applied to other specific embodiments without departing from the spirit or scope of this disclosure. The teachings of the invention provided herein can be applied to other methods and systems, and are not limited to those described above, and elements and actions of the various specific embodiments described above can be combined to provide further specific embodiments. Therefore, the novel methods and systems described herein can be implemented in many other forms; furthermore, various omissions, substitutions, and changes can be made to the form of the methods and systems described herein without departing from the spirit of this disclosure. The appended claims and their equivalents are intended to cover such forms or modifications that fall within the scope and spirit of this disclosure.
Claims
1. A method comprising: In electronic devices that include one or more processors, non-transitory memory, and displays: A virtual agent at a first location within the scene is displayed on the monitor; Obtain scene information that characterizes the scene; Obtain eye-tracking data that indicates eye behavior associated with the user's eyes in relation to the electronic device; as well as The virtual agent is guided to perform a first action based on the eye behavior associated with the user's eyes and the scene information.
2. The method according to claim 1, further comprising: Simultaneously displayed on the display are a user avatar associated with the user at a second location within the scene and a virtual agent associated with a first view frustum relative to the scene, wherein the user avatar includes visual representations of one or more eyes, and wherein the user avatar is within the first view frustum of the virtual agent; and The visual representation of one or more eyes of the user avatar is updated based on the eye behavior associated with the user's eyes.
3. The method of claim 1, wherein the eye behavior includes the movement of the user's eye from a first focal position to a second focal position.
4. The method of claim 3, wherein the movement includes rapid scanning.
5. The method of claim 4, wherein the rapid scan is directed to an object of interest.
6. The method of claim 3, wherein directing the virtual agent to perform the first action is in response to detecting the movement of the eye, the method further comprising: The detection process ensures that the user's eye maintains focus at the second focal position for at least a threshold duration. as well as In response to detecting that the user's eyes maintain focus for at least the threshold duration, the virtual agent is instructed to perform a second action different from the first action.
7. The method of claim 1, wherein the virtual agent comprises one or more virtual eyes, and wherein the first action comprises changing the appearance of the one or more virtual eyes based on the eye behavior.
8. The method of claim 7, wherein changing the appearance of the one or more virtual eyes comprises moving the one or more virtual eyes based on movement of the user's eyes, wherein, Following the movement of one or more virtual eyes, the virtual agent is associated with a second viewing frustum, which is different from the first viewing frustum.
9. The method of claim 7, wherein the eye behavior includes a change from a first focal position to a second focal position, wherein the first action includes changing the one or more virtual eyes from a third focal position to a fourth focal position, and wherein the fourth focal position satisfies a proximity threshold relative to the second focal position.
10. The method of claim 1, wherein the first action includes moving the virtual agent from a first location within the operating environment to a second location within the operating environment.
11. The method of claim 1, wherein the first action includes a change in the head pose of the virtual agent.
12. The method of claim 1, wherein the first action includes the virtual agent emitting an audible sound.
13. The method according to claim 1, wherein: When the scene information indicates a first environment type, the first action corresponds to a first action type; and When the scene information indicates a second environment type that is different from the first environment type, the first action corresponds to the second action type, wherein the second action type is different from the first action type.
14. The method according to claim 1, further comprising: The eye behavior indication activation request is determined based on the eye tracking data; as well as The virtual agent is activated in response to the determination that the eye behavior indicates the activation request.
15. The method of claim 1, wherein the electronic device includes one or more environmental sensors that output environmental data, and the method further includes determining the scene information based on the environmental data.
16. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by an electronic device including a display, cause the electronic device to: A virtual agent at a first location within the scene is displayed on the monitor; Obtain scene information that characterizes the scene; Obtain eye-tracking data that indicates eye behavior associated with the user's eyes in relation to the electronic device; as well as The virtual agent is guided to perform a first action based on the eye behavior associated with the user's eyes and the scene information.
17. The non-transitory computer-readable storage medium of claim 16, wherein the eye behavior includes the movement of the user's eye from a first focal position to a second focal position.
18. The non-transitory computer-readable storage medium of claim 17, wherein the movement comprises a rapid scan directed to an object of interest.
19. The non-transitory computer-readable storage medium of claim 17, wherein directing the virtual agent to perform the first action is in response to detecting the movement of the eye, the method further comprising: The detection process ensures that the user's eye maintains focus at the second focal position for at least a threshold duration. as well as In response to detecting that the user's eyes maintain focus for at least the threshold duration, the virtual agent is instructed to perform a second action different from the first action.
20. An electronic device, comprising: One or more processors; Non-transitory memory; monitor; as well as One or more programs stored in the non-transitory memory, which, when executed by the one or more processors, cause the electronic device to: A virtual agent at a first location within the scene is displayed on the monitor; Obtain scene information that characterizes the scene; Obtain eye-tracking data that indicates eye behavior associated with the user's eyes in relation to the electronic device; as well as The virtual agent is guided to perform a first action based on the eye behavior associated with the user's eyes and the scene information.