Neutral avatar

By using neutral virtual characters that automatically determine visual characteristics in virtual reality and augmented reality technologies, the problem of protecting user privacy while interacting and communication in a multi-user environment is solved, and the balance between privacy protection and effective interaction is achieved.

JP7676416B2Active Publication Date: 2025-05-14MAGIC LEAP INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2022543678
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-01-27
Filing Date
2021-01-25
Publication Date
2025-05-14
Estimated Expiration
2041-01-25

AI Technical Summary

Technical Problem

Existing virtual reality (VR), augmented reality (AR) and mixed reality (MR) technologies are difficult to effectively interact and communicate in a multi-user environment while maintaining user privacy.

Method used

Neutral avatars are used to automatically determine their visual characteristics, making them unique but not specific identity characteristics in a multi-user environment, thereby achieving the protection of user privacy.

Benefits of technology

It realizes interaction and communication in a multi-user environment, while protecting users' privacy and personal characteristics, and avoiding unnecessary personal information leakage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007676416000001
    Figure 0007676416000001
  • Figure 0007676416000002
    Figure 0007676416000002
  • Figure 0007676416000003
    Figure 0007676416000003
Patent Text Reader

Abstract

Neutral avatars are vague with respect to the reference physical characteristics of the corresponding user, such as weight, ethnicity, gender, or even identity. Thus, neutral avatars may be desirable for use in various co-presence environments where users wish to maintain privacy related to the above-mentioned characteristics. Neutral avatars may be configured to convey the actions and behaviors of the corresponding user in real time without using factual forms of the user's actions and behaviors.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present disclosure relates to virtual reality and augmented reality imaging and visualization systems, including mixed reality, and more particularly to animating virtual characters such as avatars. [Background technology]

[0002] Modern computing and display technologies have facilitated the development of systems for so-called "virtual reality", "augmented reality" and "mixed reality" experiences, in which digitally reproduced images are presented to a user in a manner that appears or can be perceived as real. Virtual reality (VR) scenarios typically involve the presentation of computer-generated virtual image information without transparency to other actual real-world visual inputs. Augmented reality (AR) scenarios typically involve the presentation of virtual image information as an augmentation to the visualization of the real world around the user. Mixed reality (MR) is a type of augmented reality in which physical and virtual objects can coexist and interact in real time. The systems and methods disclosed herein address various challenges associated with VR, AR and MR technologies. Summary of the Invention [Means for solving the problem]

[0003] Details of one or more implementations of the subject matter described herein are set forth in the accompanying drawings and the following description. Other features, aspects, and advantages will become apparent from the description, drawings, and claims. Neither this summary nor the following detailed description purports to define or limit the scope of the inventive subject matter.

[0004] The embodiments of the present disclosure are directed to devices, systems, and methods for facilitating virtual or augmented reality interactions. As an exemplary embodiment, one or more user input devices may be used to interact in a VR, AR, or MR session. Such a session may include virtual elements or objects in a three-dimensional space. The one or more user input devices may further be used to point, select, annotate, and draw, among other actions, on virtual objects, real objects, or voids in an AR or MR session. For ease of reading and understanding, certain systems and methods discussed herein refer to an augmented reality environment or other "augmented reality" or "AR" components. These descriptions of "augmented reality" or "AR" should be construed to include "mixed reality," "virtual reality," "VR," "MR," and the like, as well as each of those "real environments" also specifically mentioned.

[0005] As disclosed herein, a "neutral avatar" is an avatar that is vague with respect to other characteristics, such as the ethnicity, gender, or even identity of the user, that may be determined based on a combination of the characteristics listed above and the physical characteristics of the avatar. Thus, these neutral avatars may be desirable to use in various co-existence environments where users wish to maintain privacy related to the characteristics described above. Neutral avatars may be configured to convey the actions and behaviors of a corresponding user without using a factual form of the user's actions and behavior in real time. The present invention provides, for example, the following: (Item 1) 1. A computing system comprising: A hardware computer processor; A non-transitory computer readable medium having software instructions stored thereon, the software instructions being executable by the hardware computer processor, the computing system comprising: providing co-presence environment data usable by multiple users to interact within an augmented reality environment; For each of multiple users, determining a visual distinctiveness of one or more neutral avatars for the user, the visual distinctiveness being different from the visual distinctiveness of neutral avatars of others of the plurality of users; updating the coexistence environment data to include the determined visual peculiarities of a neutral avatar; and A non-transitory computer readable medium for performing operations including: A computing system comprising: (Item 2) 2. The computing system of claim 1, wherein the visual distinctiveness comprises a color, texture, or shape of the neutral avatar. (Item 3) The operation further comprises: 2. The computing system of claim 1, further comprising: storing determined visual idiosyncrasies for a particular user; and determining a neutral avatar for the user comprising selecting a stored visual idiosyncrasy associated with the user. (Item 4) 2. The computing system of claim 1, wherein determining the visual distinctiveness of a neutral avatar for a user is performed automatically, regardless of personal characteristics of the user. (Item 5) 1. A computing system comprising: A hardware computer processor; A non-transitory computer readable medium having software instructions stored thereon, the software instructions being executable by the hardware computer processor, the computing system comprising: determining a neutral avatar associated with a user within the augmented reality environment, the neutral avatar does not include an indication of the user's gender, ethnicity, and identity; the neutral avatar is configured to represent input cues from the user with changes to visual elements of the neutral avatar that are unfactual indications of corresponding input cues; providing real-time rendering updates to the neutral avatar that are visible by each of a plurality of users within a shared augmented reality environment; A non-transitory computer readable medium for performing operations including: A computing system comprising: (Item 6) 6. The computing system of claim 5, wherein the first visual element is associated with two or more input cues. (Item 7) 7. The computing system of claim 6, wherein the input cues include one or more of gaze direction, speech, head pose, and hand pose. (Item 8) 8. The computing system of claim 7, wherein the input cues include one or more of shared attention, shared gaze, and mutual gestures. (Item 9) 6. The computing system of claim 5, further comprising a signal mapping component that stores a mapping between input cues and corresponding output signals, the real-time rendering updates being determined based on the output signals. (Item 10) Item 6. The computing system of item 5, wherein the neutral avatar includes a visual element that is deformable in response to audio input cues. (Item 11) Item 11. The computing system of item 10, wherein the visual element is otherwise deformable in response to input cues indicative of particular gaze activity. (Item 12) Item 6. The computing system of item 5, wherein the neutral avatar includes a visual element that changes size in response to audio input cues. (Item 13) Item 6. The computing system of item 5, wherein the neutral avatar includes a visual element that changes shading of a portion of the neutral avatar in response to an audio input cue. (Item 14) Item 14. The computing system of item 13, wherein the portion of the neutral avatar is not associated with a mouth area of ​​the neutral avatar. (Item 15) Item 6. The computing system of item 5, wherein the neutral avatar comprises one or more geometric shapes. [Brief description of the drawings]

[0006] Details of one or more implementations of the subject matter described herein are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages will become apparent from the description, drawings, and claims.

[0007] [Figure 1] FIG. 1 depicts an illustration of a mixed reality scenario involving a virtual reality object and a physical object viewed by a person.

[0008] [Diagram 2] FIG. 2 illustrates an example AR device that can be configured to provide an AR / VR / MR scene.

[0009] [Diagram 3] FIG. 3 diagrammatically illustrates example components of an AR device.

[0010] [Figure 4] FIG. 4 is a block diagram of another example of an AR device that may include an avatar processing and rendering system in a mixed reality environment.

[0011] [Figure 5A] FIG. 5A illustrates an exemplary avatar processing and rendering system.

[0012] [Figure 5B] FIG. 5B is a block diagram illustrating an example of components and signals associated with implementing a neutral avatar.

[0013] [Figure 6] 6A, 6B, and 6C illustrate an example neutral avatar with different visual features indicative of untrue input signals from one or more user sensors.

[0014] [Figure 7] 7A, 7B, and 7C illustrate another example neutral avatar, whose visual characteristics are adjusted based on one or more of a variety of input signals.

[0015] [Figure 8] 8A and 8B illustrate another example neutral avatar, whose visual characteristics may be modified based on one or more input signals.

[0016] [Figure 9] 9A, 9B, 9C, and 9D illustrate another exemplary neutral avatar, where the visual features include portions (eg, a ring and a circle) that respond differently to different input signals.

[0017] [Figure 10] 10A-10F illustrate six example neutral avatars, where modulation in visual characteristics can be tied to one or more different input signals.

[0018] [Figure 11] 11A-11I illustrate another example neutral avatar with various forms of visual characteristics that can be dynamically updated based on one or more input signals.

[0019] [Figure 12] 12A-12B illustrate another example neutral avatar, where morphing, movement, and / or other visual changes to parts of the neutral avatar may be mapped to one or more input signals.

[0020] Throughout the drawings, reference numbers may be reused to indicate correspondence between referenced elements. The drawings are provided to illustrate example embodiments described herein and are not intended to limit the scope of the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0021] Detailed Description overview A virtual avatar may be a virtual representation of a real or fictional person in an AR environment. For example, during a telepresence session in which two or more AR users interact with each other, a viewer may perceive another user's avatar in the viewer's environment, thereby creating a tangible sense of the other user's presence in the viewer's environment. Avatars may also provide a way for users to interact with each other and do things together in a shared virtual environment. For example, a student attending an online class may perceive and interact with avatars of other students or teachers in the virtual classroom. As another example, a user playing a game in an AR / VR / MR environment may see and interact with avatars of other players in the game.

[0022] Avatars may be modeled to mimic the appearance and personality of a human user, mirroring the movements of the user's body, head, eyes, lips, etc., in making the avatar's movements as lifelike as possible. These "factual avatars" may thus convey characteristics such as body shape, gender, height, weight, hair color, hair length, hairstyle, eye color, skin color, etc., to other users. In addition, such factual avatars may allow for direct mapping of user actions to avatar animations or sounds. For example, when a user speaks, the avatar may move its mouth. In some cases, an avatar that represents a user's factual appearance and actions may be desirable, while in other circumstances avatar neutrality is desired to protect the user's privacy with respect to these factual characteristics.

[0023] As disclosed herein, a "neutral avatar" is an avatar that is neutral with respect to other characteristics, such as the user's ethnicity, gender, or even identity, which may be determined based on a combination of the characteristics listed above, and the avatar's physical characteristics. Thus, these neutral avatars may be desirable to use in a variety of coexistence environments where a user wishes to maintain privacy related to the characteristics described above.

[0024] In some embodiments, a neutral avatar may be configured to convey the actions and behaviors of a corresponding user without using a factual form of the user's actions and behavior in real time. These behaviors may include, for example: Eye gaze (e.g. the direction a person is looking) Audio activity (e.g. who is speaking) Head position (i.e. where the person's attention is directed) Hand orientation (e.g., what a person is pointing at, holding, or discussing as part of the activity during a conversation). Advantageously, the neutral avatar may be animated in a manner that conveys communication, behavior, and / or social cues to others. For example, the user's actions (e.g., gaze direction change, head movement, speech, etc.) may be mapped to the visual cues of the neutral avatar that non-factually represent the user's actions. The neutral avatar may use geometric shapes, forms, and shapes to represent the user's behavior instead of factual human features. For example, input signals from sensors of an AR device worn by a user may be mapped to abstracted geometric forms to represent the user's behavior in real time using non-specific body parts. Because of their abstraction and minimization, these geometric forms avoid connoting specific gender, ethnicity, and can be easily shared among different users.

[0025] Some additional benefits of using a neutral avatar may include: Easier and faster setup. For example, rather than a user going through the process of selecting infinite characteristics that may be exhibited in different variants of the avatar, a neutral avatar has very limited customization options. In some embodiments, an avatar is automatically assigned to a user, eliminating the requirement for a user to select custom avatar characteristics altogether. Reduces computational resources (e.g., processor cycles, storage, etc.) used in avatar rendering. Because the focus is on conveying certain behavioral, social, and communication cues in a uniform manner for all users, complex avatar graphics specific to a particular user are not necessary. · Because a neutral avatar does not represent the specific characteristics of a corresponding user (e.g., specific identity, ethnicity, gender, etc.), it can be more easily shared among multiple users using a single AR device. · Not allowing users to disclose personal information via their avatars, as may be desired for business collaboration, for example. Allows for concealment of the user's visual form while still allowing collaboration and movement within the AR environment. Users do not need to make aesthetic choices regarding their avatar that may be conspicuous, distracting, or convey unintended messages, such as in an enterprise context. Representing real-time user behavior and actions. Example of 3D display for AR device

[0026] An AR device (also referred to herein as an augmented reality (AR) system) can be configured to present 2D or 3D virtual images to a user. The images may be still images, frames of video, or videos, in combination or the like. At least a portion of the AR device can be implemented on a wearable device, which may present a VR, AR, or MR environment, alone or in combination, for user interaction. A wearable device can be used synonymously with an AR device. Additionally, for purposes of this disclosure, the term "AR" is used synonymously with the terms "MR" and "VR."

[0027] Figure 1 depicts an illustration of a mixed reality scenario with certain virtual reality objects and certain physical objects viewed by a person. In Figure 1, an MR scene 100 is depicted in which a user of the MR technology sees a real-world park-like setting 110 featuring people, trees, buildings in the background, and a concrete platform 120. In addition to these items, the user of the MR technology also perceives that they "see" a robotic figure 130 standing on the real-world platform 120, and a flying cartoon-like avatar character 140 that appears to be an anthropomorphic bumblebee, although these elements do not exist in the real world.

[0028] In order for a 3D display to produce a true depth sensation, and more specifically, a simulated sensation of surface depth, it may be desirable to generate, for each point in the display's field of view, an accommodation response that corresponds to that point's virtual depth. If the accommodation response to a display point does not correspond to that point's virtual depth as determined by the binocular depth cues of convergence and stereopsis, the human eye may experience accommodation conflicts, resulting in unstable imaging, deleterious eye strain, headaches, and, in the absence of accommodative information, a near-complete lack of surface depth.

[0029] The AR experience can be provided by a display system having a display in which images corresponding to multiple depth planes are provided to the viewer. The images may be different for each depth plane (e.g., providing a slightly different presentation of a scene or object) and may be focused separately by the viewer's eyes, thereby serving to provide depth cues to the user based on the accommodation of the eyes required to focus on different image features of the scene located on different depth planes, or based on observing different image features on different depth planes that are out of focus. As discussed elsewhere herein, such depth cues provide a believable perception of depth.

[0030] FIG. 2 illustrates an example of an AR device 200, which can be configured to provide an AR scene. The AR device 200 can also be referred to as an AR system 200. The AR device 200 includes a display 220 and various mechanical and electronic modules and systems to support the functionality of the display 220. The display 220 can be coupled to a frame 230 that is wearable by a user, wearer, or viewer 210. The display 220 can be positioned in front of the eyes of the user 210. The display 220 can present AR content to the user. The display 220 can comprise a head-mounted display (HMD) that is worn on the user's head.

[0031] In some embodiments, a speaker 240 is coupled to the frame 230 and positioned adjacent to the user's ear canal (in some embodiments, another speaker, not shown, is positioned adjacent to the user's other ear canal to provide stereo / shapeable sound control). The display 220 can include an audio sensor (e.g., a microphone) 232 to detect audio streams from the environment and capture ambient sounds. In some embodiments, one or more other audio sensors, not shown, are positioned to provide stereo sound reception. Stereo sound reception can be used to determine the location of a sound source. The AR device 200 can perform voice or speech recognition on the audio stream.

[0032] The AR device 200 may include an outwardly facing imaging system that observes the world in the environment around the user. The AR device 200 may also include an inwardly facing imaging system that may track the eye movements of the user. The inwardly facing imaging system may track either the movement of one eye or the movement of both eyes. The inwardly facing imaging system may be mounted on the frame 230 and may be in electrical communication with a processing module 260 or 270 that may process image information acquired by the inwardly facing imaging system and determine, for example, the pupil diameter or orientation of the eyes of the user 210, the eye movement, or the eye posture. The inwardly facing imaging system may include one or more cameras. For example, at least one camera may be used to image each eye. The images acquired by the cameras may be used to determine the pupil size or eye posture for each eye separately, thereby allowing the presentation of image information to each eye to be dynamically adjusted for that eye.

[0033] As an example, the AR device 200 can use an outward-facing or inward-facing imaging system to obtain an image of the user's pose. The image may be a still image, a frame of a video, or a video.

[0034] The display 220 can be operably coupled (250) to a local data processing module 260, which can be mounted in a variety of configurations, such as fixedly attached to the frame 230, by wired leads or a wireless connection, fixedly attached to a helmet or hat worn by the user, built into headphones, or otherwise removably attached to the user 210 (e.g., in a backpack configuration, in a belt-attached configuration).

[0035] The local processing and data module 260 may comprise a hardware processor and digital memory such as non-volatile memory (e.g., flash memory), both of which may be utilized to aid in processing, caching, and storing data. The data may include a) data captured from sensors (e.g., that may be operatively coupled to the frame 230 or otherwise attached to the user 210), such as image capture devices (e.g., cameras in an inward-facing and / or outward-facing imaging system), audio sensors (e.g., microphones), inertial measurement units (IMUs), accelerometers, compasses, global positioning system (GPS) units, wireless devices, or gyroscopes, or b) data obtained or processed using the remote processing module 270 or remote data repository 280, possibly for passing to the display 220 after processing or retrieval. The local processing and data module 260 may be operatively coupled to a remote processing module 270 or a remote data repository 280, such as via a wired or wireless communication link, over a communication link 262 or 264, such that these remote modules are available as resources to the local processing and data module 260. In addition, the remote processing module 280 and the remote data repository 280 may be operatively coupled to each other.

[0036] In some embodiments, the remote processing module 270 may comprise one or more processors configured to analyze and process the data or image information. In some embodiments, the remote data repository 280 may comprise a digital data storage facility, which may be available through the Internet or other networking configurations in a "cloud" resource configuration. In some embodiments, all data is stored and all calculations are performed in the local processing and data module, allowing for fully autonomous use from the remote module. Exemplary Components of an AR Device

[0037] FIG. 3 diagrammatically illustrates exemplary components of an AR device. FIG. 3 illustrates an AR device 200, which may include a display 220 and a frame 230. A blowup 202 diagrammatically illustrates various components of the AR device 200. In some implementations, one or more of the components illustrated in FIG. 3 may be part of the display 220. The various components, alone or in combination, may collect various data (e.g., auditory or visual data, etc.) associated with a user of the AR device 200 or the user's environment. In some embodiments, the AR device 200 may have additional or fewer components depending on the application for which the AR device is used. It should be noted that FIG. 3 provides a basic idea of ​​some of the various components and the types of data that may be collected, analyzed, and stored through the AR device.

[0038] In the embodiment of FIG. 3, the display 220 includes a display lens 226 that may be mounted to the user's head or housing or frame 230. The display lens 226 may comprise one or more transparent mirrors positioned in front of the user's eyes 302, 304 by the housing 230 and may be configured to bounce the projected light 338 into the eyes 302, 304 to facilitate beam shaping while also allowing transmission of at least some light from the local environment. The wavefront of the projected light beam 338 may be bent or focused to match a desired focal length of the projected light. As shown, two wide-field machine vision cameras 316 (also referred to as world cameras) are coupled to the housing 230 and may image the environment around the user. These cameras 316 may be dual capture visible / non-visible (e.g., infrared) light cameras. The cameras 316 may be part of an outward-facing imaging system. Images acquired by the world camera 316 may be processed by a pose processor 336. For example, the pose processor 336 may implement one or more object recognizers to identify the pose of the user or another person in the user's environment, or to identify physical objects in the user's environment.

[0039] Continuing with reference to FIG. 3, a pair of scanning laser shaped wavefront (e.g., for depth) light projector modules with display mirrors and optics configured to project light 338 into the eyes 302, 304 are shown. The depicted diagram also shows two miniature infrared cameras 324 paired with infrared lights (such as light emitting diodes "LEDs") configured to track the user's eyes 302, 304 and support rendering and user input. The AR device 200 may further feature a sensor assembly 339, which may include X, Y, and Z axis accelerometer capabilities and magnetic compass and X, Y, and Z axis gyroscope capabilities, and may preferably provide data at a relatively high frequency, such as 200 Hz. The head pose processor 336 may include an ASIC (application specific integrated circuit), FPGA (field programmable gate array), or ARM processor (advanced reduced instruction set machine), which may be configured to calculate real-time or near real-time user head pose from the wide field of view image information output from the capture device 316. In some embodiments, head position information sensed from one or more sensors (e.g., 6DOF sensors) in a wearable headset is used to determine avatar characteristics and movement. For example, the user's head position may drive the avatar head position and provide an estimated location for the avatar torso and for the avatar's walking movements around space.

[0040] The AR device may also include one or more depth sensors 234. The depth sensor 234 may be configured to measure a distance between an object in the environment and the wearable device. The depth sensor 234 may include a laser scanner (e.g., lidar), an ultrasonic depth sensor, or a depth-sensing camera. In an implementation in which the camera 316 has depth-sensing capabilities, the camera 316 may also be considered a depth sensor 234.

[0041] Also shown is a processor 332 configured to perform digital or analog processing and derive attitude from gyroscope, compass, or accelerometer data from the sensor assembly 339. The processor 332 may be part of the local processing and data module 260 shown in FIG. 2. The AR device 200 may also include a positioning system, such as, for example, a GPS 337 (Global Positioning System), as shown in FIG. 3, to aid in attitude and positioning analysis. In addition, the GPS may further provide remote-based (e.g., cloud-based) information about the user's environment. This information may be used to recognize objects or information within the user's environment.

[0042] The AR device may combine data obtained by the GPS 337 and a remote computing system (e.g., the remote processing module 270, another user's AR device, etc.), which can provide more information about the user's environment. As an example, the AR device can determine the user's location based on the GPS data and retrieve a world map (e.g., by communicating with the remote processing module 270) that includes virtual objects associated with the user's location. As another example, the AR device 200 can monitor the environment using the world camera 316. Based on the images obtained by the world camera 316, the AR device 200 can detect objects in the environment (e.g., by using one or more object recognizers).

[0043] The AR device 200 may also include a rendering engine 334, which may be configured to provide local rendering information to the user for the user's view of the world and facilitate the operation of the scanner and imaging into the user's eye. The rendering engine 334 may be implemented by a hardware processor (e.g., a central processing unit or a graphics processing unit, etc.). In some embodiments, the rendering engine is part of the local processing and data module 260. The rendering engine 334 may be communicatively coupled (e.g., via wired or wireless links) to other components of the AR device 200. For example, the rendering engine 334 may be coupled to the eye camera 324 via communication link 274 and to the projection subsystem 318 (which may project light into the user's eyes 302, 304 via a scanning laser array in a manner similar to a retinal scanning display) via communication link 272. The rendering engine 334 may also communicate with other processing units, such as, for example, the sensor pose processor 332 and the image pose processor 336, via links 276 and 294, respectively.

[0044] A camera 324 (e.g., a small infrared camera) may be utilized to track eye pose and support rendering and user input. Some example eye poses may include where the user is looking or the depth at which they are focused (which may be estimated using eye convergence and divergence). A GPS 337, gyroscope, compass, and accelerometer 339 may be utilized to provide a rough or fast pose estimate. One or more of the cameras 316 may obtain images and poses, which, together with data from associated cloud computing resources, may be utilized to map the local environment and share the user view with others.

[0045] The exemplary components depicted in FIG. 3 are for illustrative purposes only. Multiple sensors and other functional modules are shown together for ease of illustration and description. Some embodiments may include only one or a subset of these sensors or modules. Furthermore, the locations of these components are not limited to the locations depicted in FIG. 3. Some components may be mounted or stored within other components, such as belt-mounted components, handheld components, or helmet components. As an example, the image pose processor 336, the sensor pose processor 332, and the rendering engine 334 may be located within a beltpack and configured to communicate with other components of the AR device via wireless communication, such as ultra-wideband, Wi-Fi, Bluetooth, or via wired communication. The depicted housing 230 is preferably head-mountable and wearable by a user. However, some components of the AR device 200 may be worn on other parts of the user's body. For example, the speaker 240 may be inserted into the user's ear to provide sound to the user.

[0046] With regard to projecting light 338 into the user's eyes 302, 304, in some embodiments, the camera 324 may be utilized to measure where the center of the user's eyes is geometrically converged, which generally coincides with the location of the eye's focus or "depth of focus." The three-dimensional surface of all points at which the eyes converge may be referred to as the "monocular locus." The focal distance may take on a finite number of depths, or may vary infinitely. Light projected from the vergence distance appears focused on the subject's eyes 302, 304, while light in front of or behind the vergence distance is blurred. Examples of wearable devices and other display systems of the present disclosure are also described in U.S. Patent Publication No. 2016 / 0270656, entitled "Methods and systems for diagnosing and treating health ailments," filed March 16, 2016, which is incorporated herein by reference in its entirety and for all purposes.

[0047] The human visual system is complex and providing a realistic perception of depth is difficult. A viewer of an object may perceive the object as three-dimensional due to a combination of vergence and accommodation. Vergence movements of the two eyes relative to one another (e.g., rotational movement of the pupils toward or away from one another to converge the gaze of the eyes and fixate on an object) are closely coupled to the focusing of the eye's lenses (or "accommodation"). Under normal conditions, a change in the focus of the eye's lenses or accommodation of the eye to change focus from one object to another at a different distance will automatically produce a matching change in vergence at the same distance, a relationship known as the "accommodation-vergence reflex." Similarly, a change in vergence will induce a matching change in accommodation under normal conditions. A display system that provides a better match between accommodation and vergence may produce a more realistic and comfortable simulation of a three-dimensional image.

[0048] Furthermore, spatially coherent light with a beam diameter of less than about 0.7 millimeters can be correctly resolved by the human eye regardless of where the eye is focused. Thus, to create the illusion of proper depth of focus, eye vergence and divergence may be tracked using the camera 324, and the rendering engine 334 and projection subsystem 318 may be utilized to render all objects on or near the monocular locus in focus and all other objects variable degrees out of focus (e.g., using intentionally created blur). Preferably, the system 220 renders to the user at a frame rate of about 60 frames per second or greater. As explained above, preferably the camera 324 may be utilized for eye tracking, and software may be configured to take up not only vergence and divergence geometry, but also focus location cues to serve as user input. Preferably, such a display system is configured with suitable brightness and contrast for daytime or nighttime use.

[0049] In some embodiments, the display system preferably has a latency of less than about 20 milliseconds for visual object alignment, an angular alignment of less than about 0.1 degrees, and a resolution of about 1 arc minute, which is believed to be approximately the limit of the human eye, without being limited by theory. The display system 220 may be integrated with a localization system, which may involve a GPS element, optical tracking, a compass, an accelerometer, or other data sources to aid in position and attitude determination. The localization information may be utilized to facilitate accurate rendering within the user's view of the relevant world (e.g., such information would facilitate the glasses knowing their location relative to the real world).

[0050] In some embodiments, the AR device 200 is configured to display one or more virtual images based on the accommodation of the user's eyes. Unlike traditional 3D display approaches that force the user to focus on where the image is projected, in some embodiments, the AR device is configured to automatically vary the focus of the projected virtual content to allow a more comfortable viewing of the one or more images presented to the user. For example, if the user's eyes have a current focus of 1m, the image may be projected to match the user's focus. If the user shifts the focus to 3m, the image is projected to match the new focus. Thus, rather than forcing the user to a predetermined focus, the AR device 200 of some embodiments allows the user's eyes to function in a more natural manner.

[0051] Such an AR device 200 may eliminate or reduce the incidence of eye strain, headaches, and other physiological symptoms typically observed with virtual reality devices. To achieve this, various embodiments of the AR device 200 are configured to project virtual images at variable focal distances through one or more variable focus elements (VFEs). In one or more embodiments, 3D perception may be achieved through a multi-plane focus system that projects images onto a fixed focal plane from the user. Other embodiments employ a variable plane focus, where the focal plane is moved back and forth in the z-direction to match the user's current state of focus.

[0052] In both the multi-plane focus system and the variable plane focus system, the AR device 200 may employ eye tracking to determine the convergence-divergence movement of the user's eyes, determine the user's current focus, and project the virtual image at the determined focus. In other embodiments, the AR device 200 includes a light modulator that variably projects a variable-focus light beam in a raster pattern across the retina through a fiber scanner or other light generating source. Thus, the ability of the display of the AR device 200 to project images at variable focal distances may not only facilitate accommodation for the user to view objects in 3D, but may also be used to compensate for the user's ocular abnormalities, as further described in U.S. Patent Publication No. 2016 / 0270656, which is incorporated herein by reference in its entirety. In some other embodiments, the spatial light modulator may project the image to the user through various optical components. For example, as further described below, the spatial light modulator may project the image onto one or more wave guides, which then transmit the image to the user. An example of avatar rendering in mixed reality

[0053] The AR device may employ various mapping-related techniques to achieve a high depth of field in the rendered light field. When mapping a virtual world, it is advantageous to capture all features and points in the real world and accurately depict virtual objects in relation to the real world. To this end, FOV images captured from a user of the AR device can be added to the world model by including new photos that convey information about various points and features of the real world. For example, the AR device can collect a set of map points (such as 2D or 3D points), find new map points, and render a more accurate version of the world model. The world model of the first user can be communicated (e.g., via a network such as a cloud network) to a second user so that the second user can experience the world surrounding the first user.

[0054] FIG. 4 is a block diagram of another embodiment of an AR device, which may include an avatar processing and rendering system 690 into an augmented reality environment. In this embodiment, the AR device 600 may include a map 620, which may include at least a portion of the data in a map database. The map may reside partially locally on the AR device and partially in a networked storage location (e.g., in a cloud system) accessible by a wired or wireless network. An attitude process 610 may run on the wearable computing architecture (e.g., the processing module 260 or the controller 460) and utilize data from the map 620 to determine the position and orientation of the wearable computing hardware or the user. The attitude data may be calculated from data collected on the fly as the user experiences the system and moves within the world. The data may include images, data from sensors (such as inertial measurement units, which generally include accelerometer and gyroscope components), and surface information associated with objects in the real or virtual environment.

[0055] The sparse representation may be the output of a simultaneous localization and mapping (e.g., SLAM or vSLAM, which refers to configurations where the input is image / vision only) process. The system can be configured to find not only where in the world various components are located, but also what the world consists of. Pose can be a building block that accomplishes many goals, including capturing maps and using data from maps.

[0056] In one embodiment, the sparse point locations may not be entirely adequate by themselves, and further information may be required to produce a multi-focal AR, VR, or MR experience. A dense representation, generally referring to depth map information, may be utilized to fill this gap, at least in part. Such information may be calculated from a process referred to as stereo 640, where depth information is determined using techniques such as triangulation or time-of-flight sensing. Image information and active patterns (such as infrared patterns created using an active projector), images obtained from an image camera, or hand gestures / totems 650 may serve as inputs to the stereo process 640. A significant amount of depth map information may be fused together, some of which may be summarized with a surface representation. For example, a mathematically definable surface may be an efficient (e.g., compared to a large point cloud) and easy-to-understand input to other processing devices such as a game engine. Thus, the outputs of the stereo process (e.g., depth map) 640 may be combined in a fusion process 630. The pose 610 may also be an input to this fusion process 630, the output of which becomes the input to the map capture process 620. The sub-surfaces may interconnect to form larger surfaces, as in topographic mapping, and the map becomes a large-scale hybrid of points and surfaces.

[0057] Various inputs may be utilized to resolve various aspects in the mixed reality process 660. For example, in the embodiment depicted in FIG. 4, game parameters may be input to determine that a user of the system is playing a monster battle game with one or more monsters in various locations, dying or fleeing monsters under various conditions (such as when the user shoots the monsters), walls or other objects in various locations, and the like. A world map may contain information about the location of objects or semantic information of objects (e.g., classification of whether an object is flat or round, horizontal or vertical, table or lamp, etc.), and may be another useful input to mixed reality. Attitude relative to the world is also an input and plays an important role for almost any interactive system.

[0058] A control or input from a user is another input to the AR device 600. As described herein, user input can include visual input, gestures, totems, audio input, sensory input, and the like. To move around or play a game, for example, the user may need to command the AR device 600 as to what he or she wants. There are various forms of user control that can be utilized beyond simply moving oneself in space. In one embodiment, an object such as a totem (e.g., a user input device) or a toy gun may be held by the user and tracked by the system. The system will preferably be configured to know that the user is holding an item and understand the type of interaction the user is having with the item (e.g., if the totem or object is a gun, the system may be configured to understand the location and orientation and whether the user is clicking a trigger or other sensed button or element (which may be equipped with sensors such as an IMU that can help determine what is occurring even when the activity is not within the field of view of any of the cameras).

[0059] Tracking or recognition of hand gestures may also provide input information. The AR device 600 may be configured to track and interpret hand gestures in terms of button presses, left or right, stop, grasp, hold gestures, and the like. For example, in one configuration, a user may want to flip an email or calendar in a non-gaming environment, or "fist bump" with another person or player. The AR device 600 may be configured to utilize a minimal amount of hand gestures, which may or may not be dynamic. For example, the gestures may be simple static gestures, such as opening the hand for stop, thumbs up for OK, thumbs down for not OK, or flipping the hand to the right or left or up / down for directional commands.

[0060] Eye tracking is another input (e.g., tracking where the user is looking, controlling display technology, rendering to a specific depth or range). In one embodiment, eye vergence may be determined using triangulation, and then accommodation may be determined using a vergence / accommodation model developed for that particular person. Eye tracking is performed by an eye camera, which can determine eye gaze (e.g., direction or orientation of one or both eyes). Other techniques can be used for eye tracking, such as, for example, measurement of electrical potentials with electrodes placed near the eyes (e.g., electro-oculography).

[0061] Speech tracking may be another input that may be used alone or in combination with other inputs (e.g., totem tracking, eye tracking, gesture tracking, etc.). Speech tracking may include speech recognition, voice recognition, alone or in combination. The system 600 may include an audio sensor (e.g., a microphone) that receives an audio stream from the environment. The AR device 600 may incorporate voice recognition technology to determine who is speaking (e.g., whether the speech is from the wearer of the ARD or another person or voice (e.g., recorded voice transmitted by a loudspeaker in the environment)) and to determine what is being spoken. The local data and processing module 260 or the remote processing module 270 can process the audio data from the microphone (or audio data in another stream, such as, for example, a video stream being viewed by a user) and identify the content of the speech by applying various speech recognition algorithms, such as, for example, hidden Markov models, dynamic time warping (DTW) based speech recognition, neural networks, deep learning algorithms such as deep feedforward and recurrent neural networks, end-to-end automatic speech recognition, machine learning algorithms, or other algorithms that use acoustic or language modeling.

[0062] The local data and processing module 260 or the remote processing module 270 can also apply speech recognition algorithms, which can identify the identity of the speaker, such as whether the speaker is the user 210 of the AR device 600 or another person with whom the user is conversing. Some example speech recognition algorithms can include frequency estimation, hidden Markov models, Gaussian mixture models, pattern matching algorithms, neural networks, matrix representations, vector quantization, speaker diarization, decision trees, and dynamic time warping (DTW) techniques. Speech recognition techniques can also include anti-speaker techniques such as cohort models and world models. Spectral features may be used in representing speaker characteristics. The local data and processing module or the remote data processing module 270 can perform speech recognition using various machine learning algorithms.

[0063] An implementation of the AR device can use these user controls or inputs via the UI. UI elements (e.g., controls, pop-up windows, callouts, data entry fields, etc.) can be used, for example, to dismiss the display of information, e.g., graphical or semantic information of an object.

[0064] With respect to the camera system, the illustrated exemplary AR device 600 may include three paired cameras: a pair of relatively wide FOV or passive SLAM cameras arranged on the sides of the user's face, and a different paired camera oriented in front of the user to handle the stereo imaging process 640 and also capture hand gestures and totem / object tracking in front of the user's face. The paired cameras for the FOV camera and stereo process 640 may be part of an outward-facing imaging system. The AR device 600 may include an eye-tracking camera oriented toward the user's eyes to triangulate eye vectors and other information. The AR device 600 may also include one or more textured light projectors (such as infrared (IR) projectors) to inject texture into the scene.

[0065] The AR device 600 may include an avatar processing and rendering system 690. The avatar processing and rendering system 690 may be configured to generate, update, animate, and render avatars based on context information. Some or all of the avatar processing and rendering system 690 may be implemented as part of the local processing and data module 260 or the remote processing modules 262, 264, alone or in combination. In various embodiments, multiple avatar processing and rendering systems 690 (e.g., as implemented on different wearable devices) may be used to render the virtual avatar 670. For example, a first user's wearable device may be used to determine the first user's intention, while a second user's wearable device may determine avatar characteristics and render the first user's avatar based on the intention received from the first user's wearable device. The first user's wearable device and the second user's wearable device (or other such wearable devices) may communicate over a network.

[0066] 5A illustrates an exemplary avatar processing and rendering system 690. The exemplary avatar processing and rendering system 690 may comprise, either alone or in combination, a 3D model processing system 680, a context information analysis system 688, an avatar auto-scaler 692, an intent mapping system 694, an anatomy adjustment system 698, and a stimulus response system 696. The system 690 is intended to illustrate, and not to limit, functionality for avatar processing and rendering. For example, in an implementation, one or more of these systems may be part of another system. For example, parts of the context information analysis system 688 may be parts of the avatar auto-scaler 692, the intent mapping system 694, the stimulus response system 696, or the anatomy adjustment system 698, either alone or in combination.

[0067] The context information analysis system 688 can be configured to determine environment and object information based on one or more device sensors, described with reference to FIGS. 2 and 3. For example, the context information analysis system 688 can use images acquired by an outward facing imaging system of a viewer of the user or the user's avatar to analyze the environment and objects (including physical or virtual objects) of the user's environment or the environment in which the user's avatar is rendered. The context information analysis system 688 can analyze such images, alone or in combination with data acquired from location data or a world map, to determine the location and layout of objects in the environment. The context information analysis system 688 can also access biological characteristics of the user or humans in general to realistically animate the virtual avatar 670. For example, the context information analysis system 688 can generate an incongruity curve, which can be applied to the user's avatar such that a part of the user's avatar's body (e.g., the head) is not in an incongruent (or unrealistic) position relative to other parts of the user's body (e.g., the avatar's head is not turned 270 degrees). In some implementations, one or more object recognizers may be implemented as part of the context information analysis system 688.

[0068] The avatar auto-scaler 692, the intent mapping system 694, the stimulus response system 696, and the anatomy adjustment system 698 can be configured to determine characteristics of the avatar based on the context information. Some example characteristics of the avatar can include size, appearance, position, orientation, movement, posture, expression, etc. The avatar auto-scaler 692 can be configured to automatically scale the avatar so that the user does not have to view the avatar in an uncomfortable position. For example, the avatar auto-scaler 692 can increase or decrease the size of the avatar and bring the avatar to the user's eye level so that the user does not have to look down or up at the avatar, respectively. The intent mapping system 694 can determine the user's interaction intent based on the environment in which the avatar is rendered and map the intent (rather than the exact user interaction) to the avatar. For example, the intent of a first user can be to communicate with a second user in a telepresence session. Typically, two people face each other when communicating. The intent mapping system 694 of the first user's AR device can determine such a facing intent that exists during the telepresence session and can cause the first user's AR device to render the second user's avatar facing the first user. If the second user attempts to physically turn around, instead of rendering the second user's avatar in a turned-around position (which would cause the back of the second user's avatar to be rendered towards the first user), the first user's intent mapping system 694 can continue to render the second avatar's face towards the first user, which is the inferred intent of the telepresence session (e.g., in this example, a facing intent).

[0069] The stimulus response system 696 can identify objects of interest in the environment and determine the avatar's response to the objects of interest. For example, the stimulus response system 696 can identify a sound source in the avatar's environment and automatically reorient the avatar to look at the sound source. The stimulus response system 696 can also determine threshold exit conditions. For example, the stimulus response system 696 can cause the avatar to return to its original posture after the sound source disappears or after a period of time has passed.

[0070] The anatomical structure adjustment system 698 can be configured to adjust the user's posture based on the biological characteristics. For example, the anatomical structure adjustment system 698 can be configured to adjust the relative position between the user's head and the user's torso or between the user's upper body and lower body based on the discomfort curve.

[0071] The 3D model processing system 680 can be configured to animate the virtual avatar 670 and render it on the display 220. The 3D model processing system 680 can include a virtual character processing system 682 and a movement processing system 684. The virtual character processing system 682 can be configured to generate and update a 3D model of the user (to create and animate the virtual avatar). The movement processing system 684 can be configured to animate the avatar, for example, by changing the avatar's pose, by moving the avatar in the user's environment, or by animating the avatar's facial expressions, etc. As will be further described herein, the virtual avatar can be animated using rigging techniques. In some embodiments, the avatar is represented in two parts: a surface representation (e.g., a deformable mesh) that is used to render the virtual avatar's outward appearance, and a hierarchical set of joints (e.g., a core skeleton) that are interconnected to animate the mesh. In some implementations, the virtual character processing system 682 can be configured to edit or generate surface representations, while the movement processing system 684 can be used to animate the avatar by moving the avatar, deforming meshes, etc. An Exemplary Neutral Avatar Mapping System

[0072] FIG. 5B is a block diagram illustrating an example of components and signals associated with implementing a neutral avatar. In this example, several user sensor components 601-604 provide input signals 605 (including 605A, 605B, 605C, and 605D) to a signal mapping component 606. The signal mapping component 606 is configured to analyze the input signals 605 and determine updates to the neutral avatar, which are then transmitted as one or more output signals 607 to an avatar renderer 608 (e.g., part of the avatar processing and rendering system 690 of FIG. 5A). In the embodiment of FIG. 5B, the user sensors include gaze tracking 601, speech tracking 602, head pose tracking 603, and hand pose tracking 604. Each of these user sensors may include one or more sensors of the same type or multiple different types. The type of input signals 605 may vary from embodiment to embodiment to include fewer or additional user sensors. In some embodiments, the input signals 605 may also be processed (e.g., prior to or simultaneously with transmission to the signal mapping component 606) to determine additional input signals for use by the signal mapping component. In the example of FIG. 5B, a derived signal generator 609 may also receive each of the input signals 605A-605D and generate one or more input signals 605E that are transmitted to the signal mapping component 606. The derived signal generator may create input signals 605E that are indicative of user intent, behavior, or action that is not directly linked to one of the input signals 605A-605D.

[0073] The signal mapping component 606 may include mapping tables in various forms, such as look-up tables that allow one-to-one, one-to-many, and many-to-many mappings between input and output signals. Similarly, rule lists, pseudocode, and / or any other logic may be used by the signal mapping component 606 to determine an appropriate output signal 607 to be mapped to a current input signal 605. Advantageously, the signal mapping component 606 operates in real-time to map input signals 605 to one or more output signals 607, such that updates to neutral avatars (as implemented by the avatar renderer 608) are applied simultaneously with triggering user activity.

[0074] In some embodiments, the signal mapping component is configured to 1) measure a parameter of the user associated with a part of the user's body, and then 2) map the measured parameter to a feature of a neutral avatar, which does not represent a part of the user's body. The measured parameter may be an input signal 605, and the mapped feature of the neutral avatar may be indicated in a corresponding output signal 607 generated by the signal mapping component 606. As an example of this mapping, the user's eye rotation may be a parameter of the user, which is associated with the eyes of the user's body. As mentioned above, this eye rotation by the user may be mapped to a line or other geometric feature that is located outside the eye area of ​​the neutral avatar and therefore does not represent the user's eyes. This type of mapping of an input signal associated with one body part and a visual indicator of a neutral avatar of a second body part may be referred to as a non-factual mapping.

[0075] In some embodiments, a non-factual mapping may also be performed on the features of a neutral avatar that do not represent the user's physical actions from which the parameters were measured (e.g., animation, color, texture, sound, etc.). For example, a line feature of a neutral avatar may change color in response to a speech input signal measured from a user. This change in color does not represent a speech action performed by the user to provide speech input (e.g., opening and closing the user's mouth). Thus, this mapping may also be considered a non-factual mapping.

[0076] In some embodiments, an input signal associated with a particular body part and / or activity of a particular body part of a user may be mapped to a disparate, unrelated, unassociated, distinct, and / or different feature and / or activity of a neutral avatar. For example, the signal mapping component 606 may map the input signal to a non-factual output signal associated with a neutral avatar. For example, in response to a user speaking, the input signal 605B may be transmitted to the signal mapping component 606, which may then map the speech to a color or shading adjustment output signal 607 that is applied to the neutral avatar. Thus, the shading of some or all of the neutral avatar may be dynamically adjusted as speech input is received. Shading may be applied to parts of the avatar that are not directly associated with speech, such as non-mouth areas of the neutral avatar's face or geometric features (e.g., not directly associated with a particular facial feature of the user). For example, the upper portion of the neutral avatar may be shaded differently when the user is speaking. The shading may be dynamically updated, such as with adjustments to the shading level and / or area, as the audio input (e.g., speech volume, tone, pattern, etc.) changes. This is in contrast to typical avatar behavior, where user speech is indicated to the avatar by the avatar's mouth moving in a speech pattern. Thus, neutral avatars are configured to provide a visualization of a vague identity, with behavioral, social, and communication cues expressed in a manner that does not directly map to corresponding user actions.

[0077] In another example, rather than mapping eye gaze input signals (e.g., measured by one or more sensors in the AR device) in a one-to-one or direct manner to control the rotation of the avatar's eyes, as would be done under a factual mapping, an indirect (or non-factual) mapping may map the pupil tracking of the user's eyes to changes in the shape of a feature of the neutral avatar (e.g., the neutral avatar's head, body, or geometric shape), the shading of a feature (e.g., part or all of the neutral avatar's head, body, or geometric shape), the color of a feature, and / or any other feature of the avatar that is not the avatar's pupil. Thus, the user's eye movements may be mapped to variations in auxiliary features of the neutral avatar, such as the color or shading of the neutral avatar or even the background or objects near the neutral avatar.

[0078] In some embodiments, multiple input signals 605 may be associated with a single visual element of a neutral avatar. For example, eye gaze direction and voice may be mapped to the same visual element of a neutral avatar, such as a horizon line or other geometric shape. That single visual element may be configured to sway to represent voice activity and shift (e.g., left and right) or deform to represent gaze direction. Thus, multiple user actions may be conveyed in a more refined visual manner without the distraction of highly customized avatar visual characteristics. In some implementations, the mapping of multiple input signals to a single visual element of a neutral avatar may increase the salience and / or visual complexity of these simple neutral features. Because real human faces are capable of many subtle movements, a simple visual element that responds to a single cue (e.g., a single input signal) may be less believable in representing this complex human behavior than a single visual element with more complex behavior that responds to multiple cues (e.g., multiple input signals). Mapping multiple input signals to the same visual element of a neutral avatar may provide further visual abstraction rather than factual visual familiarity.

[0079] The visual elements of the neutral avatar are configured to convey human behaviors for communication and collaboration (eg, in a remote co-presence environment).

[0080] In some embodiments, the user's head position (e.g., from head pose tracking 603) may be mapped to changes in larger elements of the neutral avatar (e.g., shading, moving, morphing). In some embodiments, eye gaze and eye tracking information (e.g., from eye tracking 601) may be mapped to smaller elements, such as geometric shapes that move, translate, and animate and correspond to the eye tracking signals. In some embodiments, audio signals may be mapped to particle shaders or geometric elements that transition, transform, and / or animate according to audio amplitude and / or audio phonemes.

[0081] 6-12 illustrate several neutral avatar examples along with example input signals that may be mapped to output signals that trigger updates to the neutral avatar. Starting with FIG. 6, the neutral avatar is shown in each of FIGS. 6A, 6B, 6C with different visual features 603 that indicate non-factual input signals from one or more user sensors. In some implementations, as the user produces certain sounds, they are correlated with the display and / or characteristics of certain visual features, such as the opacity of visual features 603, 611, or 612. In examples, the features may become more or less opaque depending on what the user is speaking, for example, based on the volume, pitch, tone, rate, patterns, and / or speech recognized in the voice input. Thus, the visual characteristics of these visual features may be mapped to be proportional to the opacity of such visual features. For example, the shading or color of the visual features may dynamically increase and decrease to track the increase and decrease of the corresponding one or more voice features.

[0082] In one embodiment, visual feature 603 of the neutral avatar in FIG. 6A may appear when the user makes an "ah" sound (or another front vowel sound), visual feature 611 of the neutral avatar in FIG. 6B may appear when the user makes an "oo" sound (or another back vowel sound), and visual feature 612 of the neutral avatar in FIG. 6C may appear when the user makes a "th" sound. Thus, visual features 603, 611, 612 may be displayed on the neutral avatar in alternating fashion as the user makes different corresponding sounds. In some embodiments, the shapes of the user's mouth (e.g., which may be referred to as visemes) may be mapped to various output signals that may affect the look and / or behavior of the neutral avatar. For example, the relative size of the visemes, e.g., the size of the opening of the user's mouth, may be mapped to amplitude or sound pressure characteristics. This amplitude characteristic may then be used as a driver of some aspects of the neutral avatar, such as the size or shape of the neutral avatar's visual indicators. Additionally, the color, texture, shading, opacity, etc. of the visual features may be adjusted based on other factors of the audio input and / or other input signals.

[0083] In some embodiments, the position and / or translation of a visual indicator of a neutral avatar may be mapped to an eye gaze direction input signal. For example, a visual indicator that is mapped to a viseme shape (e.g., increases / decreases in size as the mouth open area in a viseme increases / decreases) may be moved based on the user's eye gaze.

[0084] In some embodiments, other transformations of the visual indicator, such as closing the eyes, shrinking, etc. may be mapped to an eye blink event.

[0085] 7A, 7B, and 7C illustrate another example neutral avatar, where visual feature 702 is adjusted based on one or more of various input signals. In this example, the deformation in visual feature 702 may be associated with the user's head pose. For example, an input signal indicating a rightward head pose may result in a deformation of visual feature 702B, while a leftward head pose may result in a deformation of visual feature 702C. The direction of the deformation (e.g., upward or downward) may be based on another input signal, such as the user's attention, or based on a specific user intent (e.g., may be determined based on multiple input signals). A single line visual feature 702 may reflect multiple input signals, such as deforming in a first manner to indicate gaze and in a second manner to indicate voice input.

[0086] In one embodiment, visual feature 702A indicates a user idle state. In Figure 7B, visual feature 702B is deformed in response to an input signal indicating a particular viseme (e.g., an "Ahh" sound). The deformation area of ​​visual feature 702B may shift to the left, for example, in response to gaze direction. Thus, visual feature 702C may indicate that the user's gaze direction has shifted to the left, and the updated deformation may indicate another viseme or silent input.

[0087] In some embodiments, a transformation (e.g., deformation) of visual feature 702 (or other simple geometric shape) may be mapped to an eye gaze shift, while head pose may be mapped to other visual features, such as a rotation of the entire hemispherical shape that contains visual feature 702. In one exemplary embodiment, visual feature 702 provides a visual reference for the overall head orientation (e.g., like face orientation), and deformation of feature 702 (e.g., as in FIG. 7C) may be mapped to an eye gaze shift toward the location of eyebrow lift and / or deformation (e.g., visual feature 702C may indicate an eye gaze shift to the top and left).

[0088] In some embodiments, the geometry of visual feature 702A, etc., may be mapped to an input signal, indicative of lip synchronization or voice animation, producing changes in visual feature 702 in a different pattern than that used for other input signals. For example, the visual feature may ripple or wobble in response to detection of a specific viseme. The radius of the line transform and / or the smoothness of the line may be adjusted (e.g., dynamically) according to the particular viseme detected, the speech amplitude, the speech pitch, and / or any other input signal derived from the user. As another example, the position of visual feature 702A on a neutral avatar may be translated vertically to represent a drop or rise in eye gaze. As another example, the length of visual feature 702 (or other visual features) may be scaled / shortened / increased to represent voice amplitude.

[0089] 8A and 8B illustrate another exemplary neutral avatar, where visual features 802 and 804 may be modified based on one or more input signals. For example, in one implementation, the upper line 802A may be mapped to a user's gaze such that changes in the user's gaze may be indicated by various changes in the upper line 802A (e.g., deformations similar to those in FIG. 7, color changes, shading changes, size or thickness changes, etc.). In this example, the lower line 804 may be mapped to changes in the audio signal such that the lower line 804B deforms as the user provides audio input. Other exemplary mappings of input signals and output signals that cause changes to a neutral avatar (e.g., the neutral avatar of FIG. 8 and any other neutral avatar with corresponding visual features) are as follows: The length of a visual feature (e.g., one or both of lines 802A, 804A) may shorten or lengthen in response to an eyeblink event. Visual features can translate left or right in response to changes in eye gaze direction. · Visual features 804 may respond to viseme changes with wiggling or deformation in shape and / or definition. · The overall length and / or position may vary with respect to amplitude.

[0090] 9A, 9B, 9C, and 9D illustrate another exemplary neutral avatar, where visual feature 904 includes a ring and a circle. In one implementation, the size of the circle is adjusted based on one or more input signals. For example, as shown in FIGS. 9B and 9C, the size of the circle may be adjusted to indicate a change in the input signal, such as to indicate when voice input is received from a user. In this example, a larger circle of visual feature 904B may indicate that an active voice signal is being received, while a smaller circle of visual feature 904C may indicate that an active voice signal is not being received. Thus, the change in size of the visual indicator may be dynamically adjusted in real time to reflect a change in the voice signal from the user. For example, the circle portion of visual feature 904C may pulse in conjunction with a user providing voice input. The same visual feature 904 may move in other manners to reflect other input signals. For example, visual features 904 may rotate between the orientations shown in Figures 9A and 9D in response to changes in the user's head pose and / or eye pose. Thus, visual features 904 may be responsive to multiple input signals, which are provided within a neutral avatar that is easily understandable and of low complexity. Other example mappings of input signals and output signals that cause a change to a neutral avatar (e.g., the neutral avatar of Figure 9 and any other neutral avatar with corresponding visual features) are as follows: The size of the circle 903 may pulse (or otherwise change) based on the audio amplitude. For example, the size may not just indicate whether the user is making a sound, but may dynamically adjust to indicate multiple levels of audio pressure. The circular shape 903 may translate left and right along the ring element to indicate eye gaze direction, which may be independent of head pose position. The circular shape 903 may flatten or stretch into a squashed cylinder shape for an eyeblink.

[0091] In the example of Figures 10A-10F, six example visualizations of a neutral avatar are illustrated, where adjustments in visual features 1002 may be tied to one or more different input signals. In one embodiment, the position, shape, and / or animation of visual features 1002 may be mapped to speech input. For example, a user's visemes (e.g., the shape of the user's mouth indicating a corresponding sound) may be mapped to variations in visual features 1002. In the example of Figure 10A, visual feature 1002A (or other substantially rounded shape) may be mapped to a viseme indicating an "Ooo" speech input, visual feature 1002B (or other shape with a generally square edge) may be mapped to a viseme indicating a "Thh" speech input, and visual feature 1002D (or other shape with sharper edges) may be mapped to a viseme indicating an "Ahh" speech input. In FIG. 10C, visual feature 1002C shows the user's gaze direction shifting to look down and to the right.

[0092] For any of these animations of visual feature 1002, the amplitude of the audio input may be visually indicated by the distance between two lines of visual feature 1002. For example, visual feature 1002E may represent a loud (e.g., high amplitude) audio input, while visual feature 1002F represents a quieter (e.g., low amplitude) audio input. In other embodiments, other visemes and / or audio or other input signals may be mapped to similar adjustments in visual feature 1002 (or other visual features of a neutral avatar).

[0093] 11A-11H illustrate another exemplary neutral avatar with various forms of visual features 1102 that may be dynamically updated based on one or more input signals. In some embodiments, the visual features 1102 (which may appear, for example, as visors or floating sunglasses in some implementations) may move from side to side in response to changes in eye and / or head pose (e.g., as shown in FIGS. 11B and 11C). FIGS. 11F and 11G illustrate visual features 1102F changing to a smaller size, such as visual feature 1102G, which may reflect a change in the user's attention or may be animated to correspond with voice input from the user. For example, larger visual feature 1102F may be shown when an input signal indicating voice input is being received, while smaller visual feature 1102G is shown when no voice input is being received (e.g., between spoken words or other pauses in speech). Visual feature 1102H may indicate changes in two input signals, one mapped to the left of visual feature 1102H and one to the right of visual feature 1102H. Thus, the size, shape, color, shading, etc. of portions of visual feature 1102H (e.g., left and right sides of visual feature 1102H) may independently indicate user behavior or social cues. In some embodiments, texture effects may be shown relative to visual feature 1102, indicating voice activity. For example, a backlight effect as shown in FIG. 11I may pulse, move, etc. according to voice input. In one embodiment, the pulse shape (and / or other characteristics) may change in a manner similar to that discussed above with reference to FIGS. 10A-10F (e.g., to reflect a particular viseme and / or amplitude). In this example, other characteristics of visual feature 1102I may remain mapped to other signal inputs, such as those discussed above. FIG. 11E illustrates another example visual feature on visual features 1102E that may be mapped to various audio input and / or other input signals.

[0094] 12A and 12B illustrate another example neutral avatar, where morphing, moving, and / or other visual changes to portions of the neutral avatar may be mapped to one or more input signals. In this example, dynamic scaling of the sphere 1202 may indicate changes in audio amplitude, where larger sphere sizes (e.g., 1202A) may be associated with higher amplitude audio input, while smaller sphere sizes (e.g., 1202B) may be associated with lower amplitude audio input. In one example, the color, texture, or other attributes of the sphere 1202 may be mapped to concrete sounds or other audio attributes, such as visemes. In one embodiment, the horizontal element 1204 may be stretched, scaled, or otherwise morphed to indicate the user's gaze or attention direction.

[0095] In any of the above examples, the links between the input and output signals may be combined, separated, and / or mapped to changes in other visual features. As noted above, in some embodiments, shading of a visual feature may indicate a change in one or more input signals. Additionally, shading of other parts of the neutral avatar, such as the avatar's face or parts of its body, may also indicate a change in the input signal. Example Implementation

[0096] The systems, methods, and devices described herein each have several aspects, no single one of which is solely responsible for its desirable attributes. Without limiting the scope of the present disclosure, some non-limiting features will now be briefly discussed. The following paragraphs describe various exemplary implementations of the devices, systems, and methods described herein. One or more computer systems can be configured to perform specific operations or actions by having software, firmware, hardware, or a combination thereof installed on the system that, when in operation, causes the system to perform an action. One or more computer programs can be configured to perform specific operations or actions by including instructions that, when executed by a data processing device, cause the device to perform an action.

[0097] Example 1: A computing system comprising a hardware computer processor and a non-transitory computer readable medium having software instructions stored thereon, the software instructions being executable by the hardware computer processor to cause the computing system to perform operations including: providing co-presence environment data usable by a plurality of users to interact within an augmented reality environment; determining, for each of the plurality of users, one or more visual idiosyncrasies of a neutral avatar for the user, the visual idiosyncrasies being different from the visual idiosyncrasies of neutral avatars of others of the plurality of users; and updating the co-presence environment data to include the determined visual idiosyncrasies of the neutral avatar.

[0098] Example 2: The computing system of example 1, wherein the visual specificity comprises a color, texture, or shape of the neutral avatar.

[0099] Example 3: The computing system of example 1, wherein the operations further include storing the determined visual idiosyncrasies for a particular user, and determining a neutral avatar for the user includes selecting the stored visual idiosyncrasies associated with the user.

[0100] Example 4: The computing system of example 1, wherein determining the visual distinctiveness of the neutral avatar for the user is performed automatically, regardless of the user's personal characteristics.

[0101] Example 5: A computing system comprising: a hardware computer processor;

[0102] 1. A computing system comprising: a non-transitory computer-readable medium having software instructions stored thereon, the software instructions being executable by a hardware computer processor to cause the computing system to perform operations including: determining a neutral avatar to be associated with a user within an augmented reality environment, the neutral avatar not including an indication of the user's gender, ethnicity, and identity, the neutral avatar being configured to represent input cues from the user along with changes to visual elements of the neutral avatar that are non-factual indications of corresponding input cues; and providing real-time rendering updates to the neutral avatar that are viewable by each of a plurality of users within the shared augmented reality environment.

[0103] Example 6: The computing system of example 5, wherein the first visual element is associated with two or more input cues.

[0104] Example 7: The computing system of example 6, wherein the input cues include one or more of gaze direction, speech, head pose, and hand pose.

[0105] Example 8: The computing system of Example 7, wherein the input cues include one or more of shared attention, shared gaze, and mutual gestures.

[0106] Example 9:

[0107] The computing system of example 5 further comprises a signal mapping component that stores a mapping between the input queues and corresponding output signals, and the real-time rendering updates are determined based on the output signals.

[0108] Example 10: The computing system of example 5, wherein the neutral avatar includes a visual element that is deformable in response to audio input cues.

[0109] Example 11: The computing system of example 10, wherein the visual element is capable of deforming in a different manner in response to input cues indicative of particular gaze activity.

[0110] Example 12: The computing system of example 5, wherein the neutral avatar includes a visual element that changes size in response to audio input cues.

[0111] Example 13: The computing system of example 5, wherein the neutral avatar includes a visual element that changes shading of a portion of the neutral avatar in response to audio input cues.

[0112] Example 14: The computing system of example 13, wherein the portion of the neutral avatar is not associated with a mouth area of ​​the neutral avatar.

[0113] Example 15: The computing system of example 5, wherein the neutral avatar comprises one or more geometric shapes.

[0114] As mentioned above, implementations of the illustrated embodiments provided above may include hardware, methods or processes, and / or computer software on a computer-accessible medium. Other considerations

[0115] Each of the processes, methods, and algorithms described herein and / or depicted in the accompanying figures may be embodied in code modules executed by one or more physical computing systems, hardware computer processors, application-specific circuits, and / or electronic hardware configured to execute specific and particular computer instructions, and may be fully or partially automated thereby. For example, the computing system may include a general-purpose computer (e.g., a server) or a special-purpose computer programmed with specific computer instructions, special-purpose circuits, etc. The code modules may be compiled and linked into an executable program, installed in a dynamic link library, or written in an interpreted programming language. In some implementations, certain operations and methods may be performed by circuitry specific to a given function.

[0116] Furthermore, certain implementations of the functionality of the present disclosure may be sufficiently mathematically, computationally, or technically complex that special purpose hardware (utilizing appropriate specialized executable instructions) or one or more physical computing devices may be required to implement the functionality, for example, due to the amount or complexity of the calculations involved, or to provide results in substantially real time. For example, an animation or video may contain many frames, each frame may have millions of pixels, and specifically programmed computer hardware is required to process the video data to provide the desired image processing task or application in a commercially reasonable amount of time. As another example, calculating weight maps, rotation, and translation parameters for a skinning system by solving a constrained optimization problem for these parameters is highly computationally intensive (see, for example, example process 1400 described with reference to FIG. 14).

[0117] The code modules or any type of data may be stored on any type of non-transitory computer readable medium, such as physical computer storage devices, including hard drives, solid state memory, random access memory (RAM), read only memory (ROM), optical disks, volatile or non-volatile storage devices, combinations of the same, and / or the like. The methods and modules (or data) may also be transmitted (e.g., as part of a carrier wave or other analog or digital propagated signal) generated on a variety of computer readable transmission media, including wireless-based and wired / cable-based media, and may take various forms (e.g., as part of a single or multiplexed analog signal, or as multiple discrete digital packets or frames). The results of the disclosed processes or process steps may be stored persistently or otherwise in any type of non-transitory tangible computer storage device, or communicated via a computer readable transmission medium.

[0118] Any process, block, state, step, or functionality in the flow diagrams described herein and / or depicted in the accompanying figures should be understood as potentially representing a code module, segment, or portion of code that includes one or more executable instructions for implementing a specific function (e.g., logical or arithmetic) or step in the process. Various processes, blocks, states, steps, or functionality can be combined, rearranged, added, deleted, modified, or otherwise altered from the illustrative examples provided herein. In some embodiments, additional or different computing systems or code modules may perform some or all of the functionality described herein. The methods and processes described herein are also not limited to any particular sequence, and the blocks, steps, or states associated therewith can be performed in other sequences as appropriate, for example, serially, in parallel, or in some other manner. Tasks or events may be added to or removed from the disclosed exemplary embodiments. Furthermore, the separation of various system components in the implementations described herein is for illustrative purposes and should not be understood as requiring such separation in all implementations. It should be understood that the program components, methods, and systems described may generally be integrated together in a single computer product or packaged in multiple computer products. Many implementation variations are possible.

[0119] The process, method, and system may be implemented in a network (or distributed) computing environment. Network environments include enterprise-wide computer networks, intranets, local area networks (LANs), wide area networks (WANs), personal area networks (PANs), cloud computing networks, crowdsourced computing networks, the Internet, and the World Wide Web. The network may be a wired or wireless network or any other type of communication network.

[0120] The systems and methods of the present disclosure each have several innovative aspects, none of which is solely responsible for or required for the desirable attributes disclosed herein. The various features and processes described above may be used independently of one another or may be combined in various ways. All possible combinations and subcombinations are intended to fall within the scope of the present disclosure. Various modifications of the implementations described in the present disclosure may be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other implementations without departing from the spirit or scope of the present disclosure. Thus, the claims are not intended to be limited to the implementations shown herein, but should be accorded the widest scope consistent with the disclosure, principles, and novel features disclosed herein.

[0121] Certain features described herein in the context of separate implementations can also be implemented in combination in a single implementation. Conversely, various features described in the context of a single implementation can also be implemented separately in multiple implementations or in any suitable subcombination. Furthermore, although features may be described above as acting in a combination and may even be initially claimed as such, one or more features from the claimed combination may in some cases be deleted from the combination, and the claimed combination may be directed to a subcombination or a variation of the subcombination. No single feature or group of features is necessary or essential to every embodiment.

[0122] In particular, conditional statements used herein, such as "can," "could," "might," "may," "eg," and the like, are generally intended to convey that certain embodiments include certain features, elements, or steps, while other embodiments do not, unless specifically stated otherwise or understood otherwise within the context as used. Thus, such conditional statements are generally not intended to imply that features, elements, and / or steps are in any way required for one or more embodiments, or that one or more embodiments necessarily include logic for determining whether those features, elements, and / or steps are to be included or performed in any particular embodiment, with or without authorial input or prompting. The terms "comprising," "including," "having," and the like, are synonymous and used inclusively in a non-limiting manner and do not exclude additional elements, features, acts, operations, etc. Also, the term "or" is used in its inclusive sense (and not its exclusive sense), thus, for example, when used to connect a list of elements, the term "or" means one, some, or all of the elements in the list. In addition, the articles "a," "an," and "the," as used in this application and the appended claims, unless otherwise specified, should be interpreted to mean "one or more" or "at least one."

[0123] As used herein, a phrase referring to a list of items "at least one of" refers to any combination of those items, including single elements. As an example, "at least one of A, B, or C" is intended to cover A, B, C, A and B, A and C, B and C, and A, B, and C. Transitive phrases such as "at least one of X, Y, and Z" are generally understood differently in the context in which they are used to convey that an item, term, etc. may be at least one of X, Y, or Z, unless specifically stated otherwise. Thus, such transitive phrases are generally not intended to suggest that an embodiment requires that at least one of X, at least one of Y, and at least one of Z, respectively, be present.

[0124] Similarly, although operations may be depicted in the figures in a particular order, it should be appreciated that such operations need not be performed in the particular order depicted, or in sequential order, or that all of the depicted operations need not be performed to achieve desirable results. Additionally, the figures may diagrammatically depict one or more exemplary processes in the form of a flow chart. However, other operations not depicted may also be incorporated within the diagrammatically depicted exemplary methods and processes. For example, one or more additional operations may be performed before, after, simultaneously with, or during any of the depicted operations. Additionally, operations may be rearranged or reordered in other implementations. In some circumstances, multitasking and parallel processing may be advantageous. Additionally, the separation of various system components in the implementations described above should not be understood as requiring such separation in all implementations, and it should be understood that the described program components and systems may generally be integrated together in a single software product or packaged in multiple software products. Additionally, other implementations are within the scope of the following claims. In some cases, the actions recited in the claims can be performed in a different order and still achieve desirable results.

Claims

1. 1. A computing system, comprising: A hardware computer processor; Non-transitory computer readable medium Equipped with the non-transitory computer-readable medium has software instructions stored thereon, the software instructions being executable by the hardware computer processor to cause the computing system to perform a number of operations; The plurality of operations include: providing co-presence environment data usable by multiple users to interact within an augmented reality environment; For each of multiple users, determining one or more visual idiosyncrasies of a neutral avatar for the user, the visual idiosyncrasies being different from visual idiosyncrasies of neutral avatars of other users of the plurality of users; Periodically, Detecting eye movements of the user; mapping the detected eye movements to visual changes in geometric features located outside an eye region of the neutral avatar; updating the coexistence environment data to include the determined visual change of the neutral avatar; Including, The visual change to the geometric feature comprises a change to a color or shading of the neutral avatar.

2. The computing system of claim 1 , wherein the visual idiosyncrasies include a color, a texture, or a shape of the neutral avatar.

3. 2. The computing system of claim 1, wherein the actions further comprise storing visual idiosyncrasies determined for a particular user, and wherein determining the neutral avatar for the user comprises selecting a stored visual idiosyncrasy associated with the user.

4. The computing system of claim 1 , wherein determining the visual distinctiveness of a neutral avatar for a user is performed automatically, regardless of personal characteristics of the user.

5. 1. A computing system, comprising: A hardware computer processor; Non-transitory computer readable medium Equipped with the non-transitory computer-readable medium has software instructions stored thereon, the software instructions being executable by the hardware computer processor to cause the computing system to perform a number of operations; The plurality of operations include: determining a neutral avatar associated with a user within an augmented reality environment, the neutral avatar not including an indication of the user's gender, ethnicity, or identity, the neutral avatar configured to represent detected eye movements of the user along with changes to a plurality of visual elements of the neutral avatar that are outside an eye region of the neutral avatar; providing real-time rendering updates for the neutral avatar that is visible by each of a plurality of users within a shared augmented reality environment; Including, The changes to the plurality of visual elements of the neutral avatar include changes to a color or shading of the neutral avatar.

6. The computing system of claim 5, wherein a first visual element of the plurality of visual elements is associated with two or more input queues.

7. The computing system of claim 6 , wherein the input cues include one or more of gaze direction, speech, head pose, and hand pose.

8. The computing system of claim 5, further comprising a signal mapping component that stores a mapping between input queues and corresponding output signals, and the real-time rendering updates are determined based on the output signals.

9. A computing system, comprising: A hardware computer processor; Non-transitory computer readable medium Equipped with the non-transitory computer-readable medium has software instructions stored thereon, the software instructions being executable by the hardware computer processor to cause the computing system to perform a number of operations; The plurality of operations include: determining a neutral avatar associated with a user within an augmented reality environment, the neutral avatar not including an indication of the user's gender, ethnicity, or identity, the neutral avatar configured to represent detected eye movements of the user along with changes to a plurality of visual elements of the neutral avatar that are outside an eye region of the neutral avatar; providing real-time rendering updates for the neutral avatar that is visible by each of a plurality of users within a shared augmented reality environment; Including, The neutral avatar includes visual elements that are deformable in response to audio input cues.

10. The computing system of claim 9 , wherein the visual elements are otherwise deformable in response to input cues indicative of particular gaze activity.

11. A computing system, comprising: A hardware computer processor; Non-transitory computer readable medium Equipped with the non-transitory computer-readable medium has software instructions stored thereon, the software instructions being executable by the hardware computer processor to cause the computing system to perform a number of operations; The plurality of operations include: determining a neutral avatar associated with a user within an augmented reality environment, the neutral avatar not including an indication of the user's gender, ethnicity, or identity, the neutral avatar configured to represent detected eye movements of the user along with changes to a plurality of visual elements of the neutral avatar that are outside an eye region of the neutral avatar; providing real-time rendering updates for the neutral avatar that is visible by each of a plurality of users within a shared augmented reality environment; Including, The neutral avatar includes a visual element that changes size in response to audio input cues.

12. A computing system, comprising: A hardware computer processor; Non-transitory computer readable medium Equipped with the non-transitory computer-readable medium has software instructions stored thereon, the software instructions being executable by the hardware computer processor to cause the computing system to perform a number of operations; The plurality of operations include: determining a neutral avatar associated with a user within an augmented reality environment, the neutral avatar not including an indication of the user's gender, ethnicity, or identity, the neutral avatar configured to represent detected eye movements of the user along with changes to a plurality of visual elements of the neutral avatar that are outside an eye region of the neutral avatar; providing real-time rendering updates for the neutral avatar that is visible by each of a plurality of users within a shared augmented reality environment; Including, The neutral avatar includes a visual element that changes shading of a portion of the neutral avatar in response to an audio input cue.

13. The computing system of claim 12 , wherein a portion of the neutral avatar is not associated with a mouth area of ​​the neutral avatar.

14. The computing system of claim 5 , wherein the neutral avatar has one or more geometric shapes.

15. A computing system, comprising: A hardware computer processor; Non-transitory computer readable medium Equipped with the non-transitory computer-readable medium has software instructions stored thereon, the software instructions being executable by the hardware computer processor to cause the computing system to perform a number of operations; The plurality of operations include: providing co-presence environment data usable by multiple users to interact within an augmented reality environment; For each of multiple users, determining one or more visual idiosyncrasies of a neutral avatar for the user, the visual idiosyncrasies being different from visual idiosyncrasies of neutral avatars of other users of the plurality of users; Periodically, Detecting eye movements of the user; mapping the detected eye movements to visual changes in geometric features located outside an eye region of the neutral avatar; updating the coexistence environment data to include the determined visual change of the neutral avatar; Including, The geometric features include an outer line of the eye region of the neutral avatar.

16. A computing system, comprising: A hardware computer processor; Non-transitory computer readable medium Equipped with the non-transitory computer-readable medium has software instructions stored thereon, the software instructions being executable by the hardware computer processor to cause the computing system to perform a number of operations; The plurality of operations include: providing co-presence environment data usable by multiple users to interact within an augmented reality environment; For each of multiple users, determining one or more visual idiosyncrasies of a neutral avatar for the user, the visual idiosyncrasies being different from visual idiosyncrasies of neutral avatars of other users of the plurality of users; Periodically, Detecting eye movements of the user; mapping the detected eye movements to visual changes in geometric features located outside an eye region of the neutral avatar; updating the coexistence environment data to include the determined visual change of the neutral avatar; Including, The visual changes to the geometric features include changes to a color or shading of a background of the neutral avatar.

Citation Information

Patent Citations

  • Mixed-reality arena

    JP2015116336A

  • System, method and program

    JP2017060611A