Transmodal Input Fusion for Multiuser Group Intention Processing in Virtual Environments

The system addresses the challenge of determining user intent in shared virtual spaces by using transmodal input fusion, enhancing efficiency and safety through real-time feedback and task adjustments.

JP7763837B2Active Publication Date: 2025-11-04MAGIC LEAP INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2023528241
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-11-13
Filing Date
2021-11-09
Publication Date
2025-11-04
Estimated Expiration
2041-11-09

AI Technical Summary

Technical Problem

Existing virtual, augmented, and mixed reality systems struggle to accurately determine and respond to the collective intent of multiple users in a shared space, leading to inefficiencies and potential safety hazards due to incomplete or inaccurate feedback and task allocation.

Method used

A system that utilizes transmodal input fusion to analyze user inputs such as gaze, hand movements, and direction through wearable devices to identify individual and group intents, generating output data for real-time adjustments and feedback.

Benefits of technology

Enhances task efficiency and safety by providing accurate feedback and reallocating users to optimize group dynamics and task completion, improving user interaction and system responsiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007763837000001
    Figure 0007763837000001
  • Figure 0007763837000002
    Figure 0007763837000002
  • Figure 0007763837000003
    Figure 0007763837000003
Patent Text Reader

Abstract

This document describes an imaging and visualization system in which the intent of a group of users in a shared space is determined and acted upon accordingly. In one aspect, a method includes, for a group of users in a shared virtual space, identifying individual objectives for each of two or more users in the group of users. For each of the two or more users, a determination of the users' individual intent is made based on input from a plurality of sensors having different input modalities. At least some of the plurality of sensors are sensors of the users' devices that enable the users to participate in the shared virtual space. Based on the individual intent, a determination is made whether the users are performing an individual objective for the users. Output data is generated and provided based on the individual objectives and the individual intents.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to virtual reality, augmented reality, and mixed reality imaging and visualization systems, and more particularly to using transmodal input fusion to determine the intent of a group of users within a shared virtual space and act accordingly. [Background technology]

[0002] Modern computing and display technologies have facilitated the development of systems for so-called "virtual reality," "augmented reality," or "mixed reality" experiences in which digitally reproduced images, or portions thereof, are presented to a user in a manner that appears or can be perceived as real. Virtual reality or "VR" scenarios typically involve the presentation of digital or virtual image information without transparency to other actual real-world visual input. Augmented reality or "AR" scenarios typically involve the presentation of digital or virtual image information as an augmentation to the user's visualization of the real world around them. Mixed reality or "MR" refers to the merging of the real and virtual worlds to produce new environments in which physical and virtual objects coexist and interact in real time. Summary of the Invention [Means for solving the problem]

[0003] This specification generally describes an imaging and visualization system in which the intent of a group of users in a shared space is determined and acted upon accordingly. The shared space can include, for example, a real space or environment using augmented reality, or a virtual space using avatars, game players, or other icons or figures representing real people.

[0004] The system can determine a user's intent based on multiple inputs, including the user's gaze, the user's hand movements, and / or the direction the user is moving. For example, a combination of these inputs can be used to determine that a user is about to reach for an object, make a gesture toward another user, focus on a particular user or object, or interact with another user or object. The system can then act on the user's intent by, for example, displaying the user's intent to one or more other users, generating group metrics or other aggregate group information based on the intents of multiple users, alerting the user or another user or providing recommendations, or reassigning users to different tasks.

[0005] In general, one innovative aspect of the subject matter described herein can be embodied in a method that includes, with respect to a group of users in a shared virtual space, identifying individual goals for each of two or more users in the group of users. For each of the two or more users, a determination of the users' individual intentions is made based on input from a plurality of sensors having different input modalities. At least some of the plurality of sensors are sensors of the users' devices that enable the users to participate in the shared virtual space. Based on the individual intentions, a determination is made whether the users are performing the individual goal for the users. Output data is generated with respect to the group of users based on the individual goals for each of the two or more users and the individual intentions for each of the two or more users. The output data is provided to each individual device of one or more users in the group of users for presentation at each of the one or more users' individual devices.

[0006] Other embodiments of this aspect include corresponding computer systems, apparatus, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the method. One or more computer systems can be configured to perform particular operations or actions by having software, firmware, hardware, or a combination thereof installed on the system that, when in operation, causes the system to perform the actions. One or more computer programs can be configured to perform particular operations or actions by including instructions that, when executed by a data processing device, cause the device to perform the actions.

[0007] The foregoing and other embodiments can each optionally include, alone or in combination, one or more of the following features: In some aspects, the individual objective for at least one user of the two or more users includes at least one of (i) a task to be performed by the at least one user, or (ii) a target object for which the at least one user should be looking. Identifying the individual objective for each of the two or more users includes determining, as the target object, a target for which at least a threshold amount of users in the group of users are looking.

[0008] In some aspects, each user's individual device includes a wearable device. Determining the user's individual intention can include receiving, from the user's wearable device, gaze data defining the user's line of sight, gesture data defining the user's hand gestures, and direction data defining a direction the user is moving, and determining the user's intention with respect to the target object based on the gaze data, gesture data, and direction data as the user's individual intention.

[0009] In some aspects, generating output data based on the individual objective for each of the two or more users and the individual intention for each of the two or more users includes determining that the particular user is not performing the individual objective for the particular user, and providing output data to each device of one or more users in the group of users regarding presentations at each device of the one or more users includes providing to a device of a leader user data indicative of the particular user and data indicative that the particular user is not performing the individual objective for the particular user.

[0010] In some aspects, the output data includes a heat map indicating the amount of users in the group of users who are performing their individual objectives. Some aspects may include performing an action based on the output data. The action may include reallocating one or more users to a different objective based on the output data.

[0011] The subject matter described herein can be implemented in particular embodiments and may provide one or more of the following advantages: Using multimodal input fusion to determine the intent of a group of users in a shared space or one or more users within a group allows the visualization system to provide feedback to users, determine and refine metrics related to users, adjust users' actions, reassign users to different roles (e.g., as part of real-time rebalancing of load), and show hidden group dynamics. By determining users' intent, the system can predict users' future actions and adjust them before they occur. This can increase the efficiency with which tasks are completed and improve safety by preventing users from performing unsafe actions. The use of multiple inputs, such as gaze direction and making gestures, allows the visualization system to more accurately determine a user's intent toward objects and / or other users relative to other techniques, such as overhead monitoring of the user or an avatar representing the user.

[0012] The details of one or more implementations of the subject matter described herein are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, drawings, and claims. The present invention provides, for example, the following. (Item 1) 1. A method implemented by one or more data processing devices, said method comprising: With respect to a group of users in a shared virtual space, identifying a distinct objective for each of two or more of the users in the group of users; With respect to each of said two or more users: determining an individual intent of the user based on input from a plurality of sensors having different input modalities, at least some of the plurality of sensors being sensors of the user's device that enables the user to participate in the shared virtual space; determining whether the user is performing the individual purpose for the user based on the individual intent; generating, with respect to the group of users, output data based on the individual goals for each of the two or more users and the individual intentions for each of the two or more users; providing, to an individual device of each of one or more users in the group of users, the output data related to presentations at the individual device of each of the one or more users; A method comprising: (Item 2) Item 10. The method of claim 1, wherein the individual objective for at least one of the two or more users comprises at least one of (i) a task to be performed by the at least one user, or (ii) a target object for which the at least one user should be looking. (Item 3) 3. The method of claim 2, wherein identifying the individual objectives for each of the two or more users includes determining as the target object a target to which at least a threshold amount of the users in the group of users are looking. (Item 4) the individual device of each user comprises a wearable device; Determining the individual intent of the user includes: receiving, from the wearable device of the user, gaze data defining a gaze of the user, gesture data defining a hand gesture of the user, and direction data defining a direction in which the user is moving; determining an intention of the user with respect to a target object based on the gaze data, the gesture data, and the direction data as the individual intention of the user; The method according to item 1, comprising: (Item 5) generating the output data based on the individual objective for each of the two or more users and the individual intention for each of the two or more users includes determining that a particular user has not performed the individual objective for the particular user; providing, to the device of each of one or more users in the group of users, the output data related to presentations at the device of each of the one or more users includes providing, to a device of a leader user, data indicative of the particular user and data indicating that the particular user has not performed the individual objective related to the particular user; The method according to item 1. (Item 6) Item 10. The method of item 1, wherein the output data comprises a heat map indicating an amount of users within the group of users who are performing the individual objectives of the users. (Item 7) Item 10. The method of item 1, further comprising: performing an action based on the output data. (Item 8) 8. The method of claim 7, wherein the action includes reallocating one or more users to different purposes based on the output data. (Item 9) 1. A computer-implemented system comprising: one or more computers; One or more computer memory devices interoperably coupled to the one or more computers and having a tangible non-transitory machine-readable medium storing one or more instructions, which when executed by the one or more computers: With respect to a group of users in a shared virtual space, identifying a distinct objective for each of two or more of the users in the group of users; With respect to each of said two or more users: determining an individual intent of the user based on input from a plurality of sensors having different input modalities, at least some of the plurality of sensors being sensors of the user's device that enables the user to participate in the shared virtual space; determining whether the user is performing the individual purpose for the user based on the individual intent; generating, with respect to the group of users, output data based on the individual goals for each of the two or more users and the individual intentions for each of the two or more users; providing, to an individual device of each of one or more users in the group of users, the output data related to presentations at the individual device of each of the one or more users; one or more computer memory devices that perform operations including 1. A computer-implemented system comprising: (Item 10) Item 10. The computer-implemented system of item 9, wherein the individual objective for at least one of the two or more users comprises at least one of (i) a task to be performed by the at least one user, or (ii) an object of interest that the at least one user should be looking at. (Item 11) Item 11. The computer-implemented system of item 10, wherein identifying the individual objectives for each of the two or more users includes determining as the target object a target to which at least a threshold amount of the users in the group of users are looking. (Item 12) the individual device of each user comprises a wearable device; Determining the individual intent of the user includes: receiving, from the wearable device of the user, gaze data defining a gaze of the user, gesture data defining a hand gesture of the user, and direction data defining a direction in which the user is moving; determining an intention of the user with respect to a target object based on the gaze data, the gesture data, and the direction data as the individual intention of the user; Item 10. The computer-implemented system of item 9, comprising: (Item 13) generating the output data based on the individual objective for each of the two or more users and the individual intention for each of the two or more users includes determining that a particular user has not performed the individual objective for the particular user; providing, to the device of each of one or more users in the group of users, the output data related to presentations at the device of each of the one or more users includes providing, to a device of a leader user, data indicative of the particular user and data indicating that the particular user has not performed the individual objective related to the particular user; Item 10. The computer-implemented system of item 9. (Item 14) 10. The computer-implemented system of claim 9, wherein the output data comprises a heat map indicating an amount of users within the group of users who are performing the individual objectives of the users. (Item 15) Item 10. The computer-implemented system of item 9, wherein the operation includes performing an action based on the output data. (Item 16) Item 16. The computer-implemented system of item 15, wherein the action includes reallocating one or more users to different purposes based on the output data. (Item 17) A non-transitory computer-readable medium having stored thereon one or more instructions, the one or more instructions comprising: With respect to a group of users in a shared virtual space, identifying a distinct objective for each of two or more users in the group of users; With respect to each of two or more users, determining an individual intent of the user based on input from a plurality of sensors having different input modalities, at least some of the plurality of sensors being sensors of the user's device that enables the user to participate in the shared virtual space; determining whether the user is performing the individual purpose for the user based on the individual intent; generating, with respect to the group of users, output data based on the individual goals for each of the two or more users and the individual intentions for each of the two or more users; providing, to an individual device of each of one or more users in the group of users, the output data related to presentations at the individual device of each of the one or more users; A non-transitory computer-readable medium executable by a computer system to perform operations including: (Item 18) Item 18. The non-transitory computer-readable medium of item 17, wherein the individual objective for at least one of the two or more users comprises at least one of (i) a task to be performed by the at least one user, or (ii) an object of interest that the at least one user should be looking at. (Item 19) Item 19. The non-transitory computer-readable medium of item 18, wherein identifying the distinct objectives for each of the two or more users includes determining as the target object a target to which at least a threshold amount of the users in the group of users are looking. (Item 20) the individual device of each user comprises a wearable device; Determining the individual intent of the user includes: receiving, from the wearable device of the user, gaze data defining a gaze of the user, gesture data defining a hand gesture of the user, and direction data defining a direction in which the user is moving; determining an intention of the user with respect to a target object based on the gaze data, the gesture data, and the direction data as the individual intention of the user; Item 18. The non-transitory computer-readable medium of item 17, comprising: [Brief explanation of the drawings]

[0013] [Figure 1A] FIG. 1A is an example of an environment in which a visualization system determines the intentions of a group of users in a shared space and acts accordingly.

[0014] [Figure 1B] FIG. 1B is an example of a wearable system.

[0015] [Figure 2] 2A-2C are exemplary attention models illustrating the attention of a group of users to a single object.

[0016] [Figure 3] 3A and 3B are exemplary attention models illustrating a group of users' attention to content and people, or interaction with content.

[0017] [Figure 4] 4A-4C are vector diagrams depicting the attention of a group of users.

[0018] [Figure 5] 5A-5C are heatmap diagrams of user attention corresponding to the vector diagrams of FIGS. 4A-4C, respectively.

[0019] [Figure 6] FIG. 6 is an exemplary attention model illustrating two users' attention to each other and to a common object.

[0020] [Figure 7] FIG. 7 is another exemplary attention model illustrating two users' attention to each other and to a common object.

[0021] [Figure 8] FIG. 8 is an exemplary attention model illustrating mutual user-to-user attention.

[0022] [Figure 9] FIG. 9 is a flowchart of an exemplary process for determining the intent of one or more users in a group of users and acting accordingly.

[0023] [Figure 10]FIG. 10 is a block diagram of a computing system that may be used in connection with the computer-implemented methods described herein.

[0024] Like reference numbers and designations in the various drawings indicate like elements. DETAILED DESCRIPTION OF THE INVENTION

[0025] Detailed Description This specification generally describes an imaging and visualization system in which the intent of a group of users in a shared space is determined and acted upon accordingly.

[0026] 1A is an example of an environment 100 in which a visualization system 120 determines the intent of a group of users in the shared space and acts accordingly. The shared space can include, for example, a real space or environment using augmented reality, or a virtual space using avatars, game players, or other icons or figures representing real people. Augmented reality, virtual reality, or mixed reality spaces are also referred to herein as shared virtual spaces.

[0027] The visualization system 120 can be configured to receive input from the user system 110 and / or other sources. For example, the visualization system 120 can be configured to receive visual input 131 from the user system 110, stationary input 132 from stationary devices, such as images and / or video from a room camera, and / or sensory input 133 from various sensors, such as gestures, totems, eye tracking, or user input.

[0028] In some implementations, the user system 110 is a wearable system that includes a wearable device, such as a wearable device 107, worn by the user 105. The wearable system (also referred to herein as an augmented reality (AR) system) can be configured to present 2D or 3D virtual images to the user. The images may be still images, frames of video, or videos, in combination or the like. The wearable system can include a wearable device that can present VR, AR, or MR content, alone or in combination, within an environment for user interaction. The wearable device can be a head-mounted device (HMD), which may include a head-mounted display.

[0029] VR, AR, and MR experiences can be provided by a display system having a display in which images corresponding to multiple rendering planes are provided to a viewer. The rendering planes can correspond to a depth plane or multiple depth planes. The images may be different for each rendering plane (e.g., providing slightly different presentations of a scene or objects) and may be focused separately by the viewer's eyes, thereby serving to provide depth cues to the user based on the ocular accommodation required to focus on different image features of a scene located on different rendering planes, or based on observing different image features on different rendering planes that are out of focus.

[0030] The wearable system can determine the location and various other attributes of the user's environment using various sensors (e.g., accelerometers, gyroscopes, temperature sensors, movement sensors, depth sensors, GPS sensors, inward-facing imaging systems, outward-facing imaging systems, etc.). This information may be further supplemented with information from stationary cameras in the room, which may provide images from different perspectives or various cues. Image data obtained by cameras (such as the indoor cameras and / or the cameras of the outward-facing imaging system) can be reduced to a set of mapping points.

[0031] FIG. 1B illustrates an exemplary wearable system 110 in more detail. Referring to FIG. 1B, the wearable system 110 includes a display 145 and various mechanical and electronic modules and systems to support the functionality of the display 145. The display 145 can be coupled to a frame 150, which is wearable by a user, wearer, or viewer 105. The display 145 can be positioned directly in front of the eyes of the user 105. The display 145 can present AR / VR / MR content to the user. The display 145 can include a head-mounted display (HMD) worn on the user's head. In some embodiments, a speaker 160 is coupled to the frame 150 and positioned adjacent to the user's ear canal (in some embodiments, another speaker, not shown, is positioned adjacent to the user's other ear canal to provide stereo / shapeable sound control). The display 145 can include an audio sensor 152 (e.g., a microphone) to detect audio streams from the environment for performing voice recognition.

[0032] The wearable system 110 may include an outward-facing imaging system that observes the world in the user's surrounding environment. The wearable system 110 may also include an inward-facing imaging system that can track the user's eye movements. The inward-facing imaging system may track the movements of either one eye or both eyes. The inward-facing imaging system may be mounted on the frame 150 and may be in electrical communication with a processing module 170 or 180 that may process image information obtained by the inward-facing imaging system and determine, for example, pupil diameter or orientation of the user's 105 eyes, eye movement, or eye posture.

[0033] As an example, the wearable system 110 can obtain images of the user's posture (e.g., gestures) using an outward-facing or inward-facing imaging system. The images may be still images, frames of video, videos, combinations thereof, or the like. The wearable system 110 can include other sensors, such as electromyography (EMG) sensors, that sense signals indicative of the action of muscle groups.

[0034] The display 145 can be operably coupled to a local data processing module 170, which can be mounted in a variety of configurations, such as fixedly attached to the frame 150, fixedly attached to a helmet or hat worn by the user, built into headphones, or otherwise removably attached to the user 105 (e.g., in a backpack-style configuration, in a belt-coupled configuration), such as by wired or wireless connection.

[0035] The local processing and data module 170 can include a hardware processor and digital memory, such as non-volatile memory (e.g., flash memory), both of which may be utilized to aid in processing, caching, and storing data. The data may include (a) data captured from environmental sensors (e.g., which may be operably coupled to the frame 150 or otherwise attached to the user 105), audio sensors 152 (e.g., microphones), or (b) data obtained or processed using the remote processing module 180 or remote data repository 190, possibly for processing or retrieval and subsequent passage to the display 145. The local processing and data module 170 may be operably coupled to the remote processing module 180 or remote data repository 190 by a communications link, such as via a wired or wireless communications link, so that these remote modules are available as resources to the local processing and data module 170. Additionally, the remote processing module 270 and the remote data repository 190 may be operably coupled to each other.

[0036] In some embodiments, remote processing module 180 may include one or more processors configured to analyze and process data and / or image information. In some embodiments, remote data repository 190 may include a digital data storage facility, which may be available through the Internet or other networking configuration in a "cloud" resource configuration.

[0037] In some embodiments, remote processing module 180 may include one or more processors configured to analyze and process data and / or image information. In some embodiments, remote data repository 190 may include a digital data storage facility, which may be available through the Internet or other networking configuration in a "cloud" resource configuration.

[0038] The environmental sensors may also include various physiological sensors. These sensors may measure or estimate the user's physiological parameters, such as heart rate, respiratory rate, galvanic skin response, blood pressure, brainwave state, etc. The environmental sensors may also include emitting devices configured to receive signals, such as lasers, visible light, light of invisible wavelengths, or sound (e.g., audible sound, ultrasound, or other frequencies). In some embodiments, one or more environmental sensors (e.g., cameras or light sensors) may be configured to measure the ambient light (e.g., brightness) of the environment (e.g., to capture the lighting conditions of the environment). Physical contact sensors, such as strain gauges, curb detectors, or the like, may also be included as environmental sensors.

[0039] 1A , visualization system 120 includes one or more object recognizers 121 that can recognize objects, recognize or map points, tag images, and associate semantic information with objects using map database 122. Map database 122 can include various points and their corresponding objects collected over time. The various devices and map database 122 can be interconnected through a network (e.g., a LAN, a WAN, etc.) and accessible to the cloud. In some implementations, part or all of visualization system 120 is implemented on one of user systems 110, and user systems 110 can communicate data with each other via a network, e.g., a LAN, a WAN, or the Internet.

[0040] Based on this information and the collection of points in map database 122, object recognizer 121 can recognize objects in an environment, e.g., a shared virtual space for a group of users. For example, object recognizer 121 can recognize faces, people, windows, walls, user input devices, televisions, other objects in the user's environment, etc. One or more object recognizers may be specialized for objects with certain characteristics. For example, an object recognizer may be used to recognize faces, while another object recognizer may be used to recognize totems, while another object recognizer may be used to recognize hands, fingers, arms, or body gestures.

[0041] Object recognition may be performed using various computer vision techniques. For example, the wearable system may analyze images acquired by an outward-facing imaging system and perform scene reconstruction, event detection, video tracking, object recognition, object pose estimation, learning, indexing, motion estimation, or image restoration, etc. One or more computer vision algorithms may be used to perform these tasks. Non-limiting examples of computer vision algorithms include Scale Invariant Feature Transform (SIFT), Speed-Up Robust Features (SURF), Orientation FAST and Rotation BRIEF (ORB), Binary Robust Invariant Scalable Keypoints (BRISK), Fast Retinal Keypoints (FREAK), Viola Jones algorithm, Eigenfaces approach, Lucas-Kanade algorithm, Horn-Schunk algorithm, Mean-shift algorithm, visual simultaneous localization and mapping (vSLAM) techniques, sequential Bayes estimators (e.g., Kalman filter, extended Kalman filter, etc.), bundle adjustment, adaptive thresholding (and other thresholding techniques), iterative nearest neighbor (ICP), semi-global matching (SGM), semi-global block matching (SGBM), feature point histograms, various machine learning algorithms (e.g., support vector machines, k-nearest neighbor algorithms, naive Bayes, neural networks (including convolutional or deep neural networks), or other supervised / unsupervised models, etc.), etc.

[0042] Object recognition can additionally or alternatively be performed by various machine learning algorithms. Once trained, the machine learning algorithms can be stored by the HMD. Some examples of machine learning algorithms can include supervised or unsupervised machine learning algorithms, including regression algorithms (e.g., ordinary least squares regression, etc.), instance-based algorithms (e.g., learning vector quantization, etc.), decision tree algorithms (e.g., classification and regression trees, etc.), Bayesian algorithms (e.g., naive Bayes, etc.), clustering algorithms (e.g., k-means clustering, etc.), association rule learning algorithms (e.g., a priori algorithm, etc.), artificial neural network algorithms (e.g., Perceptron, etc.), deep learning algorithms (e.g., Deep Boltzmann Machine, i.e., deep neural network, etc.), dimensionality reduction algorithms (e.g., principal component analysis, etc.), ensemble algorithms (e.g., stacked generalization, etc.), and / or other machine learning algorithms. In some embodiments, individual models can be customized for individual datasets. For example, the wearable device can generate or store a base model. The base model may be used as a starting point to generate additional models specific to a data type (e.g., a particular user in a telepresence session), a data set (e.g., a set of additional images acquired of a user in a telepresence session), a conditional situation, or other variations. In some embodiments, the wearable HMD can be configured to generate models for analysis of aggregated data using multiple techniques. Other techniques may include using predefined thresholds or data values.

[0043] Based on this information and the set of points in the map database, the object recognizer 121 can recognize objects, complement them with semantic information, and bring them to life. For example, if the object recognizer 121 recognizes that a set of points is a door, the visualization system 120 may associate some semantic information (e.g., a door has a hinge and 90 degrees of movement around the hinge). If the object recognizer 121 recognizes that a set of points is a mirror, the visualization system 120 may associate the semantic information that a mirror has a reflective surface that can reflect images of objects in the room. Over time, the map database 122 grows as the visualization system 120 (which may reside locally or be accessible over a wireless network) accumulates more data from the world. Once an object is recognized, the information may be transmitted to one or more wearable systems, such as the user system 110.

[0044] For example, an MR environment may contain information about a scene occurring in California. The environment may be transmitted to one or more users in New York. Based on data received from the FOV camera and other inputs, the object recognizer 121 and other software components can map points collected from various images, recognize objects, etc., so that the scene can be accurately "passed" to a second user who may be in a different part of the world. The environment may also use a topology map for localization purposes.

[0045] The visualization system 120 can generate a virtual scene for each of one or more users within the shared virtual space. For example, the wearable system may receive input from the user and other users within the shared virtual space regarding the user's environment. This may be achieved through various input devices and knowledge already held in a map database. The user's FOV camera, sensors, GPS, eye tracking, etc. communicate information to the visualization system 120. Based on this information, the visualization system 120 can determine sparse points. The sparse points may be used to determine pose data (e.g., head pose, eye pose, body pose, or hand gestures) that can be used in displaying and understanding the orientation and position of various objects in the user's surroundings. The object recognizer 121 can crawl through these collected points and recognize one or more objects using the map database. This information may then be communicated to the user's individual wearable system, and the desired virtual scene may be displayed to the user accordingly. For example, a desired virtual scene (eg, a user in CA) may be displayed in the proper orientation, position, etc., relative to various objects and other surroundings of the user in New York.

[0046] In another example, the shared virtual space can be an instruction room, such as a classroom, lecture hall, or training room. The visualization system 120 can similarly generate a virtual scene for each user in the instruction room. In yet another example, the shared virtual space can be a gaming environment in which each user is participating in a game. In this example, the visualization system 120 can generate a virtual scene for each player in the game. In another example, the shared virtual space can be a work environment, and the visualization system 120 can generate a virtual scene for each worker in the environment.

[0047] The visualization system 120 also includes a user intent detector 124, a swarm intent analyzer 125, and a swarm intent feedback generator 126. The user intent detector 124 can determine, or at least predict, the intent of one or more users within the shared virtual space. The user intent detector 124 can determine the user intent based on user input, e.g., visual input, gestures, totems, audio input, etc.

[0048] The wearable system can be programmed to receive various input modes. For example, the wearable system can receive two or more of the following types of input modes: voice commands, head pose, body pose (which may be measured, for example, by an IMU in a beltpack or a sensor external to the HMD), eye gaze (also referred to herein as eye pose), hand gestures (or gestures by other body parts), signals from a user input device (e.g., a totem), environmental sensors, etc.

[0049] The user intent detector 124 can use one or more of the inputs to determine the user's intent. For example, the user intent detector 124 can use one or more of the inputs to determine a target object toward which the user is directing their focus and / or intending to interact with. Additionally, the user intent detector 124 can determine whether the user is viewing the object or is in the process of interacting with the object, and if so, the type of interaction that is about to occur.

[0050] The user intent detector 124 can use transmodal input fusion techniques to determine the user's intent. For example, the user intent detector can aggregate direct and indirect user input from multiple sensors to produce a multimodal interaction for an application. Examples of direct input may include gestures, head pose, voice input, totem, eye gaze direction (e.g., eye gaze tracking), other types of direct input, etc. Examples of indirect input may include environmental information (e.g., environmental tracking), what other users are doing, and geolocation.

[0051] The wearable system can use an outward-facing imaging system to track and report gestures to the visualization system 120. For example, the outward-facing imaging system can obtain images of the user's hands and map the images to corresponding hand gestures. The visualization system 120 can also detect the user's head gestures using an object recognizer 121. In another example, the HMD can use an IMU to recognize head pose.

[0052] The wearable system can perform eye gaze tracking using an inward-facing camera. For example, the inward-facing imaging system can include an eye camera configured to capture images of the user's eye region. The wearable system can also receive input from a totem.

[0053] The user intent detector 124 can use various inputs and techniques to determine a target object for the user. The target object can be, for example, an object to which the user is paying attention (e.g., viewing for at least a threshold duration), moving toward, or with which the user is about to interact. The user intent detector 124 can derive a given value from an input source and produce a grid of possible values ​​for candidate virtual objects with which the user may potentially interact. In some embodiments, the value can be a confidence score. The confidence score can include a ranking, rating, evaluation, quantitative or qualitative value (e.g., a number in the range of 1 to 10, a percentage or percentile, or qualitative values ​​of “A,” “B,” “C,” etc.), etc.

[0054] Each candidate object may be associated with a confidence score, and in some cases, the candidate object with the highest confidence score (e.g., higher than the confidence scores of the other objects or higher than a threshold score) is selected as the target object by the user intent detector 124. In other cases, objects with confidence scores below the threshold confidence score are eliminated from consideration as target objects by the system, which can improve computational efficiency.

[0055] As an example, the user intent detector 124 can use eye tracking and / or head pose to determine that the user is looking at the candidate object. The user intent detector 124 can also use data from a GPS sensor on the user's wearable device to determine whether the user is approaching the candidate object or moving in a different direction. The user intent detector 124 can also use gesture detection to determine whether the user is reaching out for the candidate object. Based on this data, the user intent detector 124 can determine the user's intent, for example, to interact with the candidate object or not interact with the object (e.g., simply looking at the object).

[0056] The swarm intent analyzer 125 can analyze the intent of multiple users in a swarm, e.g., a group of users in a shared real or virtual space. For example, the swarm intent analyzer 125 can analyze the intent of an audience, e.g., a class of students, people viewing a presentation or demonstration, or people playing a game. The swarm intent analyzer 125 can, for example, generate swarm metrics based on the analysis, reallocate user tasks based on the analysis, recommend actions for particular users, and / or take other actions, as described below.

[0057] For instruction or presentation, the user group intent analyzer 125 can determine which users are paying attention to the instruction. In this example, the group intent analyzer 125 can either receive data defining the focus of instruction, e.g., a leader user, a whiteboard, a display screen, or a car or other object that is the subject of instruction, or determine the focus of instruction based on the target object to which the user is looking. For example, if all or at least a threshold number or percentage of the users in the group are looking at the same object, the group intent analyzer 125 can determine that the object is the subject of instruction.

[0058] Throughout the instruction, the user group analyzer 125 can monitor users to determine which users are paying attention to the instruction and calculate metrics for the instruction, such as average group attention, the average amount of time users spent looking at the instructional object, etc. For example, the user group analyzer 125 can determine the percentage of users paying attention to each particular object and the average amount of time each user paid attention to the object over a given period of time. When users are given a task, the user group analyzer 125 can determine, for each user, whether the user is performing the task, the percentage of time the user spent performing the task, and aggregate measures for the group, such as the percentage of users performing those tasks, the average amount of time users in the group spent performing those tasks, etc.

[0059] The user group analyzer 125 can also determine users following a display or visual inspection, users looking at a leader (e.g., a presenter or leader in a game), users looking at the environment, group responsiveness to instruction or questions, or potential for task failure (e.g., based on the number of users not following an instruction subject or not paying attention to the leader).

[0060] The target object for each user in the swarm can be used to identify distractors in the shared virtual space. For example, if most users are paying attention to the instructional focus, but several other users are looking at another object, the swarm intent analyzer 125 can determine that the other object is a distractor.

[0061] The group intent analyzer 125 can calculate the average group movement and the direction of group movement. The user group analyzer 125 can compare this movement data, for example, in real time, to a target path and provide feedback to the instructor using the group intent feedback generator 126 (described below). For example, if the instructor is teaching a fitness or dance class, the user group analyzer 125 can compare the user's movement to the target movement and determine a score for the degree to which the user's movement matches the target movement.

[0062] In some embodiments, each user may be assigned a task, for example, as part of a group project or team game. The group intent analyzer 125 can monitor each user's intent, compare the intent to their tasks, and determine whether the user is performing those tasks. The group intent analyzer 125 can use this data to calculate task efficiency and / or potential for task failure.

[0063] The group intent analyzer 125 can determine hidden group dynamics based on the user's intent. For example, the group intent analyzer 125 can find local visual attention within a group of users. In particular embodiments, the group intent analyzer 125 can determine that one subgroup is looking at the leader, while another subgroup is looking at an object being discussed by the leader.

[0064] The swarm intent analyzer 125 can use this data to determine swarm imbalances. For example, the swarm intent analyzer 125 can identify inefficient clustering of users, interference from multiple competing leaders, and / or slowdowns or errors in task handover during production line coordination tasks.

[0065] In some implementations, the group intent analyzer 125 can use multi-user transmodal convergence within the group to determine various characteristics or metrics about the group. Transmodal convergence can include gaze attention (e.g., distractions or level of user immersion), eye-hand intent (e.g., user reaching, pointing, grasping, blocking, pushing, or throwing), eye-foot intent for following a path (e.g., walking, running, jumping, sidestepping, leaning, or turning), level of physical activity (e.g., ground speed, manual effort, or fatigue), and / or member and group cognition (e.g., harsh noises, confusion, or elevated cognitive load).

[0066] The swarm intent analyzer 125 can determine swarm and subgroup statistics and / or metrics. For example, the swarm intent analyzer 125 can determine swarm motion (e.g., stop and start counts along a path), swarm physical breakup and recombination rates, swarm radius, subgroup counts (e.g., the number of users performing a particular task or focusing on a particular object), subgroup size, main group size, subgroup breakup and recombination rates, average subgroup membership rates, and / or subgroup characteristics (e.g., gender, age, role, conversation rate, communication rate, question rate, etc.).

[0067] In some implementations, a secondary sensor, such as a world camera, can face the group of users (if the user provides permission). In this example, the group intent analyzer 125 can determine additional group characteristics, such as an estimate of the group's emotional state, the level of group tension or relaxation, an estimate of group discussion intent (e.g., based on information flow, contributions from group members, or nonverbal qualities of linguistic expression), stylistic group dynamics (e.g., leadership style, presentation style, instruction style, or behavior style), and / or a visual translation of body language (e.g., indications of intent communicated nonverbally from body posture or level of agitation).

[0068] Data collected and / or generated by the group intent analyzer 125 can be shared over a local network or stored in the cloud and queried for the measurements and characteristics described above, or for secondary characteristics. Secondary characteristics can include, for example, identifying, locating, or tracking obscured group members who are not wearing wearable devices. Secondary characteristics can also include ghost group members in a room (e.g., people not wearing wearable devices) and / or ghost group members operating on an augmented team in an augmented virtual space. Secondary characteristics can also include identifying, locating, or tracking people and objects that affect group dynamics but are not part of the group, and / or identifying group movement patterns.

[0069] The group intent feedback generator 126 can provide feedback to a user or a specific user based on the results produced by the group intent analyzer 125. For example, the group intent feedback generator 126 can generate and present various group metrics for a group of users, such as the number of users paying attention to instruction, the number of users looking at the leader, etc., on a specific user's display. This visualization can be in the form of a vector diagram or heat map showing the users' intent or focus. For example, the heat map can show a specific color for each user that indicates the user's level of attention to the leader or leader. In another example, the user's color can represent the efficiency with which the user is performing their assigned tasks. In this way, a leader viewing the visualization can regain the attention of distracted users and / or return users to their individual tasks. In another example, as shown in FIGS. 5A-5C, a heat map can show the areas of user attention and the relative number of users paying attention to those areas.

[0070] In some implementations, the swarm intent feedback generator 126 can take action based on the results produced by the swarm intent feedback generator 126. For example, if a subgroup is performing a task inefficiently or is distracted, the swarm intent feedback generator 126 can reassign users to a different task or subgroup. In particular examples, the swarm intent feedback generator 126 can reassign a leader of a well-performing subgroup to an inefficient subgroup to improve the performance of the inefficient subgroup. In another example, if a task is resource-starved, e.g., does not have enough members to perform the task, the swarm intent feedback generator 126 can reassign users to that subgroup from a subgroup with an overabundance of members, e.g., causing human interference in performing the task.

[0071] With respect to individual users, the swarm intent feedback generator 126 can generate alerts or recommend actions for the individual user to the leader or other users. For example, if a user's intent deviates from the task assigned to the user, the swarm intent feedback generator 126 can generate alerts to notify the leader and / or recommend actions for the individual user, such as a new task or a correction to the way the user is performing a current task.

[0072] The swarm intent feedback generator 126 can also build and update profiles or models of users based on their activity within the swarm. For example, the profiles or models can describe the behavior of users within the swarm with information such as average attention level, task efficiency, distraction level, objects that tend to distract users, etc. This information can be used by the swarm intent feedback generator 126 during future sessions to predict how users will react to various tasks or potential distractions, proactively generate alerts to leaders, assign appropriate tasks to users, and / or determine when to reassign users to different tasks.

[0073] 2A-2C are exemplary attention models 200, 220, and 240, respectively, illustrating the attention of a group of users to a single object. Referring to FIG. 2A, model 200 includes a group of users 201-204, all looking at the same object 210. As shown for user 201, in FIG. 2A, each user 201-204 has a solid arrow 206 indicating the user's head pose direction, a dashed arrow 207 indicating the user's gaze direction, and a line 208 with circles on either end indicating the user's arm gesture direction. The same types of arrows / lines are used to indicate the same information as in FIGS. 2A-8.

[0074] The visualization system can recognize head pose direction using an IMU to recognize head pose. The visualization system can also recognize each user's gaze direction using eye tracking. The visualization system can also be a user gesture detection technique to determine the direction the users are moving their arms. Using this information, the visualization system can determine that, with respect to model 200, all of users 201-204 are looking at object 210 and that all of users 201-204 are reaching for the same object.

[0075] The visualization system can use this information to take action, generate alerts, or generate data to present to one or more users. For example, the visualization system can assign tasks to users 201-204. In particular examples, the visualization system might assign users 201-204 the task of picking up an object. Based on each user's gaze direction and arm direction combined with their relative position to object 210, the visualization system can determine that users 201-204 are reaching for object 210, but that user 202 is farther from the object than the other users. In response, the visualization system can instruct the other users to wait for a particular period of time or until a countdown provided to each user 201-204 has completed. In this manner, the visualization system can synchronize users' tasks based on their collective intent.

[0076] Referring to FIG. 2B , model 220 includes a group of users 221-225 who are all looking at the same object 230. In this example, one user 223 is gesturing toward object 230. This user 223 may be the group leader, or a presenter or instructor who is talking about object 230. In another example, object 230 may be a table that holds other objects that user 223 is describing to the other users. This model can be used to show user 223 information about the focus of the other users. In this example, all of the other users are looking at object 230, but in other examples, some users may be looking at user 223 or elsewhere. In such cases, having information about the users' attention can help user 223 return such users' focus to object 230 or, if appropriate, to user 223.

[0077] 2C, model 240 includes a group of users 241-244 who are all looking at the same object 250. In this example, one user 242 is gesturing toward object 250. For example, user 242 may be a group leader or instructor, and object 250 may be a whiteboard or display to which the other users are assumed to be looking. As with model 220, information about what the other users are looking at can help user 242 focus the other users appropriately.

[0078] 3A and 3B are example attention models 300 and 350 illustrating a group of users' attention to, or interaction with, content and people. Referring to FIG. 3A, a leader user 310 is speaking to a group of users 320, with reference to an object 305, e.g., a whiteboard, display, or other object. In this example, the group exhibits sole attention to object 305, as indicated by the dashed arrow representing the user's gaze.

[0079] 3B, model 350 represents divided or scattered audience attention. In this example, a leader user 360 stands near object 355, the subject of discussion. Some users in a group of users 370 look at leader user 360, while others look at object 355. If users should be paying attention to leader user 360 or object 355, the visualization system can alert users who are paying attention to the wrong object, or alert leader user 360 so that leader user 360 can correct the other users.

[0080] Figures 4A-4C are vector diagrams 400, 420, and 440, respectively, depicting the attention of a group of users. Figures 5A-5C are heat map diagrams 500, 520, and 540 of user attention, respectively, corresponding to the vector diagrams of Figures 4A-4C. Vector diagram 400 and heat map diagram 500 are based on the attention of a group of users in model 300 of Figure 3A. Similarly, vector diagram 420 and heat map diagram 520 are based on the attention of a group of users in model 320 of Figure 3B. Vector diagram 440 and heat map diagram 540 are based on the attention of a divergent audience, including users who are focused on many different areas, rather than devotion to one or two specific objects.

[0081] Heatmap diagrams 500, 520, and 540 are shown in two dimensions but represent three-dimensional heatmaps. The heatmap diagrams include ellipsoids, shown as ellipsoids, representing the attentional areas of users within a group of users. Smaller ellipses presented over larger ellipses represent taller or more elongated ellipsoids when three dimensions are shown. The height of the ellipsoids can represent the level of attention that users within a group are paying to the area represented by the ellipsoids, e.g., taller ellipsoids have more or less attention. The areas of the ellipsoids shown in Figures 5A-5B represent the attentional areas of users, e.g., wider areas represent larger areas where users have focused their attention. Because the ellipsoids represent ellipsoids, the following description will refer to the ellipsoids as ellipsoids.

[0082] 4A and 5A, a vector diagram 400 and a heatmap diagram 420 represent the attention of the group of users 320 of FIG. 3A. The vector diagram 400 includes a set of vectors 410 that represent the attention of the users in the group. In this example, the vector diagram 400 represents the exclusive attention of the group of users to the same object 305. As described above with reference to FIG. 3A, each user in the group is looking at the same object.

[0083] 5A includes multiple ellipsoids, each representing an area in front of a user. For example, object 305 may be on a stage or table in front of a user. Heatmap illustration 500 may include an ellipsoid for each portion of the area that represents that area and that at least one user views over a period of time. The area covered by each ellipsoid may correspond to the area in front of the user.

[0084] Heatmap chart 500 may represent the user's fluctuating average attention over a given period of time, such as the previous 5, 10, 30 minutes, or another suitable period of time. In this manner, the ellipsoids in heatmap chart 500 may move and change size as the user's attention changes.

[0085] In this example, a user, for example, user 320, can view heatmap view 500 and determine that all users are converging toward user 320 or object 305. Therefore, the user may not need to take any action to return the user's focus to user 320 or object 305.

[0086] 4B and 5B, a vector diagram 420 and a heat map diagram 520 represent the divided attention among the group of users 370 in FIG. 3B. The vector diagram 420 includes a set of vectors 430 representing the attention of the users in the group. Some of the users 430 are looking at object 325, while others are looking at user 360, who is, for example, discussing object 325. The heat map diagram includes a first set of ellipsoids 531 representing the area where user 360 is located and a second set of ellipsoids 532 representing the area where object 325 is located. The smaller-area ellipsoids at the top of ellipsoids 531 may represent the location where user 360 spent the most time because they are higher than the other ellipsoids and represent more attention to the location represented by the ellipsoid. The ellipsoids with larger areas represent a lower level of intensive attention associated with those ellipsoids and may represent larger areas where user 360 was not present but that users in the group may have looked. The area of ​​the largest ellipsoid 532 is smaller than those in the set of ellipsoids 531 because the object 325 may not have moved at all.

[0087] 4C and 5C, a vector diagram 420 and a heatmap diagram 520 represent divergent attention among a group of users. The vector diagram 420 includes a set of vectors 430 that represent the attention of users within the group. In this example, some of the users are looking at an object 445, while other users are looking at a user 480 who is discussing the object 445.

[0088] Heatmap 550 includes ellipsoid 550, which has a large area and represents all of the areas where users focused their attention over a given period of time. In addition, heatmap 540 includes ellipsoids 551-553, which represent smaller areas where users focused more attention, e.g., areas where more users focused their attention or areas where users focused their attention for a longer period of time. A user viewing this heatmap can learn that the user is not sufficiently focused on either user 480 or object 445 and can interrupt presentation or instruction to recapture the user's attention. In another example, the visualization system can determine that users are not focused on the same object based on the collective attention and the disparity between the ellipsoids and generate a warning for user 480 or a group of users.

[0089] 6 shows exemplary attention models 600 and 650 in which two users' attention is on a common object. Model 600 represents two users 621 and 622 looking at and gesturing toward object 610. Model 650 represents two users 671 and 672 looking at and gesturing toward each other, rather than object 610.

[0090] 7 is an example attention model 700 in which the attention of two users 721 and 722 is on a common object 710. In this example, user 722's attention is on one side of object 710, and the confidence level that user 722 is looking at object 710 may be lower than that of user 721 using gaze or eye tracking alone. However, if user 722 makes a gesture toward object 710, this may increase the confidence.

[0091] 8 is an example attention model 800 in which there is mutual user-to-user attention. In particular, a user 811 is looking at another user 822, and user 822 is looking at user 811.

[0092] 9 is a flowchart of an example process 900 for determining the intent of one or more users in a group of users and acting accordingly. This process can be implemented, for example, by visualization system 120 of FIG. 1A.

[0093] The system identifies distinct objectives for each of two or more users in the group of users in the shared virtual space (902). The shared virtual space may include, for example, a virtual space that uses augmented reality, a real space or environment, or other icons or figures representing avatars, game players, or real people.

[0094] The objective for a user can be a task to be performed by the user. For example, a leader can assign tasks to individual users within a group or to subgroups of users. In another example, the system can assign tasks to users randomly, pseudo-randomly, or based on a profile for the user (e.g., based on previous performance, such as task efficiency or level of distractions, when performing previous tasks).

[0095] The user's intent can be to pay attention to a target object, e.g., a physical object or a virtual object. The target object can be a person, e.g., a teacher or presenter, a display or whiteboard, a person performing surgery, an object being demonstrated or repaired, or another type of object. In this example, the target object can be defined by the reader or by the system. For example, the system can determine the target object based on the amount of users looking at the target object, e.g., at least a threshold percentage, such as 50%, 75%, etc.

[0096] The system determines (904) individual intents for each of two or more users in the group based on multiple inputs. The multiple inputs can be from multiple sensors having different input modalities. For example, the sensors can include one or more imaging systems, e.g., an outward-facing imaging system and an inward-facing imaging system, environmental sensors, and / or other suitable sensors. At least some of the sensors can be part of the users' devices, e.g., wearable systems as described above, that enable the users to participate in the shared virtual space. Other sensors can include a map of the virtual space and / or an object recognizer, to name a few examples.

[0097] An intent can define a target object with which the user is likely to interact and the user's interaction with the target object. For example, if the user is walking toward and looking at the target object, the system can determine that the user is likely to interact with the target object.

[0098] The multiple inputs may include gaze data defining the user's line of sight, gesture data defining the user's hand gestures, and direction data defining the direction the user is moving. For example, this data may be received from the user's wearable device, as described above. Using such data allows the system to more accurately determine the user's intent.

[0099] For each of the two or more users, the system determines whether the user is performing a user goal (906). The system may make this determination based on a comparison between the determined intent for the user and the user's goal. For example, if the determined intent (e.g., picking up a particular object) matches the user's goal (also picking up the target object), the system may determine that the user is performing, e.g., completing or achieving, the user goal. If the user, for example, moves away from the target object with the intent to interact with a different object, the system may determine that the user is not performing the user goal.

[0100] In another example, a user's goal may be to pay attention to the demonstration or the instructor. The system may determine the target object each user is looking at and, in effect, determine the number of users who are paying attention to the demonstration, the number of users who are paying attention to the instructor or presenter, and / or the number of users who are paying attention to the object that is the subject of the instruction or presentation. In this example, the goal for the users may be to pay more attention to the target object rather than the instructor or presenter. The system may determine the level of attention of a group of users to the target object and / or the instructor or presenter, for example, based on the number of users looking at each and / or the proportion of time each user is looking at each. The system may then determine whether the users, individually or as a group, are performing this goal based on the attention level for the user or group of users.

[0101] The system generates output data regarding the group of users (908). The output data can include characteristics, statistics, measurements, and / or status of individual users or groups of users. For example, if a particular user is not performing a user objective, the output data can indicate who the particular user is, the objective that the user is not performing, and the objective. The output data can indicate the number of users performing those objectives, such as the number or percentage of users performing those objectives (e.g., looking at the presenter or target object), average group movement, the potential for task failure, etc.

[0102] In some implementations, the output data can include a graph or chart. For example, the system can generate a heat map that shows the amount of users, for a group, who are performing their objectives. For example, the heat map can include a range of colors or shades that indicate the level at which the users are performing their objectives. For each user, the heat map can include an element representing the user, which can be presented in a color that matches the level at which the user is performing their objectives.

[0103] The system provides 910 output data for each of the one or more users regarding the presentations at the device. For example, the system can provide the output data to a group leader. In this example, the user can take action based on the data, for example, to correct the intent of the one or more users.

[0104] In some implementations, the system can take action based on the output data. For example, the system can take corrective or remedial action, such as reallocating tasks from users who are not performing their objectives to other users who have already performed their objectives. In another example, the system can determine that some group tasks are resource-starved and reallocate other users to that subgroup.

[0105] In a presentation or instruction environment, the system may determine that at least a threshold amount (e.g., number or percentage) of users are distracted or otherwise not paying attention to the target object. In this example, the system may take action to get the users to pay attention, for example, by presenting a notification on their display to pay attention to the target object.

[0106] Embodiments of the subject matter and functional operations described herein can be implemented in digital electronic circuitry, tangibly embodied computer software or firmware, computer hardware, or a combination of one or more of these, including the structures disclosed herein and their structural equivalents. Embodiments of the subject matter described herein can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible, non-transitory program carrier device for execution by or to control the operation of a data processing apparatus. Alternatively, or in addition, the program instructions can be encoded on an artificially generated propagated signal, such as a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to a receiver device suitable for execution by a data processing apparatus. The computer storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of these.

[0107] The term "data processing apparatus" refers to data processing hardware and includes all kinds of apparatus, devices, and machines for processing data, including, by way of example, a programmable processor, a computer, or multiple processors or computers. An apparatus can also be or include special purpose logic circuitry, such as an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit). An apparatus optionally includes, in addition to hardware, code that creates an execution environment for a computer program, such as code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of these.

[0108] A computer program, which may also be referred to or described as a program, software, software application, module, software module, script, or code, can be written in any form of programming language, including a compiled or interpreted language, or a declarative or procedural language, and can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program may, but need not, correspond to a file in a file system. A program can be stored in one or more scripts, part of a file that holds other programs or data, e.g., stored in a markup language document, in a single file dedicated to the program, or in multiple cooperating files, e.g., files that store one or more modules, subprograms, or portions of code. A computer program can be deployed to run on one computer or on multiple computers located at one facility or distributed across multiple facilities and interconnected by a communications network.

[0109] The processes and logic flows described herein may be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows may also be performed by, and apparatus may also be implemented as, special purpose logic circuitry, such as an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit).

[0110] A computer suitable for executing a computer program includes, by way of example, a general-purpose or special-purpose microprocessor, or both, or any other type of central processing unit. Typically, the central processing unit will receive instructions and data from a read-only memory or a random-access memory, or both. The essential elements of a computer are a central processing unit for performing or executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices, such as magnetic, magneto-optical, or optical disks, for storing data, or be operatively coupled to receive data from or transfer data to them, or both. However, a computer need not have such devices. Furthermore, a computer can be embedded within another device, such as a mobile phone, a personal digital assistant (PDA), a portable audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device, such as a universal serial bus (USB) flash drive, to name a few.

[0111] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, by way of example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices, magnetic disks such as internal hard disks or removable disks, magneto-optical disks, and CD-ROM and DVD-ROM disks. The processor and memory may be supplemented by, or incorporated in, special purpose logic circuitry.

[0112] To provide for interaction with a user, embodiments of the subject matter described herein can be implemented on a computer having a display device, e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user, and a keyboard and pointing device, e.g., a mouse or trackball, by which the user may provide input to the computer. Other types of devices can likewise be used to provide for interaction with a user; for example, feedback provided to the user can be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback, and input from the user can be received in any form, including acoustic, speech, or tactile input. Additionally, the computer can interact with the user by sending documents to and receiving documents from a device used by the user, for example, by sending a web page to a web browser on the user's device in response to a request received from the web browser.

[0113] Embodiments of the subject matter described herein can be implemented within a computing system that includes a back-end component, e.g., a data server, or includes a middleware component, e.g., an application server, or includes a front-end component, e.g., a client computer having a graphical user interface or web browser through which a user may interact with an implementation of the subject matter described herein, or any combination of one or more such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communications network. Examples of communications networks include local area networks (LANs) and wide area networks (WANs), e.g., the Internet.

[0114] A computing system may include clients and servers. Clients and servers are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. In some embodiments, the server transmits data, e.g., HTML pages, to a user device for the purpose of displaying data to and receiving user input from a user interacting with the user device, e.g., acting as a client. Data generated at the user device, e.g., the result of a user interaction, can be received from the user device at the server.

[0115] An example of one such type of computer is shown in FIG. 10, which shows a schematic diagram of a general-purpose computer system 1000. According to one implementation, the system 1000 can be used for the operations described in connection with any of the computer-implemented methods described above. The system 1000 includes a processor 1010, a memory 1020, a storage device 1030, and an input / output device 1040. Each of the components 1010, 1020, 1030, and 1040 are interconnected using a system bus 1050. The processor 1010 is capable of processing instructions for execution within the system 1000. In one implementation, the processor 1010 is a single-threaded processor. In another implementation, the processor 1010 is a multi-threaded processor. The processor 1010 is capable of processing instructions stored in the memory 1020 or on the storage device 1030 and displaying graphical information for a user interface on the input / output device 1040.

[0116] The memory 1020 stores information within the system 1000. In one implementation, the memory 1020 is a computer-readable medium. In one implementation, the memory 1020 is a volatile memory unit. In another implementation, the memory 1020 is a non-volatile memory unit.

[0117] The storage device 1030 is capable of providing mass storage for the system 1000. In one implementation, the storage device 1030 is a computer-readable medium. In various different implementations, the storage device 1030 may be a floppy disk device, a hard disk device, an optical disk device, or a tape device.

[0118] The input / output device(s) 1040 provide input / output operations for the system 1000. In one implementation, the input / output device(s) 1040 include a keyboard and / or a pointing device. In another implementation, the input / output device(s) 1040 include a display unit for displaying a graphical user interface.

[0119] While this specification contains many specific implementation details, these should not be construed as limitations on the scope of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments. Certain features described herein in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Furthermore, while features may be described above as operative in a combination and may even be initially claimed as such, one or more features from a claimed combination can, in some cases, be deleted from the combination, and the claimed combination may be directed to subcombinations or variations of subcombinations.

[0120] Similarly, while operations may be depicted in the figures in a particular order, this should not be understood as requiring such operations to be performed in the particular order shown, or in sequential order, or that all of the illustrated operations be performed, to achieve desirable results. In some situations, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system modules and components in the above-described embodiments should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems may generally be integrated together in a single software product or packaged within multiple software products.

[0121] Specific embodiments of the present subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve desirable results. As an example, the processes depicted in the accompanying figures do not necessarily require the particular order or sequential order shown to achieve desirable results. In some cases, multitasking and parallel processing may be advantageous.

[0122] The claims are as follows:

Claims

1. 1. A method implemented by one or more data processing devices, said method comprising: With respect to a group of users in a shared virtual space, identifying a distinct objective for each of two or more of the users in the group of users; For each of the two or more users, determining an individual intent of the user based on input from a plurality of sensors having different input modalities, at least some of the plurality of sensors being sensors of the user's device that enables the user to participate in the shared virtual space, and determining the individual intent of the user includes: determining one or more candidate objects with which the user may potentially interact based on the inputs from the plurality of sensors; and determining a target object for the user based on the one or more candidate objects; and and determining whether the user is performing the individual purpose for the user based on the individual intent; generating, with respect to the group of users, output data based on the individual goals for each of the two or more users and the individual intentions for each of the two or more users; providing, to an individual device of each of one or more users in the group of users, the output data related to presentations at the individual device of each of the one or more users; A method comprising:

2. 2. The method of claim 1, wherein the individual objective for at least one of the two or more users comprises at least one of (i) a task to be performed by the at least one user, or (ii) an object of interest for which the at least one user should be looking.

3. 3. The method of claim 2, wherein identifying the distinct objectives for each of the two or more users includes determining as the target object a target to which at least a threshold amount of the users in the group of users are looking.

4. the individual device of each user comprises a wearable device; Determining the individual intent of the user includes: receiving, from the wearable device of the user, gaze data defining a gaze of the user, gesture data defining a hand gesture of the user, and direction data defining a direction in which the user is moving; determining an intention of the user with respect to a target object based on the gaze data, the gesture data, and the direction data as the individual intention of the user; The method of claim 1 , comprising:

5. generating the output data based on the individual goal for each of the two or more users and the individual intention for each of the two or more users includes determining that a particular user has not performed the individual goal for the particular user; 2. The method of claim 1, wherein providing the device of each of one or more users in the group of users with the output data regarding presentations at the device of each of the one or more users includes providing a device of a leader user with data indicating the particular user and data indicating that the particular user has not performed the individual objective for the particular user.

6. The method of claim 1 , wherein the output data comprises a heat map indicating an amount of users within the group of users who are performing the individual objectives of the users.

7. The method of claim 1 , further comprising performing an action based on the output data.

8. The method of claim 7 , wherein the action includes reallocating one or more users to different purposes based on the output data.

9. The method of claim 8, wherein determining the one or more candidate objects comprises: determining one or more confidence scores for the one or more candidate objects based on the inputs from the plurality of sensors, wherein a respective confidence score of the one or more confidence scores is determined for each of the one or more candidate objects; determining the one or more candidate objects based on the one or more confidence scores; and The method according to any one of claims 1 to 8, comprising:

10. The method described in claim 9, wherein determining the target object for the user based on the one or more candidate objects includes selecting the candidate object with the highest confidence score among the one or more candidate objects as the target object.

11. A method as described in claim 9 or 10, wherein determining the target object for the user based on the one or more candidate objects includes eliminating individual candidate objects from the one or more candidate objects that have a confidence score below a predetermined threshold.

12. A method according to any one of claims 9 to 11, wherein the one or more confidence scores include one or more of a ranking, rating, evaluation, quantitative value, qualitative value, percentage, or percentile.

13. A method according to any one of claims 1 to 12, wherein the input includes at least one direct input and at least one indirect input.

14. The method of claim 13, wherein the at least one direct input includes one or more of gaze data defining the user's gaze, gesture data defining the user's hand gestures, head posture data defining the user's head posture, and voice input data defining the user's voice input.

15. A method as described in claim 13 or 14, wherein the at least one indirect input includes one or more of environmental data defining an environment, geolocation data defining the geolocation of the user, and other user data defining data of another user.

16. 1. A computer-implemented system comprising: one or more computers; 16. A computer-implemented system comprising one or more computer memory devices interoperably coupled to the one or more computers and having a tangible, non-transitory, machine-readable medium storing one or more instructions that, when executed by the one or more computers, perform the method of any one of claims 1 to 15.

17. 16. A non-transitory computer readable medium having stored thereon one or more instructions executable by a computer system to perform the method of any one of claims 1 to 15.

Citation Information

Patent Citations

  • Attention calling apparatus and method and information processing system

    JP2006146871A

  • Information processing system, information processing method, and program

    JP2007079647A

  • Availability calculation system, availability calculation method and availability calculation program

    JP2016122272A

  • Attention calling apparatus and method and information processing system

    US20060083409A1

  • Mixed reality display system and mixed reality display terminal

    WO2018225149A1