Attention decoding to enable visual breakthrough

By using computer vision and machine learning to analyze communication cues, the system optimizes the display of virtual content representations based on user attention and intention, enhancing user engagement and reducing distractions in augmented or virtual reality environments.

WO2026064163A1PCT designated stage Publication Date: 2026-03-26APPLE INC
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-09
Publication Date
2026-03-26

Smart Images

  • Figure US2025045516_26032026_PF_FP_ABST
    Figure US2025045516_26032026_PF_FP_ABST
Patent Text Reader

Abstract

Various implementations disclosed herein include devices, systems, and methods that control breakthrough of virtual content. For example, a process may present a view of an extended-reality (XR) environment via a head mounted device (HMD). The view may include virtual content positioned within a three-dimensional (3D) space of an XR environment and a person or people other than a user wearing the HMD in a physical environment is not depicted in the view. The process may further obtain sensor data representing a characteristic of the person. The process may further determine to replace a portion of the virtual content of the view by determining an attention or intention of the person to provide a depiction of the person based on the characteristic. The process may further update the view of the XR environment in which the portion of the virtual content is replaced with the depiction of the person.
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Docket No. 097425-01472(P67210W01)ATTENTION DECODING TO ENABLE VISUAL BREAKTHROUGHTECHNICAL FIELD

[0001] The present disclosure generally relates to systems, methods, and devices that control virtual content presentation to display a representation of a person or people based on interpreting a communication characteristic of the person or people.BACKGROUND

[0002] Existing techniques for presenting a person within an environment to a user of a device may be improved with respect to accuracy of potential communications with the user of the device to provide accurate and desirable viewing experiences.SUMMARY

[0003] Various implementations disclosed herein include devices, systems, and methods that control breakthrough of virtual content to display a depiction (e.g., a virtual depiction) of a person (or people) to a device user (e.g., of a head mounted device (HMD)). In some implementations, controlling breakthrough of virtual content may include allowing display of a depiction of a person for the device user based on determining that the person is communicating (e.g., talking to, looking at, etc.) with the device user. In some implementations, controlling breakthrough of virtual content may include preventing display of a depiction of a person to the device user based on determining that the person is not communicating (e.g., not talking to, not looking at, etc.) with the device user.

[0004] In some implementations, breakthrough of the virtual content may be enabled based on interpreting a characteristic(s) of the person and / or the device user to identify specified circumstances in which breakthrough is appropriate. A characteristic(s) may include a gaze direction(s), mouth movements, audio from a microphone, body language, etc. of the person and / or the device user. Likewise, characteristic(s) may be determined via usage of spatial techniques to predict attention by utilizing facial landmarks predicted by a specialized computer vision model. For example, a facial landmark detection processAttorney Docket No. 097425-01472(P67210W01) may include a computer vision task that comprises detecting and localizing specific points or landmarks on a face, such as, inter alia, eyes, nose, mouth, chin, etc.

[0005] In some implementations, subsequent to attention (indicating communications) of the person and / or the device user being identified (e.g., by either the person and / or the device user looking at or speaking to each other), an identity tracking process may be implemented via use of a computer vision model such that breakthrough of the virtual content may persist for a specified amount of time during communications between the person and the device user. For example, when the device user looks away from the person for a few seconds causing the person to be out of a field-of-view of the device user and subsequently looks back at the person, the identity tracking process would verify that the person was the same person from the previous communication and breakthrough will continue by default.

[0006] The characteristic(s) may be used to decode (e.g., determine) an attention and / or intention of the person and / or the device user based on determining that: (a) the person is looking at the device user; (b) the person is speaking to the device user; (c) the person and device user are looking at each other (e.g., joint attention); (d) the person and device user are speaking to each other; (e) the person and device user are looking at a same object(s) such as, inter alia, a pencil or a computer on a desk, (f) facial landmarks of the person and the device user, (g) facial expressions and / or inferred emotional states of the person and / or the device user, etc.

[0007] In some implementations, decoding user attention and / or intention may be performed via use of an existing gaze detection model (e.g., a computer vision model that predicts gaze direction of a person that is not the device user) that is configured to identify if or when a person is looking at a camera or not, for example, via use of bounding boxes place surrounding a face of the person to limit a search for gaze characteristics.

[0008] Some implementations may decode user attention and / or intention via the use of a custom machine learning (ML) model.

[0009] Some implementations may decode user attention and / or intention via the use of a rule-based model / algorithm. For example, a computer vision algorithm may be used to determine that the eye / pupil position satisfies a specified positional criteria to determine that a person is looking towards the device user. Likewise, a computer visionAttorney Docket No. 097425-01472(P67210W01) algorithm may be used to detect a bounding box (defining the eye region of the person) size to decode user attention and / or intention.

[0010] Some implementations may decode user attention and / or intention via the use of a spatial specificity technique such as, for example, looking at eye and / or mouth movement, facial landmarks, etc.

[0011] Some implementations may decode user attention and / or intention via the use of a spatial specificity technique such as, for example, looking at eye and / or mouth movement, facial landmarks, etc. to determine facial expressions and / or inferred emotional states. For example, detected mouth movement may indicate that a person and / or the device user is smiling thereby indicating a positive emotional state (e.g., happy) which may be used to decode user attention and / or intention.

[0012] Some implementations may decode user attention and / or intention via the use of a temporal technique such as, for example, detecting a person looking at a device user for a time period exceeding a threshold.

[0013] Some implementations may decode user attention and / or intention via the use of gesture and / or body language analysis.

[0014] In some implementations, an amount of breakthrough and / or an appearance (within a view of a device user) of a depiction of the person may be adjusted in accordance with a determined level of attention or intention of the person and / or the device user thereby indicating a confidence that circumstances are appropriate for breakthrough.

[0015] Some implementations, deactivate breakthrough of the person (e.g., removing the depiction of the person from a view) when both the person and the device user stop looking at each other for specified time period such as, for example, 5 seconds, etc.

[0016] In some implementations, adjusting an appearance of a depiction of the person may include, inter alia, adjusting a tint level of the depiction of the person, adjusting a transparency level of the depiction of the person, adjusting a focus level of the depiction of the person, etc.

[0017] In some implementations, differing states associated with a process for controlling display or breakthrough of multiple user representations within virtual content may be activated based on interpreting interactions of the users. For example,Attorney Docket No. 097425-01472(P67210W01) differing states may include a background state, a hint state, an active state, and an interactive state each enabled or disabled in response to various configurations of a user / device gaze state and or presence in a physical environment.

[0018] In some implementations, a background state may include a state associated with no user being present (e.g., in the physical environment) thereby causing a user interaction detection system to be placed in an inactive state.

[0019] In some implementations, a hint state may be triggered when a user(s) is present in the physical environment thereby triggering a subtle visual cue (to indicate to a device user (e.g., an HMD user) that a user(s) has entered the physical environment. Likewise, the hint state may be disabled (e.g., going back to a background state) when the user(s) exits the physical environment.

[0020] In some implementations, an active state may be configured to cause the subtle visual cue enabled via the hint state to transition to a partially transparent view of the user(s).

[0021] In some implementations, a subtle visual cue enabled via a hint state may cause a device user to look towards the user(s) that in combination with the user(s) looking at the device user may trigger interactive state configured to cause the subtle visual cue to transition to a fully visible view of the user(s). In some implementations, the interaction state may be disabled when the device user and the user(s) all look away from each other for a specified time period that exceeds a threshold such as 4 or 5 seconds, etc.

[0022] In some implementations, when both a device user and a user(s) look at each other, breakthrough (e.g., a transparency effect) may be activated faster than when only the device user looks at the user or the user looks at the device user. In this instance, more evidence may be beneficial (e.g., detection of one party speaking to the other) when only one party looks at the other to prevent false positives. Likewise, when both parties look at each other, there may be more evidence of an interaction and therefore breakthrough may occur faster.

[0023] In some implementations, an HMD being worn by a first user in a physical environment has a processor (e.g., one or more processors) that executes instructions stored in a non-transitory computer-readable medium to perform a method. The methodAttorney Docket No. 097425-01472(P67210W01) performs one or more steps or processes. In some implementations, the HMD presents a view of an extended-reality (XR) environment via the HMD. The view includes virtual content positioned within a three-dimensional (3D) space of the XR environment and a person other than the first user in the physical environment is not depicted in the view. In some implementations, the HMD obtains first sensor data from at least one sensor on the HMD. The first sensor data may represent a characteristic of the person. In some implementations, the HMD determines to replace at least a portion of the virtual content of the view to provide a depiction of the person based on the characteristic. In some implementations, the HMD updates the view of the XR environment in which the at least a portion of the virtual content is replaced with the depiction of the person.

[0024] In some implementations, an HMD being worn by a first user in a physical environment has a processor (e.g., one or more processors) that executes instructions stored in a non-transitory computer-readable medium to perform a method. The method performs one or more steps or processes. In some implementations, the HMD presents a view of an extended-reality (XR) environment via the HMD. The view includes virtual content positioned within a three-dimensional (3D) space of the XR environment and a first person and a second person other than the first user in the physical environment are not depicted in the view. In some implementations, first sensor data is obtained from at least one sensor on the HMD. The first sensor data may represent a first characteristic of the first person and a second characteristic of the second person. In some implementations, the HMD determines to replace at least a first portion of the virtual content of the view to provide a depiction of the first person based on the first characteristic. Determining the to replace the at least a first portion of the virtual content may include determining an attention or intention of the first person based on determining that the first person is looking at the first user. In some implementations, the HMD determines to replace at least a second portion of the virtual content of the view to provide a depiction of the second person based on the second characteristic. Determining to replace the at least a second portion of the virtual content may include determining an attention or intention of the second person based on determining that the second person is looking at the first user. In some implementations, the HMD updates the view of the XR environment in which the at least a first portion of the virtual content is replaced withAttorney Docket No. 097425-01472(P67210W01) the depiction of the first person and the at least a second portion of the virtual content is replaced with the depiction of the second person.

[0025] In accordance with some implementations, a device includes one or more processors, a non-transitory memory, and one or more programs; the one or more programs are stored in the non-transitory memory and configured to be executed by the one or more processors and the one or more programs include instructions for performing or causing performance of any of the methods described herein. In accordance with some implementations, a non-transitory computer readable storage medium has stored therein instructions, which, when executed by one or more processors of a device, cause the device to perform or cause performance of any of the methods described herein. In accordance with some implementations, a device includes: one or more processors, a non-transitory memory, and means for performing or causing performance of any of the methods described herein.BRIEF DESCRIPTION OF THE DRAWINGS

[0026] So that the present disclosure can be understood by those of ordinary skill in the art, a more detailed description may be had by reference to aspects of some illustrative implementations, some of which are shown in the accompanying drawings.

[0027] Figures 1A-B illustrate exemplary electronic devices operating in a physical environment in accordance with some implementations.

[0028] Figure 2 illustrates is a view an XR environment depicting virtual content and a depiction of a person within a field of view of a user wearing an HMD, in accordance with some implementations.

[0029] Figure 3 illustrates a process implementing the use of using bounding boxes and a gaze detection model for enabling breakthrough control, in accordance with some implementations.

[0030] Figures 4A-4C illustrate views representing various configurations of camera video / frames / images of people within a physical environment and a device user view of an XR environment via a display of a device, in accordance with some implementations.Attomey Docket No. 097425-01472(P67210W01)

[0031] Figure 5 illustrates a joint attention view representing a camera image and a device user view illustrating a representation of an XR environment presented via a device being worn by a device user, in accordance with some implementations.

[0032] Figures 6A-6B illustrate views representing various configurations of camera video / frames / images of multiple people within a physical environment and a device user view of an XR environment via a display of a device such as an HMD, in accordance with some implementations.

[0033] Figure 7 is a state diagram representing differing states associated with a process for controlling display or breakthrough of bystander representations within virtual content based on interpreting user interactions, in accordance with some implementations.

[0034] Figure 8 is a block diagram of an example a system illustrating a user device for controlling virtual content presentation to display a representation of a person based on interpreting a communication characteristic of the person, in accordance with some implementations.

[0035] Figure 9 is a flowchart representation of an exemplary method that controls virtual content presentation to display a representation of a person based on interpreting a communication characteristic of a person, in accordance with some implementations.

[0036] Figure 10 is a flowchart representation of an exemplary method that controls virtual content presentation to display at least two people in an environment based on interpreting a communication characteristic of each of the people, in accordance with some implementations.

[0037] Figure 11 is a block diagram of an electronic device of in accordance with some implementations.

[0038] In accordance with common practice the various features illustrated in the drawings may not be drawn to scale. Accordingly, the dimensions of the various features may be arbitrarily expanded or reduced for clarity. In addition, some of the drawings may not depict all of the components of a given system, method or device. Finally, likeAttorney Docket No. 097425-01472(P67210W01) reference numerals may be used to denote like features throughout the specification and figures.DESCRIPTION

[0039] Numerous details are described in order to provide a thorough understanding of the example implementations shown in the drawings. However, the drawings merely show some example aspects of the present disclosure and are therefore not to be considered limiting. Those of ordinary skill in the art will appreciate that other effective aspects and / or variants do not include all of the specific details described herein. Moreover, well-known systems, methods, components, devices and circuits have not been described in exhaustive detail so as not to obscure more pertinent aspects of the example implementations described herein.

[0040] Figures 1A-B illustrate exemplary electronic devices 105 and 110 operating in a physical environment 100. In the example of Figures 1 A-B, the physical environment 100 is a room that includes a desk 120. The electronic devices 105 and 110 may include one or more cameras, microphones, depth sensors, or other sensors that can be used to capture information about and evaluate the physical environment 100 and the objects within it, as well as information about the user 102 of electronic devices 105 and 110. The information about the physical environment 100 and / or user 102 may be used to provide visual and audio content and / or to identify the current location of the physical environment 100 and / or the location of the user within the physical environment 100.

[0041] In some implementations, views of an extended reality (XR) environment may be provided to one or more participants (e.g., user 102 and / or other participants not shown) via electronic devices 105 (e.g., a wearable device such as an HMD) and / or 110 (e.g., a handheld device such as a mobile device, a tablet computing device, a laptop computer, etc.). Such an XR environment may include views of a 3D environment that is generated based on camera images and / or depth camera images of the physical environment 100 as well as a representation of user 102 based on camera images and / or depth camera images of the user 102. Such an XR environment may include virtual content that is positioned at 3D locations relative to a 3D coordinate system associated with the XR environment, which may correspond to a 3D coordinate system of the physical environment 100.Attorney Docket No. 097425-01472(P67210W01)

[0042] Various implementations disclosed herein include devices, systems, and methods that implement breakthrough control of virtual content to display a person to a user wearing an HMD based on interpreting a characteristic of the user to identify circumstances in which breakthrough is appropriate.

[0043] In some implementations, a view of an XR environment is presented (via a display of an HMD) to a user wearing the HMD in a physical environment. The view may include virtual content positioned within a 3D space of the XR environment. In some implementations, a person (in the physical environment) differing from the user wearing the HMD is not depicted in the view.

[0044] In some implementations, first sensor data may be obtained from a sensor(s) on the HMD. For example, sensor data may include, inter alia, images of the person and / or the HMD user retrieved from an outward and / or inward facing camera(s), audio of the person and / or the HMD user retrieved from a microphone, depth data, etc. The first sensor data may represent a characteristic(s) of the person and / or the HMD user. In some implementations, a characteristic(s) may include a gaze direction(s), mouth movement(s), body language, gestures (e.g., hand waving, etc.), facial landmarks, facial expressions and / or inferred emotional states, etc.

[0045] In some implementations, it may be determined that a portion of the virtual content of the view should be replaced with a depiction of the person. The determination may be made based on interpreting a characteristic(s) of the person and / or the HMD user to identify specified circumstances in which breakthrough is appropriate. A characterise c(s) may include a gaze direction(s), mouth movements, body language, facial landmarks, facial expressions and / or inferred emotional states, etc. of the person and / or the HMD user.

[0046] In some implementations, subsequent to attention (indicating communications) of the person and / or the device user being identified (e.g., by either the person and / or the device user looking at or speaking to each other), an identity tracking process may be implemented via use of a computer vision model such that breakthrough of the virtual content may persist for a specified amount of time during communications between the person and the device user. For example, when the device user looks away from the person for a few seconds causing the person to be out of a field-of-view of theAttorney Docket No. 097425-01472(P67210W01) device user and subsequently looks back at the person, the identity tracking process would verify that the person was the same person from the previous communication and breakthrough will continue by default.

[0047] The characteristic(s) may be used to determine an attention and / or intention of the person and / or the HMD user based on determining that: (a) the person is looking at the HMD user; (b) the person is speaking to the HMD user; (c) the person and HMD user are looking at each other (e.g., joint attention); (d) the person and HMD user are speaking to each other; (e) the person and HMD user are looking at a same object(s) such as, inter alia, a pencil or a computer on a desk, (f) facial landmarks of the person and the device user, facial expressions and / or inferred emotional states, etc.

[0048] In some implementations, determining user attention and / or intention may be performed via use of an existing gaze detection model (e g., a computer vision model that predicts gaze direction of a person that is not the device user) configured to identify if or when a person is looking at a camera, for example, via use of bounding boxes place surrounding a face of the person to limit a search for gaze characteristics.

[0049] Some implementations may determine user attention and / or intention via the use of a custom machine learning (ML) model, a rule-based model / algorithm, a spatial specificity technique such as, for example, looking at eye and / or mouth movement, facial landmarks, facial expressions and / or inferred emotional states, a temporal technique such as, for example, detecting a person looking at a device user for a threshold time, body language analysis, etc.

[0050] In some implementations, a rule-based model / algorithm such as a computer vision algorithm may be used to determine that an eye / pupil position satisfies a specified positional criteria to determine that a person is looking towards the device user. Likewise, a computer vision algorithm may be used to detect a bounding box (defining the eye region of the person) size to decode user attention and / or intention.

[0051] In some implementations, an amount of breakthrough and / or an appearance (within a view of an HMD user) of a depiction of the person may be adjusted in accordance with a determined level of attention or intention of the person and / or the HMD user thereby indicating a confidence that circumstances are appropriate for breakthrough.Attorney Docket No. 097425-01472(P67210W01)

[0052] In some implementations, adjusting an appearance of a depiction of the person may include, inter alia, adjusting a tint level of the depiction of the person, adjusting a transparency level of the depiction of the person, adjusting a focus level of the depiction of the person, etc.

[0053] Figure 2 illustrates is a view of an XR environment 200 depicting virtual content 202 (e.g., comprising a scene representing mountains, water, clouds, etc.) and a depiction 210a of a person 210 within a field of view of a user 208 wearing an HMD, in accordance with some implementations.

[0054] In some implementations, depiction 210a of person 210 may be presented to user 208 (via a view presented via an HMD) based on interpreting a characteristic of the person 210 to identify conditions that may be favorable for performing a breakthrough of person 210 (e.g., depiction 210a) with respect to presentation within XR environment 200. Identifying favorable conditions for performing a breakthrough of person 210 may include determining an attention attribute of person 210 and / or user 208 with respect to communication attempts between each other. For example, attention attributes may be determined based on determining or detecting that person 210 is currently looking at (via gaze detection) or speaking to (via mouth movement detection, audio of speech, etc.) user 208 and / or user 208 is currently looking at (via gaze detection) or speaking to (via mouth movement detection, audio of speech) person 210. Alternatively, attention attributes may be determined based on determining or detecting that person 210 and user 208 are currently looking at a same object and / or are both associated with specified special landmarks.

[0055] In some implementations, attention attributes may be determined use of an existing gaze detection model (e.g., a computer vision model that predicts gaze direction of a person that is not the device user) that identifies whether a user is looking or not looking at a camera via use of bounding boxes a face a person (e.g., person 210) to limit a facial feature search as further described with respect figure 3, infra.

[0056] Figure 3 illustrates a process 300 implementing the use of bounding boxes and a gaze detection model 308 (e.g., a computer vision model that predicts gaze direction of a person that is not the device user) for enabling breakthrough control of a representation of a user, in accordance with some implementations. Process 300 acceptsAttorney Docket No. 097425-01472(P67210W01) as input a video frame 302 comprising a representation of a person 303 and a representation of a person 304.

[0057] At block 305, bounding box representations are applied to a facial region of the representation of a person 303 and a facial region of the representation of a person 304 to generate a representation 306 of video frame 302 comprising a bounding box representation 307 surrounding a facial region of person 303 and a bounding box representation 308 surrounding a facial region of person 304. Bounding box representation 307 and bounding box representation 308 are used to limit a search of video frame 302 to analyze facial regions (e.g., gaze, mouth movements, facial landmarks, facial expressions and / or inferred emotional states, etc.) of person 303 and / or person 304.

[0058] At block 309, a gaze detection model is implemented to determine whether person 303 and / or person 304 is currently looking at a camera of a device (e.g., an HMD) of a user. The gaze detection model analyzes the facial regions within bounding boxes 307 and 308 to determine gaze direction and / or mouth movements of person 303 and / or person 304.

[0059] At block 310, a decision model (e.g., a heuristic model, a machine learning model, rule-based model / algorithm, etc.) is used to determine whether person 303 and / or person 304 is currently looking at the device user (e.g., of a device such as an HMD) based on the analysis of gaze detection model (of block 309) of the facial regions within bounding box representations 307 and 308. Subsequently, the decision model is configured to determine if the gaze direction and / or mouth movements of person 303 and / or person 304 are associated with an attention of person 303 and / or person 304 being directed towards the device user.

[0060] In the example illustrated with respect to process 300, a representation 312 of video frame 302 comprises a bounding box representation 307a surrounding a facial region of person 303 and a bounding box representation 308a surrounding a facial region of person 304. Bounding box representation 307a indicates that the decision model has determined that an attention of person 303 being is not being directed towards the device user (e.g., via gaze direction analysis). Likewise, bounding box representation 308aAttorney Docket No. 097425-01472(P67210W01) indicates that the decision model has determined that an attention of person 304 being is being directed towards the device user (e.g., via gaze direction analysis) thereby enabling block 314 to control a depiction and appearance of user 304 for presentation to the device user via a display of the device such as an HMD.

[0061] Figure 4A-4C illustrate views 400a-400c representing various configurations of camera video / frames / images of people within a physical environment and a device user view of an XR environment via a display of a device such as an HMD, in accordance with some implementations.

[0062] Figure 4A illustrates view 400a representing a camera image 402 (e.g., from an outward-facing camera of an HMD) and a device user view 404 illustrating an XR environment 412 presented via a device 411 (e.g., an HMD) being worn by a device user 403 in a physical environment.

[0063] Camera image 402 comprises a representation of physical environment 401 (e.g., office space) that includes a representation of a user 406 and a user 408. Device user view 404 comprises a representation of an XR environment 412 that includes a representation of an interface 414.

[0064] The representation of user 406 includes a bounding box representation 410 surrounding a facial region of user 406. Likewise, the representation of user 408 includes a bounding box representation 409 surrounding a facial region of user 408.

[0065] The example illustrated in view 400a represents an analysis of the facial region within bounding box representation 410 (e.g., via execution of a gaze detection model) indicating that user 406 is not looking at or talking to (e.g., an attention of user 406 is not directed towards) the device user 403. Likewise, the example illustrated in view 400a represents an analysis of the facial region within bounding box representation 409 (e.g., via execution of a gaze detection model) indicating that user 408 is not looking at or talking to (e.g., an attention of user 409 is not directed towards) the device user 403. The example illustrated in view 400a additionally represents a gaze direction 433 of the user 403 directed towards an icon of interface 414 thereby indicating that the user 403 is looking at interface 414 and not looking at user 408 or user 408. Accordingly, a depictionAttorney Docket No. 097425-01472(P67210W01) and appearance of user 406 and user 408 are disabled from being presented within the representation of XR environment 412.

[0066] Figure 4B illustrates view 400b representing a camera image 415 (e.g., from an outward-facing camera of an HMD) and a device user view 416 illustrating XR environment 412 presented via the device 411 (e.g., an HMD) being worn by the device user 403.

[0067] Camera image 415 comprises a representation of the physical environment 401 that includes a representation of user 406. Device user view 416 comprises a representation of XR environment 412 that includes a representation of interface 414.

[0068] The representation of user 406 includes a bounding box representation 407 surrounding a facial region of user 406.

[0069] The example illustrated in view 400b represents an analysis of the facial region within bounding box 407 (e.g., via execution of a gaze detection model) indicating that user 406 is looking at and / or talking to (e.g., an attention of user 406 is directed towards) the device user 403. Accordingly, a depiction and appearance 422 of user 406 is presented within the representation of XR environment 412. The depiction and appearance 422 of user 406 may be presented with respect to varying levels of tint, transparency, focus, etc. based on a determined level of attention of user 406 thereby indicating a level of confidence that circumstances are appropriate for breakthrough of depiction and appearance 422 of user 406 for presentation within XR environment 412.

[0070] Figure 4C illustrates an alternative view 400c with respect to the example illustrated in view 400b of figure 4B. Alternative view 400c represents a joint attention view of a camera image 415 (e.g., from an outward-facing camera of an HMD) and a device user view 416a illustrating XR environment 412 presented via the device 411 (e.g., an HMD) being worn by the user 403.

[0071] In contrast with the example illustrated in view 400b of figure 4B, the example illustrated in view 400c of figure 4C represents user 406 and the device user 403 engaged in joint attention activities. For example, view 400c represents that user 406 and the device user 403 are looking at and / or speaking with each other.Attorney Docket No. 097425-01472(P67210W01)

[0072] View 400c illustrates an analysis of the facial region within bounding box 407 (e.g., via execution of a gaze detection model) indicating that user 406 is looking at and / or talking to (e.g., an attention of user 406 is directed towards) the device user 403. Likewise, view 400c illustrates that a gaze direction and / or mouth movements 428 of the device user 403 are directed towards user 406 thereby indicating that the device user 403 is looking at and / or talking to (e.g., an attention of the device user 403 is directed towards) user 406.

[0073] Accordingly, based on the joint attention determination, a depiction and appearance 422 of user 406 is presented within the representation of XR environment 412. The depiction and appearance 422 of user 406 may be presented with respect to varying levels of tint, transparency, focus, etc. based on a determined level of attention of user 406 and the device user thereby indicating a level of confidence that circumstances are appropriate for breakthrough of depiction and appearance 422 of user 406 for presentation within XR environment 412.

[0074] Figure 5 illustrates a joint attention view 500 representing a camera image 515 (e.g., from an outward-facing camera of an HMD) and a device user view 516 illustrating a representation of an XR environment 512 presented via a device 511 (e.g., an HMD) being worn by a device user 503 in a physical environment 501, in accordance with some implementations.

[0075] Camera image 515 comprises a representation of physical environment 501 that includes a representation of a user 506. Device user view 516 comprises the representation of XR environment 512 that includes a representation of interface 514 and a depiction and appearance 522 of user 506.

[0076] The example illustrated in view 500 represents user 506 and the device user 503 engaged in joint attention activities. For example, view 500 represents that user 406 and the device user 503 are both looking at a computer 521 and / or are both detected with specified facial landmarks or expressions that may indicate an inferred emotional state.

[0077] View 500 illustrates that a gaze direction 523 of the user 506 is directed towards computer 521 thereby indicating that the user 506 is looking at (e.g., an attention of the user 506 is directed towards) computer 521. Likewise, view 500 illustrates that aAttorney Docket No. 097425-01472(P67210W01) gaze direction 528 the device user 503 is directed towards computer 521thereby indicating that the device user 503 is looking at (e.g., an attention of the device user 503 is directed towards) computer 528.

[0078] Accordingly, based on the joint attention determination (i.e., both users looking at a same object), a depiction and appearance 522 of user 506 is presented within the representation of XR environment 512. The depiction and appearance 522 of user 506 is presented with respect to varying levels of tint, transparency, focus, etc. based on a determined level of attention of user 506 and the device user thereby indicating a level of confidence that circumstances are appropriate for breakthrough of depiction and appearance 522 of user 506 for presentation within XR environment 512.

[0079] Figure 6A-6B illustrate views 600a and 600b representing various configurations of camera video / frames / images of multiple people within a physical environment and a device user view of an XR environment via a display of a device such as an HMD, in accordance with some implementations.

[0080] Figure 6A illustrates view 600a representing a camera image 615 (e.g., from an outward-facing camera of an HMD) and a device user view 616a illustrating an XR environment 612 presented via a device 611 (e.g., an HMD) being worn by a device user 603 in a physical environment.

[0081] Camera image 615 comprises a representation of a physical environment 601 (e.g., office space) that includes a representation of a user 606 and a user 607. Device user view 616a comprises a representation of an XR environment 612 that includes a representation of an outdoor scene 614 (e.g., illustrating clouds and a mountain range).

[0082] The representation of user 606 includes a bounding box representation 617a surrounding a facial (or eye) region of user 606. Likewise, the representation of user 607 includes a bounding box representation 617b surrounding a facial (or eye) region of user 607.

[0083] The example illustrated in view 600a represents an analysis of the facial region within bounding box representation 617a (e.g., via execution of: a gaze detection model, a computer vision algorithm detecting an eye position satisfying a positional criteria, a bounding box size detection algorithm, etc.) indicating that user 606 has started to lookAttorney Docket No. 097425-01472(P67210W01) at (e.g., an attention of user 606 is initially directed towards) the device user 603. Accordingly, an active state is initialized to indicate possible attention of user 606 directed towards device user 603 thereby enabling a partially transparent initial depiction and appearance 622a of user 606 presented within the representation of XR environment 612. The depiction and appearance 622a of user 606 may be presented with respect to varying levels of tint, transparency, focus, etc. based on a determined level of attention of user 606 thereby indicating a level of confidence that circumstances are appropriate for initial breakthrough of depiction and appearance 622a of user 606 for presentation within XR environment 612.

[0084] The example illustrated in view 600a further represents that although an analysis of the facial region within bounding box representation 617b (e.g., via execution of a gaze detection model, a computer vision algorithm detecting an eye position satisfying a positional criteria, a bounding box size detection algorithm, etc.) indicates that user 607 is not currently looking at (e.g., an attention of user 607 is not directed towards) the device user 603, an interactive state is currently enabled as a gaze direction 633 of the user 603 is directed at user 607 thereby allowing breakthrough comprising a fully opaque (i.e., a complete) depiction and appearance 624a of user 607 for presentation within XR environment 612.

[0085] Figure 6B illustrates a subsequent view 600b with respect to view 600a of figure 6A. View 600b represents a view of a camera image 615 (e.g., from an outwardfacing camera of an HMD) indicating a multiple attention scenario between users 606 and 606 and device user 603 enabling a view 600b of XR environment 612 to be presented via the device 611 (e.g., an HMD) being worn by the user 603.

[0086] In contrast with the example illustrated in view 600a of figure 6A, the example illustrated in view 600b of figure 6B represents user 606, user 607, and the device user 603 engaged in joint attention activities. For example, view 600b represents that user 606, user 607, and the device user 603 are looking at and / or speaking with each other.

[0087] The example illustrated in view 600b represents an analysis of the facial region within bounding box representation 617a (e.g., via execution of: a gaze detection model, a computer vision algorithm detecting an eye position satisfying a positionalAttorney Docket No. 097425-01472(P67210W01) criteria, a bounding box size detection algorithm, etc.) indicating that user 606 is currently to looking at (e.g., an attention of user 606 is directed towards) the device user 603. Accordingly, an interactive mode has been initialized to indicate that attention of user 606 is directed towards device user 603 thereby enabling a fully opaque depiction and appearance 622b of user 606 presented within the representation of XR environment 612. The depiction and appearance 622a of user 606 may be alternatively presented with respect to varying levels of tint, transparency, focus, etc. based on a determined level of attention of user 606 thereby indicating a level of confidence that circumstances are appropriate for breakthrough of depiction and appearance 622b of user 606 for presentation within XR environment 612.

[0088] In some implementations, when device user 603 and user 606 and / or 607 look at each other, breakthrough (e.g., a transparency effect) may be activated faster than when only device user 603 looks at user 606 and / or 607 or user 606 and / or 607 looks at the device user 603. In this instance, more evidence may be beneficial (e.g., detection of device user 603 speaking to the user 606) when only one party looks at the other to prevent false positives. Likewise, when device user 603 looks at user 606 and user 606 looks at the device user 603, there may be more evidence of an interaction and therefore breakthrough may occur faster.

[0089] The example illustrated in view 600b additionally, represents an analysis of the facial region within bounding box representation 617b (e.g., via execution of: a gaze detection model, a computer vision algorithm detecting an eye position satisfying a positional criteria, a bounding box size detection algorithm, etc.) indicating that user 607 is currently to looking at (e.g., an attention of user 607 is directed towards) the device user 603. Accordingly, an interactive mode has been initialized to indicate that attention of user 607 is directed towards device user 603 thereby enabling a fully opaque depiction and appearance 624b of user 606 presented within the representation of XR environment 612. The depiction and appearance 624b of user 607 may be alternatively presented with respect to varying levels of tint, transparency, focus, etc. based on a determined level of attention of user 607 thereby indicating a level of confidence that circumstances are appropriate for breakthrough of depiction and appearance 624b of user 606 for presentation within XR environment 612.Attorney Docket No. 097425-01472(P67210W01)

[0090] Accordingly, a joint attention between user 606, user 607, and device user 603 has been detected thereby enabling full breakthrough (e.g., both user depictions are presented as fully opaque) of both users 606 and 607.

[0091] Alternatively, a joint attention determination may be detected (for enabling breakthrough) based on user 606, user 607, and device user 603 all looking at a same object such as, for example, a computer as described with respect to figure 5, supra.

[0092] Figure 7 is a state diagram 700 representing differing states associated with a process for controlling display or breakthrough of bystander representations within virtual content based on interpreting user interactions, in accordance with some implementations. For example, a bystander representation may be a user representation such as user representations 622b and 624b of users 606 and 607 in a physical environment 601 as illustrated in figure 6B.

[0093] State diagram 700 includes a background state 702, a hint state 704, an active state 708, and an interactive state 710 each enabled or disabled in response to various configurations of a user / device gaze 706 state and or presence in the physical environment.

[0094] Background state 702 comprises a state associated with no bystander being present (e.g., in the physical environment) thereby causing the user interaction detection system to be disabled (e.g., placed in an inactive state). Hint state 704 may be triggered when a bystander(s) is present in the physical environment thereby triggering a subtle visual cue (e.g., a light tint or very lightly tinted user representation) to indicate to a device user (e.g., an HMD user) that a bystander has entered the physical environment. Likewise, hint state 704 may be disabled (going back to background state 702) when the bystander(s) exits the physical environment.

[0095] In some implementations, the subtle visual cue enabled via the hint state 704 may cause the device user to look towards the bystander that in combination with detection of the bystander looking at the device user may trigger gaze state 706 to trigger interaction state 710 based on a detected gaze direction of the device user and the bystander. Interaction state 710 may be configured to cause the subtle visual cue enabled via the hint state 704 to transition to a fully visible view (e.g., an opaque version of userAttorney Docket No. 097425-01472(P67210W01) representation 624b as illustrated in figure 6B) of the bystander (e.g., user 607 of figure 6B) within virtual content of a virtual environment presented to the device user via the device (e.g., HMD) as illustrated in figure 6B. In some implementations, the interaction state 710 may trigger the hint state 70 when the bystander and the device user both look away from each other. In some implementations, the interaction state 710 may trigger the background state when the bystander exits the physical environment.

[0096] In some implementations, the subtle visual cue enabled via the hint state 704 may cause the device user to look towards the bystander (with or without the bystander looking at the device user) thereby triggering an active state 708. Active state 708 may be configured to cause the subtle visual cue enabled via the hint state 704 to transition to a partially transparent view (e.g., user representation 622a as illustrated in figure 6A) of the bystander (e.g., user 606 of figure 6A) within virtual content of a virtual environment presented to the device user via the device (e.g., HMD) as illustrated in figure 6A. In some implementations, the hint state is again triggered when the bystander looks away.

[0097] In some implementations, interaction state 710 is disabled when the device user and the bystander(s) both look away from each other for a specified time period that exceeds a threshold (e.g., 4 or 5 seconds). The interaction state 710 may include a persistence timer (to determine the specified time period) to account for abrupt gaze shifts. Accordingly, if the device user and the bystander both look away, the breakthrough of the bystander may persist for the specified time period (e.g., 5 seconds) prior to disabling the breakthrough. Likewise, if the bystander looks back at the device user within the time period window, the persistence timer may be configured to reset itself thereby maintaining breakthrough of the bystander(s). If no interaction resumes (i.e., the device user and the bystander do not look at each other within the time period), the bystander may fade out of the breakthrough and the hint state 704 is again triggered. If the bystander exits the physical environment, the background state is triggered.

[0098] Figure 8 is a block diagram of an example system illustrating a user device 100 for controlling virtual content presentation to display a representation of a person based on interpreting a communication characteristic of the person, in accordance with some implementations. The user device 800 (e.g., an HMD) may include a non-AI-basedAttorney Docket No. 097425-01472(P67210W01) interface 802, computer vision process(es) 804, a non- Al based interface 806, and a light weight engine 808.

[0099] In some embodiments, non-AI-based interface 802 is configured to accept input data (e.g., from sensors) associated with user characteristic(s) such as, inter alia, a gaze direction(s), mouth movements, audio from a microphone, body language, etc. of a person(s) in a physical environment and / or device user. For example, input data may include images from an outward-facing camera and / or inward facing camera, audio from a microphone, depth data from a depth sensor, etc.

[0100] In some embodiments, non-AI-based interface 802 obtains input data from various sources (e.g., sensors, cameras, microphones) on user device 800 indicating environmental conditions. Additional signal processing filters can combine the obtained data with additional contextual information derived from user configurations.

[0101] Inputs 810 from non-AI-based interface 802 can be fed to a computer vision process 804. A computer vision process 804 can include one or more learning-based and / or non-learning-based models for perceiving, synthesizing, and inferring information. Persons skilled in the art will appreciate that the computer vision process 804 can include any suitable number of computer vision processes to generate output 812 based on input 810.

[0102] In some embodiments, the computer vision process 804 may be used in combination with, for example, spatial techniques to predict attention by utilizing facial landmarks and / or facial expressions by detecting and localizing specific points or landmarks and expressions on a face, such as, inter alia, eyes, nose, mouth, chin, etc.

[0103] In some embodiments, user device 800 may optionally elect to use a light weight engine 808 instead of computer vision process 804 to process data in situations where only one modality of data is being processed. The light weight engine 808 may be a non-learning network.

[0104] In some embodiments, non-AI-based interface 806 is configured to accept results data from computer vision process 804 and / or a light weight engine 808 to provide instructions or recommendations for updating a view of an XR environment in which at least a portion of virtual content is replaced with a depiction of the person.Attorney Docket No. 097425-01472(P67210W01)

[0105] For example, output from computer vision process 804 and / or a light weight engine 808 is provided to non- Al based interface 806, where interface 806 disambiguates the instructions or recommendations. Based on current context and default settings provided by a user, interface 806 may select a subset of recommendations to surface to the user. For example, computer vision process 804 and / or a light weight engine 808 may be configured to determine an attention and / or intention of a user and interface 806 disambiguates the instructions or recommendations to update the view of the XR environment.

[0106] Persons of ordinary skill in the art will appreciate that computer vision process 804 can include any suitable machine learning models that are well-known or widely available such as regression techniques, classification techniques, neural networks, and deep learning networks. For instance, computer vision process 804 can include neural networks such as Artificial Neural Network (ANN), Convolutional Neural Network (CNN), Recurrent Neural Network (RNN), Adversarial Network (GAN), Reinforcement Learning Model (RLM), Encoder / Decoder Networks, and / or Transformer-Based Models (e.g., Bidirectional Encoder Representations from Transformers (BERT), Generative Pre-trained Transformer (GPT), and / or a multi-modal large language model (LLM)). Additionally or alternatively, persons of ordinary skill in the art will appreciate that computer vision process 804 can be any suitable non-learning processes such as rulebased systems, heuristics, decision trees, knowledge-based systems, statistical or stochastic systems, and expert systems.

[0107] In instances where computer vision process 804 is a machine- learning based model, computer vision process 804 can be trained to determine an attention or intention of person using one or more well-known or widely available training techniques such as supervised learning, semi-supervised learning, unsupervised learning, and / or reinforcement learning techniques. The training data can include historical e-mails.

[0108] In some embodiments, computer vision process 804 can be trained with both synthetic and non-synthetic data.

[0109] In some embodiments, computer vision process 804 can be deployed as one or more generative models, where content is automatically generated by one or moreAttorney Docket No. 097425-01472(P67210W01) computers in response to a request to generate the content. The automatically-generated content is optionally generated on-device (e.g., generated at least in part by a computer system at which a request to generate the content is received) and / or generated off-device (e.g., generated at least in part by one or more nearby computers that are available via a local network or one or more computers that are available via the internet). This automatically-generated content optionally includes visual content (e.g., images, graphics, and / or video), audio content, and / or text content.

[0110] In some embodiments, novel automatically-generated content is referred to as generative content (e.g., generative images, generative graphics, generative video, generative audio, and / or generative text). Generative content is typically generated based on a prompt input 810 to the computer vision process 804. Non-AI-based interface 802 optionally includes one or more pre-processing steps to adjust the input before it is used by an Al model to generate an output (e.g., adjustment to a user-provided prompt, creation of a system-generated prompt, and / or Al model selection). Non-AI-based interface 806 optionally includes one or more post-processing steps to adjust the output 812 by the computer vision process 804 (e.g., passing the Al model output to a different Al model, upscaling, downscaling, cropping, formatting, and / or adding or removing metadata) before the output 812 of the computer vision process 804 used for other purposes such as being provided to a different software process for further processing or being presented (e.g., visually or audibly) to a user.

[0111] A prompt input 810 for generating generative content can include one or more of: one or more words (e.g., a natural language prompt that is written or spoken), one or more images, one or more drawings, and / or one or more videos. Generative pre-trained transformer models are a type of LLM that can be effective at generating novel generative content based on a prompt input 810. In some embodiments, the computer vision process 804 uses a prompt input 810 that includes text to generate either different generative text, generative audio content, and / or generative visual content. In other embodiments, the computer vision process 804 uses a prompt input 810 that includes visual content and / or an audio content to generate generative text (e.g., a transcription of audio and / or a description of the visual content). In yet other embodiments, the computer vision process 804 uses a prompt input 810 that includes multiple types of content (e.g., text, images,Attorney Docket No. 097425-01472(P67210W01) audio, video, and / or other sensor data) to generate generative content. A prompt input 810 sometimes also includes values for one or more parameters indicating an importance of various parts of the prompt. Some prompt inputs 810 include a structured set of instructions that can be provided to the computer vision process 804 that include phrasing, a specified style, relevant context (e.g., starting point content and / or one or more examples), and / or a role for the computer vision process 804.

[0112] Generative content is generally based on the prompt but is not deterministically selected from pre-generated content and is, instead, generated using the prompt as a starting point. In some embodiments, pre-existing content (e.g., audio, text, and / or visual content) is used as part of the prompt for creating generative content (e.g., the pre-existing content is used as a starting point for creating the generative content). For example, a prompt input 810 could request that a block of text be summarized or rewritten in a different tone, and the output 812 would be generative text that is summarized or written in the different tone. Similarly a prompt input 810 could request that visual content be modified to include or exclude content specified by a prompt (e.g., removing an identified feature in the visual content, adding a feature to the visual content that is described in a prompt, changing a visual style of the visual content, and / or creating additional visual elements outside of a spatial or temporal boundary of the visual content that are based on the visual content). In some embodiments, a random or pseudo-random seed is used as part of the prompt input 810 for creating generative content (e.g., the random or pseudo-random seed content is used as a starting point for creating the generative content). For example, when generating an image from a diffusion model, a random noise pattern is iteratively denoised based on the prompt input 810 to generate an image. While specific types of computer vision processes 804 have been described herein, it should be understood that a variety of different computer vision processes could be used to generate generative content based on a prompt.

[0113] In instances where computer vision process 804 is a non-learning-based system, computer vision process 804 can use a pre-defined set of rules or a pre-defined structure to make decisions based on the inputs that the process sees. For example, the computer vision process 804 can be used to determine that an eye / pupil position satisfies a specified positional criteria to determine that a person is looking towards the deviceAttorney Docket No. 097425-01472(P67210W01) user thereby determining attention or intention of a person in a physical environment to identify circumstances in which breakthrough is appropriate. Likewise, a computer vision algorithm may be used to detect a bounding box (defining the eye region of the person) size to decode user attention and / or intention of a person in a physical environment to identify circumstances in which breakthrough is appropriate.

[0114] Some embodiments described herein can include use of learning and / or nonlearning-based process(es). The use can include collecting, pre-processing, encoding, labeling, organizing, analyzing, recommending and / or generating data. Entities that collect, share, and / or otherwise utilize user data should provide transparency and / or obtain user consent when collecting such data. The present disclosure recognizes that the use of the data in the computer vision processes can be used to benefit users. For example, the data can be used to train models that can be deployed to improve performance, accuracy, and / or functionality of applications and / or services. Accordingly, the use of the data enables the computer vision processes to adapt and / or optimize operations to provide more personalized, efficient, and / or enhanced user experiences. Such adaptation and / or optimization can include tailoring content, recommendations, and / or interactions to individual users, as well as streamlining processes, and / or enabling more intuitive interfaces. Further beneficial uses of the data in the computer vision processes are also contemplated by the present disclosure.

[0115] The present disclosure contemplates that, in some embodiments, data used by computer vision processes includes publicly available data. To protect user privacy, data may be anonymized, aggregated, and / or otherwise processed to remove or to the degree possible limit any individual identification. As discussed herein, entities that collect, share, and / or otherwise utilize such data should obtain user consent prior to and / or provide transparency when collecting such data. Furthermore, the present disclosure contemplates that the entities responsible for the use of data, including, but not limited to data used in association with computer vision processes, should attempt to comply with well-established privacy policies and / or privacy practices.

[0116] For example, such entities may implement and consistently follow policies and practices recognized as meeting or exceeding industry standards and regulatory requirements for developing and / or training computer vision processes. In doing so,Attorney Docket No. 097425-01472(P67210W01) attempts should be made to ensure all intellectual property rights and privacy considerations are maintained. Training should include practices safeguarding training data, such as personal information, through sufficient protections against misuse or exploitation. Such policies and practices should cover all stages of the computer vision processes development, training, and use, including data collection, data preparation, model training, model evaluation, model deployment, and ongoing monitoring and maintenance. Transparency and accountability should be maintained throughout. Such policies should be easily accessible by users and should be updated as the collection and / or use of data changes. User data should be collected for legitimate and reasonable uses of the entity and not shared or sold outside of those legitimate uses. Further, such collection and sharing should occur through transparency with users and / or after receiving the informed consent of the users. Additionally, such entities should consider taking any needed steps for safeguarding and securing access to such data and ensuring that others with access to the data adhere to their privacy policies and procedures. Further, such entities should subject themselves to evaluation by third parties to certify, as appropriate for transparency purposes, their adherence to widely accepted privacy policies and practices. In addition, policies and / or practices should be adapted to the particular type of data being collected and / or accessed and tailored to a specific use case and applicable laws and standards, including jurisdiction-specific considerations.

[0117] In some embodiments, computer vision processes may utilize models that may be trained (e.g., supervised learning or unsupervised learning) using various training data, including data collected using a user device. Such use of user-collected data may be limited to operations on the user device. For example, the training of the model can be done locally on the user device so no part of the data is sent to another device. In other implementations, the training of the model can be performed using one or more other devices (e.g., server(s)) in addition to the user device but done in a privacy preserving manner, e.g., via multi-party computation as may be done cryptographically by secret sharing data or other means so that the user data is not leaked to the other devices.

[0118] In some embodiments, the trained model can be centrally stored on the user device or stored on multiple devices, e.g., as in federated learning. Such decentralized storage can similarly be done in a privacy preserving manner, e.g., via cryptographicAttorney Docket No. 097425-01472(P67210W01) operations where each piece of data is broken into shards such that no device alone (i.e., only collectively with another device(s)) or only the user device can reassemble or use the data. In this manner, a pattern of behavior of the user or the device may not be leaked, while taking advantage of increased computational resources of the other devices to train and execute the ML model. Accordingly, user-collected data can be protected. In some implementations, data from multiple devices can be combined in a privacy-preserving manner to train an ML model.

[0119] In some embodiments, the present disclosure contemplates that data used for computer vision processes may be kept strictly separated from platforms where the computer vision processes are deployed and / or used to interact with users and / or process data. In such embodiments, data used for offline training of the computer vision processes may be maintained in secured datastores with restricted access and / or not be retained beyond the duration necessary for training purposes. In some embodiments, the computer vision processes may utilize a local memory cache to store data temporarily during a user session. The local memory cache may be used to improve performance of the computer vision processes. However, to protect user privacy, data stored in the local memory cache may be erased after the user session is completed. Any temporary caches of data used for online learning or inference may be promptly erased after processing. All data collection, transfer, and / or storage should use industry- standard encryption and / or secure communication.

[0120] In some embodiments, as noted above, techniques such as federated learning, differential privacy, secure hardware components, homomorphic encryption, and / or multi-party computation among other techniques may be utilized to further protect personal information data during training and / or use of the computer vision processes. The computer vision processes should be monitored for changes in underlying data distribution such as concept drift or data skew that can degrade performance of the computer vision processes over time.

[0121] In some embodiments, the computer vision processes are trained using a combination of offline and online training. Offline training can use curated datasets to establish baseline model performance, while online training can allow the computer vision processes to continually adapt and / or improve. The present disclosure recognizesAttorney Docket No. 097425-01472(P67210W01) the importance of maintaining strict data governance practices throughout this process to ensure user privacy is protected.

[0122] In some embodiments, the computer vision processes may be designed with safeguards to maintain adherence to originally intended purposes, even as the computer vision processes adapt based on new data. Any significant changes in data collection and / or applications of computer vision process use may (and in some cases should) be transparently communicated to affected stakeholders and / or include obtaining user consent with respect to changes in how user data is collected and / or utilized.

[0123] Despite the foregoing, the present disclosure also contemplates embodiments in which users selectively restrict and / or block the use of and / or access to data. That is, the present disclosure contemplates that hardware and / or software elements can be provided to prevent or block access to data. For example, in the case of some services, the present technology should be configured to allow users to select to “opt in” or “opt out” of participation in the collection of data during registration for services or anytime thereafter. In another example, the present technology should be configured to allow users to select not to provide certain data for training the computer vision processes and / or for use as input during the inference stage of such systems. In yet another example, the present technology should be configured to allow users to be able to select to limit the length of time data is maintained or entirely prohibit the use of their data for use by the computer vision processes. In addition to providing “opt in” and “opt out” options, the present disclosure contemplates providing notifications relating to the access or use of personal information. For instance, a user can be notified when their data is being input into the computer vision processes for training or inference purposes, and / or reminded when the computer vision processes generate outputs or make decisions based on their data.

[0124] The present disclosure recognizes computer vision processes should incorporate explicit restrictions and / or oversight to mitigate against risks that may be present even when such systems having been designed, developed, and / or operated according to industry best practices and standards. For example, outputs may be produced that could be considered erroneous, harmful, offensive, and / or biased; such outputs may not necessarily reflect the opinions or positions of the entities developing or deployingAttorney Docket No. 097425-01472(P67210W01) these systems. Furthermore, in some cases, references to or failures to cite third-party products and / or services in the outputs should not be construed as endorsements or affiliations by the entities providing the computer vision processes. Generated content can be filtered for potentially inappropriate or dangerous material prior to being presented to users, while human oversight and / or ability to override or correct erroneous or undesirable outputs can be maintained as a failsafe.

[0125] The present disclosure further contemplates that users of the computer vision processes should refrain from using the services in any manner that infringes upon, misappropriates, or violates the rights of any party. Furthermore, the computer vision processes should not be used for any unlawful or illegal activity, nor to develop any application or use case that would commit or facilitate the commission of a crime, or other tortious, unlawful, or illegal act including misinformation, disinformation, misrepresentations (e.g., deepfakes), deception, impersonation, and propaganda. The computer vision processes should not violate, misappropriate, or infringe any copyrights, trademarks, rights of privacy and publicity, trade secrets, patents, or other proprietary or legal rights of any party, and appropriately attribute content as required. Further, the computer vision processes should not interfere with any security, digital signing, digital rights management, content protection, verification, or authentication mechanisms. The computer vision processes should not misrepresent machine-generated outputs as being human-generated.

[0126] Figure 9 is a flowchart representation of an exemplary method 900 that controls virtual content presentation to display a representation of a person based on interpreting a communication characteristic of the person, in accordance with some implementations. In some implementations, the method 900 is performed by a device, such as a mobile device, desktop, laptop, HMD, or server device. In some implementations, the device has a screen for displaying images and / or a screen for viewing stereoscopic images such as a head-mounted display (HMD such as e.g., device 105 of Figure 1). In some implementations, the method 900 is performed by processing logic, including hardware, firmware, software, or a combination thereof. In some implementations, the method 900 is performed by a processor executing code stored in aAttorney Docket No. 097425-01472(P67210W01) non-transitory computer-readable medium (e g., a memory). Each of the blocks in the method 900 may be enabled and executed in any order.

[0127] At block 902, the method 900 presents a view of an XR environment to a first user via an HMD. The view comprises virtual content positioned within 3D space of the XR environment. In some implementations, a person other than the first user in the physical environment is not depicted in the view. For example, user 406 is not depicted in device user view 404 as described with respect to figure 4A.

[0128] At block 904, the method 900 obtains first sensor data from at least one sensor on the HMD. In some implementations, the first sensor data represents a characteristic of the person. For example, the first sensor data may represent a characteristic of a person 210 to identify conditions that may be favorable for performing a breakthrough of a person 210 with respect to presentation within an XR environment 200 as described with respect to figure 2.

[0129] In some implementations, the first sensor data may include images from an outward-facing camera or an inward facing camera.

[0130] In some implementations, the first sensor data may include audio data from a microphone.

[0131] In some implementations, the first sensor data may include depth data.

[0132] In some implementations, the characteristic may include a gaze direction of the person, a mouth movement of the person, body language of the person, facial landmarks, facial expressions and / or inferred emotional states, of the person, etc. The characteristic may be used to indicate that the person is communicating with the first user.

[0133] At block 906, the method 900 determines to replace (e.g., breakthrough) at least a portion of the virtual content of the view to provide a depiction of the person based on the characteristic. For example, depiction and appearance 422 of user 406 is presented to a device user 403 within representation of an XR environment 412 presented via an HMD as described with respect to figure 4B.Attorney Docket No. 097425-01472(P67210W01)

[0134] In some implementations, determining to replace at least a portion of the virtual content may include determining an attention or intention of the person.

[0135] In some implementations, determining the attention or intention of the person is based on determining that the person is looking at the first user and / or speaking to the first user.

[0136] In some implementations, determining the attention or intention of the person is based on determining that the person and the first user are looking at and / o speaking to each other thereby indicating joint attention.

[0137] In some implementations, determining the attention or intention of the person is based on determining that the person and the first user are looking at a same object such as, for example, a computer 528 and described with respect to figure 5.

[0138] In some implementations, determining the attention or intention of the person may performed using a gaze detection model.

[0139] In some implementations, determining the attention or intention of the person may performed using a custom machine learning model or a rule-based model / algorithm.

[0140] In some implementations, determining the attention or intention of the person may performed using a spatial technique.

[0141] In some implementations, determining the attention or intention of the person may performed using a body language analysis technique.

[0142] In some implementations, determining the attention or intention of the person may be based on determining a facial landmark or expression of the person and a facial landmark or expression of the first user. For example, determining the user attention and / or intention may be based on, for example, looking at eye and / or mouth movement, facial landmarks, etc. to determine facial expressions and / or inferred emotional states. For example, detected mouth movement may indicate that a person and / or the device user is smiling thereby indicating a positive emotional state (e.g., happy) which may be used to decode user attention and / or intention.

[0143] In some implementations, the first sensor data includes an image of an eye of the person and determining the attention or intention of the person may be based on:Attorney Docket No. 097425-01472(P67210W01) identifying a position of the eye in the image; identifying a position of a pupil within the eye; determining that the position of a pupil satisfies a specified positional criteria; and based on determining that the position of a pupil satisfies a specified positional criteria, determining that the person is looking towards the least one sensor on the HMD.

[0144] In some implementations, determining the attention or intention of the person may be performed by: detecting, at a first time, a first instance of a bounding box defining an eye region of the person; and detecting, at a second time occurring subsequent to the first time, a second instance of the bounding box defining the eye region of the person, wherein the second instance of the bounding box is larger than the first instance of the bounding box.

[0145] At block 908, the method 900 updates the view of the XR environment in which the at least a portion of the virtual content is replaced with the depiction of the person.

[0146] In some implementations, updating the view of the XR environment may continue during a specified time period while an attention or intention of the person or the first user has been interrupted.

[0147] In some implementations, the depiction of the person comprises an opaque depiction of the person.

[0148] In some implementations, the depiction of the person comprises a specified transparency level.

[0149] In some implementations, the depiction of the person comprises a specified tint level.

[0150] In some implementations, the depiction of the person comprises a specified focus level.

[0151] In some implementations, the depiction of the person initially comprises a partially transparent depiction of the person followed by an opaque depiction of the person subsequent to a specified time period.

[0152] In some implementations, it may be determined that the person and the first user have stopped looking at each other and based on determining that the person and theAttorney Docket No. 097425-01472(P67210W01) first user have stopped looking at each other. In response, the view of the XR environment may be further updated such that the at least a portion of the virtual content is replaced with the depiction of the person. In some implementations, further updating the view of the XR environment may be based on determining that the person and the first user have stopped looking at each other for a specified time period such as 5 seconds, etc.

[0153] Figure 10 is a flowchart representation of an exemplary method 1000 that controls virtual content presentation to display at least two people in an environment based on interpreting a communication characteristic of each of the people, in accordance with some implementations. In some implementations, the method 1000 is performed by a device, such as a mobile device, desktop, laptop, HMD, or server device. In some implementations, the device has a screen for displaying images and / or a screen for viewing stereoscopic images such as a head-mounted display (HMD such as e.g., device 105 of Figure 1). In some implementations, the method 1000 is performed by processing logic, including hardware, firmware, software, or a combination thereof. In some implementations, the method 1000 is performed by a processor executing code stored in a non-transitory computer-readable medium (e.g., a memory). Each of the blocks in the method 1000 may be enabled and executed in any order.

[0154] At block 1002, the method 1000 presents a view of an XR environment to a first user via an HMD. The view comprises virtual content positioned within 3D space of the XR environment. In some implementations, a first person and a second person other than the first user in the physical environment are not depicted in the view. For example, user 406 is not depicted in device user view 404 as described with respect to figure 4A.

[0155] At block 1004, the method 1000 obtains first sensor data from at least one sensor on the HMD. In some implementations, the first sensor data represents a first characteristic of the first person and a second characteristic of the second person. For example, the first sensor data may represent a characteristic of user 606 and user 607 to identify conditions that may be favorable for performing a breakthrough of a person user 606 and user 607 with respect to presentation within an XR environment 612 as described with respect to figure B.Attorney Docket No. 097425-01472(P67210W01)

[0156] In some implementations, the first sensor data may include images from an outward-facing camera or an inward facing camera.

[0157] In some implementations, the first sensor data may include audio data from a microphone.

[0158] In some implementations, the first sensor data may include depth data.

[0159] In some implementations, the characteristic may include a gaze direction of the users, a mouth movement of the users, body language of the users, facial landmarks, facial expressions to determine inferred emotional states of the users, etc. For example, detected mouth movement may indicate that a person and / or the device user is smiling thereby indicating a positive emotional state (e.g., happy) which may be used to decode user attention and / or intention.

[0160] The characteristic may be used to indicate that the users are communicating with the first user.

[0161] At block 1006, the method 1000 determines to replace (e.g., breakthrough) at least a first portion of the virtual content of the view to provide a depiction of the first person based on the characteristic. For example, depiction and appearance 622b of user 606 is presented to a device user 603 within representation of an XR environment 612 presented via an HMD as described with respect to figure 6B.

[0162] In some implementations, determining to replace at least a first portion of the virtual content may include determining an attention or intention of the first person based on determining that the first person is looking at the first user.

[0163] In some implementations, the first sensor data includes an image of an eye of the first person or the second person and determining the attention or intention of the first or second person may be based on: identifying a position of the eye in the image; identifying a position of a pupil within the eye; determining that the position of a pupil satisfies a specified positional criteria; and based on determining that the position of a pupil satisfies a specified positional criteria, determining that the first or second person is looking towards the least one sensor on the HMD.Attorney Docket No. 097425-01472(P67210W01)

[0164] In some implementations, determining the attention or intention of the first person or the second person may be performed by: detecting, at a first time, a first instance of a bounding box defining an eye region of the first person or the second person and detecting, at a second time occurring subsequent to the first time, a second instance of the bounding box defining the eye region of the first person or the second person. In some implementations, the second instance of the bounding box is larger than the first instance of the bounding box.

[0165] At block 1008, the method 1000 determines to replace (e.g., breakthrough) at least a second portion of the virtual content of the view to provide a depiction of the second person based on the second characteristic. For example, depiction and appearance 624b of user 607 is presented to a device user 603 within representation of an XR environment 612 presented via an HMD as described with respect to figure 6B.

[0166] In some implementations, determining to replace at least a second portion of the virtual content may include determining an attention or intention of the second person based on determining that the second person is looking at the first user.

[0167] In some implementations, the first user is only looking at the first person or the second person.

[0168] In some implementations, determining the attention or intention of the first person is further based on determining that the first person and the first user are looking at each other. Likewise, in some implementations determining the attention or intention of the second person is further based on determining that the second person and the first user are looking at each other.

[0169] In some implementations, it may be determined that the first person and the first user have stopped looking at each other. In response, the view of the XR environment may be further updated such that the depiction of the first person is replaced with at least a portion of the virtual content.

[0170] In some implementations, further updating the view of the XR environment may be further based on said determining that the first person and the first user have stopped looking at each other for a specified time period such as, for example, 5 seconds, etc.Attorney Docket No. 097425-01472(P67210W01)

[0171] In some implementations, determining the attention or intention of the first person and determining the attention or intention of the second person may be further based on determining that the first person, the second person and the first user are all looking at a same object such as, for example, a pencil on a desk, a computer, another person, etc.

[0172] At block 1010, the method 1000 updates the view of the XR environment in which at least a first portion of the virtual content is replaced with the depiction of the first person and at least a second portion of the virtual content is replaced with the depiction of the second person.

[0173] In some implementations, the depiction of the first person or the second person initially comprises a partially transparent depiction followed by an opaque depiction subsequent to a specified time period.

[0174] Figure 11 is a block diagram of an example device 1100. Device 1100 illustrates an exemplary device configuration for electronic devices 105 and 110 of Figure 1. While certain specific features are illustrated, those skilled in the art will appreciate from the present disclosure that various other features have not been illustrated for the sake of brevity, and so as not to obscure more pertinent aspects of the implementations disclosed herein. To that end, as a non-limiting example, in some implementations the device 1100 includes one or more processing units 1102 (e.g., microprocessors, ASICs, FPGAs, GPUs, CPUs, processing cores, and / or the like), one or more input / output (I / O) devices and sensors 1106, one or more communication interfaces 1108 (e.g., USB, FIREWIRE, THUNDERBOLT, IEEE 802.3x, IEEE 802.1 lx, IEEE 802.14x, GSM, CDMA, TDMA, GPS, IR, BLUETOOTH, ZIGBEE, SPI, I2C, and / or the like type interface), one or more programming (e.g., I / O) interfaces 1110, output devices (e.g., one or more displays) 1112, one or more interior and / or exterior facing image sensor systems 1114, a memory 1120, and one or more communication buses 1104 for interconnecting these and various other components.

[0175] In some implementations, the one or more communication buses 1104 include circuitry that interconnects and controls communications between system components. In some implementations, the one or more I / O devices and sensors 1106 include at least one of an inertial measurement unit (IMU), an accelerometer, a magnetometer, a gyroscope, aAttorney Docket No. 097425-01472(P67210W01) thermometer, one or more physiological sensors (e.g., blood pressure monitor, heart rate monitor, blood oxygen sensor, blood glucose sensor, etc.), one or more microphones, one or more speakers, a haptics engine, one or more depth sensors (e.g., a structured light, a time-of-flight, or the like), one or more IR cameras (e.g., inward facing cameras and outward facing cameras of an HMD), one or more infrared sensors, one or more heat map sensors, and / or the like.

[0176] In some implementations, the one or more displays 1112 are configured to present a view of a physical environment, a graphical environment, an extended reality environment, etc. to the user. In some implementations, the one or more displays 1112 are configured to present content (determined based on a determined user / object location of the user within the physical environment) to the user. In some implementations, the one or more displays 1112 correspond to holographic, digital light processing (DLP), liquid-crystal display (LCD), liquid-crystal on silicon (LCoS), organic light-emitting field-effect transitory (OLET), organic light-emitting diode (OLED), surface-conduction electron-emitter display (SED), field-emission display (FED), quantum-dot lightemitting diode (QD-LED), micro-electromechanical system (MEMS), and / or the like display types. In some implementations, the one or more displays 1112 correspond to diffractive, reflective, polarized, holographic, etc. waveguide displays. In one example, the device 1100 includes a single display. In another example, the device 1100 includes a display for each eye of the user.

[0177] In some implementations, the one or more image sensor systems 1114 are configured to obtain image data that corresponds to at least a portion of the physical environment 100. For example, the one or more image sensor systems 1114 include one or more RGB or IR cameras (e.g., with a complimentary metal-oxide-semiconductor (CMOS) image sensor or a charge-coupled device (CCD) image sensor), monochrome cameras, IR cameras, depth cameras, event-based cameras, and / or the like. In various implementations, the one or more image sensor systems 1114 further include illumination sources that emit light, such as a flash. In various implementations, the one or more image sensor systems 1114 further include an on-camera image signal processor (ISP) configured to execute a plurality of processing operations on the image data.Attorney Docket No. 097425-01472(P67210W01)

[0178] In some implementations, sensor data may be obtained by device(s) (e.g., devices 105 and 110 of Figure 1) during a scan of a room of a physical environment. The sensor data may include a 3D point cloud and a sequence of 2D images corresponding to captured views of the room during the scan of the room. In some implementations, the sensor data includes image data (e.g., from an RGB or IR camera), depth data (e.g., a depth image from a depth camera), ambient light sensor data (e.g., from an ambient light sensor), and / or motion data from one or more motion sensors (e.g., accelerometers, gyroscopes, IMU, etc.). In some implementations, the sensor data includes visual inertial odometry (VIO) data determined based on image data. The 3D point cloud may provide semantic information about one or more elements of the room. The 3D point cloud may provide information about the positions and appearance of surface portions within the physical environment. In some implementations, the 3D point cloud is obtained over time, e.g., during a scan of the room, and the 3D point cloud may be updated, and updated versions of the 3D point cloud obtained over time. For example, a 3D representation may be obtained (and analyzed / processed) as it is updated / adjusted over time (e.g., as the user scans a room).

[0179] In some implementations, sensor data may be positioning information, some implementations include a VIO to determine equivalent odometry information using sequential camera images (e.g., light intensity image data) and motion data (e.g., acquired from the IMU / motion sensor) to estimate the distance traveled. Alternatively, some implementations of the present disclosure may include a simultaneous localization and mapping (SLAM) system (e.g., position sensors). The SLAM system may include a multidimensional (e.g., 3D) laser scanning and range-measuring system that is GPS independent and that provides real-time simultaneous location and mapping. The SLAM system may generate and manage data for a very accurate point cloud that results from reflections of laser scanning from objects in an environment. Movements of any of the points in the point cloud are accurately tracked over time, so that the SLAM system can maintain precise understanding of its location and orientation as it travels through an environment, using the points in the point cloud as reference points for the location.

[0180] In some implementations, the device 1100 includes an eye tracking system for detecting eye position and eye movements (e.g., eye gaze detection). For example, an eyeAttorney Docket No. 097425-01472(P67210W01) tracking system may include one or more infrared (IR) light-emitting diodes (LEDs), an eye tracking camera (e.g., near-IR (NIR) camera), and an illumination source (e.g., an NIR light source) that emits light (e.g., NIR light) towards the eyes of the user. Moreover, the illumination source of the device 1100 may emit NIR light to illuminate the eyes of the user and the NIR camera may capture images of the eyes of the user. In some implementations, images captured by the eye tracking system may be analyzed to detect position and movements of the eyes of the user, or to detect other information about the eyes such as pupil dilation or pupil diameter. Moreover, the point of gaze estimated from the eye tracking images may enable gaze-based interaction with content shown on the near-eye display of the device 1100.

[0181] The memory 1120 includes high-speed random-access memory, such as DRAM, SRAM, DDR RAM, or other random-access solid-state memory devices. In some implementations, the memory 1120 includes non-volatile memory, such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. The memory 1120 optionally includes one or more storage devices remotely located from the one or more processing units 1102. The memory 1120 includes a non- transitory computer readable storage medium.

[0182] In some implementations, the memory 1120 or the non-transitory computer readable storage medium of the memory 1120 stores an optional operating system 1130 and one or more instruction set(s) 1140. The operating system 1130 includes procedures for handling various basic system services and for performing hardware dependent tasks. In some implementations, the instruction set(s) 1140 include executable software defined by binary information stored in the form of electrical charge. In some implementations, the instruction set(s) 1140 are software that is executable by the one or more processing units 1102 to carry out one or more of the techniques described herein.

[0183] The instruction set(s) 1140 includes a virtual content instruction set 1142 and an updated view instruction set 1144. The instruction set(s) 1140 may be embodied as a single software executable or multiple software executables.Attorney Docket No. 097425-01472(P67210W01)

[0184] The virtual content instruction set 1142 is configured with instructions executable by a processor to presenting a view of an XR environment such that virtual content is positioned within a 3D space of the XR environment.

[0185] The updated view instruction set 1144 is configured with instructions executable by a processor to update the view of the XR environment such that at least a portion of the virtual content is replaced with a depiction of a person(s) within a physical environment.

[0186] Although the instruction set(s) 1140 are shown as residing on a single device, it should be understood that in other implementations, any combination of the elements may be located in separate computing devices. Moreover, Figure 11 is intended more as functional description of the various features which are present in a particular implementation as opposed to a structural schematic of the implementations described herein. As recognized by those of ordinary skill in the art, items shown separately could be combined and some items could be separated. The actual number of instructions sets and how features are allocated among them may vary from one implementation to another and may depend in part on the particular combination of hardware, software, and / or firmware chosen for a particular implementation.

[0187] Those of ordinary skill in the art will appreciate that well-known systems, methods, components, devices, and circuits have not been described in exhaustive detail so as not to obscure more pertinent aspects of the example implementations described herein. Moreover, other effective aspects and / or variants do not include all of the specific details described herein. Thus, several details are described in order to provide a thorough understanding of the example aspects as shown in the drawings. Moreover, the drawings merely show some example embodiments of the present disclosure and are therefore not to be considered limiting.

[0188] While this specification contains many specific implementation details, these should not be construed as limitations on the scope of any inventions or of what may be claimed, but rather as descriptions of features specific to particular embodiments of particular inventions. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a singleAttorney Docket No. 097425-01472(P67210W01) embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Moreover, although features may be described above as acting in certain combinations and even initially claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a subcombination or variation of a subcombination.

[0189] Similarly, while operations are depicted in the drawings in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Moreover, the separation of various system components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

[0190] Thus, particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. In some cases, the actions recited in the claims can be performed in a different order and still achieve desirable results. In addition, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve desirable results. In certain implementations, multitasking and parallel processing may be advantageous.

[0191] Embodiments of the subject matter and the operations described in this specification can be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, e.g., one or more modules of computer program instructions, encoded on computer storage medium for execution by, or to control the operation of, data processing apparatus. Alternatively, or additionally, the program instructions can be encoded on an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information forAttorney Docket No. 097425-01472(P67210W01) transmission to suitable receiver apparatus for execution by a data processing apparatus. A computer storage medium can be, or be included in, a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination of one or more of them. Moreover, while a computer storage medium is not a propagated signal, a computer storage medium can be a source or destination of computer program instructions encoded in an artificially generated propagated signal. The computer storage medium can also be, or be included in, one or more separate physical components or media (e.g., multiple CDs, disks, or other storage devices).

[0192] The term “data processing apparatus” encompasses all kinds of apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, a system on a chip, or multiple ones, or combinations, of the foregoing. The apparatus can include special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit). The apparatus can also include, in addition to hardware, code that creates an execution environment for the computer program in question, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, a crossplatform runtime environment, a virtual machine, or a combination of one or more of them. The apparatus and execution environment can realize various different computing model infrastructures, such as web services, distributed computing and grid computing infrastructures. Unless specifically stated otherwise, it is appreciated that throughout this specification discussions utilizing the terms such as “processing,” “computing,” “calculating,” “determining,” and “identifying” or the like refer to actions or processes of a computing device, such as one or more computers or a similar electronic computing device or devices, that manipulate or transform data represented as physical electronic or magnetic quantities within memories, registers, or other information storage devices, transmission devices, or display devices of the computing platform.

[0193] The system or systems discussed herein are not limited to any particular hardware architecture or configuration. A computing device can include any suitable arrangement of components that provides a result conditioned on one or more inputs. Suitable computing devices include multipurpose microprocessor-based computerAttorney Docket No. 097425-01472(P67210W01) systems accessing stored software that programs or configures the computing system from a general purpose computing apparatus to a specialized computing apparatus implementing one or more implementations of the present subject matter. Any suitable programming, scripting, or other type of language or combinations of languages may be used to implement the teachings contained herein in software to be used in programming or configuring a computing device.

[0194] Implementations of the methods disclosed herein may be performed in the operation of such computing devices. The order of the blocks presented in the examples above can be varied for example, blocks can be re-ordered, combined, and / or broken into sub-blocks. Certain blocks or processes can be performed in parallel. The operations described in this specification can be implemented as operations performed by a data processing apparatus on data stored on one or more computer-readable storage devices or received from other sources.

[0195] The use of “adapted to” or “configured to” herein is meant as open and inclusive language that does not foreclose devices adapted to or configured to perform additional tasks or steps. Additionally, the use of “based on” is meant to be open and inclusive, in that a process, step, calculation, or other action “based on” one or more recited conditions or values may, in practice, be based on additional conditions or value beyond those recited. Headings, lists, and numbering included herein are for ease of explanation only and are not meant to be limiting.

[0196] It will also be understood that, although the terms “first,” “second,” etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first node could be termed a second node, and, similarly, a second node could be termed a first node, which changing the meaning of the description, so long as all occurrences of the “first node” are renamed consistently and all occurrences of the “second node” are renamed consistently. The first node and the second node are both nodes, but they are not the same node.

[0197] The terminology used herein is for the purpose of describing particular implementations only and is not intended to be limiting of the claims. As used in theAttorney Docket No. 097425-01472(P67210W01) description of the implementations and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the term “and / or” as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. It will be further understood that the terms “comprises” and / or “comprising,” when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0198] As used herein, the term “if’ may be construed to mean “when” or “upon” or “in response to determining” or “in accordance with a determination” or “in response to detecting,” that a stated condition precedent is true, depending on the context. Similarly, the phrase “if it is determined [that a stated condition precedent is true]” or “if [a stated condition precedent is true]” or “when [a stated condition precedent is true]” may be construed to mean “upon determining” or “in response to determining” or “in accordance with a determination” or “upon detecting” or “in response to detecting” that the stated condition precedent is true, depending on the context.

Claims

Attorney Docket No. 097425-01472(P67210W01)What is claimed is:

1. A method comprising: at a processor of a head mounted device (HMD) being worn by a first user in a physical environment: presenting a view of an extended-reality (XR) environment via the HMD, the view comprising virtual content positioned within a three-dimensional (3D) space of the XR environment, wherein a person other than the first user in the physical environment is not depicted in the view; obtaining first sensor data from at least one sensor on the HMD, the first sensor data representing a characteristic of the person; determining to replace at least a portion of the virtual content of the view to provide a depiction of the person based on the characteristic, wherein said determining to replace at least a portion of the virtual content comprises determining an attention or intention of the person based on determining that the person is looking at the first user; and updating the view of the XR environment in which the at least a portion of the virtual content is replaced with the depiction of the person.

2. The method of claim 1, wherein said determining the attention or intention of the person is further based on determining that the person is speaking to the first user.

3. The method of any of claims 1-2, wherein said determining the attention or intention of the person is further based on determining that the person and the first user are looking at each other.

4. The method of claim 3, further comprising: determining that the person and the first user have stopped looking at each other; and based on said determining that the person and the first user have stopped looking at each other, further updating the view of the XR environment in which the depiction of the person is replaced with at least a portion of the virtual content.Attorney Docket No. 097425-01472(P67210W01)5. The method of claim 4, wherein said further updating the view of the XR environment is further based on said determining that the person and the first user have stopped looking at each other for a specified time period, (e.g., 5 seconds)6. The method of any of claims 1-5, wherein said determining the attention or intention of the person is further based on determining that the person and the first user are speaking to each other.

7. The method of any of claims 1-6, wherein said determining the attention or intention of the person is further based on determining that the person and the first user are looking at a same object.

8. The method of any of claims 1-7, wherein said determining the attention or intention of the person is performed using a computer vision model that predicts gaze direction of the person.

9. The method of any of claims 1-7, wherein said determining the attention or intention of the person is performed using a custom model.

10. The method of any of claims 1-7, wherein said determining the attention or intention of the person is performed using a spatial technique.

11. The method of any of claims 1-7, wherein said determining the attention or intention of the person is performed using a body language analysis technique.

12. The method of any of claims 1-7, wherein the first sensor data comprises an image of an eye of the person, and wherein said determining the attention or intention of the person comprises: identifying a position of the eye in the image; identifying a position of a pupil within the eye; determining that the position of a pupil satisfies a specified positional criteria; and based on said determining that the position of a pupil satisfies a specified positional criteria, determining that the user is looking towards the least one sensor on the HMD.Attorney Docket No. 097425-01472(P67210W01)13. The method of any of claims 1-7, wherein said determining the attention or intention of the person is performed by: detecting, at a first time, a first instance of a bounding box defining an eye region of the person; and detecting, at a second time occurring subsequent to the first time, a second instance of the bounding box defining the eye region of the person, wherein the second instance of the bounding box is larger than the first instance of a bounding box.

14. The method of any of claims 1-13, wherein the depiction of the person comprises an opaque depiction of the person.

15. The method of any of claims 1-13, wherein the depiction of the person initially comprises a partially transparent depiction of the person followed by an opaque depiction of the person subsequent to a specified time period.

16. The method of any of claims 1-15, wherein the depiction of the person comprises a specified tint level.

17. The method of any of claims 1-16. wherein the depiction of the person comprises a specified focus level.

18. The method of any of claims 1-17, wherein the first sensor data comprises images from an outward-facing camera or an inward facing camera.

19. The method of any of claims 1-18, wherein the first sensor data comprises audio data from a microphone.

20. The method of any of claims 1-19. wherein the first sensor data comprises depth data.

21. The method of any of claims 1-20, wherein the characteristic comprises a gaze direction of the person.Attorney Docket No. 097425-01472(P67210W01)22. The method of any of claims 1-21. wherein the characteristic comprises mouth movement of the person.

23. The method of any of claims 1-22, wherein the characteristic comprises body language of the person.

24. The method of any of claims 1-23, wherein the characteristic comprises a facial landmark or facial expression of the person.

25. The method of any of claims 1-24. further comprising: enabling said updating to continue during a specified time period while an attention or intention of the person or the first user has been interrupted.

26. The method of any of claims 1-25. wherein said determining the attention or intention of the person is based on determining a facial landmark of the person and a facial landmark of the first user.

27. The method of any of claims 1-26, wherein said determining the attention or intention of the person is based on determining a facial expression indicating an inferred emotional state of the person and a facial expression indicating an inferred emotional state of the first user.

28. A system comprising: a processor; a computer readable medium storing instructions that when executed by the processor cause the processor to perform operations comprising any one of the methods of claims 1-27.

29. A non-transitory computer-readable medium comprising instructions that when executed by a processor cause the processor to perform operations comprising any one of the methods of claims 1-27.Attorney Docket No. 097425-01472(P67210W01)30. A method comprising: at a processor of a head mounted device (HMD) being worn by a first user in a physical environment: presenting a view of an extended-reality (XR) environment via the HMD, the view comprising virtual content positioned within a three-dimensional (3D) space of the XR environment, wherein a person other than the first user in the physical environment is not depicted in the view; obtaining first sensor data from at least one sensor on the HMD, the first sensor data representing a characteristic of the person; determining to replace at least a portion of the virtual content of the view to provide a depiction of the person based on the characteristic, wherein said determining to replace at least a portion of the virtual content comprises determining an attention or intention of the person based on determining that the person is speaking to the first user; and updating the view of the XR environment in which the at least a portion of the virtual content is replaced with the depiction of the person.

31. A method comprising: at a processor of a head mounted device (HMD) being worn by a first user in a physical environment: presenting a view of an extended-reality (XR) environment via the HMD, the view comprising virtual content positioned within a three-dimensional (3D) space of the XR environment, wherein a person other than the first user in the physical environment is not depicted in the view; obtaining first sensor data from at least one sensor on the HMD, the first sensor data representing a characteristic of the person; determining to replace at least a portion of the virtual content of the view to provide a depiction of the person based on the characteristic, wherein said determining to replace at least a portion of the virtual content comprises determining an attention or intention of the person based on a body language of the person; and updating the view of the XR environment in which the at least a portion of the virtual content is replaced with the depiction of the person.Attorney Docket No. 097425-01472(P67210W01)32. A method comprising: at a processor of a head mounted device (HMD) being worn by a first user in a physical environment: presenting a view of an extended-reality (XR) environment via the HMD, the view comprising virtual content positioned within a three-dimensional (3D) space of the XR environment, wherein a person other than the first user in the physical environment is not depicted in the view; obtaining first sensor data from at least one sensor on the HMD, the first sensor data representing a characteristic of the person; determining to replace at least a portion of the virtual content of the view to provide a depiction of the person based on the characteristic, wherein said determining to replace at least a portion of the virtual content comprises determining a joint attention or intention of the person and the first user; and updating the view of the XR environment in which the at least a portion of the virtual content is replaced with the depiction of the person.

33. The method of claim 32, wherein said determining the attention or intention of the person is based on determining that the person and the first user are looking at each other (e.g., joint attention).

34. The method of any of claims 32-33, wherein said determining the attention or intention of the person is based on determining that the person and the first user are speaking to each other.

35. The method of any of claims 32-34, wherein said determining the attention or intention of the person is based on determining that the person and the first user are looking at a same object.Attorney Docket No. 097425-01472(P67210W01)36. A method comprising: at a processor of a head mounted device (HMD) being worn by a first user in a physical environment: presenting a view of an extended-reality (XR) environment via the HMD, the view comprising virtual content positioned within a three-dimensional (3D) space of the XR environment, wherein a first person and a second person other than the first user in the physical environment are not depicted in the view; obtaining first sensor data from at least one sensor on the HMD, the first sensor data representing a first characteristic of the first person and a second characteristic of the second person; determining to replace at least a first portion of the virtual content of the view to provide a depiction of the first person based on the first characteristic, wherein said determining to replace the at least a first portion of the virtual content comprises determining an attention or intention of the first person based on determining that the first person is looking at the first user; determining to replace at least a second portion of the virtual content of the view to provide a depiction of the second person based on the second characteristic, wherein said determining to replace the at least a second portion of the virtual content comprises determining an attention or intention of the second person based on determining that the second person is looking at the first user; and updating the view of the XR environment in which the at least a first portion of the virtual content is replaced with the depiction of the first person and the at least a second portion of the virtual content is replaced with the depiction of the second person.

37. The method of claim 36, wherein the first user is only looking at the first person or the second person.

38. The method of any of claims 36-37, wherein said determining the attention or intention of the first person is further based on determining that the first person and the first user are looking at each other, and wherein said determining the attention or intention of the second person is further based on determining that the second person and the first user are looking at each other.Attorney Docket No. 097425-01472(P67210W01)39. The method of claim 38, further comprising: determining that the first person and the first user have stopped looking at each other; and based on said determining that the first person and the first user have stopped looking at each other, further updating the view of the XR environment in which the depiction of the first person is replaced with at least a portion of the virtual content.

40. The method of claim 39, wherein said further updating the view of the XR environment is further based on said determining that the first person and the first user have stopped looking at each other for a specified time period, (e g., 5 seconds)41. The method of claim 36, wherein said determining the attention or intention of the first person and said determining the attention or intention of the second person is further based on determining that the first person, the second person and the first user are looking at a same object.

42. The method of any of claims 36-41, wherein the first sensor data comprises an image of an eye of the first person or the second person, and wherein said determining the attention or intention of the first person or the second person comprises: identifying a position of the eye in the image; identifying a position of a pupil within the eye; determining that the position of a pupil satisfies a specified positional criteria; and based on said determining that the position of a pupil satisfies a specified positional criteria, determining that the first person or the second person is looking towards the least one sensor on the HMD.Attorney Docket No. 097425-01472(P67210W01)43. The method of any of claims 36-41, wherein said determining the attention or intention of the first person or the second person is performed by: detecting, at a first time, a first instance of a bounding box defining an eye region of the first person or the second person; and detecting, at a second time occurring subsequent to the first time, a second instance of the bounding box defining the eye region of the first person or the second person, wherein the second instance of the bounding box is larger than the first instance of the bounding box.

44. The method of any of claims 36-43, wherein the depiction of the first person or the second person initially comprises a partially transparent depiction followed by an opaque depiction subsequent to a specified time period.

45. The method of any of claims 36-44, wherein said determining the attention or intention of the first person and said determining the attention or intention of the second person is further based on determining a facial expression indicating an inferred emotional state of the first person and a facial expression indicating an inferred emotional state of the second person.

Citation Information

Patent Citations

  • Method for display of information from real world environment on a virtual reality (VR) device and VR device thereof

    US20180232056A1

  • Visual indicators of user attention in ar / VR environment

    US20200209624A1

  • Transmodal input fusion for multi-user group intent processing in virtual environments

    US20240004464A1