VISUAL TREATMENT OF USER FACE DEMONSTRATION WHEN CONCEALED

The system addresses inaccurate facial representations by using previous data to generate user appearances during obstructions, ensuring a realistic and smooth transition by blending live and past data and applying visual treatments to indicate non-live areas.

DE102025148243A1Pending Publication Date: 2026-05-28APPLE INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
APPLE INC
Filing Date
2025-11-20
Publication Date
2026-05-28

AI Technical Summary

Technical Problem

Existing techniques fail to adequately represent the appearance of users in electronic devices when sensor data is incomplete due to obstruction, such as when a user's mouth is obscured by hands, pens, cups, or food, leading to inaccurate facial representations.

Method used

The system detects when a portion of the user's face, such as the mouth, is obscured and uses previous user data to generate a representation during the obscured period, combining live and past data to maintain a realistic appearance, and applies visual treatments to indicate the non-live nature of the obscured area.

Benefits of technology

Ensures a continuous and realistic user representation by blending live and past data, while indicating non-live areas, thus maintaining a smooth and accurate visual experience during obstructions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The various implementations disclosed herein include devices, systems and methods that detect that a section of a user's face (e.g. the user's mouth) is obscured or about to be obscured in the sensor data and accordingly determine to use previous user data to generate at least a section of a user representation during the period in which the facial section is obscured.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL AREA

[0001] The present disclosure relates generally to electronic devices and in particular to systems, methods and devices for displaying the appearance of users based on images and other sensor data. STATE OF THE ART

[0002] Existing techniques may fail to adequately represent the appearance of users of electronic devices under various circumstances. For example, user representations may exhibit undesirable visual characteristics if the sensor data on which the representations are based is incomplete, such as when the user's hand, a pen, a cup, food, etc., obscures the user's mouth in the image sensor data, resulting in an inaccurate representation of the user's actual mouth appearance in the current image data. SUMMARY

[0003] The various implementations disclosed herein include devices, systems and methods that detect that a section of a user's face (e.g. the user's mouth) is obscured or about to be obscured in the sensor data and accordingly determine to use previous user data to generate at least a section of a user representation during the period in which the facial section is obscured.

[0004] In general, an innovative aspect of the subject matter described in this patent can be implemented in a method carried out by a processor executing instructions contained on a non-transitory, computer-readable medium. The method may involve determining a sensor data condition corresponding to the obscuring of a portion of a user's face in the sensor data, for example, detecting that a portion of a user's face (e.g., the user's mouth) is obscured or about to be obscured in the sensor data. The method may further, based on the determination of the sensor data condition, involve determining whether to use prior user data to generate at least one portion of a user representation corresponding to the portion of the user's face that is obscured in the sensor data during a specific period of time.The process may also include generating the user representation that corresponds to the portion of the user's face that is obscured in the sensor data during the period.

[0005] According to some implementations, instructions are stored in a non-transitory, computer-readable storage medium, which are computer-executable to perform or cause to be performed any of the procedures described herein. According to some implementations, a device includes one or more processors, a non-transitory memory, and one or more programs; the one or more programs being stored in the non-transitory memory and configured to be executed by the one or more processors, and the one or more programs including instructions to perform or cause to be performed any of the procedures described herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0006] In order to make the present disclosure understandable to the average person skilled in the art, a more detailed description is provided with reference to aspects of some illustrative implementations, some of which are shown in the accompanying drawings. Fig. Figure 1 illustrates a device that, according to some implementations, receives sensor data from a user. Fig. Figure 2 illustrates exemplary electronic devices that are operated in different physical environments during a communication session according to some implementations. Fig. Figures 3A-C illustrate how a section of a user's face is obscured in the sensor data at exemplary times during a period in which, according to some implementations, a representation of the user's face is to be generated based on the sensor data. Fig. 4A-C illustrate the representation of the user's face. Fig. 3A-C, which was generated for the times during the period, according to some implementations. Fig. Figure 5 illustrates an exemplary visual treatment used to provide an indication that the user's face depicted in a representation may not correspond to the user's current appearance, according to some implementations. Fig. Figures 6A-B illustrate how, according to some implementations, a section of a user's face in the sensor data is no longer obscured at exemplary times during a period in which a representation of the user's face is to be generated based on the sensor data. Fig. Figures 7A-B illustrate the representation of the user's face. Fig. 6AB, which was generated for the times during the period, according to some implementations. Fig. Figure 8 is a flowchart representation that depicts a method for generating at least one section of a user representation during a period in which a section of the face is obscured, according to some implementations. Fig. Figure 9 is a block diagram illustrating device components of an example device according to some implementations. Fig. Figure 10 is a block diagram of an exemplary head-mounted device (HMD) according to some implementations.

[0007] According to common practice, the various features illustrated in the drawings may not be drawn to scale. Accordingly, the dimensions of the various features may be enlarged or reduced for clarity. Furthermore, some of the drawings may not depict all components of a particular system, process, or device. Finally, the same reference numerals may be used to designate identical features consistently throughout the patent specification and figures. DESCRIPTION

[0008] Numerous details are described to provide a thorough understanding of the exemplary implementations shown in the drawings. However, the drawings only illustrate some exemplary aspects of the present disclosure and are therefore not to be considered limiting. The person skilled in the art will recognize that other effective aspects or variants do not include all of the specific details described herein. Furthermore, generally known systems, methods, components, devices, and circuits have not been described exhaustively in detail so as not to obscure more relevant aspects of the exemplary implementations described herein.

[0009] Fig. Figure 1 illustrates an exemplary environment 100 of an exemplary electronic device 105 operating in a physical environment 102. In some implementations, the electronic device 105 may be able to share information with another device or an intermediate device such as an information system. Additionally, the physical environment 102 includes a user 110 who wears the device 105. In some implementations, the device 105 is configured to present views of an augmented reality (XR) environment that may be based on the physical environment 102 and / or may include additional content such as virtual elements.

[0010] In the example of Fig. 1. The physical environment 102 is a space that includes physical objects such as a wall hanging 120, a plant 125, and a desk 130. The electronic device 105 may include one or more cameras, microphones, depth sensors, motion sensors, or other sensors that can be used to capture and evaluate information about the physical environment 102 and the objects in it, as well as information about the user 110.

[0011] In the example of Fig. 1. The device 105 includes one or more sensors 116 that capture light intensity images, depth sensor images, audio data, or other information about the user 110 (e.g., inward-facing sensors and / or outward-facing cameras). For example, the one or more sensors 116 can capture images of the forehead, eyebrows, eyes, eyelids, cheeks, nose, lips, chin, face, head, hands, wrists, arms, shoulders, torso, legs, or other body parts of the user (e.g., user 110). For example, inward-facing sensors can see what is inside the device 105 (e.g., the user's eyes and eye area), and other external cameras can capture the user's face outside the device 105 (e.g., egocentric cameras directed at the user 110 outside the device 105).Sensor data from a user's eye 111 can, for example, indicate various user characteristics, such as the user's gaze direction 119 over time, the user's saccadic behavior over time, the user's pupil dilation behavior over time, etc. The one or more sensors 116 can capture audio information, including the user's speech and other user-generated sounds, as well as sounds within the physical environment 100.

[0012] In some implementations, the device 105 includes an eye-tracking system for detecting the position and movements of the eyes using characteristic gaze data. For example, an eye-tracking system may include one or more infrared light-emitting diodes (IR LEDs), an eye-tracking camera (e.g., a near-infrared camera (NIR camera)), and an illumination source (e.g., an NIR light source) that emits light (e.g., NIR light) toward the eyes of the user 110. Furthermore, an illumination source of the device 105 may emit NIR light to illuminate the eyes of the user 110, and an NIR camera may capture images of the eyes of the user 110. In some implementations, the images captured by the eye-tracking system may be analyzed to detect the position and movements of the eyes of the user 110, or to obtain other information about the eyes, such as color, shape, and state (e.g., wide open, squinted, etc.).), to detect pupil dilation or pupil diameter. In addition, the gaze point estimated from the eye-tracking images can enable gaze-based interaction with content shown on the near-eye display of the device 105.

[0013] Furthermore, the one or more sensors can capture 116 images of the physical environment 100 (e.g., outward-facing sensors). For example, the one or more sensors can capture 116 images of the physical environment 100, which includes physical objects such as a wall hanging 120, a plant 125, and a desk 130. In addition, the one or more sensors can capture 116 images (e.g., light intensity images and / or depth data).

[0014] One or more sensors, for example one or more sensors 115 on the device 105, can identify user information based on proximity to or contact with a section of the user 110. For example, the one or more sensors 115 can acquire sensor data that can provide biological information relating to a user's cardiovascular status (e.g., pulse), body temperature, respiratory rate, etc.

[0015] The one or more sensors 116 or the one or more sensors 115 can acquire data from which a user orientation 121 within the physical environment can be determined. In this example, the user orientation 121 corresponds to a direction in which the user's torso 110 is pointing.

[0016] Some implementations disclosed herein determine a user understanding or a scene understanding based on sensor data obtained from a user-worn device such as the first device 105. Such a user understanding may indicate a user state related to providing user support or enabling a communication session.

[0017] Content can be visual, e.g., displayed on a screen of the device 105, or audible, e.g., audio 118 generated by a speaker of the device 105. In the case of audio content, the audio 118 can be generated in such a way that only the user 110 is likely to hear it, e.g., through a speaker near the user's ear 112, or at a volume below a threshold so that people nearby are unlikely to hear it. In some implementations, the audio mode (e.g., the volume) is determined based on whether other people are within a threshold distance or how close other people are to the user 110.

[0018] In some implementations, the content and sensor capabilities provided by Device 105 can be delivered using components, sensors, or software modules that are sufficiently small and efficient in terms of power consumption and use to fit into lightweight, battery-powered, wearable products such as wireless earbuds or other ear-mounted devices, or head-mounted devices (HMDs) such as smart glasses / augmented reality glasses (AR glasses). Features can be enabled by using a combination of multiple devices. For example, a smartphone (wirelessly connected and interoperable with wearable devices) can provide computing resources, connections to cloud or internet services, location services, and so on.

[0019] The device 105 can generate user facial representations of the user 110 for various purposes based on image and / or other sensor data. For example, the user 110 can use the device 105 (e.g., a head-mounted device (HMD)) which has image sensors that capture images of parts of the user's face (e.g., images of the user's eyes via cameras in the HMD and / or images of the user's cheeks, nose, and mouth via downward-facing cameras on the HMD). A stream of image and / or other sensor data can be maintained over time and used to animate a user facial representation, e.g., to provide a user avatar representing the user's face as the user makes facial expressions and moves their face in other ways over time.

[0020] A user face rendering can combine live and past data about the user. For example, live sensor data depicting the current appearance of parts of the user's face (e.g., images of the user's eyes via cameras in the HMD and / or images of the user's cheeks, nose, and mouth via downward-facing cameras on the HMD) can be combined with past data depicting the face at one or more earlier times (e.g., registration data depicting the face without the HMD in one or more expressions, such as a neutral expression, a smiling expression, etc.).

[0021] User facial representation data can be 3D or otherwise use information about the 3D appearance of the user's face. In some implementations, real-time sensor data corresponding to the user's current / live facial appearance (e.g., real-time images from inward- and downward-facing sensors) is combined with information about the 3D shape of the user's face to provide the user facial representation. A user facial representation can be used for numerous purposes, including, but not limited to, providing a representation of the user to one or more other users during a communication session.

[0022] The implementations disclosed herein take into account circumstances during a period in which a representation of a user's face is being captured, when the sensors that record portions of the user's face (e.g., sensors 116) are obscured, for example, when cameras that record the user's mouth 110 (e.g., downward-facing cameras on a head-mounted display, HMD) are covered by the user's hand during a FaceTime® call. In this case, the device (e.g., HMD) may not be able to accurately determine the facial expression and performs one or more processes to compensate for this lack of information. For example, the device 105 (e.g., HMD) may identify content to be displayed during the period in which the portion of the user's face is obscured.For example, information from one or more earlier times when the face was not obscured can be used (e.g., camera images available from the time immediately before the facial section was obscured, and / or camera images from a registry showing the face in a particular (e.g., neutral) configuration).

[0023] Some implementations utilize information from a previous user registration. During such a registration, the system (e.g., HMD) may have captured images or other sensor data corresponding to the user's face in one or more specific configurations (e.g., smiling, frowning, neutral, closed mouth, expressionless, etc.). Such information can be used for later periods when a user representation requires current sensor data from the user, but this sensor data is partially or completely unavailable because part of the user's face is obscured. Information captured during a live capture session can also be captured and retained for use during such periods, for example, by storing camera or other sensor data from one or more points in time prior to the current time.

[0024] In some implementations, based on the detection that a portion of the user's face (e.g., the user's mouth) is obscured or about to be obscured in the sensor data, the device 105 (e.g., HMD) or another device used to display the representation determines that previous user sensor data is used to generate a user representation for the period during which the facial portion is or will be obscured. In some implementations, the user's immediately preceding expression is retained, for example, by reusing the sensor data from the immediately preceding time. For instance, as soon as the user 110 does something that covers their mouth, the device 110 (e.g., HMD) can provide the user's previous facial expression for the obscured portion of their face, e.g.,This is achieved by resetting the user's current expression for the obscured portion of their face to the user's expression immediately before covering their mouth. Other parts of the user's face (e.g., the user's eyes, cheeks, etc.) can continue to be rendered based on current sensor data. A treatment (e.g., a soft blend) can be applied between a portion of the user's face rendered based on previous data and a portion rendered based on current data. For example, the user's previous mouth expression (based on the currently obscured mouth) can be combined with the current sensor data of the upper face, which continues to be captured, so that the user's eyes and upper face in the rendering still reflect the user's current face (e.g., live expression).

[0025] In some implementations, this representation, based on a previous expression, is gradually modified over time as the user's facial portion remains obscured. This can convey to an observer that the user's face is not frozen in the previous position and / or prevent the continued display of a representation in an unnatural or otherwise undesirable frozen pose (e.g., frozen with the mouth wide open). In some examples, this involves a gradual (e.g., over a period of time) transformation or fading of the appearance of the obscured portion of the user's face into a different expression (e.g., a neutral / expressionless or other predetermined expression).For example, the user's face can initially be displayed in its previous pose and then gradually faded / transformed back to a neutral expression using registration data. This ensures that if the user does something unusual with their mouth, they don't retain that unusual (e.g., strange / frozen) expression for an extended period. As long as the mouth is covered, the expression might remain neutral.

[0026] Once the portion of the user's face is no longer obscured, the device 105 (e.g., HMD) can crossfade from the predefined (e.g., neutral) expression back to the live animation view, for example, using live sensor data of the previously obscured portion of the user's face. In alternative implementations, once the portion of the user's face is no longer obscured, for example, in the case of a momentary obscuration, the device 105 crossfades from the display based on the previous expression (or from the current crossfade of the previous and predefined expressions) back to the animated live view.

[0027] In some implementations, one or more visual treatments are applied during the period when a portion of a user's face is not based on live, current sensor data—for example, while that portion of the face is obscured. Such treatment may blur, augment (for example, by adding a light blue tint), or otherwise modify the appearance of the obscured area of ​​the face to conceal artifacts that may arise from using a combination of live and previous sensor data. Additionally (or alternatively), such visual treatments may suggest to an observer that what they are seeing may not be the user's actual mouth—for example, that it may not represent the user's actual, current facial expression. The visual effect may convey uncertainty or some other degree of inaccuracy.The extent or other attributes of the visual effect may depend on the proportion of the user's face that is obscured; for example, the extent and / or size of the blur and / or glow effect may increase based on the extent of the user's face that is obscured.

[0028] In some implementations, Device 105 (e.g., HMD) is configured to predict that a portion of the user's face is about to be obscured (but is not yet obscured), for example, based on the detection of the user's hand approaching their face. A visual treatment can then be applied based on this prediction. In some implementations, during a period before an occlusion, when Device 105 determines that future occlusion is likely, the user's displayed face may still correspond to their actual appearance. However, an additional visual treatment (blur / light, etc.) can be applied to provide the viewer with additional context for what happens after the occlusion occurs, by linking the appearance and intensity of the effect to the proximity of the hand to the mouth.This can be particularly useful because the obscuring object (e.g., a hand) might not be displayed directly in front of the mouth, for example, if the hand is not tracked / rendered when it is in close proximity to the head / device. In some implementations, additional visual treatment (blur / light / etc.) can be applied to reduce the extent of the change that occurs once the mouth is obscured, as it can be applied partially. In these cases, with occlusion-based prediction (e.g., hand-based prediction), the device can determine not to apply a visual effect to the facial segment (e.g., the mouth) but to restrict it to the torso, thus linking the effect to the hand and not obscuring the mouth.

[0029] In some implementations, the device 105 predicts, based on its movement trajectory, that a hand is likely to cover the mouth. The device 105 may only display a representation of the hand if hand tracking is available, and this may not be available if the hand is within a threshold distance of the device / head. However, based on the prediction that the hand might cover the mouth, the device 105 may decide to display the hand for slightly longer than it otherwise would, using the predicted hand movement to display the hand once tracking is lost. This may further emphasize the connection between the hand covering the mouth and the visual treatment, which would otherwise be less apparent.

[0030] In some implementations, one or more heuristics are used to determine when the mouth should no longer be treated as covered, e.g., by requiring the system to observe a certain number of uncovered frames or a predetermined period of time before the device 105 begins to de-cover the treatment.

[0031] In some implementations, a user's facial representation includes or is otherwise based on Gaussian splats, for example, via a 3D rendering based on Gaussian splats. In such implementations, the splat-based rendering may be caused by the gradual blending of facial expressions over time. Splat-based blending can appear very realistic and thus unintentionally suggest to an observer that the user's face has an expression that does not correspond to their actual, current expression. Accordingly, visual treatments or other processes may be employed to intentionally suggest that a user's face could have a different expression.Instead of, for example, a smooth crossfade between a user's previous facial expression and a predetermined / neutral expression, the transition can be intentionally speckled and modified with a classic film crossfade effect, so that while it feels "smooth," it doesn't resemble a natural human movement. For instance, a film crossfade effect is smooth, but an external viewer can easily recognize that it's an artificial transition effect rather than the user actually closing their mouth.

[0032] Fig. Figure 2 illustrates exemplary electronic devices that, according to some implementations, are operated in various physical environments during a communication session of a first user at a first device and a second user at a second device, with the second user having a view of a 3D representation of the first device. Fig. Figure 2 illustrates in particular an exemplary operating environment 200 of electronic devices 210, 265, which are operated in different physical environments 202 and 250, respectively, during a communication session, e.g., while the electronic devices 210, 265 exchange information with each other or with an intermediate device such as a communication session system / server. In this example of Fig. 2 is the physical environment 202, a room that includes a wall hanging 212, a plant 214, and a desk 216 (e.g., physical environment 102 of Fig. 1) The electronic device 210 includes one or more cameras, microphones, depth sensors, or other sensors that can be used to capture and evaluate information about the physical environment 202 and objects within it, as well as information about the user 225 of the electronic device 210 (e.g., a handheld device). The information about the physical environment 202 and / or the user 225 can be used to provide visual content (e.g., for user representations) and audio content (e.g., for audible voice or text transcription) during the communication session. For example, a communication session can provide one or more participants (e.g., users 225, 260) with views of a 3D environment generated based on camera images and / or depth camera images of the physical environment 202, a representation of the user 225.

[0033] Furthermore, in this example of Fig. 2. The physical environment 250 is a room that includes a wall hanging 252, a sofa 254, and a coffee table 256. The electronic device 265 includes one or more cameras, microphones, depth sensors, or other sensors that can be used to capture and evaluate information about the physical environment 250 and objects within it, as well as information about the user 260 of the electronic device 265 (e.g., a user-worn device or HMD device such as Device 105). The information about the physical environment 250 and / or the user 260 can be used to provide visual and audio content during the communication session.For example, a communication session can provide views of a 3D environment generated based on camera images and / or depth camera images (from the electronic device 265) of the physical environment 250, as well as a representation of the user 260 based on camera images and / or depth camera images (from the electronic device 265) of the user 260. For example, a 3D environment from the device 210 can be sent by a communication session instruction set 280, which is in communication with the device 265, by a communication session instruction set 282 (e.g., via the information system 290 over the network connection 285).

[0034] The information system 290 can coordinate the sharing of content (e.g., data linked to user representations 240, 275) between two or more devices (e.g., the electronic devices 210 and 265). Fig. Figure 2 illustrates an example of a view 205 provided on the device 210, including a user representation 240 (e.g., a figure of at least one section of the user 260), provided that consent has been given to display each user's representations during a given communication session. In particular, the user representation 240 of user 260 is generated based on one or more user representation techniques. The generation of user representations is further explained herein.

[0035] Fig. Figure 2 further illustrates a view 266, including a representation 275 (e.g., a figure) of at least one section of the user 225 (e.g., from the center of the upper body upwards) within the 3D environment 270. The user representation 240 of the user 260 can be generated on the device 210 (e.g., the receiving / viewing device) by generating representations of the user 260 for multiple points in time within a period, based on the data received from device 265 (e.g., a frame-specific 3D representation of the user 260). Alternatively, in some embodiments, the user representation 240 of the user 260 is generated on the device 265 (e.g., the transmitting device) and sent to the device 210 (e.g., the receiving / viewing device for viewing a figure of the sender).In some embodiments, each of the representations 240 by user 260 and 275 by user 225 is generated by generating splats that correspond to the user representation data.

[0036] In the example of Fig. Figure 2 illustrates electronic devices 210 and 265 as head-mounted devices (HMDs). However, each of the electronic devices 210 and 265 can be a mobile phone, a tablet, a laptop, or any other form of portable device (e.g., a head-mounted device (glasses), headphones, an ear-mounted device, etc.). In some implementations, functions of each of the devices 210 and 265 are performed across two or more devices, for example, a mobile device and a base station, or a head-mounted device and an ear-mounted device. Various capabilities can be distributed among several devices, including, but not limited to, performance capabilities, CPU capabilities, GPU capabilities, storage capabilities, memory capabilities, visual content display capabilities, audio content generation capabilities, and the like.The multiple devices that can be used to achieve the functions of the electronic devices 210 and 265 can communicate with each other via wired or wireless communication. In some implementations, each device communicates with a separate controller or server to manage and coordinate the user experience (e.g., a communication session server). Such a controller or server may be located within or remote from the physical environment 202 and / or the physical environment 250.

[0037] In the example of Fig. 2. The 3D environments 230 and 270 can also be based on a common coordinate system that can be shared with other users (e.g., providing a virtual space for characters for a multi-person communication session). In other words, a common coordinate system can be used for the 3D environments 230 and 270. A common reference point can be used to align the coordinate systems. In some implementations, the common reference point can be a virtual object within the 3D environment that each user can view within their respective views. For example, a shared table as the center point around which the user representations (e.g., the users' characters) are positioned in the 3D environment. Alternatively, the common reference point is not visible within each view.For example, a shared coordinate system of a 3D environment can use a common reference point to position each user's representation (e.g., around a table / desk). Thus, if the common reference point is visible, each view of the device can display the "center" of the 3D environment to determine the perspective of the other user representations. Displaying the common reference point can be particularly important in a multi-user communication session, allowing each user's view to perspectively complement the location of every other user during the session.

[0038] In some implementations, the representations of individual users can be realistic or unrealistic and / or depict a user's current and / or past appearance. For example, a photorealistic representation of user 225 or 260 can be created based on a combination of live images and past images of the user. The past images can be used to generate sections of the representation for which no live image data is available (e.g., sections of a user's face that are not within the field of view of a camera or sensor of the electronic device 210 or 265, or that are obscured, for example, by the respective device and / or covered by a user's hand).In one example, the electronic devices 210 and 265 HMDs and the live image data of the user's face include a downward-facing camera capturing images of the user's cheeks and mouth, as well as inward-facing camera images of the user's eyes, which can be combined with prior image data of other parts of the user's face, head, and torso that are not currently observable by the device's sensors. Prior data about the user's appearance can be obtained earlier during the communication session, during a previous use of the electronic device, during a registration process designed to obtain sensor data about the user's appearance from different perspectives and / or under different conditions, or by other means.

[0039] In some implementations, generating one or more user representations for a communication session, as in Fig. Figure 2 illustrates (e.g., Generating User Displays 240, 275) that these renderings are based on one or more rendering techniques, such as using a 3D mesh or a 3D point cloud. Alternatively, a 3D Gaussian splat rendering approach can be used. Such an approach can use UV mapping and generate a proxy mesh display.

[0040] Fig. Figures 3A-C illustrate how a portion of a user's face is obscured in the sensor data at exemplary times during a period in which a representation of the user's face is to be generated based on the sensor data. At a first time point (shown in Fig. 3A) The user wears the device 105, but the rest of the user's face is accessible to be detected by one or more sensors via an unobstructed / unobstructed view (e.g., via one or more outward / downward-facing sensors on the device 105). After the first time point, the user wears the device at a second time point (see Fig. 3B) continues to use the device 105 and moves his hand 320 into a position that at least partially obstructs / covers the view of one or more sensors (e.g., one or more outward / downward-facing sensors on the device 105 may have a restricted view, where a section of the user's face (e.g., the mouth area) is not captured in the sensor data). After the second time point, at a third time point (see Fig. 3C) continues to use the device 105 and continues to position his hand 320 in a position such that the view of one or more sensors is at least partially obstructed / covered (e.g., the view of one or more outward / downward-facing sensors on the device 105 may continue to have a restricted view in which a section of the user's face (e.g., the mouth area) is not captured in the sensor data).

[0041] Fig. 4A-C illustrate the representation of the user's face. Fig. 3A-C, which were generated for the time points during the period. In particular, for the first time point (illustrated in Fig. 3A) a user representation 410 (illustrated in Fig. 4A) is generated based on live sensor data, which captures sensor data (e.g., via one or more outward / downward-facing sensors on device 105) of the user's face. The user display 410 shows a current / live appearance of the user's face because the face is not obscured at this initial time. The user display 410 can provide such an appearance only using live sensor data or a combination of live and previously captured sensor data (e.g., data from a previous user registration). Sections of the user's face that are not obscured in the live / current data are displayed in the user display 410 to correspond to the user's live / current appearance (e.g., if the user is currently smiling, the user's mouth area is displayed as smiling in the user display 410, etc.).In this example of user representation 410, the eye area of ​​the face is displayed based on a combination of current sensor data (e.g., internal forward-facing cameras capturing the live / current condition of the user's eye area) and previous registration data (e.g., information such as the color, about the appearance of the user's eyes, captured in sensor data when the device was not being worn by the user).

[0042] For the second point in time (illustrated in Fig. 3B) can be a user representation 420 (illustrated in Fig. 4B) based on live sensor data that captures sensor data (e.g., captured by one or more outward / downward-facing sensors on the device 105) of the user's face and / or previously captured sensor data (e.g., captured by one or more outward / downward-facing sensors on the device 105 at the first time point and / or during a previous registration process when the device 105 was not worn by the user). User representation 420 shows a non-live / non-current appearance of at least one part of the user's face because that part of the face is obscured at this second time point. Specifically, in this example, user representation 420 shows the mouth area of ​​the face based on sensor data captured at the first time point when the mouth area was not obscured; that is, the mouth area may retain its appearance from the previous first time point.In this example of user representation 420, the eye area of ​​the face is displayed based on a combination of current sensor data (e.g., internal forward-facing cameras capturing the live / current condition of the user's eye area) and previous registration data (e.g., information such as the color of the user's eyes, captured in sensor data when the device was not being worn by the user). A user representation can thus combine information from a user registration (e.g., information about the user's eye color) with current information about the user's face (e.g., information about the user's current gaze direction and gaze state, as well as information about the uncovered portions of the user's face) and / or previous information about the user's face from a recent point in time (e.g.,Information about the lower part of the user's face (currently obscured in the sensor data, but not obscured at a previous, recent time) is combined. Under certain circumstances, prior information about a user's face is obtained as part of a previous event unrelated to a prior registration. For example, at least some of the prior information about the user's face may be obtained during the same communication session as the current information about the user's face. Such information about the user's face may provide information about the appearance of the user's face, including, but not limited to, information about a recent time when the user's mouth was not obscured during a particular communication session.To ensure a smooth, continuous, or otherwise desirable transition between facial segments representing previous and current sensor data, various blending processes or visual effects can be used.

[0043] For the third point in time (illustrated in Fig. 3C) can display a user representation 430 (illustrated in Fig. 4C) based on live sensor data that captures sensor data (e.g., captured by one or more outward / downward-facing sensors on the device 105) of the user's face and / or previously captured sensor data (e.g., captured by one or more outward / downward-facing sensors on the device 105 at a previous time and / or during a previous registration process when the device 105 was not worn by the user). User representation 430 shows a non-live / non-current appearance of at least one portion of the user's face because that portion of the face remains obscured at this third time. Specifically, in this exemplary user representation 430, the mouth area of ​​the face is depicted based on sensor data captured during a registration process when the mouth area was not obscured.Such a registration may have occurred separately and / or at a time prior to the current communication session. During such a registration, one or more facial configurations / expressions may be captured in sensor data and used to provide the appearance of the portion of the face that is obscured during live capture. In this example, a neutral facial expression is generated based on sensor data captured during such a registration for the obscured portion of the face (e.g., the mouth area). In this example for User Display 430, the eye area of ​​the face is generated based on a combination of current sensor data (e.g., internal forward-facing cameras capturing the live / current condition of the user's eye area) and previous registration data (e.g.,Information such as the color and appearance of the user's eyes, captured in sensor data when the device was not being worn, is displayed. A user display can thus combine information from user registration (e.g., information about the appearance of the neutral mouth area and the user's eye color) with current information about the user's face (e.g., information about the user's current gaze direction and gaze state, as well as information about the uncovered portions of the user's face). To ensure a smooth, continuous, or otherwise desirable transition between such sections, various blending processes or visual effects can be used between facial sections displayed based on previous / registration sensor data and those displayed based on current sensor data.

[0044] Fig. Figure 5 illustrates an exemplary visual treatment used to provide an indication (530) that the user's face depicted (430) in a user representation may not correspond to the user's current appearance. For example, a user representation may be displayed to a second user during a live communication session. It may be desirable to provide the second user viewing the image with an indication of when the user representation's appearance is not live / current. Such an indication can take various forms, including, but not limited to, highlighting, colorizing, blurring, outlining, darkening, or softening the area that does not correspond to the user's live / current appearance.

[0045] Such a specification can be useful, for example, in implementations where the user representation (e.g., User Representation 430) is generated using a Gaussian splatting technique (e.g., by using points represented by parameters that include information about the Gaussian distribution, which represents the appearance of points on the surface of the face to generate views from specific angles, such as stereo viewpoints). Such techniques can generate user representations that would otherwise be overly realistic and / or easily mistaken for the user's live / actual appearance or movements. For example, it may be undesirable to specify a realistic appearance in which the user is smiling when they are not actually smiling.

[0046] Fig. Figures 6A-B illustrate how a section of a user's face is no longer obscured in the sensor data at exemplary points in time during a period in which a representation of the user's face is to be generated based on the sensor data. After the third point in time ( Fig. 3C) the user carries at a fourth time (see Fig. 6A) the device 105 and continues to position his hand 320 in a position such that the view of one or more sensors is at least partially obstructed / covered (e.g., the view of one or more outward / downward-facing sensors on the device 105 may remain restricted, so that a section of the user's face (e.g., the mouth area) is not captured in the sensor data). After the fourth time point ( Fig. 6A) the user carries at the fifth time (shown in Fig. 6B) the device 105 and has moved his hand 320 away from his face so that one or more sensors (e.g., one or more outward / downward-facing sensors on the device 105) can again detect sensor data corresponding to the part of the user's face (e.g., the mouth area) that was previously obscured in the sensor data.

[0047] Fig. Figures 7A-B illustrate the representation of the user's face. Fig. 6AB, which was generated for the time points during the period. For the fourth time point (illustrated in Fig. 6A) can display a user representation 710 (illustrated in Fig. 7A) based on live sensor data that captures sensor data (e.g., captured by one or more outward / downward-facing sensors on the device 105) of the user's face and / or previously captured sensor data (e.g., captured by one or more outward / downward-facing sensors on the device 105 at a previous time and / or during a previous registration process when the device 105 was not worn by the user). The user representation 710 shows a non-live / non-current appearance of at least one portion of the user's face because that portion of the face remains obscured at this fourth time point.Specifically, in this example for user representation 710, although the current appearance of the mouth (which is obscured by hand 320) shows an open mouth expression, the mouth area of ​​the face is still rendered based on sensor data captured during a registration process when the mouth area was not obscured (and showed a closed mouth expression). Such registration allows one or more facial configurations / expressions to be captured in sensor data and used to provide the appearance of the portion of the face that is obscured during live capture. In this example, a neutral facial expression is generated based on sensor data captured during such registration for the obscured portion of the face (e.g., the mouth area).In this example of user display for the 710, the eye area of ​​the face is displayed based on a combination of current sensor data (e.g., internal forward-facing cameras capturing the live / current condition of the user's eye area) and previous registration data (e.g., information such as the color, about the appearance of the user's eyes, captured in sensor data when the device was not being worn by the user). In this example of... Fig. 7A The eyes have a partially closed appearance based on the current partially closed state of the eyes at the fourth time point, but exhibit other characteristics, such as color, based on the user's previous registration. A user representation can combine information from a user registration (e.g., information about the appearance of the neutral mouth area and the user's eye color) with current information about the user's face (e.g., information about the user's current gaze direction and gaze state, as well as information about the uncovered portions of the user's face).To ensure a smooth, continuous, or otherwise desirable transition between such sections, various blending processes or visual effects can be used between facial sections rendered based on previous / recorded sensor data and facial sections rendered based on current sensor data.

[0048] For the fifth point in time (illustrated in Fig. 6B) a user representation 720 (illustrated in Fig. 7B) is generated based on live sensor data that captures sensor data (e.g., captured by one or more outward / downward-facing sensors on the device 105) of the user's face. User representation 720 shows the live / current appearance of the user's face because the face is no longer obscured at this fifth point in time. User representation 410 can provide such an appearance only using live sensor data or a combination of live and previously captured sensor data (e.g., data from a previous user registration). Sections of the user's face that are not obscured in the live / current data are displayed in user representation 720 to correspond to the user's live / current appearance (e.g.,In User Display 410, if the user is smiling, the user's mouth area is depicted as smiling; if the user's eyes are partially closed, the eyes are depicted as partially closed, and so on. In this example of User Display 410, the eye area of ​​the face is displayed based on a combination of current sensor data (e.g., internal forward-facing cameras capturing the live / current condition of the user's eye area – partially closed) and previous registration data (e.g., information such as the color of the user's eyes, captured in sensor data when the device was not being worn by the user).

[0049] A transition (e.g., over a period of time) can be applied to gradually change the appearance of the user's face in User Representation 710 to the appearance of the user's face in User Representation 720. Such a gradual transition in facial expression from a non-live to a live appearance can provide a better viewing experience.

[0050] Fig. Figure 8 is a flowchart illustrating an exemplary procedure 800. In some implementations, a device (e.g., device 105) performs Fig. 1) The techniques of Method 800 for generating at least a portion of a user representation during a period in which a portion of the face is obscured. In some implementations, the techniques of Method 800 are executed on a mobile device, desktop, laptop, HMD, or server device. In some implementations, Method 800 is executed on processing logic, including hardware, firmware, software, or a combination thereof. In some implementations, Method 800 is performed on a processor executing code stored in non-transitory computer-readable medium (such as memory). In some implementations, Method 800 is implemented on a processor of a device such as a viewing device that renders a representation of a user (for example, Device 210 renders a user's face). Fig. 2 a 3D representation 240 of the user 260 (of a figure) from data obtained from the device 265).

[0051] In Block 810, Method 800 involves determining a sensor data condition corresponding to a portion of a user's face that is obscured in sensor data. This may involve detecting that a portion of a user's face (e.g., the user's mouth) is obscured or about to be obscured in the sensor data. It may involve determining that the portion of the user's face is currently obscured or about to be obscured by a user's hand. For example, a user's hand may be tracked by one or more sensors on the device or another device. The movement of the hand over time may be used to predict that the hand will be in a position that prevents the sensors from obtaining sensor data about a portion of the face. In addition, information about a user and / or the physical environment (such as...) may be used.Previous hand movements of the user, the tendency to put their hand in front of their mouth in certain situations, etc.) are used to predict that the hand will be in a position that prevents the sensors from obtaining sensor data about the facial area.

[0052] In block 820, procedure 800, based on determining the sensor data condition, involves determining to use previous user data to generate at least one section of a user representation that corresponds to the portion of the user's face that was obscured in the sensor data during a period of time. As already mentioned in relation to the Fig. As explained in 3A-C, 4A-C, 6A-B and 7A-B, the previous user data may include user data representing the appearance of the portion of the user's face captured during a period immediately prior to the occurrence of an occlusion, and / or user data representing the appearance of the portion of the user's face captured during a registration period in which images of the user's face were taken in a variety of facial configurations (e.g. smiling, frowning, neutral, mouth closed, expressionless, etc.).

[0053] In block 830, procedure 800 involves generating the user representation corresponding to the portion of the user's face that is obscured in the sensor data during the period. The user representation can be generated during a live acquisition session in which sensor data from periods without obscuration are stored for use during periods of obscuration. Generating the user representation corresponding to the portion of the user's face obscured in the sensor data set may involve generating the user representation to obtain the user's immediately preceding facial expression during the period (e.g., using the last available / most recent information about the facial segment).

[0054] Other sections of the user's face can be displayed based on live sensor data, reflecting the live appearance of those other sections over a period of time. Visual processing is applied between the section of the user's face displayed based on previous data and the other sections displayed based on current data. For example, a smooth blend can be applied between a section of the user's face displayed based on previous data and a section displayed based on current data.

[0055] Generating the user appearance corresponding to the portion of the user's face that is obscured in the sensor data during the relevant period can involve generating a gradual change for that portion of the user's face. This could involve transitioning from a current appearance to a last observed appearance, and then gradually transitioning from that to a registration-based appearance. A gradual change can be applied to transform an initial appearance of the facial portion corresponding to a first expression occurring immediately before the obscuration into a second appearance of the facial portion corresponding to a second expression that differs from the first (e.g., transitioning to neutral).

[0056] Method 800 may further include determining a second sensor data condition corresponding to the portion of the user's face that is no longer obscured in the sensor data. This may involve, for example, detecting that a portion of a user's face (e.g., the user's mouth) is no longer obscured or is about to be obscured in the sensor data. Based on the determination of the second sensor data condition, Method 800 may further include determining to use live user data to generate at least the portion of the user representation corresponding to the portion of the user's face that is no longer obscured by the sensor data during a second period, and generating the user representation corresponding to the portion of the user's face that is no longer obscured by the sensor data during the second period.Generating the user representation corresponding to the portion of the user's face that is no longer obscured by the sensor data during the second period can involve generating a gradual change for that portion of the user's face. This gradual change can transform an initial appearance of the facial portion corresponding to a first expression (e.g., an immediately preceding or neutral expression) into a second appearance of the facial portion corresponding to a second expression different from the first (e.g., an animated live view).In some implementations, a triple transition effect is used during the transition from the previous frame to the neutral frame, blending between the previous / neutral frame (based on the progress of the transition) to define the mouth area, and then between the mouth area and the live frame across the entire face (so that the eyes remain alive, the mouth remains fixed, and the area in between is blurred).

[0057] As in Fig. As illustrated in Figure 5, Procedure 800 may involve the application of a visual treatment, whereas a user representation is based on non-live user data, meaning that the portion of the user's face depicted in the user representation may not reflect the user's current facial expression. A feature of the visual treatment may be based on the portion of the face that is obscured.

[0058] In some implementations, the user display is generated based on Gaussian representations with 3D positions, and then a view is generated based on these Gaussian representations, with transitions between facial expressions configured to convey unnatural changes. This can be done, for example, so that the movement feels "fluid" but not like natural human movement. For instance, a film crossfade effect is smooth, but an external viewer can easily recognize that it is an artificial transition effect rather than the user actually closing their mouth.

[0059] Fig. Figure 9 is a block diagram of an exemplary device 900. Device 900 illustrates an exemplary device configuration for the devices described herein (e.g., devices 105, 210, 265, 410, etc.). Although certain specific features are illustrated, the person skilled in the art will recognize from the present disclosure that various other features have not been illustrated for the sake of brevity, so as not to obscure more relevant aspects of the implementations disclosed herein. For this purpose, in some implementations, the device 900 includes, as a non-limiting example, one or more processing units 902 (e.g., microprocessors, ASIC, FPGA, GPU, CPU, processing cores and / or the like), one or more input / output devices (I / O devices) and sensors 906, one or more communication interfaces 908 (e.g., USB, FIREWIRE, THUNDERBOLT, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, GSM, CDMA, TDMA, GPS, IR, BLUETOOTH, ZIGBEE, SPI, I2C and / or similar interface types), one or more programming interfaces (e.g. I / O interfaces) 910, one or more displays 912, one or more inward and / or outward-facing image sensor systems 914, a memory 920 and one or more communication buses 904 for connecting these and various other components together.

[0060] In some implementations, the one or more communication buses 904 include switching logic that connects and controls the communications between system components. In some implementations, the one or more I / O devices and sensors 906 include at least one inertial measurement unit (IMU), accelerometer, magnetometer, gyroscope, thermometer, one or more physiological sensors (e.g., blood pressure measuring device, heart rate measuring device, blood oxygen sensor, blood glucose sensor, etc.), one or more microphones, one or more loudspeakers, a haptic motor, one or more depth sensors (e.g., a fringe projection sensor, a time-of-flight sensor, or the like), and / or the like.

[0061] In some implementations, the one or more displays 912 are configured to present the user with a view of a physical or graphical environment. In some implementations, the one or more displays 912 correspond to holographic digital light processing (DLP), liquid crystal display (LCD), liquid crystal on silicon (LCoS), organic light-emitting field-effect transistor (OLET), organic light-emitting diode (OLED), surface conduction electron emitter (SED), field emission display (FED), quantum dot light-emitting diode (QD-LED), microelectromechanical system (MEMS), and / or similar display types. In some implementations, the one or more displays 912 correspond to diffraction, reflection, polarized, holographic, and other waveguide displays. In one example, the device 10 includes a single display.In another example, device 10 includes a display for each of the user's eyes.

[0062] In some implementations, the one or more image sensor systems 914 are configured to obtain image data corresponding to at least one section of the physical environment 102. For example, the one or more image sensor systems 914 include one or more RGB cameras (e.g., with a complementary metal-oxide-semiconductor image sensor (CMOS image sensor) or a charge-coupled device image sensor (CCD image sensor)), monochrome cameras, IR cameras, depth cameras, event-based cameras, and / or the like. In various implementations, the one or more image sensor systems 914 further include illumination sources that emit light, such as a flash. In various implementations, the one or more image sensor systems 914 further include an in-camera image signal processor (ISP) configured to perform a variety of processing operations on the image data.

[0063] Memory 920 includes high-speed random-access memory such as DRAM, SRAM, DDR-RAM, or other solid-state random-access memory devices. In some implementations, Memory 920 includes non-volatile memory such as one or more magnetic disk storage devices, optical disk storage devices, flash storage devices, or other non-volatile solid-state storage devices. Memory 920 optionally includes one or more storage devices located remotely from the one or more Processing Units 902. Memory 920 includes a non-transient, computer-readable storage medium.

[0064] In some implementations, the memory 920 or the non-transitory, computer-readable storage medium of the memory 920 stores an optional operating system 930 and one or more instruction sets 940. The operating system 930 includes procedures for handling various basic system services and performing hardware-dependent tasks. In some implementations, the one or more instruction sets 940 include executable software defined by binary information stored in the form of electrical charge. In some implementations, the one or more instruction sets 940 are software executable by the one or more processing units 902 to perform one or more of the techniques described herein.

[0065] The instruction set(s) 940 includes a registration instruction set 942, an obfuscation detection instruction set 944, a user presentation instruction set 946, and a communication session instruction set 948. The instruction set(s) 940 may be embodied as a single executable software or as multiple executable software programs.

[0066] In some implementations, the registration instruction set 942 can be executed by the processing unit(s) 902 to generate registration data from image data. The registration instruction set 942 can be configured to provide instructions to the user to acquire image or other sensor information to generate the registration personification and to determine whether additional image information is needed to generate an accurate registration personification to be used by the figure display process. For these purposes, in various implementations, the instruction includes instructions and / or logic for this purpose, as well as heuristics and metadata.

[0067] In some implementations, the occlusion detection instruction set 944 can be executed by the processing unit(s) 902 to determine when an obstacle, such as a user's hand, prevents a sensor from capturing a live / current appearance of a portion of the user, as described here. The occlusion detection instruction set 944 may include or utilize an instruction set that performs body and / or hand tracking. For this purpose, in various implementations, the instruction includes instructions and / or logic for this, as well as heuristics and metadata for it.

[0068] In some implementations, the user presentation instruction set 946 is executable by the processing unit(s) 902 to generate a user presentation using one or more of the techniques discussed herein, or as may otherwise be appropriate. For this purpose, in various implementations, the instruction includes instructions and / or logic for it, as well as heuristics and metadata for it.

[0069] In some implementations, the communication session instruction set 948 can be executed by the processing unit(s) 902 to establish a communication session between two or more electronic devices (e.g., device 210 and device 265, as in Fig. 2 illustrated) using one or more of the techniques discussed herein, or as may otherwise be appropriate. For these purposes, the instruction in various implementations includes instructions and / or logic for it, as well as heuristics and metadata for it.

[0070] Although instruction set(s) 940 are shown to reside on a single device, it is understood that in other implementations any combination of the elements may reside on separate data processing devices. Furthermore, Fig. Section 9 is intended more as a functional description of the various features present in a given implementation, as opposed to a structural scheme of the implementations described herein. As the average person skilled in the art will recognize, elements shown separately could be combined, and some elements could be separated. The actual number of instruction sets and how features are assigned to them may vary from one implementation to another and may depend in part on the specific combination of hardware, software, and / or firmware chosen for a particular implementation.

[0071] Fig.Figure 10 illustrates a block diagram of an exemplary head-worn device 1000 according to some implementations. The head-worn device 1000 includes a housing 1001 (or encapsulation) containing various components of the head-worn device 1000. The housing 1001 includes (or is coupled to) an eye pad (not shown) located at a proximal end (towards the user 25) of the housing 1001. In various implementations, the eye pad is a piece of plastic or rubber that comfortably and precisely holds the head-worn device 1000 in the correct position on the user's face 25 (e.g., surrounding the user's eye 35).

[0072] The housing 1001 accommodates a display 1010, which shows an image emitted by a user 25 towards or onto their eye. In various implementations, the display 1010 emits the light through an eyepiece that has one or more optical elements 1005 which refract the light emitted by the display 1010, so that the display appears to the user 25 at a virtual distance greater than the actual distance from the eye to the display 1010. The optical element(s) 1005 may include, for example, one or more lenses, a waveguide, other optical diffraction elements (DOEs), and the like. To enable the user 25 to focus on the display 1010, the virtual distance is, in various implementations, at least greater than a minimum focal distance of the eye (e.g., 7 cm).To provide a better user experience, the virtual distance is also greater than 1 meter in various implementations.

[0073] The housing 1001 also incorporates a tracking system that includes one or more light sources 1022, camera 1024, camera 1032, camera 1034, and a controller 1080. The one or more light sources 1022 emit light onto the eye of the user 25, which is reflected as a light pattern (e.g., a circle of shimmering) that can be detected by the camera 1024. Based on the light pattern, the controller 1080 can determine an eye-tracking characteristic of the user 25. For example, the controller 1080 can determine the gaze direction and / or blink state (eyes open or eyes closed) of the user 25. As another example, the controller 1080 can determine a pupil center, pupil size, or viewpoint. In various implementations, the light is emitted by one or more light sources 1022, reflected by the user's eye 25, and detected by the camera 1024.In various implementations, the light from the user's eye 25 is reflected by a hot mirror or passes through an eyepiece before reaching the camera 1024.

[0074] Display 1010 emits light in a first wavelength range, and the one or more light sources 1022 emit light in a second wavelength range. Similarly, camera 1024 detects light in the second wavelength range. In various implementations, the first wavelength range is a visible wavelength range (e.g., a wavelength range within the visible spectrum of approximately 400–700 nm), and the second wavelength range is a near-infrared wavelength range (e.g., a wavelength range within the near-infrared spectrum of approximately 700–1400 nm).

[0075] In various implementations, eye tracking (or, more specifically, a particular gaze direction) is used to enable user interaction (e.g., the user selects an option on the display by looking at it), to provide foveated rendering (e.g., presenting a higher resolution in an area of ​​the display that the user is looking at and a lower resolution elsewhere on the display), or to correct distortions (e.g., for images to be displayed on the display). In various implementations, the one or more light sources emit light toward the user's eye, which is reflected in the form of a multitude of shimmers.

[0076] In various implementations, the camera 1024 is a frame / shutter-based camera that generates an image of the user's eye 35 at a specific time or multiple time points at a given frame rate. Each image includes a matrix of pixel values ​​corresponding to pixels in the image that correspond to positions in a matrix of the camera's light sensors. In some implementations, each image is used to measure or track pupil dilation by measuring changes in pixel intensities associated with one or both of the user's pupils.

[0077] In various implementations, the camera 1024 is an event camera that includes a variety of light sensors (e.g., a matrix of light sensors) at a variety of respective locations, which, in response to a specific light sensor detecting a change in light intensity, generates an event message indicating a specific location of that particular light sensor.

[0078] In various implementations, Camera 1032 and Camera 1034 are frame / shutter-based cameras that can generate an image of the user's face at one or more specific times and at a given frame rate. For example, Camera 1032 captures images of the user's face below the eyes, and Camera 1034 captures images of the user's face above the eyes. The images captured by Camera 1032 and Camera 1034 can include light intensity images (e.g., RGB) and / or depth image data (e.g., time-of-flight, infrared, etc.).

[0079] It will therefore be evident that the implementations described above are given as examples and that the present invention is not limited to what has been shown and described in detail above. Rather, its scope includes combinations and partial combinations of the various features described above, as well as variations and modifications thereof, which will be clear to a person skilled in the art upon reading the preceding description and which are not disclosed in the prior art.

[0080] As described above, one aspect of the present technology is the collection and use of physiological data to enhance a user's experience with an electronic device in relation to interaction with electronic content. This disclosure considers that, in some cases, such collected data may include personally identifiable information that can uniquely identify a particular individual or be used to identify an individual's interests, characteristics, or tendencies. Such personally identifiable information may include physiological data, demographic data, location data, telephone numbers, email addresses, home addresses, device characteristics of personal devices, or any other personally identifiable information.

[0081] This disclosure recognizes that the use of such personal data in the present technology can be to the benefit of users. For example, the personal data can be used to improve the interaction with and control capabilities of an electronic device. Accordingly, the use of such personal data enables computational control of the electronic device. Furthermore, this disclosure also considers other uses of personal data that are beneficial to the user.

[0082] This disclosure further considers that entities responsible for collecting, analyzing, disclosing, transferring, storing, or otherwise using such personal and / or physiological data should adhere to data protection best practices and / or best practices. In particular, such entities should implement and consistently use data protection policies and practices that are generally recognized as meeting or exceeding industry or regulatory requirements for maintaining and protecting the confidentiality of personal data. For example, personal data from users should be collected for legitimate and appropriate uses by the entity and should not be shared or sold outside of those legitimate uses. Furthermore, such collection should only occur after obtaining informed consent from users.Furthermore, such entities must take all necessary steps to protect and secure access to such personal data and ensure that others who have access to the personal data comply with their data protection policies and procedures. In addition, such entities may submit to a third-party evaluation to confirm that they adhere to generally accepted data protection policies and practices.

[0083] Despite the foregoing, this disclosure also considers implementations in which users selectively block the use of or access to personal data. That is to say, this disclosure considers the possibility of providing hardware or software elements to prevent or block access to such personal data. For example, in the case of personalized content delivery services, this technology could be configured to allow users to choose, during registration for services, whether to consent to ("opt in") or decline ("opt out") the collection of personal data. In another example, users could choose not to provide any personal data to targeted content delivery services.In yet another example, users can choose not to provide any personal information, but allow the transmission of anonymous information for the purpose of improving the functionality of the device.

[0084] Although the present disclosure broadly covers the use of personal information for implementing one or more different disclosed embodiments, it also provides that the different embodiments can be implemented without the need for access to such personal information. That is to say, the various embodiments of the present technology do not become inoperable due to the absence of all such personal data or any part thereof.For example, content can be selected and delivered to users by inferring preferences or settings based on non-personal data or a mere minimum of personal information, such as the content requested by the device assigned to a user, other non-personal information available to content delivery services, or publicly available information.

[0085] In some implementations, data is stored using a public / private key system that allows only the data owner to decrypt the stored data. In some other implementations, the data can be stored anonymously (e.g., without identifying and / or personally identifiable information about the user, such as a real name, username, time and location data, or the like). This prevents other users, hackers, or third parties from determining the user's identity associated with the stored data. In some implementations, a user can access their stored data from a different device than the one used to upload the data. In these cases, the user may be required to provide login credentials to access their stored data.

[0086] Numerous specific details are set forth herein to provide a thorough understanding of the claimed subject matter. The person skilled in the art will understand that the claimed subject matter can also be implemented without these specific details. In other cases, methods, devices, or systems that would be known to a person skilled in the art have not been described in detail in order to avoid obscuring the claimed subject matter.

[0087] Unless specifically stated otherwise, the discussions contained in this description, using terms such as "processing", "processing data", "calculating", "determining" and "identifying" or the like, are understood to refer to actions or processes of a data processing device such as one or more computers or similar electronic data processing device or devices that manipulate or transform data represented as physical, electronic or magnetic quantities in storage devices, registers or other information storage devices, transmission devices or display devices of the data processing platform.

[0088] The system or systems discussed here are not limited to any particular hardware architecture or configuration. A data processing device may include any suitable arrangement of components that provides a result based on one or more inputs. Suitable data processing devices include general-purpose microprocessor-based computer systems that access stored software, which programs or configures the data processing system from a general-purpose computing device to a specialized data processing device that implements one or more implementations of the subject matter here.Any suitable programming, scripting, or other type of language or combination of languages ​​may be used to implement the teachings contained herein in software to be used in programming or configuring a data processing device.

[0089] Implementations of the methods disclosed herein can be carried out during the operation of such data processing devices. For example, the order of the blocks shown in the examples above can be varied, blocks can be reordered, combined, or decomposed into sub-blocks. Certain blocks or processes can be executed in parallel.

[0090] The use of "adapted to" or "configured to" herein is intended as an open and inclusive formulation that does not exclude devices that are adapted or configured to perform additional tasks or steps. Similarly, the use of "based on" is intended to be open and inclusive in that a process, step, calculation, or other action that is "based" on one or more specified conditions or values ​​may, in practice, be based on additional conditions or a value beyond those specified. Headings, lists, and numbering contained herein are provided for convenience only and are not intended to be restrictive.

[0091] It is also understood that, although the terms "first," "second," etc., may be used here to describe different objects, these objects are not restricted by these terms. These terms are only used to distinguish one object from another. For example, a first node could be called a second node, and similarly, a second node could be called a first node without changing the meaning of the description, as long as every occurrence of "first node" is consistently renamed and every occurrence of "second node" is consistently renamed. The first node and the second node are both nodes, but they are not the same node.

[0092] The terminology used herein serves only to describe certain implementations and is not intended to limit the claims. As used in the description of the implementations and the accompanying claims, the singular forms "a," "an," "the," "a," and "a" are intended to include the plural forms unless the context clearly indicates otherwise. It is also understood that the term "or," as used here, refers to and includes any and all possible combinations of one or more of the related terms listed.It is further understood that the terms “include” or “comprehensive”, when used in this description, indicate the presence of listed features, integers, steps, operations, elements or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components or groups thereof.

[0093] As used herein, the term "if" can, depending on the context, be interpreted as "upon" or "in response to finding" or "according to a finding" or "in response to recognizing" that a stated antecedent condition is satisfied. Likewise, the phrase "when it is found [that a stated antecedent condition is satisfied]" or "when [a stated antecedent condition is satisfied]" can, depending on the context, be interpreted as "upon finding" or "in response to a finding that" or "according to a determination" or "upon recognizing" or "in response to recognizing" that the stated antecedent condition is satisfied.

[0094] The foregoing description and summary of the invention are to be understood as illustrative and exemplary in every respect, but not as limiting, and the scope of the invention disclosed herein is to be determined not only from the detailed description of illustrative implementations, but according to the full breadth permitted by patent law.

[0095] It is understood that the implementations shown and described herein are only illustrative of the basic ideas of the present invention and that various modifications can be implemented by the person skilled in the art without deviating from the scope and spirit of the invention.

Claims

[1] Procedure, encompassing: on a processor of a head-worn device (HMD): Determining a sensor data condition that corresponds to a section of a user's face that is obscured in sensor data; based on determining the sensor data condition, determining, to use previous user data to generate at least one section of a user representation that corresponds to the section of the user's face that is obscured in the sensor data during a given period; and Generating the user representation that corresponds to the portion of the user's face that is obscured in the sensor data during the specified period. [2] Method according to claim 1, wherein determining the sensor data condition includes determining that the portion of the user's face is currently covered or about to be covered by a hand of the user. [3] Method according to one of claims 1 or 2, wherein the section of the user's face includes a mouth area of ​​the user. [4] Method according to any one of claims 1 to 3, wherein the prior user data comprises user data representing the appearance of the portion of the user's face captured during a period immediately prior to the occurrence of an obscuration. [5] Method according to any one of claims 1 to 3, wherein the prior user data comprises user data representing an appearance of the section of the user's face that was captured during a registration period in which images of the user's face were captured in a plurality of facial configurations. [6] Method according to any one of claims 1 to 5, wherein the user display is generated during a live acquisition session, during which sensor data from a period without occlusion are stored for use during occlusion periods. [7] Method according to any one of claims 1 to 6, wherein generating the user representation corresponding to the part of the user's face obscured in the sensor data comprises generating the user representation to maintain the immediately preceding facial expression of the user during the period, wherein other parts of the user's face are represented based on live sensor data corresponding to the live appearance of the other parts of the user's face during a period, and wherein a visual treatment is provided between the part of the user's face and the other parts of the user's face. [8] Method according to any one of claims 1 to 7, wherein generating the user representation corresponding to the part of the user's face that is obscured in the sensor data during the period comprises generating a gradual change for the part of the user's face that is obscured in the sensor data. [9] Method according to claim 8, wherein the gradual change transforms a first appearance of the portion of the face corresponding to a first expression occurring immediately before the covering into a second appearance of the portion of the face corresponding to a second expression that differs from the first expression. [10] Method according to any one of claims 1 to 9, further comprising: Determining a second sensor data condition that corresponds to the portion of the user's face that is no longer obscured in the sensor data; based on determining the second sensor data condition, determining, to use live user data to generate at least the portion of the user representation that corresponds to the portion of the user's face that is no longer obscured in the sensor data during a second period; and Generating the user representation that corresponds to the section of the user's face that is no longer obscured in the sensor data during the second period. [11] Method according to any one of claims 1 to 10, wherein generating the user representation corresponding to the part of the user's face that is no longer obscured in the sensor data during the second period comprises generating a gradual change for the part of the user's face, wherein the gradual change transforms a first appearance of the part of the face corresponding to a first expression into a second appearance of the part of the face corresponding to a second expression different from the first expression. [12] Method according to any one of claims 1 to 11, further comprising applying a visual treatment while a user representation is based on non-live user data indicating that the portion of the user's face shown in the user representation may not reflect a current facial expression of the user, wherein an attribute of the visual treatment is based on a portion of the face that is obscured. [13] Method according to any one of claims 1 to 12, wherein the user display is generated based on generating Gaussian-based representations with 3D positions and subsequently generating a view based on the Gaussian-based representations, wherein transitions between facial expressions are configured to convey unnatural changes. [14] Device comprising: a non-transient, computer-readable storage medium; and one or more processor(s) coupled to the non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium comprises program instructions which, when executed on the one or more processor(s), cause the one or more processor(s) to perform operations comprising any one of the methods according to claims 1 to 13. [15] Non-transitory computer-readable storage medium storing program instructions that can be executed on a device to perform operations comprising any of the methods of claims 1 to 13.