Visual processing of user representations when interacting with secure UI elements
By detecting sensitive user input triggers, pausing or adjusting the transmission of virtual representation data and generating modified live frames, the problem of sensitive information leakage in shared extended reality environments is solved, and user input privacy protection is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- APPLE INC
- Filing Date
- 2025-10-16
- Publication Date
- 2026-04-21
AI Technical Summary
In shared extended reality environments, sensitive user input is exposed to the risk of potential eavesdropping or hacking attacks, and existing technologies struggle to effectively manage user input to protect privacy.
By detecting sensitive user input triggers, the transmission of virtual representation data can be paused or adjusted. For example, the capture and generation of sensor data can be paused, and modified live frames can be generated to confuse the user's eyes, ensuring that sensitive information is not leaked.
It effectively protects users' sensitive information, prevents eavesdropping and hacking attacks, and ensures users' privacy and security in a shared XR environment.
Smart Images

Figure CN121902188A_ABST
Abstract
Description
Background Technology
[0001] Some devices can generate and present extended reality (XR) environments. An XR environment can include a fully or partially simulated environment that people perceive and / or interact with through electronic systems. In XR, a subset of a person's physical motion or a representation thereof is tracked, and in response, one or more properties of one or more virtual objects simulated in the XR environment are adjusted in a manner consistent with real-world properties.
[0002] Some XR environments allow multiple users to interact with virtual objects or each other within the XR environment. For example, users can use gestures to interact with user input components in the XR environment. Furthermore, some XR environments allow multiple users to interact with each other within a shared XR environment. However, what is needed are improved techniques for managing user input in shared XR environments. Attached Figure Description
[0003] Figure 1 The flowchart illustrates a technique for adjusting the transmission of virtual representation data according to one or more implementation schemes.
[0004] Figure 2 An example diagram showing a virtual representation of a user according to one or more implementation schemes is presented.
[0005] Figure 3 A flowchart is shown of a method for managing virtual representation data based on a sensitive user input component, according to one or more embodiments.
[0006] Figure 4 An example illustration of a modified live frame representing a virtual representation of a user according to one or more implementation schemes is shown.
[0007] Figure 5 A flowchart is shown of a method for generating modified live frames of virtual representation data according to one or more embodiments.
[0008] Figure 6 A flowchart is shown of a technique for incorporating an eye portion from a reference frame into a live frame, according to one or more embodiments.
[0009] Figure 7 An example network diagram is shown for an electronic device participating in an extended reality coexistence session according to one or more implementation schemes.
[0010] Figure 8 An exemplary system for various XR technologies is shown in block diagram form according to one or more embodiments. Detailed Implementation
[0011] This disclosure relates to systems, methods, and computer-readable media for managing virtual representation data in a shared extended reality environment. Specifically, the embodiments described herein relate to techniques for improving security when using user input components in a shared extended reality environment.
[0012] For the purposes of this specification, the terms “extended reality” or “XR” refer to an environment that is fully or partially simulated.
[0013] For the purposes of this specification, the term "character" refers to a virtual, photorealistic representation of a subject that accurately reflects the subject's physical characteristics, movements, etc., generated based on the subject's tracking data.
[0014] For the purposes of this specification, the term "coexistence session" refers to a virtual communication session in which two or more users are active in a public XR environment. In some implementations, a specific user can view other users in the coexistence session in the form of roles.
[0015] For the purposes of this specification, the term "live frame" refers to a frame representing a user's virtual representation, or a frame of sensor data used to generate a user's virtual representation in real-time or near real-time (e.g., during a coexistence session). Therefore, a live frame reflects the characteristics of the user during the capture of a live frame.
[0016] For the purposes of this specification, the term "reference frame" refers to a frame of image or sensor data captured prior to a live frame. For example, a reference frame may be captured prior to a live frame during a coexistence session, or offline during a registration session, etc.
[0017] Coexistence sessions enable users to interact with each other using virtual representations such as avatars, characters, or photorealistic models, generated from local sensor data captured by electronic devices in the form of tracking data. The tracking data can be used to determine the user's visual and geometric characteristics, based on which a virtual representation of the subject is generated. The virtual representation, or data associated with it, can be sent to other electronic devices participating in the coexistence session, causing the subject to appear as a virtual representation on those other devices.
[0018] In a coexisting session, users can generate user input in several ways, such as virtual or physical user input components, gestures, and gaze. However, some user interactions may involve sensitive information, such as PIN codes, passwords, and personally identifiable information. In such cases, when an unauthorized party uses the movement of a user's virtual representation to infer user input, the transmission of virtual representation data may expose the user's sensitive information to potential eavesdropping, hacking, or keylogging attacks. The implementation described herein opportunistically obfuscates tracked user movements, allowing user input movements to be hidden relative to other users in the coexisting session, thereby providing additional privacy to the local user.
[0019] According to some implementations, sensitive input triggers can be detected based on physical and / or virtual input components that are present in the vicinity of the user and interact with the user. In some implementations, triggers can be detected based on a combination of application context and the presence of a user input component, such as if a user prompt is presented for sensitive user information. As another example, sensitive input triggers can be detected when an input component capable of receiving user input that meets sensitivity criteria, such as predefined categories including personally identifiable information, passwords, security codes, etc. Examples of user input components may include virtual or physical keyboards, keypads, text fields, or other user interface elements or devices that can be used by the user to provide sensitive information.
[0020] According to one or more embodiments, when a sensitive input component is detected or a sensitive input trigger is otherwise activated, the transmission of virtual representation data for the user can be adjusted. For example, the transmission of virtual representation data can be paused. In some embodiments, pausing the transmission of virtual representation data may involve pausing the capture of sensor data (such as camera data) used to generate the virtual representation data. For example, when a sensitive input trigger is active, one or more cameras can be turned off or deactivated.
[0021] In some implementations, when synchronization of presentation status information is paused for a local user, an additional user can continue to interact with elements in the shared session. The local device can provide an indication that synchronization is paused, allowing the additional device to indicate to its respective user that the local user is not experiencing the same representation in the multi-user communication session. Additionally, the local user can continue to receive presentation status information from the remote user and optionally update the local presentation status while synchronization is paused.
[0022] In some implementations, the transmission of virtual representation data can be adjusted by generating a modified live frame containing eye portions from a reference frame of the virtual representation data. In some implementations, the reference frame may be a frame captured or generated during the registration process. The eye portions may include left and right eye portions and may be a single region of the user's virtual representation, or may include separate regions for the left and right eyes. The modified live frame can be generated by identifying the eye portions in a live frame of virtual representation data captured by a camera or other sensor of a local device. The modified live frame can be generated by incorporating eye portions from the reference frame into the live frame based on the eye portions in the live frame. For example, eye portions from the reference frame can be mapped to eye portions in the live frame based on the user's head pose or head position in the live frame. The modified live frame can be provided for display at a remote device, such that the user's eye portions are blurred or replaced by eye portions from the reference frame.
[0023] In the following description, numerous specific details are set forth for purposes of explanation in order to provide a thorough understanding of the disclosed concepts. As part of this description, some of the accompanying drawings of this disclosure are block diagrams representing structures and devices to avoid obscuring the novel aspects of the disclosed concepts. For clarity, not all features of actual specific embodiments may be described. Additionally, some of the drawings of this disclosure are provided in the form of flowcharts as part of this specification. The blocks in any particular flowchart may be presented in a specific order. However, it should be understood that the specific order of any given flowchart is only for illustrative purposes of one embodiment. In other embodiments, any of the various elements depicted in the flowcharts may be omitted, or the illustrated sequence of operations may be performed in a different order, or even simultaneously. Furthermore, other embodiments may include additional steps not depicted as part of the flowcharts. Moreover, the language used in this disclosure has been primarily chosen for readability and instructional purposes and may not have been chosen to define or limit the subject matter of the invention, thereby resorting to the necessary claims to determine such inventive subject matter. In this disclosure, reference to “an implementation” or “implementation” means that a particular feature, structure or characteristic described in connection with the implementation is included in at least one implementation of the disclosed subject matter, and the repeated references to “an implementation” or “implementation” should not be construed as necessarily involving all of the same implementation.
[0024] It should be understood that in any actual implementation of development (as in any software and / or hardware development project), numerous decisions must be made to achieve the developer's specific goals (e.g., compliance with system and business-related constraints), and these goals may differ between different implementations. It should also be understood that such development work can be complex and time-consuming, but nevertheless, it remains routine work for those of ordinary skill in the art who design and implement graphical modeling systems in benefit from this disclosure.
[0025] Figure 1 A flowchart illustrates techniques for adjusting the transmission of virtual representation data according to one or more implementation schemes. Specifically, Figure 1 An example of a technique for adjusting the transmission of virtual representation data between a first device 100 and a second device 105 in response to user interaction with a sensitive input component, according to one embodiment of this disclosure, is illustrated. Although the flowchart illustrates various procedures executed by specific components in a particular order, it should be understood that, according to one or more embodiments, various procedures may be executed by alternative devices or modules. Furthermore, various procedures may be executed in an alternative order, and various combinations of these procedures may be executed simultaneously. Additionally, according to some embodiments, one or more procedures may be omitted, or other procedures may be added.
[0026] The flowchart begins at box 110, where a first device 100 captures local sensor data. This local sensor data may be captured by a camera, microphone, motion sensor, gaze tracker, or some combination thereof. Local sensor data may include, but is not limited to, image data, audio data, depth data, motion data, gaze data, etc. The sensor data can be any data captured by sensors that capture user characteristics, such as cameras, microphones, motion sensors, gaze trackers, or any other sensors that capture the user's current characteristics from a local electronic device. According to one or more embodiments, the local sensor data captured at box 110 may capture the motion or characteristics of the user of the first device 100. For this purpose, the first device 100 may be a head-mounted device or other wearable device, and the local sensor data may be captured by user-facing sensors on the wearable device.
[0027] At box 115, the first device 100 generates first user virtual representation data based on local sensor data collected at box 110. The first user virtual representation can be generated to reflect real-world characteristics of the user of the first device 100, such as appearance, motion, geometry, or volume. The first user virtual representation may include, but is not limited to, avatars, characters, photorealistic models, cartoons, holograms, etc. The first user virtual representation may include, but is not limited to, the first user's facial features, body features, gestures, expressions, movement, voice, clothing, accessories, or other attributes. In some embodiments, the first user virtual representation may be a photorealistic model of the user. The first user virtual representation data generated at box 115 may include the user's virtual representation, or may include data that can be used to generate or render the user's virtual representation, such as tracking data, motion data, appearance data, posture information, facial expression information, etc. In some embodiments, static and dynamic virtual representation data can be used to generate the user's virtual representation. For example, tracking data collected during a coexistence session may be combined with registration data to generate the user's virtual representation. At box 120, the first device 100 transmits the first user virtual representation data to the second device 105.
[0028] Similarly, at box 125, the second device 105 captures local sensor data. This local sensor data may be captured by a camera, microphone, motion sensor, gaze tracker, or some combination thereof. The local sensor data may include, but is not limited to, image data, audio data, depth data, motion data, gaze data, etc. According to one or more embodiments, the local sensor data captured at box 125 may capture the motion or characteristics of the user of the second device 105. For this purpose, the second device 105 may be a head-mounted device or other wearable device, and the local sensor data may be captured by user-facing sensors on the wearable device.
[0029] At block 130, second device 105 generates second user virtual representation data based on local sensor data collected at block 125. The second user virtual representation can be generated to reflect real-world characteristics of the user of second device 105, such as appearance, movement, geometry, or volume. The second user virtual representation may include, but is not limited to, the first user's facial features, body features, gestures, expressions, movements, voice, clothing, accessories, or any other suitable attributes. At block 135, second device 105 sends the second user virtual representation data to first device 100. Thus, as shown in time block 140, first device 100 and second device 105 continuously provide virtual representation data to each other. For example, this may occur when first device 100 and second device 105 are active in a public coexistence session. For example, first device 100 and second device 105 may share at least a portion of an extended reality environment.
[0030] Although virtual representation data is shared between the first device 100 and the second device 105 during time block 140, at block 145, the flowchart includes the first device 100 rendering a second user virtual representation based on second user virtual representation data received from the second device 105. This may include generating and / or rendering the avatar or persona of the user of the second device to reflect the characteristics of the user of the second device during the coexistence session. Similarly, at block 150, the second device 105 renders a first user virtual representation based on first user virtual representation data received from the first device 100. [Go to...] Figure 2 An example is shown where a first device 100 captures sensor data of a user 200A to determine the current characteristics of the user 200A and sends corresponding virtual representation data to a second device 105. The second device 105 then renders a view of a character 210A that reflects the characteristics of the user 200A.
[0031] return Figure 1 The flowchart proceeds to box 155, where the first device 100 detects a sensitive user input trigger. According to some embodiments, a sensitive user input trigger can be detected when a user input component capable of providing user input with a predefined sensitivity classification is detected or provided. For example, sensitivity classification can be applied to inputs that may include or transmit passwords, credit card numbers, personal messages, personally identifiable information, health information, or other personal or confidential data. Sensitivity classifications can be predefined, defined by the application providing the user input component, user-defined, or some combination thereof. Sensitive input components can be physical or virtual, such as physical or virtual keyboards, keypads, etc. Furthermore, sensitive user input triggers can also be detected based on context or application state (such as the state of a running application). For example, if a user input field marked as a sensitive field is presented, user interaction with the user input component to input data into the field can be a sensitive user input trigger. In some embodiments, user interaction can be based on the user's gaze targeting the input component within a predefined time period, determining that the user is currently, or predicting that the user will interact with the input component (e.g., based on hand proximity), or some combination thereof.
[0032] In response to the detection of such a trigger, the flowchart proceeds to block 160, and the first device adjusts the transmission of first user virtual representation data. Adjusting the transmission of virtual representation data may involve modifying the transmission itself, such as suspending the transmission of some or all of the virtual representation data generated by device 100, or modifying the data to be transmitted. Optionally, as shown in time block 165, adjusting the transmission may include stopping the transmission of virtual representation data. That is, virtual representation data may be generated by the first device 100, but transmission may be suspended. In some embodiments, adjusting the transmission of virtual representation data may include suspending the generation of virtual representation data by the first device 100, or suspending sensor data collection for the user of the first device 100, such that virtual representation data is not generated and therefore not transmitted to the second device 105. Thus, at block 170, the second device 105 stops receiving first user representation data or receives reduced first user representation data. This is shown in time block 165, where virtual representation data is transmitted from the second device 105 to the first device 100, but not from the first device 100 to the second device 105. Alternatively, the second device may receive the modified first user representation data. The first user's representation data can be modified so that the eye region is modified based on the actual movement of the first user's eyes.
[0033] At box 175, the second device 105 adjusts the presentation of the first user virtual representation. For example, at least a portion of the virtual representation may be paused, or may appear inconsistent with the current characteristics of the user of the first device 100. In some embodiments, the second device 105 may additionally apply visual processing to the first user virtual representation to signal that the first user virtual representation is in paused mode, or to obfuscate at least a portion of the virtual representation from which sensitive user input, such as eyes, hands, fingers, etc., can be derived or inferred.
[0034] Return to Figure 2 For example, user 200B is shown interacting with input component 215 by scanning a virtual keypad to enter a code. According to one or more embodiments, interaction with the virtual keypad may satisfy sensitive user input triggers. Therefore, the first device 100 may adjust the transmission of virtual representation data of the user of the first device 100. Thus, the second device 105 shows character 210B whose facial expression no longer reflects the facial expression of user 200B of the first device 100. This is because the virtual representation data of user 200B is in a paused mode at the first device 100. Similarly, when user 200C continues to use input component 215B, character 210C at the second device 105 remains in a paused mode. Although not shown, the second device 105 may continue or not continue sending virtual representation data to the first device 100. Furthermore, the first device 100 may or may not present the current virtual representation of the user of the second device while in a paused mode.
[0035] return Figure 1 As shown in box 180, the first device 100 can detect the completion of sensitive user input. This can be determined, for example, when the user stops interacting with the user input component (e.g., within a predefined amount of time), or when the sensitive user input component is no longer detected. As another example, the completion of sensitive user input can be determined when the sensitive input text box is no longer displayed. As yet another example, the user can definitively indicate that sensitive user input has stopped, for example, based on input to a confirmation button, submit button, gestures, voice commands, etc. Furthermore, in some implementations, the completion of sensitive user input can be determined based on a timeout.
[0036] In response to determining that sensitive user input has been completed, the flowchart proceeds to block 185, and the first device 100 resumes continuous transmission of first user virtual representation data. Resuming transmission may involve restarting the capture of sensor data of the user of the first device 100, and / or generating virtual representation data of the user of the first device 100. Therefore, as shown in time block 190, transmission between the first device 100 and the second device 105 resumes, such that the second device 105 resumes receiving virtual representation data from the first device 100. Alternatively, resuming continuous transmission of the first user virtual representation data may include adjusting the transmitted virtual representation so that the virtual representation data represents the current characteristics of the local user (such as gaze).
[0037] The flowchart ends at box 195, where the second device 105 resumes rendering the first user virtual representation based on the continuously received first user virtual representation data. That is, the second device 105 resumes rendering the user's role or other virtual representation of the first device 100 in a manner consistent with the user characteristics during the coexistence session. In some implementations, a transition effect may be presented when rendering resumes. For example, one or more intermediate frames may be generated to transition the paused role to the resumed role.
[0038] Return to Figure 2 For example, user 200D no longer interacts with the input component. Therefore, first device 100 can restart the transmission of virtual representation data. Consequently, second device 105 presents character 210D, which matches the appearance of user 200D and is generated based on the virtual representation received from first device 100 that captures sensor data of user 200D.
[0039] Figure 3This is a flowchart 300 illustrating an example of a technique for adjusting the transmission of virtual representation data in response to user interaction with a sensitive input component, according to one embodiment of this disclosure. It should be understood that the various processes described may be performed in different orders, and some processes may be performed in parallel. Furthermore, according to some embodiments, not all processes may be required. Therefore, the boxes depicted and / or described as optional only indicate that some embodiments may involve performing the actions described in the boxes, while other embodiments may not.
[0040] Flowchart 300 begins at box 305, where a coexistence session is initiated. A coexistence session can be a virtual communication session in which two or more devices share at least a portion of a public XR environment. According to some implementations, a coexistence session may include virtual components, such as a virtual representation for each user. The coexistence session can be initiated by a user's electronic device, a server, or any other suitable device.
[0041] The flowchart proceeds to box 310, where a determination is made regarding whether a sensitive input component has been detected. According to one or more embodiments, a sensitive input component can be a physical or virtual input component capable of providing data classified as sensitive data. The determination can be based on the characteristics of the input component or in combination with other factors such as an open window or other contextual information. Specific parameters used to determine whether an input component is a sensitive input component can be predefined or defined by a specific application, such that the same input component may be a sensitive input component when used in one application but not when used in another. Furthermore, an input component can be classified as a sensitive input component based on user-defined parameters, system-defined parameters, or some combination thereof.
[0042] If a sensitive input component is detected at box 310, optionally, flowchart 300 proceeds to decision box 320 and determines whether a user interaction with the sensitive user input component has been detected. The user interaction can be an action performed by the user to generate user input using the sensitive input component. In some embodiments, the user interaction can be an observed or detected user interaction, for example, based on image data or other sensor data, based on input received by the input device, etc. In some embodiments, the user interaction can be a predicted user interaction based on user tracking data. As an example, a user interaction can be detected if the user or one or both of the user's hands move within a predefined distance from and / or toward the sensitive user input component. If no user interaction is detected at box 320, or if the flowchart returns to box 310 if no sensitive input component is detected, the flowchart proceeds to box 325, and the local device continues to send virtual representation data. As described above, this may include capturing tracking data of the local user, using the tracking data to generate virtual representation data of the local user, and sending the virtual representation data to another device active in the coexistence session. The virtual representation data may include data from which a virtual representation of the user is generated or rendered.
[0043] Returning to box 310, if sensitive input is detected, and optionally, at box 320, user interaction with the input component is detected, flowchart 300 proceeds to box 330. At box 330, the transmission of virtual representation data is adjusted by the local device. Adjusting the transmission data may involve modifying the transmission itself, such as pausing transmission or modifying the data to be transmitted. Optionally, as shown in box 335, adjusting the transmission of virtual representation data may include stopping the capture of sensor data. Sensor data may be any data captured by any other sensor of the user's current characteristics by a camera, microphone, motion sensor, gaze tracker, or the capturing electronic device of the local electronic device. Additionally, optionally, as shown in box 340, adjusting the transmission of virtual representation may involve stopping the transmission of virtual representation data. That is, the virtual representation data may be generated by the local device, but the transmission may be paused.
[0044] The transmission of virtual representation data can be adjusted within a predefined time period until user interaction is complete, until user input is confirmed, or based on another criterion or a combination thereof. In one example, as shown in flowchart 300, it can be determined whether a sensitive input component is still detected at box 310, and the flowchart can continue the adjusted transmission of virtual representation data until the sensitive input component is no longer detected at box 310, or optionally, until user interaction with the sensitive input component is no longer detected at box 320. Then, flowchart 300 ends at box 325, and virtual representation data is transmitted without adjustment.
[0045] According to some implementation schemes, adjusting the transmission of virtual representation data may involve modifying the live frames of the virtual representation data to confuse at least a portion of the user, such as the eyes, mouth, etc. Figure 4 An example of a technique for adjusting the transmission of virtual representation data in response to user interaction with a sensitive input component, according to one embodiment of this disclosure, is illustrated. It should be understood that the various processes described may be executed in different orders, and some processes may be executed in parallel. Furthermore, according to some embodiments, not all processes may be required.
[0046] Flowchart 400 begins at box 405, where a coexistence session is initiated. A coexistence session can be a virtual communication session in which two or more devices share at least a portion of a public XR environment. According to some implementations, a coexistence session may include virtual components, such as a virtual representation for each user. The coexistence session can be initiated by a user's electronic device, a server, or any other suitable device.
[0047] Flowchart 400 proceeds to box 410, where sensor data of the local user is captured. Sensor data may include any data captured by sensors such as cameras, microphones, motion sensors, gaze trackers, or any other sensors that capture the user's current characteristics. This data may include image data, audio data, depth data, motion data, gaze data, or similar types of information that can be used to generate a virtual representation of the user. At box 415, a live frame of virtual representation data is generated from the sensor data. The live frame may include, for example, sensor data from which a virtual representation of the user can be generated, reflecting the current visual characteristics of the tracked user. For example, the live frame may include 2D or 3D representation data of the user, such as geometric data, texture data, image data, or other data from which a virtual representation can be generated, such as in the form of a character.
[0048] Go to Figure 5 An example is shown where a first device 100 captures sensor data of user 500A to determine the current characteristics of user 500A and sends corresponding virtual representation data to a second device 105. The second device 105 then renders a view of role 510A that reflects the characteristics of user 500A. Thus, real-time data is represented as role 510A at the second device 105.
[0049] Return to Figure 4Flowchart 400 proceeds to box 420, where a determination is made regarding whether a sensitive input component has been detected. According to one or more embodiments, a sensitive input component can be a physical or virtual input component capable of providing data classified as sensitive data. The determination can be based on the characteristics of the input component or in combination with other factors such as open windows or other contextual information. Specific parameters used to determine whether an input component is a sensitive input component can be predefined or defined by a specific application, such that the same input component may be a sensitive input component when used in one application but not when used in another. Furthermore, an input component can be classified as a sensitive input component based on user-defined parameters, system-defined parameters, or some combination thereof.
[0050] If a sensitive input component is detected at box 410, optionally, flowchart 400 proceeds to decision box 425 and determines whether a user interaction with the sensitive user input component has been detected. The user interaction can be an action performed by the user to generate user input using the sensitive input component. In some embodiments, the user interaction can be an observed or detected user interaction, for example, based on image data or other sensor data, based on input received by the input device, etc. In some embodiments, the user interaction can be a predicted user interaction based on user tracking data. As an example, a user interaction can be detected if the user or one or both of the user's hands move within a predefined distance from and / or toward the sensitive user input component. If no user interaction is detected at box 425, or if the flowchart returns to box 420 if no sensitive input component is detected, flowchart 400 proceeds to box 430, and the local device continues to send virtual representation data. As described above, this may include capturing tracking data of the local user, using the tracking data to generate virtual representation data of the local user, and sending the virtual representation data to another device active in the coexistence session. Virtual representation data may include data from which a virtual representation of the user is generated or rendered.
[0051] Returning to box 420, if sensitive input is detected, and optionally, at box 425, user interaction with the input component is detected, flowchart 400 proceeds to box 435. At box 435, the eye portion of a reference frame is retrieved. According to one or more embodiments, the reference frame may be a frame of the user's virtual representation captured prior to the current live frame. In some embodiments, the reference frame may include only the eye region, or may include more facial features from which the eye region can be retrieved. In some embodiments, the eye portion of the reference frame may be predefined and may be generated and stored prior to the coexistence session. For example, during a registration period, a local user may use their device to capture sensor data of their face to generate character data for driving the virtual representation during the coexistence session. The eye portion may be a single, continuous region of the face containing both eyes, or may include a combination of portions of the virtual representation data corresponding to the eyes, eyeballs, pupils, and irises.
[0052] Flowchart 400 proceeds to box 440, where the eye portion of the reference frame is incorporated into the live frame to generate the modified frame. The eye portion can be incorporated in various ways. For example, the eye region of the live frame can be extracted and replaced with the reference eye region. As another example, a composite frame can be generated by increasing the transparency of the eye region in the live frame and overlaying the reference eye portion, making the eye region in the live frame invisible in the adjusted frame. In some implementations, the reference eye region and the live frame eye region can be aligned, for example, based on head pose data such as head position, eye tracking data, etc. References will be discussed below. Figure 6 The various techniques used to incorporate reference eye portions into the live frame are described in more detail.
[0053] Flowchart 400 proceeds to box 445, where a modified frame providing a virtual representation of the local user is provided for rendering at a remote device. As described above, the modified frame may include data from which a 3D representation of the user can be generated and / or rendered. The modified frame may be sent to a second device, and / or made available to additional devices in a coexistence session.
[0054] Return to Figure 5For example, user 500B is shown interacting with input component 515A by scanning a virtual keypad to enter codes. According to one or more embodiments, interaction with the virtual keypad may satisfy sensitive user input triggers. Therefore, the first device 100 may adjust the transmission of virtual representation data of user 500B. Specifically, a reference eye region 525 may be obtained, for example, from reference frame 520. In some embodiments, the reference eye region 525 may be extracted from reference frame 520 during runtime. Alternatively, the reference eye region 525 may be previously extracted and stored, such as during the registration process of user 500. Device 100 may replace the eye region with a replacement eye region 530A to generate character 510B. Therefore, the replacement eye region 530A is presented to the user in a manner that blurs the true gaze direction of user 500B. Thus, the second device 105 shows character 510B whose eyes no longer reflect the eyes of user 500B of the first device 100, although other features of the user may be presented in a consistent manner, such as head orientation, mouth movement, etc. In this scenario, the eyebrows of character 510B are displayed as a reflection of the eyebrows of user 400B, despite the different gaze direction. Similarly, as user 500C continues to use input component 515B, character 510C at second device 105 continues to reflect reference eye area 525 as alternative eye area 530B, while other features of character 410C (such as eyebrows, lips, etc.) continue to reflect the movement of user 400C.
[0055] The transmission of virtual representation data can be adjusted within a predefined time period until user interaction is complete, until user input is acknowledged, or based on another criterion or a combination thereof. In one example, as shown in flowchart 400, it can be determined whether a sensitive input component is still detected at box 420, and the flowchart can continue the adjusted transmission of virtual representation data until the sensitive input component is no longer detected at box 420, or optionally, until user interaction with the sensitive input component is no longer detected at box 425. Flowchart 400 then ends at box 430, providing a live frame of virtual representation data without adjustment.
[0056] Return to Figure 5 For example, user 500D no longer interacts with the input components. Therefore, the first device 100 can restart the transmission of virtual representation data. Consequently, the second device 505 presents character 510D in a manner that matches the appearance of user 500D and is generated based on the virtual representation received from the first device 100 that captures sensor data of user 500D. Specifically, the eye area of character 510D now reflects the eye area of user 500D.
[0057] Figure 6This is a flowchart illustrating an example of a technique, according to some implementations, for generating a modified live frame of virtual representation data for a user in response to detecting user interaction with a sensitive input component. Specifically, relative to... Figure 6 The described technique involves incorporating a reference eye portion into a live frame to generate a modified frame, as described above relative to... Figure 5 Box 540 provides a general description. It should be understood that the various processes described can be executed in different orders, and some processes can be executed in parallel. Furthermore, depending on some implementation schemes, not all processes may be required.
[0058] The flowchart begins at box 605, where the electronic device detects one or more facial landmarks in a live frame of virtual representation data. The live frame of virtual representation data can be generated based on sensor data captured from the user, such as image data, depth data, motion data, gaze data, etc. Therefore, the live frame can include a visual representation of the user. Facial landmarks can include, but are not limited to, points or regions corresponding to the user's eyes, nose, mouth, eyebrows, chin, or other facial features, and can be detected in two or three dimensions. Facial landmark detection can be performed using any suitable computer vision technique, such as face detection, face alignment, face recognition, feature detection, etc.
[0059] At frame 610, the electronic device identifies the eye region in the live frame based on one or more facial landmarks. The eye region may include, for example, portions of the live frame including the user's left eye, right eye, or both eyes. Identification of the eye region can be performed using any suitable geometric or spatial techniques, such as bounding boxes, contours, masks, etc. In some embodiments, the eye region may be a continuous region or may consist of multiple distinct regions (such as left and right eye portions). In some embodiments, the region may include the eyeball, iris, and pupil, etc. Furthermore, in some embodiments, the eye region may be defined excluding eyelids, such that the virtually represented eyelids remain consistent with the live frame.
[0060] The flowchart continues at box 615, where the electronic device determines the head pose in the live frame. The head pose may include, but is not limited to, the orientation, rotation, or position of the user's head in the live frame. The determination of the head pose may be performed based on sensor data (such as image data, for example, using visual-inertial techniques) and / or motion data (such as data captured by accelerometers, IMUs, etc.).
[0061] At box 620, the electronic device maps the eye region from a reference frame of the virtual representation data to the eye region in the live frame based on head pose. The reference frame of the virtual representation data may be obtained during the registration process at the electronic device and, in some embodiments, may include data for generating or rendering a neutral or static facial expression of the user. Alternatively, the reference frame may be any previous frame of the virtual representation data and may include at least the eye region. Mapping may include, but is not limited to, aligning, transforming, distorting, or projecting the eye region from the reference frame to the eye region in the live frame such that the eye region in the reference frame matches the eye region in the live frame in terms of size, shape, position, orientation, etc.
[0062] The flowchart proceeds to box 625, where the electronic device performs an alpha blending technique on the eye region of the reference frame and the live frame based on the mapping. The alpha blending technique may include, but is not limited to, using a weighted average to combine the pixel values of the eye region in the neutral reference frame and the eye region in the live frame, such that the appearance of the eye region in the live frame is reduced and the appearance of the eye region in the neutral reference frame is increased.
[0063] The flowchart ends at box 630, where the electronic device applies a smoothing operation to the blended frame. The smoothing operation may include, but is not limited to, reducing noise, artifacts, or discontinuities in the blended frame so that the transition between the eye region in the neutral reference frame and the rest of the live frame is smooth and natural. The smoothing operation can be performed using any suitable image processing technique, such as filtering, blurring, interpolation, etc.
[0064] See Figure 7This document depicts a simplified block diagram of an electronic device 100 communicatively connected to an additional electronic device 105 via a network 715, according to one or more embodiments of the present disclosure. The electronic device 100 may be part of a multi-functional device such as a mobile phone, tablet computer, personal digital assistant, portable music / video player, wearable device, head-mounted system, projection-based system, base station, laptop computer, desktop computer, network device, or any other electronic system such as those described herein. The electronic device 100, the additional electronic device 105, and / or network storage device may additionally or alternatively include one or more additional devices (such as server devices, base stations, accessory devices, etc.), in which various functionalities may be included, or various functionalities may be distributed across these additional devices. Exemplary networks (such as network 715) include, but are not limited to, local area networks (such as Universal Serial Bus (USB) networks), organizational LANs, and wide area networks (such as the Internet). According to one or more embodiments, the electronic device 100 is used to participate in multi-user communication sessions (such as coexistence sessions) in an XR environment. It should be understood that the various components and functions within the electronic device 100, the additional electronic device 105, and the network storage device may be distributed differently on the device or on the additional device.
[0065] Electronic device 100 may include one or more processors 725, such as a central processing unit (CPU). Processor 725 may include a system-on-a-chip (such as a system-on-a-chip present in a mobile device) and may include one or more dedicated graphics processing units (GPUs). Additionally, processor 725 may include multiple processors of the same or different types. Electronic device 100 may also include memory 735. Memory 735 may include one or more different types of memory that can be used in conjunction with processor 725 to perform device functions. For example, memory 735 may include cache, ROM, RAM, or any kind of transient or non-transitory computer-readable storage medium capable of storing computer-readable code. Memory 735 may store various programming modules for execution by processor 725, including XR module 765, tracing module 770, and various other application programs 775. Electronic device 100 may also include storage device 730. Storage device 730 may include one or more non-transitory computer-readable media, including, for example, magnetic disks (fixed hard disks, floppy disks, and removable disks) and magnetic tapes, optical media (such as CD-ROMs and digital video optical discs (DVDs)), and semiconductor storage devices (such as electrically programmable read-only memory (EPROM) and electrically erasable programmable read-only memory (EEPROM)). According to one or more embodiments, storage device 730 may be configured to store virtual representation data 760. Electronic device 100 may be attached to include a network interface 750 from which additional network components can be accessed via network 715.
[0066] The electronic device 100 may also include one or more cameras 740 or other sensors 745 (such as depth sensors) to determine the depth or other characteristics of the environment. In one or more embodiments, each of the one or more cameras 740 may be a conventional RGB camera or a depth camera. Furthermore, the cameras 740 may include stereo cameras or other multi-camera systems, time-of-flight camera systems, etc. The cameras 740 may include one or more user-facing cameras, one or more scene-facing cameras, or some combination thereof.
[0067] Electronic device 100 may also include a display 755. The display device 755 may utilize digital light projection, OLED, LED, uLED, liquid crystal on silicon, laser scanning light source, or any combination of these technologies. The display device 755 can be used to present a representation of a multi-user communication session, including shared virtual elements and other XR objects within the multi-user communication session. The display 755 may have an opaque, transparent, or translucent display. The transparent or translucent display may have a medium through which light is guided to the user's eyes. Optical waveguides, optical reflectors, holographic media, optical combiners, combinations thereof, or other similar technologies may be used for the medium. In some embodiments, the transparent or translucent display may be selectively controlled to become opaque. Projection-based systems may utilize retinal projection techniques that project graphic images onto the user's retina. Projection systems may also project virtual objects into a physical environment (e.g., as a hologram or projected onto a physical surface).
[0068] Storage device 730 can be used to store various data and structures that can be used to provide state information for tracking application and session states. Storage device 730 may include, for example, a virtual representation data repository 760. Virtual representation data repository 760 can be used to store information to be used to generate virtual representations for local users (such as static virtual representation data generated during registration periods, user-specific models, etc.).
[0069] According to one or more embodiments, memory 735 may include one or more modules comprising computer-readable code executable by processor 725 to perform functions. The memory may include, for example, a tracking module 770 configured to determine characteristics of a local user based on sensor data captured by electronic device 100 (such as camera 740, sensor 745, etc.). Memory 735 may also include an XR module 765, which can be used to provide coexistence sessions in an XR environment. In some embodiments, XR module 765 may, for example, use tracking data from tracking module 770 and data from virtual representation data 760 to generate a virtual representation of the local user.
[0070] In some implementations, the transmission of virtual representation data may be paused or adjusted based on detected sensitive input components (such as virtual input components associated with application 775 and / or physical components detected, for example, by camera 740) or other signals sent to or received by electronic device 100. Virtual representation data may be sent to additional electronic device 105, allowing additional electronic device 105 to use the virtual representation data to present a virtual representation of the user of electronic device 100.
[0071] Although electronic device 100 is depicted as including the numerous components described above, in one or more embodiments, the various components may be distributed across multiple devices. Therefore, while certain calls and transmissions are described herein with respect to the specific system depicted, in one or more embodiments, various calls and transmissions may be directed differently based on the functions of different distributions. Additionally, additional components may be used, and certain combinations of the functions of any components may be combined.
[0072] Now for reference Figure 8 This document illustrates a simplified functional block diagram of an exemplary multi-functional electronic device 800 according to one embodiment. The electronic device may be a multi-functional electronic device, or may have some or all of the components described herein. The multi-functional electronic device 800 may include a processor 805, a display 810, a user interface 815, graphics hardware 820, device sensors 825 (e.g., proximity / ambient light sensors, accelerometers, and / or gyroscopes), a microphone 830, an audio codec 835, a speaker 840, communication circuitry 845, digital image capture circuitry 850 (e.g., including a camera system), a memory 860, a storage device 865, and a communication bus 870, among other combinations. The multi-functional electronic device 800 may be, for example, a mobile phone, a personal music player, a wearable device, and a tablet computer.
[0073] Processor 805 can execute necessary instructions to implement or control the operation of various functions performed by device 800. Processor 805 may, for example, drive display 810 and receive user input from user interface 815. User interface 815 allows a user to interact with device 800. For example, user interface 815 may take various forms, such as buttons, keypad, dial pad, click wheel, keyboard, display screen, or touchscreen. Processor 805 may also be a system-on-a-chip (such as those present in mobile devices) and include a dedicated graphics processing unit (GPU). Processor 805 may be based on a Reduced Instruction Set Computer (RISC) or Complex Instruction Set Computer (CISC) architecture or any other suitable architecture and may include one or more processing cores. Graphics hardware 820 may be dedicated computing hardware for processing graphics and / or assisting processor 805 in processing graphics information. In one embodiment, graphics hardware 820 may include a programmable GPU.
[0074] Image capture circuit 850 may include one or more lens assemblies, such as 880A and 880B. The lens assemblies may have various combinations of characteristics, such as different focal lengths. For example, lens assembly 880A may have a shorter focal length relative to the focal length of lens assembly 880B. Each lens assembly may have a separate associated sensor element 890. Alternatively, two or more lens assemblies may share a common sensor element. Image capture circuit 850 may capture still images, video images, and enhanced images, etc. The output from image capture circuit 850 may be processed at least in part by video codec 855 and / or processor 805 and / or graphics hardware 820 and / or a dedicated image processing unit or pipeline incorporated within circuit 845. Images captured in this way may be stored in memory 860 and / or storage device 865.
[0075] Memory 860 may include one or more different types of media used by processor 805 and graphics hardware 820 to perform device functions. For example, memory 860 may include memory cache, read-only memory (ROM), and / or random access memory (RAM). Storage device 865 may store media (e.g., audio files, image files, and video files), computer program instructions or software, preference information, device configuration file information, and any other suitable data. Storage device 865 may include one or more non-transitory computer-readable storage media, including, for example, magnetic disks (fixed hard disks, floppy disks, and removable disks) and magnetic tape, optical media (such as CD-ROMs and digital video optical discs (DVDs)), and semiconductor storage devices (such as electrically programmable read-only memory (EPROM) and electrically erasable programmable read-only memory (EEPROM)). Memory 860 and storage device 865 can be used to tangibly hold computer program instructions or computer-readable code organized into one or more modules and written in any desired computer programming language. When executed by, for example, processor 805, such computer program code can implement one or more of the methods described herein.
[0076] Humans can interact with and / or perceive the physical environment or physical world without the aid of electronic devices. The physical environment can include physical features, such as physical objects or surfaces. An example of a physical environment is a physical forest that includes physical plants and animals. Humans can directly perceive and / or interact with the physical environment through various means, such as hearing, vision, taste, touch, and smell. In contrast, humans can use electronic devices to interact with and / or perceive a fully or partially simulated extended reality (XR) environment. This XR environment can include mixed reality (MR) content, augmented reality (AR) content, virtual reality (VR) content, and so on. Using an XR system, a person's physical movements, or some of their representations, can be tracked, and in response, the characteristics of virtual objects simulated in the XR environment can be adjusted in a manner consistent with at least one physical law. For example, the XR system can detect movement of the user's head and adjust the graphical and auditory content presented to the user (similar to how such views and sounds change in a physical environment). For example, the XR system can detect movement of electronic devices (e.g., mobile phones, tablets, laptops, etc.) presenting the XR environment and adjust the graphical and auditory content presented to the user (similar to how such views and sounds change in a physical environment). In some cases, the XR system can adjust the features of the graphical content in response to other inputs such as representations of physical motion (e.g., voice commands).
[0077] Many different types of electronic systems enable users to interact with and / or perceive XR environments. A non-exclusive list of examples includes head-up displays (HUDs), head-mounted systems, projection-based systems, windows or vehicle windshields with integrated display capabilities, displays formed as lenses placed over the user's eyes (e.g., contact lenses), head-mounted receivers / headsets, input systems with or without haptic feedback (e.g., wearable or handheld controllers), speaker arrays, smartphones, tablets, and desktop / laptop computers. Head-mounted systems may have opaque displays and one or more speakers. Other head-mounted systems may be configured to accept opaque external displays (e.g., smartphones). Head-mounted systems may include one or more image sensors for capturing images or video of the physical environment and / or one or more microphones for capturing audio of the physical environment. Head-mounted systems may have transparent or semi-transparent displays instead of opaque displays. Transparent or semi-transparent displays may have a medium through which light is directed to the user's eyes. Displays can utilize various display technologies, such as uLED, OLED, LED, liquid crystal on silicon, laser scanning light sources, digital light projection, or combinations thereof. Optical waveguides, optical reflectors, holographic media, optical combiners, or combinations thereof, or other similar technologies can be used as the medium. In some implementations, transparent or translucent displays can be selectively controlled to become opaque. Projection-based systems can utilize retinal projection technology, which projects graphic images onto a user's retina. Projection systems can also project virtual objects into the physical environment (e.g., as holograms or onto physical surfaces).
[0078] The technologies defined in this document take into account options for obtaining and using users' personal information. For example, such personal information may be provided during multi-user communication sessions on electronic devices. However, with regard to the collection of such personal information, it should be obtained with the user's informed consent, so that the user is aware of their personal information and can control the use of their personal information.
[0079] Parties with access to personal information will use it solely for legitimate and reasonable purposes and will comply with privacy policies and practices that are at least in accordance with appropriate laws and regulations. Furthermore, such policies should be comprehensive, user-accessible, and considered to meet or exceed government / industry standards. In addition, personal information will not be distributed, sold, or otherwise shared except for any reasonable and legitimate purpose.
[0080] However, users can limit the extent to which parties can access their personal information. The processes and devices described herein can allow for changes to settings or other preferences that enable users to control access to their personal information. Furthermore, while some of the characteristics defined herein are described in the context of the use of personal information, aspects of these characteristics can be implemented without the need for the use of such information. As an example, a user's personal information can be masked or otherwise generalized so that the information does not identify the specific user from whom the information is obtained.
[0081] It should be understood that the above description is intended to be illustrative and not restrictive. Material has been presented to enable any person skilled in the art to make and use the disclosed subject matter protected by the claims and to provide the material in the context of specific embodiments, variations of which will be readily apparent to those skilled in the art (e.g., some embodiments of the disclosed embodiments may be used in combination with each other). Therefore, Figures 1 to 6 The specific arrangement of the steps or actions shown or Figures 7 to 8 The arrangement of elements shown should not be construed as limiting the scope of the disclosed subject matter. Therefore, the scope of the invention should be determined by reference to the appended claims and the full scope of their equivalents. In the appended claims, the terms “including” and “in which” are used as common English equivalents to the corresponding terms “comprising” and “wherein”.
Claims
1. A method, the method comprising: Detect user interactions with sensitive input components performed by a first user on a first device; as well as In response to detecting user interaction with the sensitive input component, the transmission of virtual representation data corresponding to the first user to the second device is adjusted. The first device and the second device are active in the virtual communication session.
2. The method according to claim 1, further comprising: The input component is determined to be a sensitive input component based on a predefined classification of the input component by the corresponding application.
3. The method according to claim 1, further comprising: The sensitivity of an input component is determined based on the application state of the corresponding application.
4. The method of claim 1, wherein the sensitive input component includes a virtual input component.
5. The method of claim 1, wherein the sensitive input component includes a physical input component.
6. The method according to any one of claims 4 to 5, wherein detecting the user interaction comprises: Determine that the first user's gaze is directed at the sensitive input component within a predefined time period.
7. The method of claim 6, wherein detecting the user interaction further comprises: Determine if the user interacts with the sensitive input component to generate user input.
8. A non-transitory computer-readable medium comprising computer-readable code executable by one or more processors to perform the following operations: Detecting user interactions with sensitive input components performed by a first user on a first device; and In response to detecting user interaction with the sensitive input component, the transmission of virtual representation data corresponding to the first user to the second device is adjusted. The first device and the second device are active in the virtual communication session.
9. The non-transitory computer-readable medium of claim 8, further comprising computer-readable code for performing the following operations: The input component is determined to be a sensitive input component based on a predefined classification of the input component by the corresponding application.
10. The non-transitory computer-readable medium of claim 8, further comprising computer-readable code for performing the following operations: The sensitivity of an input component is determined based on the application state of the corresponding application.
11. The non-transitory computer-readable medium of claim 8, wherein the sensitive input component includes a virtual input component.
12. The non-transitory computer-readable medium of claim 8, wherein the computer-readable code for adjusting the transmission of the virtual representation data further comprises computer-readable code for performing the following operations: Pause the capture of camera data that generates the virtual representation data.
13. The non-transitory computer-readable medium of claim 8, wherein the computer-readable code for adjusting transmission further comprises computer-readable code for performing the following operations: Suspend the transmission of at least a portion of the virtual representation data.
14. The non-transitory computer-readable medium of claim 8, wherein the virtual representation data includes data from which a photorealistic representation of the first user is generated.
15. A system comprising: One or more processors; and One or more computer-readable media, the one or more computer-readable media including computer-readable code executable by the one or more processors to perform the following operations: Detect user interactions with sensitive input components performed by a first user on a first device; as well as In response to detecting user interaction with the sensitive input component, the transmission of virtual representation data corresponding to the first user to the second device is adjusted. The first device and the second device are active in the virtual communication session.
16. The system of claim 15, further comprising computer-readable code for performing the following operations: The input component is determined to be a sensitive input component based on a predefined classification of the input component by the corresponding application.
17. The system of claim 15, further comprising computer-readable code for performing the following operations: The sensitivity of an input component is determined based on the application state of the corresponding application.
18. The system of claim 15, wherein the sensitive input component includes a virtual input component.
19. The system of claim 15, wherein the sensitive input component includes a physical input component.
20. The system of claim 15, wherein the computer-readable code for adjusting the transmission of the virtual representation data further comprises computer-readable code for performing the following operations: Pause the capture of camera data that generates the virtual representation data.