Estimation of expressions for head-mounted devices using low profile antennas and impedance specialty sensing
By using low-profile slot antennas and cross-polarized antenna systems in head-mounted devices, combined with machine learning models, users' facial and hand gestures can be captured non-contactly, solving the accuracy and privacy issues of facial expression and hand gesture expression in existing technologies, and achieving efficient gesture recognition and silent capture in virtual reality environments.
Patent Information
- Application Number
- CN202480061673.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-09-29
- Filing Date
- 2024-09-25
- Publication Date
- 2026-05-26
AI Technical Summary
Existing technologies are insufficient in terms of accuracy and efficiency to effectively track or determine the facial expressions and hand gestures of electronic device users, especially in virtual reality and augmented reality environments, where privacy issues and device accessibility limitations exist.
Employing a low-profile slot antenna and cross-polarized antenna system in a head-mounted device, the system non-contactly captures the 3D position and posture configuration of the user's face and hands by measuring the self-resonant frequency and performance changes of the RF signal, combined with a machine learning model. It also utilizes S-parameter and impedance characteristic sensing technology to reduce environmental and covering interference.
It achieves highly accurate and low-privacy facial and hand gesture capture in virtual reality and augmented reality environments, supports the recognition of silent expressions and gestures, and improves the physical robustness of the device and the user experience.
Smart Images

Figure CN122095336A_ABST
Abstract
Description
Cross-reference to related applications
[0001] This application claims the benefit of U.S. Provisional Application Serial No. 63 / 541,601, filed on September 29, 2023, the full text of which is incorporated herein by reference. Technical Field
[0002] This disclosure relates in general to systems, methods, and apparatus for determining the facial and / or hand posture configuration of a user of an electronic device, such as identifying the three-dimensional (3D) positions of points on the cheeks, chin, lips, tongue, hands in front of the mouth, etc. of a user of a head-mounted device (HMD). Background Technology
[0003] Existing technologies for tracking or otherwise determining the facial expressions, facial features, hand gestures, etc., of users of electronic devices can be improved in terms of accuracy, efficiency, and other attributes. Summary of the Invention
[0004] The various specific embodiments disclosed herein include devices, systems, and methods for identifying the 3D positions of points on the surface or tissue geometry of a user's facial deformities, hand gesture configurations (e.g., in front of the mouth), and / or the user's head / face and / or hands. Facial and / or hand gesture information can be used to deliver content, such as virtual content, within extended reality (XR) environments (e.g., VR or AR environments). The identification of a user's facial deformities, hand gesture configurations, and / or the 3D positions of points on the surface or tissue geometry of the user's head / face and / or hands can be achieved via an XR head-mounted device utilizing radio frequency (RF) signals obtained via one or more antennas. Such information can be used to predict the user's facial configurations, such as, in particular, 3D key points, such as, in particular, the user's cheeks, chin, lips, tongue, etc. Similarly, such information can be used to predict and / or identify the user's hand gesture configurations and associated expressions.
[0005] Head-mounted devices may include head-mounted displays (HMDs), heads-up display glasses, AR glasses including clear eyepieces, vision-correcting glasses, etc. One or more antennas may be located on or integrated into a portion of the head-mounted device (such as, in particular, the bottom portion of an XR head-mounted device). In some embodiments, a user's head / face / hand features (such as, in particular, the user's facial expressions and / or hand gestures) may interact dielectrically and non-contactly with at least one of the one or more antennas, such that changes in the user's facial and / or hand gesture configurations may be manifested within the self-resonant frequency and / or performance associated with at least one of the one or more antennas. In some embodiments, the values of the self-resonant frequency and / or performance may be measured via the XR head-mounted device. The resulting data associated with the user's head / face / hand gesture features may be used, for example, to configure personas, interpret user facial / hand gestures / expressions (expressions), and / or uniquely identify users. Personas may include, in particular, photorealistic representations of the user, abstract representations of the user (e.g., animated representations of the user), any type of avatar, etc.
[0006] In some embodiments, one or more antennas are configured to minimize interference from the user's hands, face coverings (e.g., beard, mask, etc.), and / or the associated environment by utilizing a directional radiation pattern associated with the angular dependence of the intensity of the RF waves emitted from the one or more antennas. Some embodiments utilize, in particular, slot antennas, which may be advantageous in some cases due to their low profile and simplified structural elements. Some embodiments utilize slot antennas constructed as folds, which can provide a location for placing secondary antennas. Some embodiments utilize a vertically polarized U-shaped slot antenna placed in the center of the XR headset and a horizontally polarized antenna placed adjacent to the side portions of the XR headset (e.g., at a specified distance from the edge of the XR headset). Some embodiments include multiple antennas utilizing scattering parameters (any type of S-parameter) that include vector signals / entities describing the electrical behavior of a linear electrical network when subjected to various steady-state stimuli from electrical signals. For example, in a 2-port antenna configuration, there can be four or more S-parameters (e.g., S11, S21, S12, and S22), and each port can be configured to inject a signal and measure the reflected signal relative to its own port, such as, in particular, the S11 parameter (input port voltage reflection coefficient indicating the power reflected back to the transmitting antenna) or the S22 parameter (output port voltage reflection coefficient), and the transmitted signal from the additional port, such as the S21 parameter (forward voltage gain indicating the transmit power between antennas) or the S12 parameter (reverse voltage gain). Specifically, the electrical network (e.g., the surface of a face or hand) is deformed (e.g., changes in facial geometry, different facial expressions, different hand gestures, etc.), and the electromagnetic impedance distribution of the face and / or hand has also changed. Subsequently, as the incident signal (e.g., from port 1) passes through the network (e.g., face, hand, etc.), the amount of signal reflected (e.g., back to port 1) and transmitted (e.g., to port 2) changes, and can be measured as S-parameters, in particular S11 for analyzing the reflected signal and S21 for analyzing the transmitted signal. In some implementations, one or more antennas / antenna arrays can be used to achieve controlled beam steering, thus modifying the phase of the input signal relative to the radiating elements that control the directivity of the antenna system's radiation pattern. Some implementations utilize multi-port antenna configurations, allowing the test point matrix to be expanded based on the number of antennas used. For example, two antennas can generate parameters S11, S22, S21, and S12, while four or more antennas can generate S11, S22, S33, S44, S12, S13, S14…Snm, etc., for a total of 16 test points for amplitude and another 16 test points for phase.
[0007] In some implementations, the inputs to the machine learning (ML) model and / or rule-based model may include, in particular, the return loss amplitudes of ports S11, S22…Sn, the phase shifts of ports S11, S22…Sn, the amplitudes of combinations of ports S21 or any other, the phase shifts of combinations of ports S21 or any other, etc. Some implementations provide additional featureizations to be utilized beyond the original measurements, such as the first derivative of the sequence, the difference from previous frame data, the standard deviation, coefficients from a polynomial fit, subtraction of available port vectors, etc. Some implementations provide outputs including a series of 3D keypoints, facial expression classification, and / or user identification. The interpretation (rule-based) algorithm or ML model and / or additional mechanisms (e.g., using motion sensor inputs) can be configured to consider the user's body (e.g., chest) and modify associated disturbances such as head / hand orientation (e.g., tilting forward as much as possible).
[0008] In some embodiments, a device has a structure, one or more antennas positioned on a portion of the structure, and a processor. The one or more antennas include at least one or a combination of antenna topologies. The processor is configured to interpret data from the one or more antennas to identify deformations in the tissue geometry of at least a portion of the user's body when the device is worn on the user's head. The data from the one or more antennas may include data associated with one or more impedance characteristics.
[0009] According to some embodiments, an apparatus includes one or more processors, non-transitory memory, and one or more programs; the one or more programs are stored in the non-transitory memory and configured to be executed by the one or more processors, and the one or more programs include instructions for performing or causing to perform any of the methods described herein. According to some embodiments, a non-transitory computer-readable storage medium stores instructions that, when executed by one or more processors of the apparatus, cause the apparatus to perform or cause to perform any of the methods described herein. According to some embodiments, an apparatus includes: one or more processors, non-transitory memory, and components for performing or causing to perform any of the methods described herein. Attached Figure Description
[0010] To enable those skilled in the art to understand this disclosure, more detailed descriptions can be made with reference to aspects of some exemplary embodiments, some of which are shown in the accompanying drawings.
[0011] Figure 1 Exemplary electronic devices operating in motion in a physical environment according to some specific implementations are illustrated.
[0012] Figure 2Examples of systems are provided that can analyze user facial poses and / or expressions and / or uniquely identify users based on certain specific implementations.
[0013] Figure 3A Examples of head-mounted devices configured to predict a user’s facial configuration using RF signals obtained via an antenna, according to some specific implementations, are illustrated.
[0014] Figure 3B Alternative examples of head-mounted devices configured to predict a user’s facial configuration using RF signals obtained via an antenna, according to some specific implementations, are illustrated.
[0015] Figure 4A Exemplary views are shown of a head-mounted device worn by a user, according to some specific implementations, capable of analyzing the user's facial posture and / or expression.
[0016] Figure 4B An exemplary view is shown, illustrating how, according to some specific implementations, facial expressions are presented to a user via a head-mounted device to generate corresponding output characters and associated key points.
[0017] Figure 4C Exemplary views are shown of a head-mounted device worn by a user according to some specific implementations, capable of analyzing the user's hand gestures and / or expressions.
[0018] Figure 4D An alternative exemplary view is shown, according to some specific implementations, of a head-mounted device worn by a user that can analyze the user's hand gestures and / or expressions.
[0019] Figure 5 This is a flowchart representation of an exemplary method for predicting the body part configuration of a user based on some specific implementations using RF signals obtained via an antenna.
[0020] Figure 6 It is a block diagram based on some specific implementations of electronic devices.
[0021] As is customary practice, various features illustrated in the accompanying drawings may not be drawn to scale. Therefore, for clarity, the dimensions of various features may be arbitrarily expanded or reduced. Furthermore, some drawings may not depict all components of a given system, method, or apparatus. Finally, similar reference numerals may be used throughout the specification and drawings to denote similar features. Detailed Implementation
[0022] Numerous details have been described to provide a thorough understanding of the exemplary embodiments illustrated in the accompanying drawings. However, the drawings illustrate only some exemplary aspects of this disclosure and should not be considered limiting. Those skilled in the art will understand that other effective aspects and / or variations do not include all the specific details described herein. Furthermore, well-known systems, methods, components, devices, and circuits have not been described exhaustively so as not to obscure further relevant aspects of the exemplary embodiments described herein.
[0023] Figure 1 An exemplary electronic device 105 operating in physical environment 100 is illustrated. Figure 1 In the example, physical environment 100 is a room. Electronic device 105 may include one or more antennas, one or more cameras, microphones, depth sensors, or other sensors that can be used to capture and evaluate information about physical environment 100 and objects within it, as well as information such as the facial features and / or hand gesture configuration (e.g., in front of the mouth) of user 102 of electronic device 105. Information about physical environment 100 and / or user 102's facial features and / or hand gesture configuration can be used to provide visual and audio content and / or identify the current location of physical environment 100 and / or the location of the user within physical environment 100. Similarly, information about the user's facial features and / or hand gesture configuration can be used to configure personas, interpret user facial poses / expressions, interpret user hand gestures / expressions, uniquely identify users, etc.
[0024] In some implementations, a view of an extended reality (XR) environment may be provided to one or more participants (e.g., user 102 and / or other participants not shown) via an electronic device 105 (e.g., a wearable device such as an HMD, XR headset, etc.). Such an XR environment may include a view of a 3D environment generated based on camera images and / or depth camera images of the physical environment 100, and a representation of user 102 based on camera images and / or depth camera images of user 102. Such an XR environment may include virtual content positioned at a 3D location relative to a 3D coordinate system (i.e., 3D space) associated with the XR environment, which may correspond to the 3D coordinate system of the physical environment 100.
[0025] In some embodiments, video (e.g., pass-through video depicting the physical environment) is received from the image sensor of a device (e.g., device 105). In some embodiments, the 3D representation of the virtual environment is aligned with the 3D coordinate system of the physical environment. The size of the 3D representation of the virtual environment can be generated, in particular, based on the scale of the physical environment or the positioning of open spaces, floors, walls, etc., such that the 3D representation is configured to align with corresponding features of the physical environment. In some embodiments, the viewpoint in the 3D coordinate system can be determined based on the position of the electronic device within the physical environment. This viewpoint can be determined, in particular, based on image data, depth sensor data, motion sensor data, etc., which can be retrieved via a virtual inertial odometry (VIO) system, a simultaneous localization and mapping (SLAM) system, etc.
[0026] In some implementations, an HMD or heads-up display (e.g., device 105) or other external device communicatively coupled to a server can be configured to capture a user's overall profile for use within an XR environment, such as providing multimodal input to generate expressive characters. Capturing a user's overall profile may include capturing facial features associated with the user's mouth, such as mouth configuration, and / or capturing hand movements / poses associated with the user's hands, to provide a channel for representing facial pose information and / or 3D hand gesture expression / interpretation. For example, emotional cues can be obtained via facial expressions, improving credibility, clarity, and congruence among users when nonverbal signals such as facial expressions are matched with contextual context and verbal content. In some implementations, open and closed mouth poses can potentially be used for diet tracking. Similarly, hand configurations / expressions can be used to obtain emotional cues associated with a specified posture (e.g., in front of the user's mouth), such as those described below. Figure 4D The illustrated silent / quiet posture.
[0027] In some implementations, HMDs or heads-up display glasses can utilize an antenna associated with or integrated into a portion of the HMD or heads-up display glasses (e.g., the lower portion) to implement an RF-based method for capturing a user's mouth configuration information. The aforementioned RF-based method for capturing mouth / hand configuration information eliminates the privacy issues associated with camera-based facial / hand feature capture methods, while supporting silent facial expression / hand gesture determination. Similarly, the RF-based method for capturing mouth and / or hand configuration information does not require devices that come into contact with the user's body and can be entirely contained within the HMD or heads-up display glasses, thus providing high physical robustness and user usability.
[0028] In some implementations, the HMD (or head-up display glasses) is configured to sense antenna impedance characteristics. For example, a dielectric material (such as human tissue) located in close proximity to the antenna can be loaded as a component of the antenna, potentially affecting the associated reference ground plane and thus the antenna characteristic impedance and radiation performance relative to a given reference frequency. Associated geometric changes within the dielectric (e.g., occurring during movement of the user's mouth or hand) can reflect changes on the antenna's sensing ground plane, and the subsequently changing antenna performance can be measured relative to S-parameters. The antenna can include a low-profile antenna topology, such as a submillimeter low-profile cross-polarized antenna system providing a suitable structure for integration into the HMD. Similarly, S21…Snm (multiple antennas) data, such as S21 port amplitude, S21 port phase shift, etc., can be used to quantify the antenna's performance gain. In some implementations, the HMD includes dedicated sensing hardware configured to unlock increased frame rates and resolutions.
[0029] Some implementations utilize electromagnetic simulation computational resources to determine the antenna system topology design. Similarly, determining the S-parameter response of a sample mouth configuration can include constructing representative mouth configurations on a 3D human head or hand model. For example, representative mouth configurations can include, in particular, a closed mouth, an open mouth, a closed smile, a toothy smile, and a mouth with the tongue protruding and open. Likewise, determining the S-parameter response of a sample hand configuration can include constructing representative hand configurations on a 3D human hand model. For example, sample hand configurations can include, in particular, a silent / quiet posture represented by fingers in front of the user's mouth. In some implementations, the properties of lossy copper materials can be used for antennas with varying design parameters. Minimizing mismatch losses between antennas can include implementing impedance matching networks for each antenna at a predefined frequency of interest.
[0030] Figure 2 An example of a system 200 according to some specific implementation is illustrated, which is capable of analyzing a user's facial and / or hand gestures and / or expressions (expressions) to configure roles and / or uniquely identify the user. System 200 includes a head-mounted device 208 (e.g., an XR head-mounted device, such as an HMD, heads-up display glasses, AR glasses including clear eyepieces, vision-correcting glasses, etc.) and a machine learning model 210 (or pipeline) configured to analyze antenna data to track user facial and / or hand features for predicting facial and / or hand configurations, such as 3D key points, particularly relative to the user's cheeks, lips, hands, and / or tongue.
[0031] System 200 enables the generation of a user's facial / hand pose configuration prediction 220 via a head-mounted device 208 using RF signals obtained via antenna 204. Optional sensors 205, such as inward-facing or outward-facing cameras, can be used to provide additional data for generating the facial / hand pose configuration prediction 220 and / or configuring subsequent roles. Antenna 204 may be located near, on, and / or integrated with the bottom portion of head-mounted device 208. Similarly, antenna 204 may be located near, on, and / or integrated with any portion of head-mounted device 208. In some specific implementations, user facial and / or hand features, such as, in particular, the user's mouth, can alter the sensing ground plane of antenna 204, such that changes in the user's facial configuration can be manifested within the self-resonant frequency and / or performance associated with at least one of the one or more antennas. The values of self-resonant frequency and / or performance can be measured via head-mounted device 208, and the resulting facial / hand pose configuration prediction 220 (e.g., generated via machine learning model 210) associated with the user's facial and / or hand pose features can be used, for example, to generate rendered content 224 (such as, in particular, characters, user facial / hand poses / expressions), identify users, etc.
[0032] Figure 3A An example of a head-mounted device 300 according to some specific embodiments is illustrated, which is configured to predict a user's facial / hand posture (in front of the face) configuration using RF signals obtained via antennas 302a and 302b. Antennas 302a and 302b may each comprise a single antenna or multiple antennas. Alternatively, head-mounted device 300 may comprise a single antenna. The example illustrated in Figure 3 shows a bottom view of head-mounted device 300, wherein antennas 302a and 302b are mounted to the bottom portion 301 of head-mounted device 300. Alternatively, antennas 302a and 302b may each be mounted to any part of head-mounted device 300 or integrated with the head-mounted device. Head-mounted device 300 may additionally include a vector network analyzer (VNA) 304, a battery 306, a communication module 308, and an impedance matching network 310. The head-mounted device 300 illustrates a vector network analyzer (VNA) 304, a battery 306, and a communication module 308 as externally mounted components. Alternatively, the VNA 304, battery 306, and communication module 308 can be configured as internally mounted components integrated within a portion of the head-mounted device 300, as described below regarding... Figure 3B exemplified.
[0033] The head-mounted device 300 can operate as a standalone device. Alternatively, the head-mounted device 300 can operate in conjunction with a connected computing device or controller.
[0034] Antennas 302a and 302b may each include a custom antenna configuration mounted to the bottom portion 301 of the head-mounted device 300 via a low-profile 3D-printed base (wedge-shaped) structure 312 slightly angled toward the user's face. Alternatively, antennas 302a and 302b may each include, in particular, a cross-polarized antenna system comprising a first vertically polarized U-slot antenna positioned on the central portion 303 of the head-mounted device 300 and a second horizontally polarized antenna (e.g., ...) positioned on the right side portion 305 of the head-mounted device 300. Figure 3A (As illustrated), such as a specified distance from the edge of the first vertically polarized U-shaped slot antenna. Regarding Figure 3A The illustrated example represents a cross-polarized antenna system in which antennas 302a and 302b are positioned orthogonally to each other. Utilizing the aforementioned orthogonally positioned cross-polarized antenna system can help reduce ambient noise and produce higher signal resolution, which can improve the capture of multiple different aspects of mouth movements (such as mouth closing movements, smiling with teeth showing, etc.) or hand movements (such as finger and joint movements, such as fingers moving towards or away from the antenna, etc.). Although Figure 3A Antennas 302a and 302b are illustrated as being orthogonally positioned relative to each other, but some specific implementations may include antennas 302a and 302b being positioned relative to each other at any type of angular location. Similarly, head-mounted device 300 may include additional antennas positioned orthogonally to antennas 302a and / or 302b (or in any angular configuration).
[0035] The antenna slots of antennas 302a and 302b can each have a specific size that optimizes size while providing sufficient sensing capability. To reduce the antenna footprint, each of antennas 302a and 302b can be physically folded (e.g., longitudinally) to maintain the same electrical length as a straightened slot.
[0036] In some embodiments, antennas 302a and 302b can each be constructed using a laser or alternative cutting equipment to form an antenna base structure 312 comprising a specified shape (for antennas 302a and 302b) from, for example, an acrylic sheet (e.g., including a specified thickness). Subsequently, a copper layer 311 (e.g., copper strip, copper sheet, etc.) can be applied to the antenna base structure 312, such that any voids present between the copper layer 311 and the antenna base structure 312 are removed during application, thereby producing antennas 302a and 302b. Ultra-miniature version A (SMA) coaxial cables can be electrically and mechanically connected (e.g., via solder) to antennas 302a and 302b. The SMA coaxial cable may include an outer conductor mesh structure that is electrically and mechanically connected (e.g., via solder) to a ground plane formed from the copper layer. In some implementations, the antenna feed line can be electrically and mechanically connected (e.g., via solder) across the antenna slots of antenna 302a and / or antenna 302b.
[0037] In the example implementation, the initial self-resonant frequencies of antennas 302a and 302b can each be approximately 2.1 GHz, with an S11 amplitude of -23 dB. Therefore, to shift the operating frequencies of antennas 302a and 302b to a short-range range of the target frequency, such as approximately 2.5 GHz, an impedance matching network 310 can be constructed by connecting a parallel inductor 310a (e.g., approximately 30 nH) between the antenna feed and the ground plane. Similarly, a capacitor 310b (e.g., approximately 0.6 pF) can be connected in series with port S11 and the antenna feed.
[0038] In some implementations, the S11 and S21 parameters can be measured using a VNA 304 mechanically attached to the front portion 300b of the head-mounted device 300. The VNA 304 can be configured to measure the S11 and S21 parameters from a specified frequency range and calculate return loss amplitude and phase shift. Combined information from the S11 and S21 parameters can be used to estimate fine-grained mouth configurations. The VNA 304 can include a dedicated single or multiple chipset-integrated VNA design and can be integrated into the motherboard of the head-mounted device 300.
[0039] In some implementations, facial expressions and / or hand gestures may not require a high frame rate, but they can benefit from higher resolution scans to distinguish them from different mouth or hand configurations. Therefore, multiple points, such as 61 points, can be sampled from a specified frequency range at 5 frames per second (FPS), generating 61 S11 parameter return loss amplitudes, 61 S11 parameter phase shifts, 61 S21 parameter return loss amplitudes, and 61 S11 parameter phase shifts, resulting in a total of 244 values. These values can be fed into a machine learning pipeline for predicting the user's facial configuration. In contrast, mouth movements occurring during speech can include rapid movements but may require less resolution to capture their full morphology. In response, 31 points can be sampled from a specified frequency range, generating, for example, 124 values at 8.5 FPS.
[0040] In some implementations, calibrations may be performed that are associated with differences between the SMA coaxial cable, connectors, fixed dielectric components (e.g., headgear, support wedges, etc.) and additional associated components, resulting in reduced signal sensitivity.
[0041] The head-mounted device 300 can be configured to operate as a standalone device via battery power and wireless functionality. For example, the head-mounted device 300 can be configured to operate as a standalone device by using the communication module 308 to control the VNA 304 and by using the battery 306 for onboard power.
[0042] The head-mounted device 300 may include a machine learning pipeline, such as Figure 2 The machine learning model 210 receives vectors as input: S11 return loss amplitude, S11 phase shift, S21 amplitude, S21 phase shift...Snm. Each of the aforementioned vectors may include values of length 31 or 61, for a total of 244 values for sensing facial expressions or 144 values for sensing speech motion. In addition to using these raw measurements as machine learning features, we also perform additional featureization. For example, for each vector, the following calculations may be performed: the first derivative of the sequence (61 or 31 features x 4), the difference from the previous data frame (60 or 30 x 4), the standard deviation (1 feature x 4), the coefficients from the third-order polynomial fit (4 features x 4), and the subtraction of the S11 to S21 vectors (61 or 31 features x 2), thereby producing a total of 870 or 450 features for machine learning.
[0043] Figure 3B Examples are given based on some specific implementations relative to Figure 3A An alternative example of the head-mounted device 330 of the head-mounted device 300. With Figure 3ACompared to the 300-head-mounted device, Figure 3B The head-mounted device 330 includes, relative to Figure 3A The external component configuration of the head-mounted device 300 is an alternative to the internal component configuration. For example, Figure 3B The head-mounted device 330 includes a VNA 304, a battery 306, and a communication module 308, which are configured as internally mounted components integrated within a portion of the head-mounted device. The VNA 304, communication module 308, and / or battery 306 may include dedicated single or multiple chipset integration designs and may be integrated into a dedicated printed circuit board (PCB) and / or motherboard within the head-mounted device 330.
[0044] Figure 4A An exemplary view 400 according to some specific implementation is illustrated, wherein a user 401 wears a head-mounted device 405 capable of analyzing the user's facial posture and / or expressions. The head-mounted device 405 may include, relative to... Figure 3A Head-mounted devices 300 or Figure 3B The head-mounted device 330 is the same as or similar to the head-mounted device 405, and therefore may include antennas 402a and 402b mounted to the bottom portion 417 of the head-mounted device 405. The head-mounted device 405 may additionally include a vector network analyzer (VNA) 404, a battery 406, a communication module 408, and an impedance matching network 410. The VNA 404, battery 406, communication module 408, and impedance matching network 410 may include, for example... Figure 3A The illustrated external mounting components. Alternatively, VNA 404, battery 406, communication module 408, and impedance matching network 410 may include, as shown below, Figure 3B The illustrated internal mounting components. Any configuration of the internal and external mounting components can be implemented. Exemplary view 400 illustrates the interaction of user 401's mouth 401a (or facial expression) with antenna 402 dielectrically and non-contactly, such that changes in mouth configuration can be represented as changes in the self-resonant frequency and performance of antenna 402. These changes can be measured by head-mounted device 405, and machine learning pipelines and / or modules can be configured to predict 11 3D keypoints of user 401's cheeks, lips, and tongue, as illustrated in phase and amplitude curves 407 for parameters S11 and S21. Phase and amplitude curves 407 can be used, for example, to configure more expressive personas for telepresence purposes, thereby reducing privacy issues inherent in camera-based systems while supporting (silent) facial expressions that are undetectable by audio-based systems.
[0045] Figure 4BAn exemplary view 419 illustrates, according to some specific implementation, the presentation of facial expressions 424 to a user 420 via a head-mounted device 428 to generate a corresponding output character 432 and associated key points 434. The generation of the output character 432 may utilize the user 420's mouth to interact with the antenna dielectric ground and non-contact ground, such that changes in the mouth configuration can be represented as changes in the antenna's self-resonant frequency and performance. Facial expressions 424 may include measurements associated with mouth widths (e.g., corner-to-corner of the mouth) recorded relative to a neutral configuration (e.g., average = 49.4 mm, SD = 3.9) used to scale a 3D head model of the user 420 within associated software.
[0046] In some implementations, the head-mounted device 428 can be calibrated by capturing two different datasets to evaluate its overall accuracy. The first dataset includes captured facial expressions, such as facial expressions 424, which include exaggerated facial movements sustained for a short period of time. The second dataset includes mouth movements 425 that occur during speech. Mouth movements 425 that occur during speech can include continuous and rapid small movements.
[0047] In some implementations, the accuracy of the head-mounted device 428 can be benchmarked by using sensors such as cameras to determine the ground-based mouth configuration and to mark key points in a high-resolution video stream of the user 420's mouth. The associated software development kit (SDK) can be configured to output a blended shape (e.g., 37) for configuring the user 420's head and to extract associated key points 434 that can be scaled to match measurements of the user 420.
[0048] View 419 illustrates a facial expression 424 (e.g., 10) to be presented to user 420 via head-mounted device 428. In response to this presentation, a request is presented to user 420 to provide a facial expression matching configuration associated with facial expression 424. The resulting facial expression (e.g., mouth movement 425) provided by user 420 is timestamped for future evaluation of discrete configuration classification accuracy.
[0049] Requesting facial expressions may include a first data collection phase, which includes requesting all facial expressions (e.g., 10) in a random sequence. The first data collection phase may include five rounds forming a data capture session (e.g., 10 facial expressions across 5 rounds). During the first data collection phase, synchronized sensor data and ground-based data may be continuously captured (e.g., 5 frames per second (FPS)) to capture various intermediate facial configurations among the ten terminal facial expressions (e.g., 10 paired combinations) for classification and accuracy evaluation.
[0050] Requests for facial expressions may include a second data collection phase, which includes collecting mouth movement data generated by speech. The mouth movement data generated by speech includes data that changes more rapidly than facial expression data. The second data collection phase may include presenting randomly selected sentences (e.g., 10) to the user one at a time. Each session of the second data collection phase may include 5 rounds of 10 random sentences (e.g., 50 sentences per session). During the second data collection phase, synchronized sensor data and ground-based data may be captured at 8.5 FPS.
[0051] In some implementations, multiple training processes can be enabled that are associated with different user calibration and training schemes. For example, calibration and training schemes may include, in particular, session accuracy calibration and training across the worn head-mounted device, accuracy calibration and training within a session of the worn head-mounted device, and accuracy calibration and training across users.
[0052] Session accuracy calibration and training across worn head-mounted devices can include training the model using persistent session data for testing relative to all but one session of the user data collection, such as... Figure 2 The machine learning model 210 can train the model for all combinations of session data and combine all results.
[0053] Conversational accuracy calibration and training within a worn head-mounted device can include training the model relative to each session of user data collection associated with facial expressions and speech movements, such as... Figure 2 The machine learning model 210. In some specific implementations, each data collection session may include 5 rounds of data collection, thereby allowing training using all combinations of data from 4 rounds of data collection, and 1 round may be used for testing to measure accuracy within the worn head-mounted device session.
[0054] Cross-user accuracy calibration and training can include training the model relative to estimated and reported cross-user accuracy by training the model with respect to data from all users except one user, and then testing it with data from one user. Figure 2 Machine learning model 210.
[0055] In some implementations, facial data associated with a user's facial features can be used to uniquely identify the user, especially for loading personalized content (e.g., system settings, game progress actions, interpupillary distance, etc.) and providing automatic content control (e.g., for children).
[0056] In some specific implementations, antennas (such as Figure 4AThe illustrated antenna 402 comprises a low-profile antenna embedded in the surface of the head-mounted device (e.g., a flat bottom surface, such as bottom portion 417), such that the angle of the antenna can affect the performance of the head-mounted device. For example, the performance of the head-mounted device may be affected because the maximum directivity of the cross-polarized slot antenna system is orthogonal to the ground plane of the antenna; therefore, a small angle of the antenna relative to the direction toward the user's mouth can allow the signal to distinguish mouth configurations and reduce generated noise. Similarly, an antenna array with controllable beam steering can form an antenna radiation pattern with beam tilt, such as at 45° relative to the user's face.
[0057] In some implementations, the orientation associated with the antenna can be configured to reduce noise from the environment and associated hand movements. Similarly, the aforementioned tilted radiating beam antenna can be configured to take into account neck orientation, such as nodding vertical neck movements that may affect the signal and prediction results. For example, the tilted radiating beam antenna can be configured to precisely target only the rigid body of the face.
[0058] Figure 4C An exemplary view 438 is illustrated according to some specific implementation, wherein a user 440 wears a head-mounted device 441 capable of analyzing the user's hand (and facial) posture and / or expression (facial expression). The head-mounted device 441 may include, relative to... Figure 3A Head-mounted devices 300 or Figure 3B Head-mounted device 330 (and / or Figure 4A The head-mounted device 441 may be a head-mounted device identical or similar to the head-mounted device 405, and therefore may include antennas 402a and 402b mounted to the bottom portion 417 of the head-mounted device 441. The head-mounted device 441 may additionally include a vector network analyzer (VNA) 404, a battery 406, a communication module 408, and an impedance matching network 410. The VNA 404, battery 406, communication module 408, and impedance matching network 410 may include, for example... Figure 3A The illustrated external mounting components. Alternatively, VNA 404, battery 406, communication module 408, and impedance matching network 410 may include, as shown below, Figure 3BThe illustrated internal mounting components. Any configuration of the internal and external mounting components can be implemented. Exemplary view 438 illustrates a hand gesture configuration / expression 449a of user 440 (e.g., in front of mouth 443), which interacts dielectrically and non-contactly with antennas 402a and 402b, such that changes in hand gesture configuration can manifest as changes in the self-resonant frequencies and performance of antennas 402a and 402b. Interaction with the dielectrically and non-contactly with antennas 402a and 402b can minimize interference from face coverings (e.g., beard, mask, etc.) and / or the associated environment by utilizing a directional radiation pattern associated with the angular dependence of the intensity of the RF waves emitted from antennas 402a and 402b. In some implementations, metal accessories (e.g., a heavy bracelet or ring on the hand of user 440) may cause interference with RF waves emitted from antennas 402a and 402b to detect changes in hand posture configuration, and therefore, head-mounted device 441 may present the resulting error message to indicate that a training process should be performed to retrain head-mounted device 441 while wearing the metal accessory.
[0059] In some implementations, changes in hand gesture configuration can be measured by a head-mounted device 405, and machine learning pipelines and / or modules can be configured to predict 3D keypoints of the user 440's hands, fingers, and joints, as illustrated in phase and amplitude curves 451a for parameters S11 and S21. Phase and amplitude curves 451a can be used, for example, to configure more expressive personas for telepresence purposes, thereby reducing privacy issues inherent in camera-based systems while supporting (silent) hand gesture expressions that are undetectable by audio-based systems.
[0060] Figure 4D Examples are given based on some specific implementations relative to Figure 4C An alternative exemplary view 457 of exemplary view 438 shows a user 440 wearing a head-mounted device 441 capable of analyzing the user's hand (and facial) posture and / or expression (facial expression). Figure 4C Compared to the exemplary view 438, Figure 4D Exemplary view 457 illustrates a hand gesture configuration / expression 449b of user 440 (e.g., in front of mouth 443) that represents an emotional cue associated with a silent / quiet posture. The hand gesture configuration / expression 449b includes changes in hand gesture configuration measured by head-mounted device 405 and a machine learning pipeline and / or module for predicting 3D keypoints of user 440's hands, fingers, and joints, as illustrated in phase and amplitude plots of parameters S11 and S21, as shown in plot 451b.
[0061] Figure 5This is a flowchart representation of an exemplary method 500 for predicting a user's body part configuration using RF signals obtained via an antenna, according to some specific implementations. In some implementations, method 500 is performed by a device (such as a mobile device, desktop computer, laptop computer, HMD, or server device). In some implementations, the device has a screen for displaying images and / or a screen for viewing stereoscopic images, such as a head-mounted display (HMD, such as... Figure 1 (Device 105). In some embodiments, method 500 is performed by processing logic components, including hardware, firmware, software, or combinations thereof. In some embodiments, method 500 is performed by a processor that executes code stored in a non-transitory computer-readable medium (e.g., memory). Each block in method 500 can be enabled and executed in any order.
[0062] Method 500 is used in structures that are positioned within a part of the structure (e.g., such as...). Figure 3A and Figure 3B The implementation takes place at one or more antennas on the illustrated bottom portion 301 and at a device coupled to a processor of a non-transitory computer-readable storage medium. The one or more antennas may include at least one or a combination of antenna topologies. The device may be, in particular, a head-mounted device (HMD) including an internal display, an augmented reality (AR) glasses kit including a transparent eyepiece, a pair of vision-correcting glasses, etc.
[0063] In some embodiments, one or more antennas may include a controllable beam-steering antenna configured to produce a directional radiation pattern. In some embodiments, one or more antennas include a slot antenna comprising a folded slot. In some embodiments, the slots of one or more antennas can be modified to change the shape or size of one or more antennas. In some embodiments, one or more antennas may utilize common polarization or cross polarization. In some embodiments, one or more antennas may include a vertically polarized U-shaped slot antenna and a horizontally polarized antenna, the horizontally polarized antenna being positioned such that the vertically polarized U-shaped antenna is closer to the central axis of the structure than the horizontally polarized antenna. In some embodiments, one or more antennas may utilize a 2-port configuration such that each port of the 2-port configuration is configured to inject a signal and measure the reflected signal to its own port and the transmitted signal from the other port. In some embodiments, one or more antennas may provide an antenna radiation pattern beam tilt relative to the user's face.
[0064] At box 504, when the device is worn on a user's head, method 500 interprets data from one or more antennas. This interpretation may include executing a rule-based model or machine learning model configured to analyze the antenna data to track at least one body part of the user (e.g., head, hands, etc.). For example, user facial features may be tracked to predict facial configurations, such as 2D or 3D keypoints, particularly relative to the user's cheeks, lips, chin, and / or tongue, as described in paragraph 0027 regarding... Figure 2 Paragraph 0038 about Figure 3A And paragraph 0041 about Figure 4A As described above. Similarly, user hand gesture characteristics can be tracked to predict and / or distinguish user hand gesture configurations and associated expressions. Data from one or more antennas may include, in particular, return loss amplitude and phase shift, additional characterization, etc.
[0065] At block 506, method 500, in response to interpreting data, identifies deformations in the tissue geometry of at least a portion (e.g., body part) of the user (e.g., head, hand, etc.) when the device is worn on the user's head. In some implementations, data from one or more antennas may be associated with one or more impedance characteristics.
[0066] In some specific implementations, interpreting the data and identifying at least some of the deformations may include, in particular, inputting data from one or more antennas into a machine learning-based or rule-based model, outputting 2D or 3D keypoints corresponding to locations on the facial surface (such as the user's mouth, tongue, chin, and / or cheeks), outputting facial expression classifications and / or user identifiers, taking into account interference from the user's chest, etc., thereby identifying deformations in the tissue geometry of the user's head, as described in paragraph 0040 regarding... Figure 4A As described.
[0067] In some specific implementations, interpreting the data and identifying at least some of the deformations may include, in particular, inputting data from one or more antennas into a machine learning model, outputting 2D or 3D keypoints corresponding to positions on the surface of the hand (such as fingers, joints, wrists, palms, or backs of the hands), outputting hand expression classifications and / or user identifiers, thereby identifying deformations of the user's hand (e.g., those located near the user's face), as described in paragraphs 0057-0059 regarding... Figure 4C and Figure 4D As described.
[0068] Figure 6 This is a block diagram of example device 600. Device 600 illustrates... Figure 1An exemplary device configuration of electronic device 105. Although certain specific features have been illustrated, those skilled in the art will understand from this disclosure that various other features have not been illustrated for the sake of brevity and to avoid obscuring more relevant aspects of the specific embodiments disclosed herein. Therefore, as a non-limiting example, in some specific implementations, device 600 includes one or more processing units 602 (e.g., microprocessors, ASICs, FPGAs, GPUs, CPUs, processing cores, etc.), one or more input / output (I / O) devices and sensors 604, one or more communication interfaces 608 (e.g., USB, FireWire, Thunderbolt, IEEE 802.3x, IEEE 802.11x, IEEE 802.14x, GSM, CDMA, TDMA, GPS, IR, Bluetooth, ZigBee, SPI, I2C, and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 610, output devices (e.g., one or more displays) 612, one or more internal and / or external image sensor systems 614, memory 620, and one or more communication buses 604 for interconnecting these components and various other components.
[0069] In some embodiments, one or more communication buses 604 include circuitry for interconnecting system components and controlling communication between system components. In some embodiments, one or more I / O devices and sensors 606 include at least one of the following: an inertial measurement unit (IMU), an accelerometer, a magnetometer, a gyroscope, a thermometer, one or more physiological sensors (e.g., a blood pressure monitor, a heart rate monitor, a blood oxygen sensor, a blood glucose sensor, etc.), one or more microphones, one or more speakers, a haptic engine, one or more depth sensors (e.g., structured light, time-of-flight, etc.), one or more cameras (e.g., an inward-facing camera and an outward-facing camera of an HMD), one or more infrared sensors, one or more thermal sensors, etc.
[0070] In some embodiments, one or more displays 612 are configured to present a view of a physical environment, a graphical environment, an extended reality environment, etc., to a user. In some embodiments, one or more displays 612 are configured to present content to a user (determined based on the user's determined user / object position within the physical environment). In some embodiments, one or more displays 612 correspond to holographic, digital light processing (DLP), liquid crystal display (LCD), liquid crystal on silicon (LCoS), organic light-emitting field-effect transistor (OLET), organic light-emitting diode (OLED), surface-conducting electron emission display (SED), field emission display (FED), quantum dot light-emitting diode (QD-LED), microelectromechanical systems (MEMS), and / or similar display types. In some embodiments, one or more displays 612 correspond to waveguide displays such as diffraction, reflection, polarization, and holography. In one example, device 600 includes a single display. In another example, device 600 includes displays for each of the user's eyes.
[0071] In some embodiments, one or more image sensor systems 614 are configured to acquire image data corresponding to at least a portion of the physical environment 100. For example, one or more image sensor systems 614 may include one or more RGB cameras (e.g., having a complementary metal-oxide-semiconductor (CMOS) image sensor or a charge-coupled device (CCD) image sensor), monochrome cameras, IR cameras, depth cameras, event-based cameras, etc. In various embodiments, one or more image sensor systems 614 may also include an illumination source emitting light, such as a flash. In various embodiments, one or more image sensor systems 614 may also include an on-camera image signal processor (ISP) configured to perform multiple processing operations on the image data.
[0072] In some specific implementations, sensor data can be obtained from devices (e.g., Figure 1The device 105 acquires the sensor data during a scan of the room's physical environment. The sensor data may include a 3D point cloud and a sequence of 2D images corresponding to views of the room captured during the scan. In some embodiments, the sensor data includes image data (e.g., from an RGB camera), depth data (e.g., depth images from a depth camera), ambient light sensor data (e.g., from an ambient light sensor), and / or motion data from one or more motion sensors (e.g., accelerometers, gyroscopes, IMUs, etc.). In some embodiments, the sensor data includes visual inertial odometry (VIO) data determined based on the image data. The 3D point cloud can provide semantic information about one or more elements of the room. The 3D point cloud can provide information about the location and appearance of surface portions within the physical environment. In some embodiments, the 3D point cloud is acquired over time (e.g., during a scan of the room) and can be updated, with updated versions of the 3D point cloud obtained over time. For example, when the 3D representation is updated / adjusted over time (e.g., when a user scans a room), the 3D representation can be obtained (and analyzed / processed).
[0073] In some embodiments, sensor data may be positioning information, and some embodiments include VIO (Vehicle Identification and Odometry) to determine equivalent odometry information to estimate travel distance using sequential camera images (e.g., light intensity image data) and motion data (e.g., acquired from an IMU / motion sensor). Alternatively, some embodiments of this disclosure may include a Simultaneous Localization and Mapping (SLAM) system (e.g., a positioning sensor). This SLAM system may include a GPS-independent, multi-dimensional (e.g., 3D) laser scanning and range measurement system that provides real-time simultaneous localization and mapping. This SLAM system can generate and manage highly accurate point cloud data produced by reflections from laser scans of objects in the environment. Accurately tracking the movement of any points in the point cloud over time allows the SLAM system to use points in the point cloud as reference points for its location, maintaining an accurate understanding of its position and orientation as it travels through the environment.
[0074] In some embodiments, device 600 includes an eye-tracking system for detecting eye positioning and eye movement (e.g., eye gaze detection). For example, the eye-tracking system may include one or more infrared (IR) light-emitting diodes (LEDs), an eye-tracking camera (e.g., a near-infrared (NIR) camera), and an illumination source (e.g., an NIR light source) that emits light (e.g., NIR light) towards the user's eyes. Furthermore, the illumination source of device 600 may emit NIR light to illuminate the user's eyes, and the NIR camera may capture images of the user's eyes. In some embodiments, the images captured by the eye-tracking system may be analyzed to detect the positioning and movement of the user's eyes, or to detect other information about the eyes such as pupil dilation or pupil diameter. Furthermore, the gaze point estimated from the eye-tracking images enables gaze-based interaction with content displayed on a near-eye display of device 600.
[0075] Memory 620 includes high-speed random access memory, such as DRAM, SRAM, DDR RAM, or other random access solid-state memory devices. In some embodiments, memory 620 includes non-volatile memory, such as one or more disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 620 optionally includes one or more storage devices remotely located to one or more processing units 602. Memory 620 includes a non-transitory computer-readable storage medium.
[0076] In some embodiments, memory 620 or a non-transitory computer-readable storage medium of memory 620 stores an optional operating system 630 and one or more instruction sets 640. Operating system 630 includes procedures for handling various basic system services and for performing hardware-related tasks. In some embodiments, instruction set 640 includes executable software defined by binary information stored in charge. In some embodiments, instruction set 640 is software executable by one or more processing units 602 to perform one or more of the techniques described herein.
[0077] Instruction set 640 includes data interpretation instruction set 642 and surface deformation tracking instruction set 644. Instruction set 640 can be embodied in a single software executable file or multiple software executable files.
[0078] The data interpretation instruction set 642 is configured with instructions that can be executed by the processor to interpret data from one or more antennas.
[0079] The surface deformation discrimination instruction set 644 is configured with instructions that can be executed by the processor to track the deformation of the tissue geometry of the user's head or hand (near the mouth) when the device is worn on the user's head.
[0080] Although instruction set 640 is shown as residing on a single device, it should be understood that in other specific implementations, any combination of elements may reside in separate computing devices. Furthermore, Figure 6 This is intended more as a functional description of various features present in a particular specific implementation than as a structural diagram of the specific implementation described herein. As will be appreciated by those skilled in the art, the items shown individually can be combined, and some items can be separated. The actual number of instruction sets and how features are allocated therein will vary depending on the specific implementation and may depend in part on the specific combination of hardware, software, and / or firmware chosen for that particular implementation.
[0081] Those skilled in the art will understand that well-known systems, methods, components, devices, and circuits have not been described exhaustively so as not to obscure more relevant aspects of the specific embodiments of the examples described herein. Furthermore, other effective aspects and / or variations do not include all the details in the specific details described herein. Therefore, several details are described to provide a thorough understanding of the exemplary aspects illustrated in the accompanying drawings. Moreover, the drawings only illustrate some exemplary embodiments of this disclosure and should not be considered limiting.
[0082] While this specification contains numerous specific details of implementation, these details should not be construed as limiting the scope of any invention or potentially claimed content, but rather as descriptions of features specific to particular embodiments of a particular invention. Certain features described in the context of different embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may be implemented individually or in any suitable sub-combination in multiple embodiments. Furthermore, while certain features may be described above as functioning in certain combinations and even initially claimed in this manner, one or more features of a claimed combination may be removed from that combination in certain circumstances, and the claimed combination may involve sub-combinations or variations thereof.
[0083] Similarly, although the operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring such operations to be performed in a sequential order or the specific order shown, or requiring all illustrated operations to achieve the desired result. In some situations, multitasking and parallel processing may be advantageous. Furthermore, the partitioning of the various system components in the above embodiments should not be construed as requiring such partitioning in all embodiments, and it should be understood that the described program components and systems may generally be integrated together in a single software product or packaged into multiple software products.
[0084] Therefore, specific embodiments of the subject matter have been described. Other embodiments are also within the scope of the following claims. In some cases, the actions described in the claims can be performed in a different order and still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific order or sequence shown to achieve the desired result. In some embodiments, multitasking and parallel processing may be advantageous.
[0085] The embodiments of the subject matter and operation described in this specification may be implemented in digital electronic circuits or in computer software, firmware, or hardware (including the structures disclosed in this specification and their structural equivalents) or in a combination thereof. Embodiments of the subject matter described in this specification may be implemented as one or more computer programs, such as one or more modules of computer program instructions encoded on a computer storage medium for execution by or control of the operation of a data processing device. Alternatively or additionally, the program instructions may be encoded on artificially generated propagating signals, such as machine-generated electrical, optical, or electromagnetic signals, which are generated to encode information for transmission to a suitable receiver device for execution by the data processing device. The computer storage medium may be or be included in a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination thereof. Furthermore, although the computer storage medium is not a propagating signal, it may be a source or destination of computer program instructions encoded in artificially generated propagating signals. The computer storage medium may also be or be included in one or more separate physical components or media (e.g., multiple CDs, disks, or other storage devices).
[0086] The term "data processing apparatus" encompasses all kinds of devices, apparatuses, and machines for processing data, including programmable processors, computers, systems-on-a-chip, or many or combinations of the foregoing. The apparatus may include special-purpose logic circuitry (e.g., FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits)). In addition to hardware, the apparatus may include code that creates an execution environment for the computer program under consideration, such as code constituting processor firmware, protocol stacks, database management systems, operating systems, cross-platform runtime environments, virtual machines, or combinations thereof. The apparatus and execution environment can implement a variety of different computing model infrastructures, such as web services, distributed computing, and grid computing infrastructures. Unless otherwise specifically stated, it should be understood that throughout this specification, discussions using terms such as "processing," "computing," "calculating," "determining," and "identifying" refer to the actions or processes of computing devices, such as one or more computers or similar electronic computing devices, that manipulate or convert data represented as physical electronic or magnetic quantities within the memory, registers, or other information storage devices, transmission devices, or display devices of a computing platform.
[0087] The one or more systems discussed herein are not limited to any particular hardware architecture or configuration. A computing device may include any suitable arrangement of components that provide results conditioned on one or more inputs. Suitable computing devices include computer systems based on multi-purpose microprocessors that access stored software that programs or configures the computing system from a general-purpose computing device to a special-purpose computing device that implements one or more specific implementations of this subject. The teachings contained herein may be implemented in the software used for programming or configuring the computing device using any suitable programming, scripting, or other type of language or combination of languages.
[0088] Specific implementations of the methods disclosed herein can be performed in the operation of such computing devices. The order of the boxes presented in the examples above can be varied; for example, the boxes can be reordered, combined, and / or divided into sub-blocks. Some boxes or processes can be executed in parallel. The operations described in this specification can be implemented as operations performed by a data processing apparatus on data stored on one or more computer-readable storage devices or received from other sources.
[0089] The use of "applies to" or "configured to" in this document implies open and inclusive language, which does not exclude applicability to or configuration for performing additional tasks or steps. Additionally, the use of "based on" implies openness and inclusivity, as processes, steps, calculations, or other actions "based on" one or more of the stated conditions or values may in practice be based on additional conditions or values beyond those stated. The headings, lists, and numbering included herein are for illustrative purposes only and are not intended to be restrictive.
[0090] It will also be understood that while terms such as "first," "second," etc., may be used in this document to describe various elements, these elements should not be limited by these terms. These terms are merely used to distinguish one element from another. For example, a first node may be called a second node, and similarly, a second node may be called a first node, changing the meaning of the description, provided that all occurrences of "first node" are consistently renamed and all occurrences of "second node" are consistently renamed. First nodes and second nodes are both nodes, but they are not the same node.
[0091] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the claims. As used in the description of these embodiments and in the appended claims, the singular forms “a,” “an,” and “the” are intended to also cover the plural forms unless the context clearly indicates otherwise. It will also be understood that the term “and / or” as used herein refers to and covers any and all possible combinations of one or more of the associated listed items. It will be further understood that the term “comprising,” when used in this specification, specifies the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0092] As used herein, the term "if" can be interpreted as meaning "when the prerequisite is true" or "when the prerequisite is true" or "in response to determination" or "according to determination" or "in response to detection" that the prerequisite is true, depending on the context. Similarly, the phrases "if it is determined [the prerequisite is true]" or "if [the prerequisite is true]" or "when [the prerequisite is true]" can be interpreted as meaning "when it is determined that the prerequisite is true" or "in response to determination" or "according to determination" that the prerequisite is true or "when it is detected that the prerequisite is true" or "in response to detection" that the prerequisite is true, depending on the context.
Claims
1. An apparatus, the apparatus comprising: structure; One or more antennas, said one or more antennas being positioned on a portion of said structure, said one or more antennas comprising at least one or a combination of antenna topologies; as well as A processor configured to interpret data from the one or more antennas to identify deformations in the tissue geometry of at least a portion of the user when the device is worn on the user's head.
2. The device of claim 1, wherein the data from the one or more antennas is associated with one or more impedance characteristics.
3. The device of claim 1, wherein the one or more antennas utilize a 2-port configuration, wherein each port in the 2-port configuration is configured to inject a signal and measure reflected signals to its own port and transmitted signals from the other port.
4. The device of claim 1, wherein at least a portion comprises the user's head.
5. The device of claim 4, wherein the one or more antennas provide an antenna radiation pattern beam tilt relative to the user's face.
6. The device of claim 5, wherein the data from the one or more antennas includes return loss amplitude and phase shift.
7. The device of claim 4, wherein the data from the one or more antennas includes additional characterization.
8. The device of claim 4, wherein interpreting data from the one or more antennas to identify deformations in the tissue geometry of the at least a portion of the user comprises inputting the data from the one or more antennas into a machine learning model.
9. The device of claim 4, wherein interpreting data from the one or more antennas to identify deformations of the tissue geometry of at least a portion of the user comprises outputting 2D or 3D key points corresponding to positions on the surface of the user's face.
10. The device of claim 9, wherein the 2D or 3D key points are on the user's mouth, tongue, chin, or cheek.
11. The device of claim 4, wherein interpreting data from the one or more antennas to identify deformations in the tissue geometry of the at least a portion of the user includes outputting a facial expression classification.
12. The device of claim 4, wherein interpreting data from the one or more antennas to identify deformations of the tissue geometry of the at least a portion of the user includes outputting a user identifier.
13. The device of claim 4, wherein interpreting data from the one or more antennas to identify deformation of the tissue geometry of the at least portion of the user includes taking into account interference from the user's chest.
14. The device of claim 1, wherein the at least portion includes at least one hand of the user located near the user's face.
15. The device of claim 14, wherein interpreting data from the one or more antennas to identify the geometry of the at least portion of the user comprises outputting 2D or 3D key points corresponding to positions on the surface of the at least one hand.
16. The device of claim 15, wherein the 2D or 3D key points are on the fingers, joints, wrist, palm, or back of the hand of the user.
17. The device of claim 14, wherein interpreting data from the one or more antennas to identify deformations in the tissue geometry of the at least a portion of the user includes outputting a hand expression classification.
18. The device of claim 14, wherein interpreting data from the one or more antennas to identify deformations of the tissue geometry of the at least a portion of the user includes outputting a user identifier.
19. A method comprising: At a device having a structure, one or more antennas comprising at least one or more antenna topologies positioned on a part of said structure, and a processor coupled to a non-transitory computer-readable storage medium: When the device is worn on a user's head, it interprets data from the one or more antennas; as well as In response to the interpretation of the data, when the device is worn on the user's head, deformation of the tissue geometry of at least a portion of the user is identified.
20. A non-transitory computer-readable storage medium storing program instructions executable by one or more processors to perform operations, the operations including: At a device having a structure, one or more antennas comprising at least one or more antenna topologies positioned on a part of said structure, and a processor coupled to a non-transitory computer-readable storage medium: When the device is worn on a user's head, it interprets data from the one or more antennas; as well as In response to the interpretation of the data, when the device is worn on the user's head, deformation of the tissue geometry of at least a portion of the user is identified.