Detection and utilization of facial micro-movements

The system uses coherent light to analyze facial micro-movements for interpreting subvocalization and neuromuscular activity, addressing the challenge of silent communication and control, achieving effective unvoiced interaction and verification.

JP2025528023APending Publication Date: 2025-08-26キュー(キュー)リミテッド
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
JP2025503196
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-02-28
Filing Date
2023-07-19
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively interpret and utilize facial micro-movements during subvocalization for communication and control, lacking efficient methods to discern and utilize neuromuscular activity beyond vocalized speech.

Method used

A system and method utilizing coherent light sources and detectors to project and analyze facial skin micro-movements, correlating these movements with stored data to identify and interpret communication, perform identity verification, and enable unvoiced conversation, among other functions.

Benefits of technology

Enables accurate interpretation of subvocalized speech and neuromuscular activity, facilitating communication, identity verification, and control operations without perceptible vocalization, enhancing interaction and security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025528023000001_ABST
    Figure 2025528023000001_ABST
Patent Text Reader

Abstract

Systems, methods, and non-transitory computer-readable media containing instructions for detecting and using facial skin micro-movements are disclosed. In some non-limiting embodiments, the detection of facial skin micro-movements occurs using a speech detection system that may include a wearable housing, a light source (either a coherent light source or a non-coherent light source), a light detector, and at least one processor. The one or more processors may be configured to analyze light reflections received from the facial region to determine facial skin micro-movements and extract meaning from the determined facial skin micro-movements. Examples of meaning that may be extracted from the determined facial skin micro-movements may include words spoken by an individual (silently or aloud), an individual's identity, an individual's emotional state, an individual's heart rate, an individual's respiratory rate, or any other biometric, emotional, or speech-related indicator value.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 390,653, filed July 20, 2022, U.S. Provisional Patent Application No. 63 / 394,329, filed August 2, 2022, U.S. Provisional Patent Application No. 63 / 438,061, filed January 10, 2023, U.S. Provisional Patent Application No. 63 / 441,183, filed January 26, 2023, and U.S. Provisional Patent Application No. 63 / 487,299, filed February 28, 2023, all of which are incorporated by reference in their entireties.

[0002]

[0002] This disclosure relates generally to the field of discerning information from neuromuscular activity. One example is discerning communication by detecting facial skin movements that occur during subvocalization. Other examples include enabling control-based neuromuscular activity and discerning changes in neuromuscular activity over time. [Background technology]

[0003]

[0003] The human brain and neural activity are complex and involve many subsystems. One of these subsystems is the facial area, which humans use to communicate with others. From birth, humans are trained to activate craniofacial muscles and produce sounds. Even before language skills are fully developed, infants use facial expressions, including microexpressions, to convey deeper information about themselves. However, once language skills are learned, speech is the primary technique humans use to communicate.

[0004]

[0004] The normal process of vocalized speech uses multiple groups of muscles and nerves from the chest and abdomen, through the throat, to the mouth and face. To produce a given phoneme, motor neurons activate muscle groups in the face, larynx, and mouth in preparation for the propulsion of airflow from the lungs, and these muscles continue to move during speech to produce words and sentences. Without this airflow, no sound would be produced from the mouth. Silent speech occurs when airflow from the lungs is absent, but the muscles of the face, larynx, and mouth move to produce or allow the necessary sounds to be produced.

[0005]

[0005] Some of the disclosed embodiments are directed to providing a new approach for extracting meaning from neuromuscular activity that detects facial skin micro-movements that occur during subvocalization, such as unvoiced speech. Summary of the Invention

[0006]

[0006] Embodiments consistent with the present disclosure provide systems, methods, and devices for detecting and using facial movements.

[0007] Some disclosed embodiments may include a system, method, and non-transitory computer-readable medium for identifying an individual using facial skin micro-movements. These embodiments may include operating a wearable coherent light source configured to project light toward a facial region of the individual's head, operating at least one detector configured to receive coherent light reflections from the facial region and output an associated reflection signal, analyzing the reflection signal to determine specific facial skin micro-movements of the individual, accessing a memory correlating multiple facial skin micro-movements with the individual, searching for a match between the determined specific facial skin micro-movements and at least one of the multiple facial skin micro-movements in the memory, initiating a first action if a match is identified, and initiating a second action different from the first action if a match is not identified.

[0008] Some disclosed embodiments may include a system, method, and non-transitory computer-readable medium for interpreting facial skin movement. These embodiments may include projecting light onto a plurality of facial regions of an individual, the plurality of regions including at least a first region and a second region, the first region being closer to at least one of the zygomaticus or the laughing muscle than the second region; receiving reflections from the plurality of regions; detecting a first facial skin movement corresponding to the reflection from the first region and a second facial skin movement corresponding to the reflection from the second region; determining, based on a difference between the first and second facial skin movements, that the reflection from the first region closer to at least one of the zygomaticus or the laughing muscle is a stronger indicator of communication than the reflection from the second region; and, based on the determination that the reflection from the first region is a stronger indicator of communication, processing the reflection from the first region to identify communication and ignoring the reflection from the second region.

[0009] Some disclosed embodiments may include a system, method, and non-transitory computer-readable medium for performing identity verification operations based on facial micro-movements. These embodiments may include receiving a reference signal for verifying in a trusted manner a correspondence between a specific individual and an account at an institution, the reference signal being derived based on reference facial micro-movements detected using first coherent light reflected from the specific individual's face, storing a correlation between the specific individual's identity and the reference signal reflecting the facial micro-movements in a secure data structure, receiving a request to authenticate the specific individual via the institution thereafter, receiving a real-time signal indicative of a second coherent light reflection derived from a second facial micro-movement of the specific individual, comparing the real-time signal to the reference signal stored in the secure data structure, thereby authenticating the specific individual, and, upon authentication, notifying the institution that the specific individual has been authenticated.

[0010]

[0010] Some disclosed embodiments may include a system, method, and non-transitory computer-readable medium for continuous authentication based on facial skin micro-motion. These embodiments may include receiving, during an ongoing electronic transaction, a first signal representing a coherent light reflex associated with a first facial skin micromovement during a first time period; using the first signal to determine the identity of a specific individual associated with the first facial skin micromovement; receiving, during the ongoing electronic transaction, a second signal representing a coherent light reflex associated with a second facial skin micromovement, the second signal being received during a second time period following the first time period; using the second signal to determine that a specific individual is also associated with the second facial skin micromovement; receiving, during the ongoing electronic transaction, a third signal representing a coherent light reflex associated with a third facial skin micromovement, the third signal being received during a third time period following the second time period; using the third signal to determine that the third facial skin micromovement is not associated with the specific individual; and initiating an action based on a determination that the third facial skin micromovement is not associated with the specific individual.

[0011] Some disclosed embodiments may include a system, method, and non-transitory computer-readable medium for performing a thresholding operation to interpret facial skin micro-movements. These embodiments may include detecting a facial micro-movement when there is a lack of perceptible vocalization associated with the facial micro-movement, determining an intensity level of the facial micro-movement, comparing the determined intensity level to a threshold, interpreting the facial micro-movement when the intensity level is above the threshold, and ignoring the facial micro-movement when the intensity level is below the threshold.

[0012] Some disclosed embodiments may include a system, method, and non-transitory computer-readable medium for establishing a non-voiced conversation. These embodiments may include establishing a wireless communication channel to enable unspoken conversation via a first wearable device and a second wearable device, both of which each include a coherent light source and a photodetector configured to detect facial skin micromovements from coherent light reflections; detecting, by the first wearable device, first facial skin micromovements occurring in the absence of perceptible vocalizations; transmitting a first communication from the first wearable device to the second wearable device via the wireless communication channel, the first communication derived from the first facial skin micromovements and transmitted for presentation via the second wearable device; receiving a second communication from the second wearable device via the wireless communication channel, the second communication derived from the second facial skin micromovements detected by the second wearable device; and presenting the second communication to a wearer of the first wearable device.

[0013] Some disclosed embodiments may include systems, methods, and non-transitory computer-readable media for initiating content interpretation operations prior to utterance of the content to be interpreted. These embodiments may include receiving a signal representative of facial skin micro-movements, determining at least one word to be uttered from the signal prior to utterance of the at least one word in a source language, providing an interpretation of the at least one word prior to utterance of the at least one word, and causing the interpretation of the at least one word to be presented as the at least one word is uttered.

[0014] Some disclosed embodiments may include a system, method, and non-transitory computer-readable medium for performing private voice assistant operations. These embodiments may include receiving a signal indicating specific facial skin micromovements reflecting a private request to the assistant, where responding to the private request requires identification of a specific individual associated with the specific facial skin micromovements, accessing a data structure maintaining correlations between the specific individual and multiple facial skin micromovements associated with the specific individual, searching the data structure for a match indicating a correlation between a stored identity of the specific individual and the specific facial skin micromovements, initiating a first action in response to determining that a match exists in the data structure, where the first action includes enabling access to information specific to the specific individual, and initiating a second action different from the first action if no match is identified in the data structure.

[0015] Some disclosed embodiments may include a system, method, and non-transitory computer-readable medium for determining subvocalized phonemes from facial skin micro-movements. These embodiments may include controlling at least one coherent light source to illuminate a first region of a face and a second region of the face, performing a first pattern analysis on light reflected from the first region of the face to determine first facial skin micro-movements in the first region of the face, performing a second pattern analysis on light reflected from the second region of the face to determine second facial skin micro-movements in the second region of the face, and ascertaining at least one subvocalized phoneme using the first facial skin micro-movements in the first region of the face and the second facial skin micro-movements in the second region of the face.

[0016] Some disclosed embodiments may include a system, method, and non-transitory computer-readable medium for generating a synthetic representation of a facial expression. These embodiments may include controlling at least one coherent light source to illuminate a portion of a face, receiving an output signal from a light detector, the output signal corresponding to reflection of the coherent light from the portion of the face, applying speckle analysis to the output signal to determine speckle analysis-based facial skin micro-movements, using the determined speckle analysis-based facial skin micro-movements to identify at least one pre-uttered or uttered word during a period of time, using the determined speckle analysis-based facial skin micro-movements to identify at least one change in facial expression during the period of time, and outputting data for causing a virtual representation of the face to mimic the at least one change in facial expression in conjunction with audio presentation of the at least one word during the period of time.

[0017] Some disclosed embodiments may include a system, method, and non-transitory computer-readable medium for performing operations for attention-related interactions based on facial skin micro-movements. These embodiments may include determining an individual's facial skin micro-movements based on coherent light reflection from a facial region of the individual, determining a particular engagement level of the individual using the facial skin micro-movements, receiving data associated with an expected interaction with the individual, accessing a data structure that correlates information reflecting alternative engagement levels with different presentation techniques, determining a particular presentation technique for the expected interaction based on the particular engagement level and the correlation information, and associating the particular presentation technique with the expected interaction for a subsequent engagement with the individual.

[0018] Some disclosed embodiments may include a system, method, and non-transitory computer-readable medium for performing a speech synthesis operation from detected facial skin micro-movements. These embodiments may include determining facial skin micro-movements of a first individual speaking with a second individual based on light reflection from a facial region of the first individual, accessing a data structure correlating facial micro-movements with words, performing a lookup in the data structure for a particular word associated with the particular facial skin micro-movement, obtaining input related to a preferred speech consumption characteristic of the second individual, adopting the preferred speech consumption characteristic, and synthesizing an audible output of the particular word using the adopted preferred speech consumption characteristic.

[0019] Some disclosed embodiments may include a system, method, and non-transitory computer-readable medium for performing operations for personal presentation of a preliminary utterance. These embodiments may include receiving a reflection signal corresponding to light reflected from a facial region of an individual, using the received reflection signal to determine a particular facial skin micro-movement of the individual in the absence of a perceptible utterance associated with the particular facial skin micro-movement, accessing a data structure correlating facial skin micro-movements with words, performing a lookup in the data structure of a particular unuttered word associated with the particular facial skin micro-movement, and causing audible presentation of the particular unuttered word to the individual prior to utterance of the particular word by the individual.

[0020] Some disclosed embodiments may include a system, method, and non-transitory computer-readable medium for interpreting a speech disorder based on facial movements. These embodiments may include receiving a signal associated with specific facial skin movements of an individual with a speech disorder that affects how the individual pronounces a plurality of words, accessing a data structure including correlations between the plurality of words and a plurality of facial skin movements corresponding to how the individual pronounces the plurality of words, identifying a specific word associated with the specific facial skin movements based on the received signal and the correlations, and generating an output of the specific word for presentation, wherein the output differs from the individual's pronunciation of the specific word.

[0021] Some disclosed embodiments may include a system, method, and non-transitory computer-readable medium for continuous verification of the authenticity of a communication based on optical reflection from facial skin. These embodiments may include generating a first data stream representing a communication by a subject, the communication having a duration, generating a second data stream for verifying the identity of the subject from facial skin optical reflections captured during the duration of the communication, transmitting the first data stream to a destination and transmitting the second data stream to the destination, wherein the second data stream, upon receipt at the destination, is correlated with the first data stream so that the second data stream is usable to repeatedly check that the communication originated from the subject during the duration of the communication.

[0022] Some disclosed embodiments may include a system, method, and non-transitory computer-readable medium for noise suppression utilizing facial skin micromovements. These embodiments may include operating a wearable coherent light source configured to project light toward a facial region of a head of a wearer, operating at least one detector configured to receive coherent light reflections from the facial region associated with the facial skin micromovements and outputting an associated reflection signal, analyzing the reflection signal to determine speech timing based on the facial skin micromovements in the facial region, receiving an audio signal from at least one microphone, the audio signal including voices of words spoken by the wearer along with ambient sounds, correlating the reflection signal with the received audio signal based on the speech timing to determine portions of the audio signal associated with the words spoken by the wearer, and outputting the determined portions of the audio signal associated with the words spoken by the wearer while omitting output of other portions of the audio signal that do not include the words spoken by the wearer.

[0023] Some disclosed embodiments may include a system, method, and non-transitory computer-readable medium for providing private answers to unvoiced questions. These embodiments may include receiving a signal indicative of a particular facial micro-movement in the absence of perceptible vocalization, accessing a data structure correlating the facial micro-movement with words, using the received signal to perform a lookup in the data structure for a particular word associated with the particular facial micro-movement, determining a query from the particular word, accessing at least one data structure to perform a lookup for an answer to the query, and generating a discrete output including the answer to the query.

[0024] Some disclosed embodiments may include a system, method, and non-transitory computer-readable medium for executing control commands based on facial skin micromovements. These embodiments may include operating at least one coherent light source to enable illumination of a lip portion of a face, receiving a specific signal representing a coherent light reflectance associated with a specific non-lip portion of the facial skin micromovement, accessing a data structure that associates a plurality of non-lip facial skin micromovements with control commands, identifying in the data structure a specific control command associated with the specific signal associated with the specific non-lip facial skin micromovement, and executing the specific control command.

[0025] Some disclosed embodiments may include a system, method, and non-transitory computer-readable medium for detecting changes in neuromuscular activity over time. These embodiments may include establishing a baseline of neuromuscular activity from coherent light reflexes associated with past skin micromovements, receiving a current signal representative of coherent light reflexes associated with an individual's current skin micromovements, identifying a deviation of the current skin micromovements from the baseline of neuromuscular activity, and outputting an indicator of the deviation.

[0026] Some disclosed embodiments may include systems, methods, and non-transitory computer-readable media for projecting graphical content and interpreting nonverbal utterances. These embodiments may include operating a wearable light source configured to project light in a graphical pattern onto a facial area of ​​an individual, the graphical pattern configured to visually convey information; receiving an output signal from a sensor corresponding to a portion of the light reflected from the facial area; determining facial skin micro-movements associated with the nonverbalization from the output signal; and processing the output signal to interpret the facial skin micro-movements.

[0027] Some disclosed embodiments may include a system, method, and non-transitory computer-readable medium for interpreting facial skin micromovements. These embodiments may include receiving coherent light reflections from a facial region associated with an individual's facial skin micromovements, outputting a reflection signal associated with the light reflections, capturing a voice produced by the individual, outputting an audio signal associated with the captured sound, and using both the reflection signal and the audio signal to generate an output corresponding to a word produced by the individual.

[0028] Some disclosed embodiments may include a system, method, and non-transitory computer-readable medium for interpreting facial skin micromovements. These embodiments may include receiving, during a first time period, a first signal representing a preliminary vocalization's facial skin micromovements, receiving, during a second time period following the first time period, a second signal representing a voice, analyzing the voice to identify words spoken during the second time period, correlating the words spoken during the second time period with the preliminary vocalization's facial skin micromovements received during the first time period, storing the correlation, receiving, during a third time period, a third signal representing a preliminary vocalization's facial skin micromovements received in the absence of vocalization, using the stored correlation to identify a language associated with the third signal, and outputting the language.

[0029] Some disclosed embodiments may include a system, method, and non-transitory computer-readable medium for operating a multi-function earpiece. These embodiments may include operating a speaker integrated with an ear-worn housing associated with the multi-function earpiece to present a voice, operating a light source integrated with the ear-worn housing to project light toward the wearer's facial skin, operating a light detector integrated with the ear-worn housing and configured to receive reflections from the skin corresponding to the wearer's facial skin micro-movements indicative of a pre-spoken word, and simultaneously presenting the voice from the speaker, projecting the light toward the skin, and detecting the received reflections indicative of the pre-spoken word.

[0030] Some disclosed embodiments may include a driver for integrating with a software program and for enabling the neuromuscular detection device to interface with the software program, the driver comprising: an input handler for receiving non-audible muscle activation signals from the neuromuscular detection device, a lookup component for mapping certain ones of the non-audible activation signals to corresponding commands within the software program, a signal processing module for receiving the non-audible muscle activation signals from the input handler, providing certain ones of the non-audible muscle activation signals to the lookup component, and receiving an output as the corresponding commands, and a communications module for communicating the corresponding commands to the software program, thereby enabling control within the software program based on the non-audible muscle activity detected by the neuromuscular detection device.

[0031] Some disclosed embodiments may include a system, method, and non-transitory computer-readable medium for performing a context-driven facial micro-motor behavior, including receiving, during a first time period, a first signal representing a first coherent light reflex associated with a first facial skin micro-motor, analyzing the first coherent light reflex to determine a first plurality of words associated with the first facial skin micro-motor, receiving first information indicative of a first contextual condition under which the first facial skin micro-motor occurred, receiving, during a second time period, a second signal representing a second coherent light reflex associated with a second facial skin micro-motor, analyzing the second coherent light reflex to determine a second plurality of words associated with the second facial skin micro-motor, and receiving first information indicative of a first contextual condition under which the first facial skin micro-motor occurred. receiving second information indicating a second contextual condition when the second information occurs; accessing a plurality of control rules correlating a plurality of actions with the plurality of contextual conditions, wherein a first control rule specifies a format of private presentation based on the first contextual condition and a second control rule specifies a format of non-private presentation based on the second contextual condition; upon receiving the first information, implementing the first control rule to privately output the first plurality of words; and upon receiving the second information, implementing the second control rule to non-privately output the second plurality of words.

[0032] Some disclosed embodiments may include a system, method, and non-transitory computer-readable medium for extracting a response to content based on facial skin micro-movements. These embodiments may include determining facial skin micro-movements of an individual during a period in which the individual consumes content based on coherent light reflections from a facial region of the individual, determining at least one specific micro-expression from the facial skin micro-movements, accessing at least one data structure including correlations between a plurality of micro-expressions and a plurality of non-verbal perceptions, determining a specific non-verbal perception of the content consumed by the individual based on the at least one specific micro-expression and the correlations in the data structure, and initiating an action associated with the specific non-verbal perception.

[0033] Some disclosed embodiments may include a system, method, and non-transitory computer-readable medium for removing noise from facial skin micro-movement signals. These embodiments may include operating a light source to illuminate a facial skin area of ​​an individual during a period in which the individual is engaged in at least one non-speech-related physical activity, receiving a signal representative of light reflections from the facial skin area, analyzing the received signal to identify a first reflection component indicative of pre-speech facial skin micro-movements and a second reflection component associated with the at least one non-speech-related physical activity, and filtering out the second reflection component from the first reflection component indicative of pre-speech facial skin micro-movements to enable interpretation of a word.

[0034]

[0034] Consistent with other disclosed embodiments, a non-transitory computer-readable storage medium can store program instructions that, when executed by at least one processing device, perform any of the methods described herein.

[0035]

[0035] The foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the scope of the claims.

[0036]

[0036] The accompanying drawings, which are incorporated in and constitute a part of this disclosure, illustrate various disclosed embodiments. [Brief explanation of the drawings]

[0037] [Figure 1] FIG. 1 is a schematic diagram of a user using a first exemplary speech detection system consistent with certain embodiments of the present disclosure. [Figure 2A] FIG. 2A is a schematic diagram of a user using a second exemplary speech detection system consistent with certain embodiments of the present disclosure. [Figure 2B] FIG. 2B is a perspective view of a user using a third exemplary speech detection system consistent with certain embodiments of the present disclosure. [Figure 3] FIG. 3 is a schematic diagram of a user using a fourth exemplary speech detection system consistent with certain embodiments of the present disclosure. [Figure 4] FIG. 4 is a block diagram illustrating some of the components of a speech detection system and a remote processing system consistent with some embodiments of the present disclosure. [Figure 5A] FIG. 5A is a schematic diagram of a portion of a speech detection system when detecting facial skin micro-movements, consistent with some embodiments of the present disclosure. [Figure 5B] FIG. 5B is a schematic diagram of a portion of a speech detection system when detecting facial skin micro-movements, consistent with some embodiments of the present disclosure. [Figure 6] FIG. 6 is a schematic illustration of a reflection image associated with light reflections received from an area of ​​a facial region associated with a single spot, consistent with certain embodiments of the present disclosure. [Figure 7] FIG. 7 is a block diagram of a memory consistent with disclosed embodiments. [Figure 8] FIG. 8 is an example alternative action utterance detection process diagram consistent with certain embodiments of the present disclosure. [Figure 9]FIG. 9 is a flowchart of an exemplary process for identifying an individual, consistent with certain embodiments of the present disclosure. [Figure 10] FIG. 10 is a flowchart of an exemplary process for identifying an individual using facial skin micro-motion, consistent with certain embodiments of the present disclosure. [Figure 11] FIG. 11 is an illustration of two exemplary use cases for interpreting facial skin motion from light reflection, consistent with certain embodiments of the present disclosure. [Figure 12] FIG. 12 is an illustration of another exemplary use case for interpreting facial skin motion from light reflection, consistent with certain embodiments of the present disclosure. [Figure 13] FIG. 13 is a flowchart of an exemplary process for interpreting facial skin movement consistent with certain embodiments of the present disclosure. [Figure 14] FIG. 14 is a schematic diagram of the operation of an example authentication service configured to provide personal identity verification based on facial micro-movements, consistent with certain embodiments of the present disclosure. [Figure 15] FIG. 15 is a simplified diagram of an exemplary system for personal identity verification using facial micro-movements, consistent with certain embodiments of the present disclosure. [Figure 16A] FIG. 16A is a simplified diagram of an exemplary system for personal identity verification using facial micro-movements, consistent with certain embodiments of the present disclosure. [Figure 16B] FIG. 16B is a simplified diagram of an exemplary system for personal identity verification using facial micro-movements, consistent with certain embodiments of the present disclosure. [Figure 17A] FIG. 17A is a flowchart of an exemplary process for verifying an individual's identity using facial micro-movements, consistent with certain embodiments of the present disclosure. [Figure 17B] FIG. 17B is a flowchart of an exemplary process for generating a reference signal for individual identity verification, consistent with certain embodiments of the present disclosure. [Figure 18]FIG. 18 is a schematic diagram of an exemplary authentication system and service configured to provide continuous authentication of an individual based on facial skin micro-movements, consistent with certain embodiments of the present disclosure. [Figure 19] FIG. 19 is a simplified diagram of an exemplary system configured to provide continuous authentication of an individual based on facial micro-movements, consistent with certain embodiments of the present disclosure. [Figure 20] FIG. 20 is a flowchart of an exemplary process for continuous authentication of an individual using facial micro-movements, consistent with certain embodiments of the present disclosure. [Figure 21] FIG. 21 is a flowchart of another exemplary process for continuous authentication of an individual using facial micro-movements, consistent with certain embodiments of the present disclosure. [Figure 22] FIG. 22 is a flowchart of another exemplary process for continuous authentication of an individual using facial micro-movements, consistent with certain embodiments of the present disclosure. [Figure 23] FIG. 23 is a flowchart of another exemplary process for continuous authentication of an individual using facial micro-movements, consistent with certain embodiments of the present disclosure. [Figure 24] FIG. 24 includes a series of displacement versus time charts including threshold levels associated with a number of facial regions, consistent with certain embodiments of the present disclosure. [Figure 25A] FIG. 25A is a schematic illustration of exemplary displacement levels of facial micro-movements for which a threshold trigger mechanism may be employed, consistent with certain embodiments of the present disclosure. [Figure 25B] FIG. 25B is a schematic illustration of exemplary displacement levels of facial micro-movements for which a threshold trigger mechanism may be employed, consistent with certain embodiments of the present disclosure. [Figure 26] FIG. 26 is a block diagram of an example speech detection system that uses a threshold and threshold adjustment as a triggering mechanism, consistent with certain embodiments of the present disclosure. [Figure 27] FIG. 27 is a displacement versus time graph including background noise, consistent with certain embodiments of the present disclosure. [Figure 28A] FIG. 28A illustrates an example of measuring potential skin differences to determine facial micromovements, consistent with certain embodiments of the present disclosure. [Figure 28B] FIG. 28B illustrates an example of measuring potential skin differences to determine facial micromovements, consistent with certain embodiments of the present disclosure. [Figure 29] FIG. 29 is a flowchart illustrating an exemplary method for interpreting or disregarding facial micro-movements using thresholds, consistent with certain embodiments of the present disclosure. [Figure 30] FIG. 30 is a schematic diagram of a system configured to enable non-vocal conversation between individuals, consistent with certain embodiments of the present disclosure. [Figure 31] FIG. 31 is a schematic illustration of an exemplary processing of detected facial skin micro-motion of an individual, consistent with certain embodiments of the present disclosure. [Figure 32] FIG. 32 is a schematic diagram of another system configured to enable non-vocal conversation between individuals, consistent with certain embodiments of the present disclosure. [Figure 33] FIG. 33 is a flowchart of an exemplary process for establishing a non-vocal conversation, consistent with certain embodiments of the present disclosure. [Figure 34] FIG. 34 is a schematic diagram of an exemplary content interpretation process that is initiated prior to the utterance of the content to be interpreted, consistent with certain embodiments of the present disclosure. [Figure 35] FIG. 35 is a flowchart of an exemplary process for initiating content interpretation prior to utterance of the content to be interpreted, consistent with certain embodiments of the present disclosure. [Figure 36] FIG. 36 illustrates an example protocol for performing private voice assistant operations using different facial skin micro-movements, consistent with embodiments of the present disclosure. [Figure 37] FIG. 37 illustrates an example of a second action that is initiated if a match is not identified within the exemplary data structure, consistent with an embodiment of the present disclosure. [Figure 38] FIG. 38 shows a flowchart of an example process for performing private voice assistant operations consistent with embodiments of the present disclosure. [Figure 39] FIG. 39 is an exemplary diagram illustrating how different areas of facial skin are used to detect subvocalized phonemes, consistent with certain embodiments of the present disclosure. [Figure 40] FIG. 40 shows three graphs illustrating exemplary alternative timing for completing a process that includes detecting sub-vocalized phonemes, consistent with embodiments of the present disclosure. [Figure 41] FIG. 41 is a flowchart of an exemplary process for determining subvocalized phonemes from facial skin micro-movements, consistent with certain embodiments of the present disclosure. [Figure 42A] FIG. 42A is a perspective view of a user wearing an exemplary headset and a resulting virtual representation of one of the user's facial expressions, consistent with certain embodiments of the present disclosure. [Figure 42B] FIG. 42B is another perspective view of a user wearing an exemplary headset and resulting virtual representations of different facial expressions of the user, consistent with certain embodiments of the present disclosure. [Figure 43] FIG. 43 is a block diagram illustrating an example operating environment for generating synthetic representations of facial expressions consistent with certain embodiments of the present disclosure. [Figure 44] FIG. 44 is a block diagram illustrating an exemplary system for generating synthetic representations of facial expressions and / or determining spoken phonemes from facial skin micro-movements, consistent with certain embodiments of the present disclosure. [Figure 45] FIG. 45 is a flowchart illustrating an exemplary method for generating a synthetic representation of a facial expression and / or determining spoken phonemes from facial skin micro-movements, consistent with certain embodiments of the present disclosure. [Figure 46]FIG. 46 is a flowchart illustrating another exemplary method for generating a synthetic representation of a facial expression and / or determining spoken phonemes from facial skin micro-movements, consistent with certain embodiments of the present disclosure. [Figure 47] FIG. 47 is a schematic diagram of an exemplary process for understanding a presentation method based on facial skin micro-motion, consistent with certain embodiments of the present disclosure. [Figure 48] FIG. 48 is a schematic illustration of a user using an exemplary system for attention-related interaction based on facial skin micro-movements, consistent with certain embodiments of the present disclosure. [Figure 49] FIG. 49 is a schematic diagram of receiving an expected interaction via a smartphone, consistent with certain embodiments of the present disclosure. [Figure 50] FIG. 50 is a flowchart of an exemplary process for understanding a presentation method based on facial skin micro-motion, consistent with certain embodiments of the present disclosure. [Figure 51] FIG. 51 illustrates a first individual wearing a speech detection system while communicating with at least one second individual, consistent with certain embodiments of the present disclosure. [Figure 52] FIG. 52 illustrates a flowchart of an exemplary process for initiating content interpretation prior to utterance of the content to be interpreted, consistent with certain embodiments of the present disclosure. [Figure 53A] FIG. 53A is a schematic illustration of an audible presentation of an unuttered word before it is uttered, consistent with certain embodiments of the present disclosure. [Figure 53B] FIG. 53B is a schematic illustration of an audible presentation of an unuttered word before utterance, consistent with certain embodiments of the present disclosure. [Figure 54] FIG. 54 is a block diagram of an exemplary speech detection system that uses received reflections to determine unspoken words from facial micro-movements and triggers audible presentation, consistent with certain embodiments of the present disclosure. [Figure 55] FIG. 55 shows an exemplary schematic diagram of synthesis translation between languages, consistent with some embodiments of the present disclosure. [Figure 56]FIG. 56 illustrates an exemplary additional feature of personal presentation of preliminary utterances consistent with certain embodiments of the present disclosure. [Figure 57] FIG. 57 is a flowchart of an exemplary method for determining an unuttered word from facial micro-movements using received reflections and triggering an audible presentation, consistent with certain embodiments of the present disclosure. [Figure 58] FIG. 58 is a perspective view of an individual using a first exemplary speech detection system, consistent with certain embodiments of the present disclosure. [Figure 59A] FIG. 59A is a schematic diagram of a portion of a speech detection system when detecting facial skin micro-movements, consistent with some embodiments of the present disclosure. [Figure 59B] FIG. 59B is a schematic diagram of a portion of a speech detection system when detecting facial skin micro-movements, consistent with certain embodiments of the present disclosure. [Figure 60] FIG. 60 is a block diagram illustrating exemplary components of a first example of a speech detection system consistent with certain embodiments of the present disclosure. [Figure 61] FIG. 61 is a flowchart of an exemplary method for determining facial skin micro-movements, consistent with certain embodiments of the present disclosure. [Figure 62] FIG. 62 illustrates an exemplary system for correcting speech impediments based on facial movements, consistent with certain embodiments of the present disclosure. [Figure 63] FIG. 63 is a flowchart of an exemplary process for correcting speech impediments based on facial movements, consistent with certain embodiments of the present disclosure. [Figure 64] FIG. 64 is a schematic diagram of an exemplary speech detection system that sends two data streams to a destination to verify the authenticity of a communication, consistent with certain embodiments of the present disclosure. [Figure 65] FIG. 65 is a schematic diagram of example functionality used to authenticate a communication at a destination, consistent with certain embodiments of the present disclosure. [Figure 66]FIG. 66 is a flowchart illustrating an example method for verifying the authenticity of a communication using received reflections, consistent with certain embodiments of the present disclosure. [Figure 67] FIG. 67 illustrates an exemplary head-mounted system for noise suppression consistent with certain embodiments of the present disclosure. [Figure 68] FIG. 68 illustrates an example of audio signal processing for noise suppression, consistent with certain embodiments of the present disclosure. [Figure 69] FIG. 69 is a flowchart of an exemplary process for noise suppression consistent with certain embodiments of the present disclosure. [Figure 70] FIG. 70 illustrates an exemplary system for providing private answers to silent questions consistent with embodiments of the present disclosure. [Figure 71] FIG. 71 illustrates an example of an image data application that may be used to provide private answers to unvoiced questions, consistent with embodiments of the present disclosure. [Figure 72] FIG. 72 illustrates a flowchart of an exemplary process for providing private answers to silent questions, consistent with embodiments of the present disclosure. [Figure 73] FIG. 73 is a schematic diagram of an individual using a first exemplary speech detection system, consistent with certain embodiments of the present disclosure. [Figure 74] FIG. 74 is a schematic diagram of two individuals each using an exemplary speech detection system, consistent with certain embodiments of the present disclosure. [Figure 75] FIG. 75 is a flowchart of an exemplary method for performing unvoiced audio control, consistent with certain embodiments of the present disclosure. [Figure 76] FIG. 76 is a schematic illustration of an exemplary timeline of the progression of a medical condition that may be detectable by measuring skin micromotion over time, consistent with some embodiments of the present disclosure. [Figure 77]FIG. 77 is a block diagram of an exemplary system capable of detecting changes in neuromuscular activity over time, consistent with certain embodiments of the present disclosure. [Figure 78] FIG. 78 is a block diagram of an example function for detecting deviations in a medical condition, consistent with certain embodiments of the present disclosure. [Figure 79] FIG. 79 is a flowchart illustrating an exemplary method for detecting changes in neuromuscular activity over time using received light reflexes, consistent with certain embodiments of the present disclosure. [Figure 80] FIG. 80 is a schematic illustration of detecting non-verbal information from an individual using projected graphic patterns, consistent with certain embodiments of the present disclosure. [Figure 81] FIG. 81 is a schematic illustration of varying a projected graphic pattern, consistent with some embodiments of the present disclosure. [Figure 82] FIG. 82 is a flowchart of an exemplary process for detecting non-verbal information from an individual using projected graphic patterns, consistent with certain embodiments of the present disclosure. [Figure 83] FIG. 83 shows an exemplary embodiment of a user wearing a head-mounted system for interpreting facial skin micro-movements. [Figure 84] FIG. 84 shows a flowchart of an exemplary method for interpreting facial skin micro-movements. [Figure 85A] FIG. 85A illustrates an exemplary embodiment of a training operation for interpreting facial skin micro-movements in first through third time periods, consistent with certain disclosed embodiments. [Figure 85B] FIG. 85B illustrates an exemplary embodiment of a training operation for interpreting facial skin micro-movements in first through third time periods, consistent with certain disclosed embodiments. [Figure 85C] FIG. 85C illustrates an exemplary embodiment of a training operation for interpreting facial skin micro-movements in first through third time periods, consistent with certain disclosed embodiments. [Figure 86]FIG. 86 is a flow diagram of an example of the first through third periods shown in FIGS. 85A-85C with an exemplary additional extension period, consistent with certain disclosed embodiments. [Figure 87] FIG. 87 is a flowchart of an exemplary method for interpreting facial skin micro-movements, consistent with certain disclosed embodiments. [Figure 88] FIG. 88 is a schematic illustration of a user wearing an exemplary headset with added facial micro-motion detection capabilities consistent with certain embodiments of the present disclosure. [Figure 89] FIG. 89 is a schematic diagram of an exemplary facial micro-movement detection process consistent with certain embodiments of the present disclosure. [Figure 90] FIG. 90 is a flowchart of an exemplary process for operating a multi-function earpiece, consistent with certain embodiments of the present disclosure. [Figure 91] FIG. 91 is a schematic diagram of a user wearing an exemplary headset of an alternative form factor, consistent with certain embodiments of the present disclosure. [Figure 92] FIG. 92 illustrates a block diagram of exemplary drivers for interfacing with software programs and devices consistent with disclosed embodiments. [Figure 93] FIG. 93 shows a schematic diagram of an exemplary driver for integrating with a software program and neuromuscular detection device, consistent with disclosed embodiments. [Figure 94] FIG. 94 shows a schematic diagram of an exemplary system for integrating with and enabling devices to interface with software programs, consistent with embodiments of the present disclosure. [Figure 95] FIG. 95 is a block diagram illustrating an exemplary operating environment for generating context-driven facial micromotor output consistent with certain embodiments of the present disclosure. [Figure 96]FIG. 96 is a block diagram illustrating an exemplary system for generating context-driven facial micromotor output consistent with certain embodiments of the present disclosure. [Figure 97] FIG. 97 is a flowchart illustrating an exemplary method for generating context-driven facial micromotor output consistent with certain embodiments of the present disclosure. [Figure 98] FIG. 98 is a flowchart illustrating another exemplary method for generating context-driven facial micromotor output consistent with certain embodiments of the present disclosure. [Figure 99] FIG. 99 is a schematic illustration of a user wearing an exemplary headset and the resulting facial micro-movement based context-driven output consistent with certain embodiments of the present disclosure. [Figure 100] FIG. 100 is a schematic diagram of an exemplary system for extracting responses to content based on facial skin micro-movements, consistent with certain embodiments of the present disclosure. [Figure 101] FIG. 101 includes block diagrams of two example use cases for triggering actions based on reactions to content, consistent with certain embodiments of the present disclosure. [Figure 102] FIG. 102 is a flowchart of an exemplary process for extracting a response to content based on facial skin micro-movements, consistent with certain embodiments of the present disclosure. [Figure 103] FIG. 103 illustrates an individual performing a first non-speech-related activity (e.g., walking) and a second non-speech-related activity (e.g., sitting) while wearing a speech recognition system consistent with an embodiment of the present disclosure. [Figure 104] FIG. 104 illustrates an exemplary expanded view of the speech detection system of FIG. 103, consistent with an embodiment of the present disclosure. [Figure 105] FIG. 105 shows an exemplary comparison between a first signal of an individual performing speech-related facial skin movements while walking and a second signal of an individual performing speech-related facial skin movements while sitting, consistent with an embodiment of the present disclosure. [Figure 106] FIG. 106 illustrates an exemplary decomposition and classification of an electronic representation of an optical signal into a first reflectance component indicative of pre-speech facial skin micro-movements and a second reflectance component associated with at least one non-speech related physical activity, consistent with an embodiment of the present disclosure. [Figure 107] FIG. 107 illustrates an exemplary second reflected component of a light signal reflecting from a facial region of an individual simultaneously engaging in a first physical activity and a second physical activity, consistent with an embodiment of the present disclosure. [Figure 108] FIG. 108 shows a flowchart of an exemplary process for removing noise from facial skin micromotion, consistent with certain embodiments of the present disclosure. [Figure 109] FIG. 109 illustrates another exemplary decomposition and classification of a representation of an optical signal to identify a first reflection component indicative of pre-phonation facial skin micro-movements, consistent with an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0038]

[0148] The following detailed description includes reference to the accompanying drawings. Wherever possible, the same reference numbers are used in the drawings and the description to refer to the same or similar parts. While several exemplary embodiments are described herein, modifications, adaptations, and other implementations are possible. For example, substitutions, additions, or modifications can be made to the components shown in the drawings, and the exemplary methods described herein can be modified by substituting, rearranging, deleting, or adding steps to the disclosed methods. Therefore, the following detailed description is not limited to the disclosed embodiments and examples. Instead, the appropriate scope is defined by the appended claims.

[0039]

[0149] Various terms used in this specification and claims may be defined or summarized differently when described in connection with different disclosed embodiments. It is understood that the definition, summary, and explanation of a term in each instance applies to all instances, even when not repeated, except where a transitional definition, explanation, or summary would render the embodiment inoperable. It is also understood that once a term is defined in this specification, that definition applies to all other uses of the term in this specification, unless there is an inherent contradiction. Furthermore, the illustrative embodiments in the drawings and their descriptions should not be considered definitions of the terms in the claims, but are non-limiting examples used to describe particular embodiments.

[0040]

[0150] Throughout, this disclosure refers to "embodiments" and "disclosed embodiments," which refer to examples of ideas, concepts, and / or implementations of the inventions described herein. Many related and unrelated embodiments are described throughout this disclosure. The fact that some "disclosed embodiments" are described as exhibiting a feature or characteristic does not mean that other disclosed embodiments necessarily share that feature or characteristic.

[0041]

[0151] This disclosure employs open-ended permissive language indicating, for example, that some embodiments may employ, may encompass, or may include a particular feature. Use of the term "may" and other open-ended terminology is intended to indicate that not all embodiments employ a particular disclosed feature, but that at least one embodiment employs a particular disclosed feature.

[0042]

[0152] Different embodiments of the present disclosure may include systems, methods, and / or computer-readable media containing instructions. A system refers to at least two interconnected or interrelated components or parts that work together to achieve a common purpose, function, or subfunction. A method refers to at least two steps, actions, or techniques to be followed to complete a task or subtask, reach a goal, or reach a next step. A computer-readable medium containing instructions refers, for example, to any storage mechanism containing program code instructions to be executed by a computer processor. Examples of computer-readable media are further described elsewhere in this disclosure. The instructions can be written in any type of computer programming language, such as an interpreted language (e.g., scripting languages ​​such as HTML and JavaScript), a procedural or functional language (e.g., C or Pascal that can be compiled to convert to executable code), an object-oriented programming language (e.g., Java or Python), a logic programming language (e.g., Prolog or answer set programming), and / or any other programming language. The instructions executed by the at least one processor may include implementing one or more program code instructions in hardware, software (including one or more signal processing and / or application specific integrated circuits), firmware, or any combination thereof, as described above. Having a processor perform an operation may include causing the processor to calculate, execute, or perform one or more arithmetic, mathematical, logical, reasoning, or inference steps.

[0043]

[0153] Some disclosed embodiments may include detecting facial skin micromovements. The term "facial skin micromovements" broadly refers to facial skin movements that may be detectable using sensors but may not be easily detectable with the naked eye. Facial skin micromovements include various types of movements, including involuntary movements caused by muscle recruitment and other types of small skin deformations, ranging from micrometers to millimeters and lasting fractions of a second to several seconds. In some cases, facial skin micromovements are part of larger-scale skin movements visible to the naked eye (e.g., a smile may involve many facial skin micromovements). In other cases, facial skin micromovements are not part of larger-scale skin movements visible to the naked eye. Such micromovements may occur over a facial area of ​​several square millimeters, but they may occur in facial skin surface areas of less than 1 square centimeter, less than 1 square millimeter, less than 0.1 square millimeters, less than 0.01 square millimeters, or even smaller areas. In some embodiments, facial skin micromovements correspond to one or more muscle recruitments in the facial region of an individual's head. The facial region may include specific anatomical areas, such as a portion of the cheek above the mouth, a portion of the cheek below the mouth, a portion of the midchin, a portion of the cheek below the eyes, the neck, the chin, and other areas associated with specific muscle recruitment that may cause facial skin micromovements. In some embodiments, the specific muscle may be connected to cutaneous tissue and not connected to any bone. In particular, the specific muscle may be located in subcutaneous tissue associated with cranial nerve V or cranial nerve VII. As described in more detail herein, the first facial skin micromovement 522A and the second facial skin micromovement 522B of FIG. 5A are non-limiting examples of facial skin micromovements consistent with the present disclosure.

[0044]

[0154] When certain muscles contract, they pull on the facial skin, causing facial skin movements. Some of the movements that occur when certain muscles contract can be micro-movements. By way of example, in the context of the present disclosure, certain muscles that can cause facial skin micro-movements can be broadly divided into four groups: orbital, nasal, oral, and lingual. The orbital group of facial muscles includes two muscles associated with the orbit. These muscles control eyelid movements, which are important for protecting the cornea from damage. Both of these muscles are innervated by cranial nerve VII. The nasal group of facial muscles is associated with the movement of the nose and the skin surrounding it. This group includes three muscles, all of which are innervated by cranial nerve VII. The oral group is the most important group for facial expression and is responsible for moving the mouth and lips. Such movements are necessary when singing, whistling, or emphasizing vocal communication. The oral group of muscles consists of the orbicularis oris, buccinator, and various smaller muscles. In certain embodiments, the disclosed system can monitor facial skin micro-movements corresponding to recruitment of the buccinator muscles. The buccinator muscle is located relatively deep between the mandible and maxilla compared to other facial muscles. The tongue group of muscles consists of four intrinsic muscles (e.g., superior longitudinal, inferior longitudinal, and transverse) used to change the shape of the tongue and four extrinsic muscles (e.g., genioglossus, hyoglossus, styloglossus, and palatoglossus) used to change the position of the tongue. Any of the tongue muscles listed above can cause tongue movements that can be detected by analyzing detected facial skin micromovements. As described in more detail herein, muscle fibers 520 in Figures 5A and 5B are non-limiting examples of facial muscles that cause facial skin micromovements consistent with the present disclosure.

[0045]

[0155] Consistent with the present disclosure, facial skin micromovements may be detected during subvocalization. The term "subvocalization" refers to any speech-related activity that occurs without speech, before speech, or imperceptibly preceding speech. In one embodiment, speech-related activity may include silent speech (i.e., when airflow from the lungs is absent but facial muscles produce the desired sound). In another embodiment, speech-related activity may include silent speaking (i.e., when words are produced with some airflow from the lungs but imperceptibly using an audio sensor). In yet another embodiment, speech-related activity may include pre-phonatory muscle recruitment (i.e., subvocalization that occurs before the onset of speech is sometimes referred to herein as pre-phonation). In some cases, pre-phonatory facial skin micromovements may be triggered by voluntary muscle recruitment that occurs when certain craniofacial muscles begin to produce a word. In other cases, pre-vocalization facial skin micromovements may be triggered by involuntary facial muscle recruitment by an individual as certain craniofacial muscles prepare to utter a word. For example, involuntary facial muscle recruitment may occur between 0.1 and 0.5 seconds before actual utterance. In some cases, the proposed system can use facial skin micromovements that occur during detected subvocalization to identify the word being uttered. Determining the word a user intends to say before the user actually utters it can have many advantages, as the system can begin processing the word without waiting for the user to articulate and enunciate it. In one example, the disclosed system can generate subtitles for live broadcasts without delay. In another example, the disclosed system can translate what a user is saying into different languages ​​in real time. Furthermore, because the disclosed system can detect words before they are uttered, the actual utterance of these words is not a prerequisite. Therefore, facial skin micromovements that occur during subvocalization can be detected even in the absence of perceptible utterance. Movements of the facial skin or muscles that are not vocalized but convey speech-related information are referred to herein as unvoiced speech.Detecting unvoiced speech may have a variety of uses, including, but not limited to, enabling silent communication with other users, initiating commands, or enabling interaction with virtual personal assistance. As described in more detail herein, the subvocalization decoding module 708 of FIG. 7 is a non-limiting example of a software module used to decode facial skin micromovements of some subvocalizations.

[0046]

[0156] In some embodiments, the detection of facial skin micro-movements is performed using a speech detection system. While the abbreviation "speech detection system" is employed, it should be understood that the system may alternatively or additionally be configured to detect non-spoken commands, facial expressions, or emotions. The system may also be used for user authentication. The speech detection system may include any device of a group of devices operatively coupled to one another. As used herein, the term "system" includes any device or group of devices operatively connected together and configured to perform a function. In some embodiments, the system may include a computer (e.g., a desktop computer, a laptop computer, a server, a smartphone, a portable digital assistant (PDA), or similar device) or multiple computers or servers operatively connected to one another (e.g., using wires or wirelessly) to share information and / or data. The computer(s) may include a dedicated computer (e.g., hardwired and coded to perform a desired function) or a general-purpose computer (e.g., using software to perform any desired function). In some embodiments, the system may include a cloud server. As described elsewhere in this disclosure, the cloud server may be a computer platform that provides services over a network such as the Internet. In one embodiment, the speech detection system may include a wearable housing, a coherent or non-coherent light source, a photodetector, and a processor. However, the specific list of components above is not intended to limit the systems covered by this disclosure. Numerous variations and / or modifications can be made to the exemplary speech detection system, as will be understood by those skilled in the art with the benefit of this disclosure. For example, not all components are essential for detecting facial skin micro-movements in all cases.Additionally, components may be rearranged into various configurations while providing the functionality of the various disclosed embodiments. In some cases, speech detection systems according to some embodiments of the present disclosure need not be wearable but can be aimed at the skin from a location not connected to the human body. A wearable or non-wearable system can project coherent light toward a user's facial area, analyze the reflected light, and determine facial skin micro-movements. Alternatively, in other cases, speech detection systems according to some embodiments of the present disclosure need not include a coherent light source. Specifically, the light detector can be an ultra-high-resolution image sensor (e.g., greater than 120 megapixels) or any other sensor capable of detecting facial skin micro-movements, and the detection of facial skin micro-movements can be achieved using one or more image processing algorithms. As described in more detail herein, speech detection system 100 of FIGS. 1-3 is a non-limiting example of a speech detection system consistent with the present disclosure. As shown in these examples, the system includes a wearable housing 110, a light source 410, a light detector 412, and a processing device 400.

[0047]

[0157] Some disclosed embodiments include a wearable housing configured to be worn on an individual's head. The term "wearable housing" broadly includes any structure or enclosure designed to connect to a human head, such as configured to be worn by a user. Such a wearable housing may be configured to house or support one or more electronic components or sensors. In one example, the wearable housing is configured to couple with eyeglasses. In another example, the wearable housing is coupled with wireless earbuds. The wearable housing may have a cross-section that is button-shaped, P-shaped, square, rectangular, rounded rectangular, or any other regular or irregular shape that can be worn by a user. Such a structure may allow the wearable housing to be worn on, in, or around a body part associated with the user's head (e.g., over the ear, in the ear, around the neck). The wearable housing may be made from plastic, metal, composite material, a combination of two or more of plastic, metal, and composite material, or other suitable material. Consistent with disclosed embodiments, the housing may be worn on the ear. There are several ways in which the housing can be attached to the ear. 1. In-the-ear (ITE): The housing is inserted directly into the ear canal and may be held in place by the shape of the ear. Examples include wireless earbuds and earplugs. In some cases, the housing may be custom-made to fit the specific shape of an individual's ear and sit on the concha. 2. Behind-the-ear (BTE): The housing may sit behind the ear and have a small tube that extends into the ear canal. Examples include hearing aids and Bluetooth headsets. 3. Over-the-ear (OTE): The housing sits on the ear and may be held in place by a headband or other support. Examples include headphone and earmuff-like structures. 4. Over-the-head (OTH): The housing may be held in place by a headband that covers the top of the head.In other embodiments, the wearable housing may be attached to a secondary device, such as eyeglasses (sunglasses or vision-correcting glasses), a hat, a helmet, a visor, or any other type of head-wearable device. In some cases, the wearable housing may be attached to the secondary device using at least one adapter. Specifically, the at least one adapter may be configured to allow an individual to wear the speech detection system in two or more different ways. For example, a single adapter may allow the wearable housing to be attached to eyeglasses and wireless earphones. As described in more detail herein, the wearable housing 110 of FIGS. 1 and 2A is a non-limiting example of a wearable housing consistent with the present disclosure.

[0048]

[0158] Some embodiments include a coherent light source configured to project light toward a user's facial region. Other embodiments include a non-coherent light source configured to project light toward a user's facial region. As used herein, the term "light source" broadly refers to any device configured to emit light. The term "coherent light" includes light that is highly ordered and exhibits a high degree of spatial and temporal coherence. This can occur, for example, when light waves are in phase with each other and have uniform frequencies and wavelengths, resulting in a light beam that is highly directional and limited in its outward spread as it travels. Alternatively, coherent light can include scenarios when light waves have a constant phase difference. In some examples, coherent light can be generated by coherent light sources such as lasers and other types of light sources with narrow spectral ranges and high monochromaticity (i.e., the light consists of a single wavelength). In contrast, non-coherent light can be generated by non-coherent light sources such as incandescent light bulbs and natural sunlight, which have wide spectral ranges and low monochromaticity.

[0049]

[0159] For example, coherent light can contain many waves of the same frequency with different phases and amplitudes, not necessarily at the same time and place. Controlling interference may require prior knowledge of the phase information of the light. In one embodiment, the coherent light source may be an alternative light source, such as a solid-state laser, a laser diode, a high-power laser, a laser such as a quantum cascade laser (QCL), or a light-emitting diode (LED)-based light source. Additionally, the coherent light source can emit light in various formats, such as pulsed light, continuous wave (CW), or quasi-CW. For example, one type of light source that can be used is a vertical-cavity surface-emitting laser (VCSEL). Another type of light source that can be used is an external cavity diode laser (ECDL). In some examples, the light source may include a laser diode configured to emit light at a wavelength between approximately 650 nm and 1150 nm. Alternatively, the coherent light source may include a laser diode configured to emit light at a wavelength between about 800 nm and about 1020 nm, between about 850 nm and about 950 nm, or between about 1300 nm and about 1700 nm. Unless otherwise specified, the terms "about" and "substantially the same" with respect to numerical values ​​may include variations of up to 5% relative to the stated value. As described in more detail herein, the light source 410 in FIGS. 4, 5A, and 5B is a non-limiting example of a light source consistent with the present disclosure. It should be recognized that, in the context of the present disclosure, the use of a coherent light source is intended as a non-limiting exemplary implementation in the context of speech detection systems, methods, and computer-readable media. Many of the embodiments described herein can be implemented with coherent or non-coherent light, and reference to either as an example herein is not intended to be limiting.For example, even when not explicitly stated, the described and claimed speech detection systems, methods, and computer program products may be configured to measure incoherent light reflection to detect facial skin micro-movements.

[0050]

[0160] Some embodiments include at least one detector configured to receive light reflections from a user's facial region. The terms "light detector" or simply "detector" broadly refer to any device, element, or system capable of measuring one or more characteristics of electromagnetic waves (e.g., power, frequency, phase, pulse timing, pulse duration, or other characteristics) and generating an output related to the measured one or more characteristics. Examples of detectors consistent with this disclosure include a photosensitive sensor, an imaging sensor, a phase detector, a MEMS sensor, a wave meter, a spectrometer, a spectrophotometer, a homodyne detector, or a heterodyne detector. In some embodiments, at least one detector may be configured to detect coherent light reflections. Additionally or alternatively, at least one detector may be configured to detect incoherent light reflections. At least one detector may include multiple detectors composed of multiple detection elements. At least one detector may include different types of light detectors. At least one detector may include multiple detectors of the same type that may differ in other characteristics (e.g., sensitivity, size). Combinations of several types of detectors may be used for different reasons. Consistent with some embodiments, the at least one detector can measure any form of light reflection and scattering, including secondary speckle patterns, different types of specular reflection, diffuse reflection, speckle interference, and any other form of light scattering. In some embodiments, the at least one detector is configured to output a reflection signal associated with the detected coherent light reflection. In the context of this disclosure, the term "reflection signal" broadly refers to any form of data retrieved from the at least one photodetector in response to light reflection from a facial region. The reflection signal may be any electronic representation of a property determined from the light reflection or a raw measurement signal detected by the at least one photodetector. As described in more detail herein, photodetector 412 of FIGS. 4, 5A, and 5B is a non-limiting example of a photodetector consistent with this disclosure.

[0051]

[0161] Some embodiments include at least one processor configured to determine specific facial skin micromovements using reflected signals from the detector. The term "at least one processor" may include any physical device or group of devices having electrical circuitry that performs logical operations on one or more inputs. For example, the at least one processor may include one or more integrated circuits (ICs), including application-specific integrated circuits (ASICs), microchips, microcontrollers, microprocessors, all or part of a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a field-programmable gate array (FPGA), a server, a virtual server, or other circuitry suitable for executing instructions or performing logical operations. The instructions executed by the at least one processor may be, for example, preloaded into a memory integrated with or incorporated into the controller, or stored in a separate memory. The memory may include random access memory (RAM), read-only memory (ROM), a hard disk, an optical disk, a magnetic medium, flash memory, other persistent, fixed, or volatile memory, or any other mechanism capable of storing instructions. In some embodiments, the at least one processor may include two or more processors. Each processor may have a similar structure, or the processors may be of different structures that are electrically connected or decoupled from each other. For example, the processors may be separate circuits or may be integrated into a single circuit.When two or more processors are used, the processors may be configured to operate independently or cooperatively, and may be co-located or remotely located from one another. The processors may be coupled electrically, magnetically, optically, acoustically, mechanically, or by other means that allow them to interact. As described in more detail herein, processing unit 112 of FIG. 1 and processing device 400 of FIG. 4 are non-limiting examples of at least one processor consistent with this disclosure.

[0052]

[0162] In some embodiments, at least one processor can determine facial skin micromotion by applying optical reflectance analysis. The term "optical reflectance analysis" includes evaluating surface properties by analyzing patterns of light scattered from a surface. When light strikes a surface (e.g., facial skin), some of it is absorbed, some is transmitted, and some is reflected. The amount and type of light reflected depends on the surface properties and the angle at which the light strikes. In one example, when a non-coherent light source is used, optical reflectance analysis can include scattering analysis, which involves measuring the scattering of light from a surface (e.g., facial skin). In another example, when a non-coherent light source is used, optical reflectance analysis can include speckle analysis or any pattern-based analysis. For example, coherent light illuminating a rough, uneven, or textured surface can be reflected or scattered in various different directions, resulting in a pattern of bright and dark areas called "speckle." Such analysis can be performed using a computer (e.g., including a processor) to identify speckle patterns and derive information about the surface (e.g., facial skin) represented in at least the reflected signals received from the photodetector. Speckle patterns can arise as a result of overlapping interference of coherent light waves, resulting in changes in intensity. The detected speckle patterns, or any other detected patterns, can then be processed to generate reflected image data. As described in more detail herein, the optical reflectance processing module 706 shown in FIG. 7 is a non-limiting example of a software module used to determine facial skin micromotion by applying optical reflectance analysis.

[0053]

[0163] Consistent with the present disclosure, the reflected image data may be processed by any image processing algorithm, including classical and / or artificial neural network (ANN)-based algorithms, such as convolutional neural networks (CNNs) and recurrent neural networks (RNNs). In some examples, the reflected image data may be preprocessed by transforming the image data using a transformation function to obtain a transformed speckle image. For example, the transformed reflected image data may include one or more convolutions of the speckle image. The transformation function may include one or more image filters, such as a low-pass filter, a high-pass filter, a band-pass filter, an all-pass filter, etc. In some examples, the transformation function may include a nonlinear function. In some examples, the reflected image data may be preprocessed by smoothing at least a portion of the reflected image data, for example, using a Gaussian convolution, a median filter, etc. In some examples, the reflected image data may be preprocessed to obtain different representations of the reflected image data. For example, the reflected image data may include a representation of at least a portion of the reflected image data in the frequency domain, a discrete Fourier transform of at least a portion of the reflected image data, a discrete wavelet transform of at least a portion of the reflected image data, a time / frequency representation of at least a portion of the reflected image data, a reduced-dimensional representation of at least a portion of the reflected image data, a lossy representation of at least a portion of the reflected image data, a lossless representation of at least a portion of the reflected image data, a time-ordered sequence of any of the above, or any combination of the above. In some examples, the reflected image data may be preprocessed to extract edges, and the preprocessed reflected image data may include information based on and / or related to the extracted edges. In some examples, the reflected image data may be preprocessed to extract features from the reflected image data. Some examples of such features may include information related to edges, corners, blobs, ridges, Scale Invariant Feature Transform (SIFT) features, temporal features, etc.

[0054]

[0164] In some embodiments, performing the light reflectance analysis may include evaluating the reflected image data and / or pre-processed reflected image data using one or more rules, functions, procedures, artificial neural networks, object detection algorithms, visual event detection algorithms, action detection algorithms, motion detection algorithms, background subtraction algorithms, inference models, etc. Some non-limiting examples of such inference models may include pre-programmed manually programmed inference models, classification models, regression models, results of training a training algorithm, such as a machine learning algorithm and / or a deep learning algorithm, on training examples, where the training examples may include examples of data instances, and in some cases, the data instances may be labeled with corresponding desired labels and / or results, etc. In some embodiments, performing the speckle analysis may include analyzing pixels, voxels, point clouds, range data, etc. included in the reflected image data.

[0055]

[0165] Some embodiments may include analyzing the reflected image data to decode speech. The process of decoding speech from the reflected image data may include identifying patterns or recognizing signatures within the reflected image data. For example, the recognition data, patterns, or signatures may be associated with certain phonemes, phoneme combinations, words, word combinations, or any other speech-related components. Recognizing information within such reflected image data may allow the speech to be decoded. Such recognition and / or decoding may be assisted by machine learning. For example, machine learning models or algorithms may be employed to recognize and / or understand utterances or commands. Some non-limiting examples of machine learning algorithms that may be used include classification algorithms, data regression algorithms, image segmentation algorithms, visual detection algorithms (e.g., object detectors, motion detectors, edge detectors, etc.), speech recognition algorithms, mathematical embedding algorithms, natural language processing algorithms, support vector machines, random forests, nearest neighbor algorithms, deep learning algorithms, artificial neural network algorithms, convolutional neural network algorithms, recurrent neural network algorithms, linear machine learning models, nonlinear machine learning models, ensemble algorithms, etc. For example, the trained machine learning algorithms may include predictive models, classification models, regression models, clustering models, segmentation models, artificial neural networks (e.g., deep neural networks, convolutional neural networks, recurrent neural networks, etc.), random forests, support vector machines, and other inference models. In some examples, the training examples may include example inputs and desired outputs corresponding to the example inputs. Further, in some examples, training a machine learning algorithm using training examples can produce a trained machine learning algorithm, which can be used to estimate outputs for inputs not included in the training examples.In some examples, engineers, scientists, processes, and machines that train machine learning algorithms can further use validation examples and / or test examples. For example, the validation examples and / or test examples can include example inputs along with desired outputs corresponding to the example inputs, and the trained machine learning algorithm and / or intermediately trained machine learning algorithm can be used to estimate outputs for example inputs of the validation examples and / or test examples, and the estimated outputs can be compared with the corresponding desired outputs, and the trained machine learning algorithm and / or intermediately trained machine learning algorithm can be evaluated based on the results of the comparison. In some examples, the machine learning algorithm can have parameters and hyperparameters, and the hyperparameters are set manually by a human or automatically by a process external to the machine learning algorithm (e.g., a hyperparameter search algorithm), and the parameters of the machine learning algorithm are set by the machine learning algorithm according to the training examples. In some implementations, the hyperparameters are set according to the training examples and the validation examples, and the parameters are set according to the training examples and the selected hyperparameters.

[0056]

[0166] In some examples, decoding speech from reflection image data may include a trained machine learning algorithm used as an inference model that generates an inference output when provided with input. For example, the trained machine learning algorithm may include a classification algorithm, where the input may include samples and the inference output may include a classification of the samples. In another example, the trained machine learning algorithm may include a regression model, where the input may include samples and the inference output may include an inferred value for the samples. In yet another example, the trained machine learning algorithm may include a clustering model, where the input may include samples and the inference output may include an assignment of the samples to at least one cluster. In an additional example, the trained machine learning algorithm may include a classification algorithm, where the input may include an image and the inference output may include a classification of an item depicted in the image. In yet another example, the trained machine learning algorithm may include a regression model, where the input may include an image and the inference output may include an inferred value (e.g., estimated facial skin movement) of an item depicted in the image. In an additional example, the trained machine learning algorithm may include an image segmentation model, where the input may include an image and the inference output may include a segmentation of the image. In yet another example, the trained machine learning algorithm may include an object detector, the input may include an image, and the inference output may include one or more detected objects in the image and / or one or more locations of the objects in the image.In some examples, the trained machine learning algorithm can include one or more equations and / or one or more functions and / or one or more rules and / or one or more procedures, the inputs can be used as inputs to the equations and / or functions and / or rules and / or procedures, and the inference output can be based on the output of the equations and / or functions and / or rules and / or procedures (e.g., selecting one of the outputs of the equations and / or functions and / or rules and / or procedures, using statistical measures of the outputs of the equations and / or functions and / or rules and / or procedures, etc.) As described in more detail herein, the reflection image 600 of FIG. 6 is a non-limiting example of a visualization of reflection image data consistent with the present disclosure.

[0057]

[0167] In some embodiments, an artificial neural network may be configured to analyze an input and generate a corresponding output. Some non-limiting examples of such artificial neural networks include shallow artificial neural networks, deep artificial neural networks, feedback artificial neural networks, feedforward artificial neural networks, autoencoder artificial neural networks, probabilistic artificial neural networks, time-delay artificial neural networks, convolutional artificial neural networks, recurrent artificial neural networks, long / short-term memory artificial neural networks, etc. In some examples, the artificial neural network may be manually configured. For example, the structure of the artificial neural network may be manually selected, the types of artificial neurons of the artificial neural network may be manually selected, parameters of the artificial neural network (e.g., parameters of the artificial neurons of the artificial neural network) may be manually selected, etc. In some examples, the artificial neural network may be configured using machine learning algorithms. For example, a user may select hyperparameters for an artificial neural network and / or a machine learning algorithm, and the machine learning algorithm may use the hyperparameters and training examples to determine the parameters of the artificial neural network, e.g., using backpropagation, using gradient descent, using stochastic gradient descent, using mini-batch gradient descent, etc. In some examples, an artificial neural network may be created from two or more other artificial neural networks by combining two or more other artificial neural networks into a single artificial neural network.

[0058]

[0168] The disclosed embodiments may include and / or access data structures or data. Data structures consistent with the present disclosure may include any collection of data values ​​and relationships between them. For example, the data structure may include correlations between facial micromovements and words or phonemes, and at least one processor may perform a lookup within the data structure for a particular word or phoneme associated with the detected facial skin micromovement. Data may be stored linearly, horizontally, hierarchically, relationally, non-relationally, unidimensionally, multidimensionally, operationally, ordered, unordered, object-oriented, centralized, decentralized, distributed, custom, or in any manner that allows data access. By way of non-limiting example, data structures may include arrays, associative arrays, linked lists, binary trees, balanced trees, heaps, stacks, queues, sets, hash tables, records, tagged unions, ER models, and graphs. For example, the data structure may include an XML database, an RDBMS database, an SQL database, or NoSQL alternatives for data storage / retrieval, such as MongoDB, Redis, Couchbase, Datastax Enterprise Graph, Elastic Search, Splunk, Solr, Cassandra, Amazon DynamoDB, Scylla, HBase, and Neo4J. The data structure may be a component of the disclosed system or a component of remote computing (e.g., a cloud-based data structure). The data in the data structure may be stored in contiguous or non-contiguous memory. Furthermore, a data structure as used herein does not require that the information be co-located. It may be distributed across multiple servers, for example, servers that may be owned or operated by the same or different entities. Thus, the term "data structure" as used herein in the singular includes multiple data structures. As described in more detail herein, data structure 124 of FIG. 1 and data structures 422 and 464 of FIG. 4 are non-limiting examples of data structures consistent with the present disclosure.

[0059]

[0169] Consistent with the present disclosure, the at least one processor can generate an output associated with the determined facial skin micro-movements. The term "generating an output" broadly refers to issuing a command, issuing data, and / or initiating an action in any type of electronic device. In some embodiments, the output may be voice (e.g., delivered via a speaker configured to fit the user's ear), and the voice may be an audible presentation of words associated with the silent or pre-utterance. In one example, the audible presentation of words may include an answer to a question the user silently asks the virtual personal assistant. In another example, the audible presentation of words may include synthetic speech (e.g., an artificial production of human speech). According to other disclosed embodiments, the output may be directed to a display (e.g., a visual display such as a computer monitor, a television, a mobile communication device, VR or XR glasses, or any other device enabling visual perception), and the generated output may include a graphic, image, or textual presentation (e.g., subtitles) of words associated with the pre-utterance or pre-utterance. The textual presentation of words may be presented simultaneously as the words are spoken. In other embodiments, the output may be directed to a communication device associated with the user, and the generated output may be any data exchanged with the communication device. The term "communication device" is intended to include all possible types of devices capable of exchanging data using a network configured to transmit the data. In some examples, the communication device may include a smartphone, tablet, smartwatch, personal digital assistant, desktop computer, laptop computer, Internet of Things (IoT) device, dedicated terminal, wearable communication device, and any other device capable of data communication. As described in more detail herein, the output determination module 712 of FIG. 7 is a non-limiting example of a software module used to generate an output related to the determined facial skin micromovements.

[0060]

[0170] The disclosed embodiments may include exchanging data (e.g., text data) using a network. The term "communications network" or simply "network" may include any type of physical or wireless computer networking arrangement used to exchange data. For example, the network may be the Internet, a private data network, a virtual private network using a public network, a Wi-Fi network, a LAN or WAN network, a combination of one or more of the foregoing, and / or other suitable connections that may enable information exchange between various components of the system. In some embodiments, the network may include one or more physical links used to exchange data, such as Ethernet, coaxial cable, twisted pair cable, optical fiber, or any other suitable physical medium for exchanging data. The network may also include a public switched telephone network ("PSTN") and / or a wireless cellular network. The network may be a secure or unsecure network. In other embodiments, one or more components of the system may communicate directly via a dedicated communications network. Direct communication may use any suitable technology, including, for example, BLUETOOTH™, BLUETOOTH LE™ (BLUETOOTH LE: BLE), Wi-Fi, near-field communications (NFC), or other suitable communication method that provides a medium for exchanging data and / or information between separate entities. As described in more detail herein, communication network 126 of FIG. 1 is a non-limiting example of a communication network consistent with the present disclosure.

[0061]

[0171] As used herein, a non-transitory computer-readable storage medium (or similar configurations such as non-transitory computer-readable media) refers to any type of physical memory capable of storing information or data readable by at least one processor. Examples include random access memory (RAM), read-only memory (ROM), volatile memory, non-volatile memory, hard drives, CD-ROMs, DVDs, flash drives, disks, any other optical data storage media, any physical media with a pattern of holes, markers, or other readable elements, PROMs, EPROMs, FLASH-EPROMs, or any other flash memory, NVRAMs, caches, registers, any other memory chips or cartridges, and networked versions thereof. The terms "memory" and "computer-readable storage medium" can refer to multiple structures, such as multiple memories or computer-readable storage media located within a wearable device or at a remote location. Furthermore, one or more computer-readable storage media can be utilized in practicing a computer-implemented method. Thus, the term computer-readable storage medium should be understood to include tangible items and exclude carrier waves and transient signals.

[0062]

[0172] Reference is now made to FIG. 1 , which illustrates a schematic diagram of an individual 102 using a speech detection system consistent with some embodiments of the present disclosure. It should be understood that FIG. 1 is a single exemplary representation, and that some illustrated elements may be omitted and other elements may be added within the scope of the present disclosure. In the illustrated exemplary implementation, the speech detection system 100 may be a head-mounted system for the user 102. Specifically, the speech detection system 100 (also referred to herein simply as the “system”) may have the form and appearance of an over-the-ear clip-on headset. Alternatively, the system may be head-mounted in one of many other ways within the scope of the present disclosure, including in-ear earphones, integrated into or connectable to eyeglass temples, a headband, or any other mechanism capable of securing the system or portions thereof to a human head. The speech detection system 100 may be configured to direct projected light 104 (e.g., coherent light) toward respective locations on the face of the user 102, thus generating an array of light spots 106 that span a facial region 108 of the face. The facial region 108 may be at least 1 cm 2 , at least 2 cm 2 , at least 4 cm 2 , at least 6 cm 2 , or at least 8 cm 2 In some embodiments, the size of the face region 108 may be determined to allow for detection of movement of different portions of the facial muscles. Although only one beam of projected light 104 is shown in the illustrated example, it is understood that all spots projected toward the face region 108 are associated with a corresponding light beam or beams of light. In other embodiments, the light source may project light in a manner other than as an array of spots. For example, the face region may be illuminated uniformly or non-uniformly.

[0063]

[0173] For head-wearable embodiments, speech detection system 100 may include a wearable housing 110 configured to be worn on the head of user 102. Wearable housing 110 may include or be associated with a processing unit 112 configured to interpret facial skin micromovements, an output unit 114 configured to fit over the user's ear and provide an audible and / or vibratory output, and a light-sensing unit 116 configured to project light toward portions of the user's 102's face other than the lips and detect reflections of the projected light. In the illustrated example, light-sensing unit 116 may be connected to output unit 114 by an arm 118 and may thus be held in a position proximate to and / or facing the user's face. According to some disclosed embodiments, light-sensing unit 116 does not contact the user's skin in facial region 108; rather, light-sensing unit 116 may be held a certain distance away from the skin surface in facial region 108. The distance of the light-sensing unit 116 from the skin surface may be at least 5 mm, at least 7.5 mm, at least 10 mm, at least 15 mm, or at least 20 mm.

[0064]

[0174] The light sensing unit 116 may be configured to receive reflections of the light 104 from the facial region 108 and output an associated reflection signal. Specifically, the reflection signal may indicate a light pattern (e.g., a secondary speckle pattern) that may result from the reflection of coherent light from each of the spots 106 within the field of view of the speech detection system 100. To cover a sufficiently large facial region 108, the detector of the speech detection system 100 may have a wide field of view, e.g., the field of view may have an angular width of at least 60°, at least 70°, or at least 90°. Within this field of view, the speech detection system 100 may detect and process signals reflecting the light patterns of all spots 106 or only a certain subset of the spots 106. For example, the processing unit 112 may select a subset of the spots 106 determined to provide the greatest amount of useful and reliable information regarding the relative movement of the skin surface of the user 102 and may avoid processing data from other spots 106. Further details of the configuration and operation of the light sensing unit 116 are described below with reference to FIG. 5.

[0065]

[0175] Consistent with the present disclosure, the speech detection system 100 can detect facial skin micro-movements of the user 102 and extract meaning from the detected movements, even without the utterance of speech or any other vocal utterance by the user 102. The extracted meaning may be an identification of the user 102 wearing the speech detection system 100, an identification of a subvocalization by the user 102, such as a word spoken silently by the user 102, an identification of a word spoken aloud by the user 102, an identification of a phoneme spoken silently by the user 102, or an identification of a phoneme spoken aloud by the user 102. Similarly, the extracted meaning may include an identification of the heart rate of the user 102, an identification of the respiratory rate of the user 102, and / or other characteristics associated with verbal or non-verbal communication by the user 102. In one example, the speech detection system 100 can generate an output signal including data related to identification information, UI commands, a synthesized audio signal, a text transcription, or any combination thereof. In one example, the synthesized audio signal may be played back to the user 102 via a speaker in the output unit 114. This playback may be useful to provide feedback to the user 102 regarding the speech output.

[0066]

[0176] Consistent with the present disclosure, speech detection system 100 can exchange data (e.g., output signals) with various communication devices associated with a user, such as mobile communication device 120 or server 122. The term “communication device” is intended to include all possible types of devices capable of exchanging data using a digital communication network, an analog communication network, or any other network configured to convey data. In some examples, the communication device may include a wearable communication device such as a smartphone, a tablet, a smart watch, a personal digital assistant, a laptop computer, an IoT device, a dedicated terminal, industrial machinery, a vehicle, a smart house, an appliance, or any other electronic device capable of exchanging information or data with another electronic device. In other examples, the communication device may include a non-wearable communication device such as a desktop computer, a smart home hub, a router, a server, or any other network-connected equipment. In some cases, the processing device of mobile communication device 120 or server 122 may supplement or replace some functionality of processing unit 112 of speech detection system 100. In some embodiments, the output signal generated by speech detection system 100 may be transmitted to mobile communication device 120 or a cloud server via a communication link. The term "cloud server" refers to a computer platform that provides services over a network such as the Internet. In the exemplary embodiment shown in FIG. 1, server 122 may utilize one or more virtual machines that may not correspond to individual pieces of hardware. For example, computing and / or storage capabilities may be implemented by allocating an appropriate portion of desired computing / storage power from a scalable repository such as a data center or distributed computing environment. In one exemplary configuration, server 122 may be a cloud server that determines neural activity of user 102 based on facial skin micromovements.In one example, server 122 may implement the methods described herein using customized hardwired logic, one or more application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), firmware, and / or program logic that, in combination with a computer system, makes server 122 a dedicated machine.

[0067]

[0177] In some embodiments, the server 122 can access the data structure 124 to, for example, determine correlations between words and multiple facial movements. The data structure 124 can utilize volatile or nonvolatile, magnetic, semiconductor, tape, optical, removable, non-removable, or other types of storage devices or tangible or non-transitory computer-readable media, or any medium or mechanism for storing information. The data structure 124 can be part of the server 122, as shown, or can be separate from the server 122. When the data structure 124 is not part of the server 122, the server 122 can exchange data with the data structure 124 via a communications link. The data structure 124 can include one or more memory devices that store data and instructions used to execute one or more features of the disclosed methods. In one embodiment, the data structure 124 can include any of several suitable data structures, from a small data structure hosted on a workstation to a large data structure distributed across a data center. The data structure 124 can also include any combination of one or more data structures controlled by a memory controller device (e.g., a server) or software. Consistent with the present disclosure, the speech detection system 100 can communicate with the mobile communication device 120 or the server 122 using a communication network 126 as defined above.

[0068]

[0178] Reference is now made to FIG. 2A , which illustrates another exemplary implementation of a speech detection system 100 according to the present disclosure. In this example, the wearable housing 110 may be integrated with or otherwise attached to eyeglasses 200 having frames 202. In this exemplary implementation, eyeglasses 200 may include nasal electrodes 204 and temporal electrodes 206 attached to frames 202 and in contact with the user's skin surface. Electrodes 204 and 206 may receive surface electromyogram (sEMG) signals, which provide additional information regarding the activation of the user's facial muscles. Speech detection system 100 may use the electrical activity sensed by electrodes 204 and 206, along with the output of optical sensing unit 116, in generating a synthesized audio signal, for example. Additionally or alternatively, speech detection system 100 may include one or more additional optical sensing units 208, similar to optical sensing unit 116, for sensing skin movement in other areas of the user's face, such as eye movement. These additional light-sensing units may be used together with or instead of the light-sensing unit 116. In the illustrated example, the light-sensing unit 116 may illuminate the first face region 108A, and the light-sensing unit 208 may illuminate the second face region 108B. The first face region 108A and the second face region 108B may not overlap.

[0069]

[0179] In some disclosed embodiments, the speech detection system may be incorporated into, integrated with, or otherwise attached to an extended reality appliance. As used herein, the term "extended reality appliance" includes any type of device or system that enables a user to perceive and / or interact with an extended reality environment. The term "extended reality environment" refers to all types of combined real and virtual environments and human-machine interactions that are at least partially generated by computer technology. One non-limiting example of an extended reality environment may be a virtual reality (VR) environment. A virtual reality environment may be an immersive, simulated, non-physical environment that provides a user with the perception of being in a virtual environment. Another non-limiting example of an extended reality environment may be an augmented reality (AR) environment. An augmented reality environment may include a direct or indirect live view of a physical real-world environment augmented with virtual, computer-generated perceptual information, such as virtual objects, with which a user can interact. Another non-limiting example of an extended reality environment may be a mixed reality (MR) environment. A mixed reality environment may be a hybrid of a physical, real-world environment and a virtual environment, where physical and virtual objects coexist and can interact in real time. Examples of extended reality appliances include VR headsets, AR headsets, MR headsets, smart glasses, and wearable projection devices.

[0070]

[0180] Reference is now made to FIG. 2B , which illustrates another exemplary implementation of speech detection system 100 according to some embodiments of the present disclosure. In the illustrated example, speech detection system 100 may be part of extended reality appliance 250. Extended reality appliance 250 may include all of the sensors, etc., described above with reference to glasses 200, etc. For example, extended reality appliance 250 may include one or more of a gyroscope, an accelerometer, a magnetometer, an image sensor, a depth sensor, an infrared sensor, a proximity sensor, and / or any other sensor configured to measure one or more properties associated with the individual wearing extended reality appliance 250 and generate an output related to the measured properties. In some cases, speech detection system 100 may use input from any one of the sensors of extended reality appliance 250 to determine spoken or sub-vocalized words spoken by individual 102. For example, speech detection system 100 may use input from an image sensor of extended reality appliance 250 along with data from light sensing unit 116 (see FIG. 1) to extract the meaning of facial movements. In other cases, extended reality appliance 250 may generate output including visual and / or audible presentations associated with words detected by speech detection system 100. For example, individual 102 may interact with extended reality appliance 250 using silent commands.

[0071]

[0181] Reference is now made to FIG. 3 , which illustrates another exemplary implementation of a speech detection system 100 according to the present disclosure. In the implementation illustrated in FIG. 3 , the speech detection system 100 may be integrated with a mobile communication device 120. Specifically, the mobile communication device 120 may include a photodetector configured to detect light reflection 300 from a face region 108. In this example, the light projected onto the face region 108 originates from a non-wearable light source 302, which may be a coherent or non-coherent light source. In some configurations, the non-wearable light source 302 may be included in the mobile communication device 120. Alternatively, the non-wearable light source 302 may be separate from the mobile communication device 120.

[0072]

[0182] Consistent with the present disclosure, as shown in FIG. 3 , the pattern of light projected onto the facial region 108 may be a single spot 106 large enough to illuminate different portions of the facial region 108. For example, the spot 106 may include a first portion 304A associated with a first facial muscle and a second portion 304B associated with a second facial muscle. A processing device of the mobile communication device 120 can then apply light reflection analysis to the received reflection 300 to determine facial skin micro-movements. In particular, the processing device of the mobile communication device 120 can determine a first facial skin micro-movement of the first portion 304A and a second facial skin micro-movement of the second portion 304B. The processing device can use both the first and second facial skin micro-movements to extract meaning (e.g., determining whether it is a speech or a command, or authenticating the user 102) and generate an output. 3 may be used when the extracted meaning includes continuous authentication of the user 102. Specifically, the speech detection system 100 may provide an authentication service that uses facial micro-movement biometrics for continuous authentication while the mobile communication device 120 is in use.

[0073]

[0183] 4 is a block diagram of an exemplary configuration of speech detection system 100 and an exemplary configuration of remote processing system 450. It should be understood that FIG. 4 is a representation of only one embodiment, and that some illustrated elements may be omitted and other elements may be added within the scope of the present disclosure. In the illustrated embodiment, speech detection system 100 comprises a processing unit 112 including a processing device 400 and a memory device 402, an output unit 114 including a speaker 404, a light indicator 406, and a haptic feedback device 408, a light sensing unit 116 including at least one light source 410 and at least one light detector 412, an audio sensor 414, a power source 416, one or more additional sensors 418, a network interface 420, and a data structure 422. Speech detection system 100 may have direct or indirect access to a bus 424 (or any other communication mechanism) that interconnects the aforementioned subsystems and components for transferring information and commands within speech detection system 100. Some of the subsystems and components listed above are referred to herein in the singular, but may be plural in alternative configurations. For example, in some configurations, speech detection system 100 may include multiple light sources 410 or multiple light detectors 412.

[0074]

[0184] The processing device 400 shown in FIG. 4 may constitute any physical device or group of devices having electrical circuitry that performs logical operations on one or more inputs. The instructions executed by at least one processor may, for example, be integrated with or preloaded into memory incorporated into the processing device 400, or may be stored in a separate memory (e.g., memory device 402 or data structure 422). As described above, the processing device may include two or more processors. Each processor may have a similar structure, or the processors may be of different structures that are electrically connected or decoupled from each other. For example, the processors may be separate circuits or integrated into a single circuit. When two or more processors are used, the processors may be configured to operate independently or cooperatively, and may be co-located or remotely located from each other. The processors may be coupled electrically, magnetically, optically, acoustically, mechanically, or by other means that allow them to interact. Consistent with the present disclosure, at least some of the functionality described below with respect to the processing device 400 may be performed by a processing device of a remote processing system 450.

[0075]

[0185] The memory device 402 shown in FIG. 4 may include, for example, one or more magnetic disk storage devices, one or more optical storage devices, and / or high-speed random access memory and / or non-volatile memory, such as flash memory (e.g., NAND, NOR). Consistent with the present disclosure, components of the memory device 402 may be distributed across two or more units of the speech detection system 100 and / or across two or more units of memory devices. In particular, the memory device 402 may be used to store software products and / or data stored on a non-transitory computer-readable medium. As mentioned above, the terms “memory” and “computer-readable storage medium” may refer to multiple structures, such as multiple memories or computer-readable storage media located within the speech detection system 100 or at a remote location (e.g., remote processing system 450). Furthermore, one or more computer-readable storage media may be utilized in practicing a computer-implemented method. Referring now to FIG. 7, examples of software modules stored on the memory device 402 will be described.

[0076]

[0186] The output unit 114 shown in FIG. 4 can trigger output from various output devices, such as a speaker 404, a light indicator 406, and a haptic feedback device 408. Examples of the speaker 404 can include or incorporate a loudspeaker, a wireless earphone, an audio headphone, a hearing aid-type device, a bone conduction headphone, and any other device capable of converting an electrical audio signal into a corresponding voice. In some embodiments, the speaker 404 may be configured to generate an audio signal audible only to the user 102. Alternatively, the speaker 404 may be configured to emit a voice into the air so that nearby persons can hear it. The light indicator 406 can include one or more light sources, such as, for example, an LED array associated with different colors. The light indicator 406 can be used to indicate the battery status of the speech detection system 100 or to indicate its operating mode. The haptic feedback device 408 may include a vibration motor, a linear actuator, a vibration transducer, or any other force feedback device capable of providing a tactile or haptic cue or converting an electrical signal into a corresponding vibration or application of force.

[0077]

[0187] The light-sensing unit 116 shown in FIG. 4 may include a light source 410 and a light detector 412. The light source 410 may project coherent or non-coherent light onto the face region 108. As discussed above, the light source 410 may be an alternative light source, such as a solid-state laser, a laser diode, a high-power laser, or a light-emitting diode (LED)-based light source. Additionally, the light source 410 may emit light in various formats, such as pulsed light, continuous wave (CW), or quasi-CW. In one embodiment, the light source 410 may be an infrared laser diode configured to emit an input beam of coherent radiation. The light source 410 may be associated with a beam-splitting element, such as a Dammann grating or another suitable type of diffractive optical element (DOE), for splitting the input beam into multiple output beams that form respective spots 106 in a matrix of locations spread over the face region 108. In another embodiment (not shown), light source 410 may include multiple laser diodes or other emitters that generate respective groups of output beams that cover different respective sub-areas within facial region 108. In one embodiment, processing unit 112 may select and activate only a subset of the emitters, rather than activating all of the emitters. For example, to reduce power consumption of speech detection system 100, processing unit 112 may activate only one emitter or a subset of two or more emitters that illuminates a particular area on the user's face that is known to provide the most useful information for generating the desired speech output.

[0078]

[0188] The photodetector 412 shown in FIG. 4 can be used to detect reflections from the face region 108 indicative of facial skin movement. As described above, the photodetector can measure properties of coherent or incoherent light, such as power, frequency, phase, pulse timing, pulse duration, and other properties. In some embodiments, the photodetector 412 can include an array of detection elements, such as a set of charge-coupled device (CCD) sensors and / or a set of complementary metal-oxide semiconductor (CMOS) sensors, with objective optics for imaging the face region 108 onto the array. Due to the small dimensions and proximity of the light-sensing unit 116 to the skin surface, the photodetector 412 can have a field of view wide enough to detect spots 106 on the face region 108 at high angles of at least 60°, at least 70°, or at least 90°. The photodetector 412 can be configured to generate an output related to the measured properties of the detected light. Consistent with the present disclosure, the output of the photodetectors 412 may include any form of data determined in response to light reflections received from the face region 108. In some embodiments, the output may include a reflection signal that includes an electronic representation of one or more properties determined from the coherent or incoherent light reflections. In other embodiments, the output may include raw measurements detected by at least one photodetector 412.

[0079]

[0189] In some embodiments, the photodetector 412 can measure one of several optical attributes associated with skin changes. The term “skin change” refers to any detectable movement, change, or alteration occurring in the skin. Such skin changes may include changes in the epidermis (i.e., the outermost layer of skin), changes in the dermis (i.e., the middle layer of skin), changes in the subcutaneous tissue (i.e., the deepest layer of skin), and changes in deep muscle tissue. The optical attributes can be measured without contacting the skin of the individual 102. One example of several optical attributes of reflected light that can be measured by the photodetector 412 can include intensity, frequency, reflectance, angle, sharpness, bidirectional reflectance distribution function, color, brightness, glossiness, transparency, opacity, surface texture, surface relief, surface movement, and other optical attributes derivable from analysis of light reflection. The output of the photodetector 412 can be used to determine information related to skin changes. In some embodiments, information related to these skin changes may be derived from changes in the distance from the skin to the detector as the skin moves; in other embodiments, the changes may not be derived from changes in the distance of the skin from the light detector 412. For example, the determined speed or angular velocity of facial skin changes can be determined by detecting changes in non-distance measurements (e.g., image sharpness) over time. Thus, in one non-limiting example, optical attributes can be detected from random intensity fluctuations observed when coherent light interacts with a rough or scattering surface, such as human skin. In another non-limiting example, optical attributes can be detected based on the interference of light waves, such as when an interference pattern is used to measure the phase difference or amplitude change between two or more optical paths.

[0080]

[0190] In some embodiments, the light-sensing unit 116 may not require a reference for light source parameters, such as the wavelength, intensity, or coherence of the light source, and may not require a reference beam (typically used with a beam splitter) to measure one or more optical attributes of the reflected light. For example, the light-sensing unit 116 may use a single beam to illuminate the skin and then process the light reflection returned to the light detector 412. While some speech detection systems may include a single pixel sensor (e.g., a photodiode), in other embodiments, the light detector 412 may include one or more multi-pixel sensors that enable the generation of images that provide spatial information beyond a single point (e.g., each pixel sensor includes more than 4 megapixels, more than 10 megapixels, or more than 10 megapixels). For example, the reflected image shown in FIG. 6 may be generated from the output of the light detector 412. As described throughout this disclosure, the output of the light detector 412 may be analyzed using image processing to determine the pattern of light scattered from the surface. For example, secondary speckle characteristics may be determined.

[0081]

[0191] In some non-limiting examples, the light-sensing unit 116 may use a diffractive element to split the outgoing beam into multiple beams and not rely on the superposition of coherent light waves to cause interference. In some non-limiting examples, the light-sensing unit 116 may be positioned such that the light detector 412 can be positioned along a different optical axis than the light source 410. In other non-limiting examples, aligning the light source and sensor along the same optical axis may be used to maintain coherence, achieve path length matching, ensure spatial overlap, and maintain sensitivity and accuracy of the interference pattern. However, because some implementations of the light detector 412 detect reflected images and not distance to points, the light-sensing unit 116 may include a first optical axis for the outgoing light and a second optical axis for the inward light that is not aligned with the first optical axis. In some embodiments, the light detector 412 is configured to measure both sub-micron speed and depth changes in the range of 5 to 500 microns. In alternative embodiments, the light detector 412 is configured to measure changes of less than 1 micron. All of the examples provided in this paragraph are alternatives and, depending on the implementation details, can be implemented in many alternative embodiments provided herein.

[0082]

[0192] The audio sensor 414 shown in FIG. 4 may include one or more audio sensors configured to capture audio by converting voices into digital information. Some examples of audio sensors may include a microphone, a unidirectional microphone, a bidirectional microphone, a cardioid microphone, an omnidirectional microphone, an onboard microphone, a wired microphone, a wireless microphone, or any combination of the above. The audio sensor 414 may be configured to capture voices spoken by the user 102, allowing the user 102 to use the speech detection system 100 as a traditional headphone if desired. Additionally or alternatively, the audio sensor 414 may be used in conjunction with the silent speech detection capabilities of the speech detection system 100. In one embodiment, the audio signal output by the audio sensor 414 may be used to change the operating state of the speech detection system 100. For example, the processing unit 112 may generate a speech output only when the audio sensor 414 does not detect the user 102 speaking a word. In another embodiment, the audio sensor 414 may be used in a calibration procedure, in which the optical sensing unit 116 detects micro-movements of the skin while the user 102 is uttering a particular phoneme or word. The processing unit 112 can compare the reflected signal output by the optical detector 412 with the voice detected by the audio sensor 414 to calibrate the optical sensing unit 116. This calibration can include prompting the user 102 to shift the position of the optical sensing unit 116 to align the optics with a desired position relative to the facial region 108. In yet another embodiment, the audio sensor 414 enables on-the-fly training of the neural network of the speech detection system 100. For example, the speech detection system 100 can be configured to correlate facial micro-movements of the skin with words using an audio signal captured simultaneously with the micro-movements.After recognizing the recorded words, the speech detection system 100 can perform a lookback to identify the facial micro-movements that precede the articulation of those words, thereby training the speech detection system 100. Similarly, the speech detection system can be used to train facial expressions, commands, user recognition, and emotions.

[0083]

[0193] The power source 416 shown in FIG. 4 can provide electrical energy to power the speech detection system 100. The power source can include any device or system capable of storing, distributing, or transmitting electrical power, including, but not limited to, one or more batteries (e.g., lead-acid, lithium-ion, nickel-metal hydride, nickel-cadmium), one or more capacitors, one or more connections to an external power source, one or more power converters, or any combination thereof. Referring to the example shown in FIG. 4, the power source 416 can be portable, meaning that the speech detection system 100 can be wearable. The portability of the power source allows the user 102 to use the speech detection system 100 in a variety of situations. In other embodiments, the power source 416 can be associated with a connection to an external power source (such as a power grid) that can be used to charge the power source 416.

[0084]

[0194] The additional sensors 418 shown in FIG. 4 may include various sensors, such as image sensors, motion sensors, environmental sensors, electromyography (EMG) sensors, resistive sensors, ultrasonic sensors, proximity sensors, biometric sensors, or other sensing devices configured to facilitate related functionality. For example, the speech detection system 100 may include one or more image sensors configured to capture visual information from the user's 102's environment by converting light (not emitted from the light source 410) into image data. Consistent with this disclosure, the image sensor may be included in any device or system capable of detecting and converting optical signals in the near-infrared, infrared, visible, and / or ultraviolet spectrums into electrical signals. Examples of image sensors may include digital cameras, semiconductor charge-coupled devices (CCDs), or complementary metal-oxide semiconductor (CMOS) or N-type metal-oxide-semiconductor (NMOS) active pixel sensors. The electrical signals may be used to generate image data. Consistent with this disclosure, image data can include pixel data streams, digital images, digital video streams, data derived from captured images, and data that can be used to construct one or more 3D images, a sequence of 3D images, a 3D video, or a virtual 3D representation. Image data acquired by one or more image sensors can be transmitted to processing unit 112 or remote processing system 450 via wired or wireless transmission.

[0085]

[0195] The speech detection system 100 may also include one or more motion sensors configured to measure the movement of the user 102. Specifically, the motion sensor may perform at least one of detecting the movement of the user 102, measuring the velocity of the user 102, measuring the acceleration of the user 102, or measuring any other action involving movement. In some embodiments, the motion sensor may include one or more accelerometers configured to detect changes in acceleration (e.g., proper acceleration) and / or measure the acceleration of the speech detection system 100. In some embodiments, the motion sensor may include one or more gyroscopes configured to detect changes in orientation of the speech detection system 100 and / or measure information related to the orientation of the speech detection system 100. In some embodiments, the motion sensor may include one or more using an image sensor, a LIDAR sensor, a radar sensor, or a proximity sensor. For example, by analyzing captured images, the processing device 400 may determine the movement of the speech detection system 100, e.g., using an ego-motion algorithm. Additionally, the processing device may determine the movement of objects in the environment of the speech detection system 100, e.g., via object tracking.

[0086]

[0196] The speech detection system 100 can also include one or more environmental sensors of different types configured to capture data reflecting the user's 102's environment. In some embodiments, the environmental sensors can include one or more chemical sensors configured to at least one of measure chemical properties in the user's 102's environment, measure changes in chemical properties in the user's 102's environment, detect the presence of chemicals in the user's 102's environment, and / or measure concentrations of chemicals in the user's 102's environment. Examples of measurable chemical properties include pH levels, toxicity, and temperature. Examples of chemicals or phenomena that can be measured include electrolytes, certain enzymes, certain hormones, certain proteins, smoke, carbon dioxide, carbon monoxide, oxygen, ozone, hydrogen, and hydrogen sulfide. In other embodiments, the environmental sensors can include one or more temperature sensors configured to detect changes in the temperature of the user's 102's environment and / or measure the temperature of the user's 102's environment. In other embodiments, the environmental sensors can include one or more barometers configured to detect changes in the atmospheric pressure of the user's 102's environment and / or measure the atmospheric pressure of the user's 102's environment. In other embodiments, the environmental sensors may include one or more light sensors configured to detect changes in ambient light in the user's 102 environment.

[0087]

[0197] The network interface 420 shown in FIG. 4 can provide bidirectional data communication to a network, such as the communications network 126. In one embodiment, the network interface 420 can include an Integrated Services Digital Network (ISDN) card, a cellular modem, a satellite modem, or a modem for providing a data communication connection over the Internet. As another example, the network interface 420 can include a Wireless Local Area Network (WLAN) card. In another embodiment, the network interface 420 can include an Ethernet port connected to a radio frequency receiver and transmitter and / or an optical (e.g., infrared) receiver and transmitter. The specific design and implementation of the network interface 420 can depend on the communications network or networks over which the speech detection system 100 is intended to operate. For example, in some embodiments, the speech detection system 100 can include a network interface 420 designed to operate over a GSM network, a GPRS network, an EDGE network, a Wi-Fi or WiMax network, and a Bluetooth network. In any such implementation, network interface 420 may be configured to send and receive electrical, electromagnetic, or optical signals that carry digital data streams or signals representing various types of information.

[0088]

[0198] The data structure 422 shown in FIG. 4 may include any hardware, software, firmware, or combination thereof for storing information in a database and facilitating retrieval of information from the database. The term "database" may be understood to include a collection of data, which may be distributed or non-distributed. A database may include a database management system that controls the organization, storage, and retrieval of data contained in the database. As described above, the data contained in the database may be stored linearly, horizontally, hierarchically, relationally, non-relationally, unidimensionally, multidimensionally, operationally, ordered, non-ordered, object-oriented, centralized, decentralized, distributed, custom, or in any manner that enables data access. In disclosed embodiments, the data structure 422 may include correlations between facial micromovements and words, commands, emotions, facial expressions, and / or biological states. At least one processor can perform lookups within the data structure to thereby interpret the detected facial skin micromovements. According to one embodiment, at least some of the data stored in the data structure 422 may alternatively or additionally be stored in a remote processing system 450.

[0089]

[0199] Consistent with the present disclosure, the speech detection system 100 may be configured to communicate with a remote processing system 450 (e.g., a mobile communication device 120 or a server 122). The remote processing system 450 may have direct or indirect access to a bus 452 (or other communication mechanism) that interconnects subsystems and components for transferring information within the remote processing system 450. For example, the bus 452 may interconnect a memory interface 454, a network interface 456, a power supply 458, a processing device 460, one or more additional sensors 462, a data structure 464, and a memory device 466.

[0090]

[0200] The memory interface 454 shown in FIG. 4 may be used to access software products and / or data stored in a non-transitory computer-readable medium or other memory device, such as memory devices 402, 466, data structure 422, or data structure 464. Memory device 466 may include software modules for performing processes consistent with the present disclosure. In particular embodiments, memory device 466 may include a shared memory module 472, a node registration module 473, a load balancing module 474, one or more computing nodes 475, an internal communications module 476, an external communications module 477, and a database access module (not shown). Modules 472-477 may include software instructions for execution by at least one processor (e.g., processing device 460) associated with remote processing system 450. The shared memory module 472, node registration module 473, load balancing module 474, computing module 475, and external communications module 477 may cooperate to perform various operations.

[0091]

[0201] The shared memory module 472 may enable information sharing between the remote processing system 450 and one or more other devices associated with the speech detection system 100. In some embodiments, the shared memory module 472 may be configured to enable the processing device 460 to access, retrieve, and store data. For example, using the shared memory module 472, the processing device 460 may at least one of execute software programs stored in the memory device 402, 466, the data structure 422, or the data structure 464, store information in the memory device 402, 466, the data structure 422, or the data structure 464, or retrieve information from the memory device 402, 466, the data structure 422, or the data structure 464.

[0092]

[0202] The node registration module 473 can be configured to track the availability of one or more computational nodes 475. In some examples, the node registration module 473 may be implemented as a software program, such as a software program executed by one or more computational nodes 475, a hardware solution, or a combined software and hardware solution. In some implementations, the node registration module 473 can communicate with one or more computational nodes 475, for example, using the internal communication module 476. In some examples, one or more computational nodes 475 can notify the node registration module 473 of their status by sending a message, for example, upon startup, upon shutdown, at regular intervals, at selected times, in response to a query received from the node registration module 473, or at any other determined time. In some examples, the node registration module 473 can inquire about the status of one or more computational nodes 475 by sending a message, for example, upon startup, at regular intervals, at selected times, or at any other determined time.

[0093]

[0203] The load balancing module 474 can be configured to divide the workload among one or more computing nodes 475. In some examples, the load balancing module 474 may be implemented as a software program, such as a software program executed by one or more computing nodes 475, a hardware solution, or a combined software and hardware solution. In some implementations, the load balancing module 474 can interact with the node registration module 473 to obtain information regarding the availability of one or more computing nodes 475. In some implementations, the load balancing module 474 can communicate with one or more computing nodes 475 using, for example, the internal communication module 476. In some examples, one or more computing nodes 475 can notify the load balancing module 474 of their status by sending a message, for example, upon startup, upon shutdown, at regular intervals, at selected times, in response to a query received from the load balancing module 474, or at any other determined time. In some examples, the load balancing module 474 can inquire about the status of one or more computing nodes 475 by sending a message, for example, upon startup, at regular intervals, at selected times, or at any other determined time.

[0094]

[0204] Internal communications module 476 can be configured to receive and / or transmit information to and from one or more components of remote processing system 450. For example, control and / or synchronization signals can be sent and / or received via internal communications module 476. In one embodiment, input information to a computer program, output information of a computer program, and / or intermediate information of a computer program can be sent and / or received via internal communications module 476. In another embodiment, information received via internal communications module 476 can be stored in memory device 466 or data structure 464. For example, information retrieved from data structure 464 can be transmitted using internal communications module 476. In another example, a reference signal reflecting facial micro-movements of user 102 can be stored in data structure 464 and accessed using internal communications module 476.

[0095]

[0205] The external communication module 477 can be configured to receive and / or transmit information to and from one or more speech detection systems 100. For example, control signals can be sent and / or received via the external communication module 477. In one embodiment, information received via the external communication module 477 can be stored in the memory device 466, the data structure 464, and / or any memory device within the one or more speech detection systems 100. In another embodiment, information retrieved from the data structure 464 may be transmitted using the external communication module 477 to the speech detection system 100 or any entity with which the user 102 communicates. For example, when the user 102 communicates with a financial institution (e.g., a bank), the information retrieved from the data structure 464 can be transmitted to enable authentication of the user 102. In another embodiment, the external communication module 477 can be used to transmit and / or receive sensor data. Examples of such input data can include data received from the speech detection system 100, information captured from the environment of the user 102 using one or more sensors, such as the additional sensor 418 and the additional sensor 462.

[0096]

[0206] In some embodiments, aspects of modules 472-477 may be implemented in hardware, software (contained in one or more signal processing and / or application specific integrated circuits), firmware, or any combination thereof, and may be executable by one or more processors alone or in various combinations with each other. In particular, modules 472-477 may be configured to interact with each other and / or other modules of speech detection system 100 to perform functions consistent with disclosed embodiments. Memory device 466 may include additional or fewer modules and instructions.

[0097]

[0207] 4 may share functionality similar to that of the corresponding elements of speech detection system 100, as described above. The specific design and implementation of the above-described components may vary based on the implementation of remote processing system 450. Additionally, remote processing system 450 may include more or fewer components. For example, when remote processing system 450 is a mobile communication device (e.g., mobile communication device 120) associated with user 102, it may include a speaker, a microphone, and additional sensors.

[0098]

[0208] The components and arrangement of speech detection system 100 and remote processing system 450 shown in FIG. 4 are not intended to limit the disclosed embodiments. As will be understood by one of ordinary skill in the art with the benefit of this disclosure, numerous variations and / or modifications can be made to the illustrated speech detection system 100 and remote processing system 450. For example, not all components are essential to the operation of the input unit in all cases. Not all components may be located in any suitable portion of speech detection system 100 or remote processing system 450. Furthermore, components may be rearranged in various configurations while providing the functionality of the disclosed embodiments. For example, some speech detection systems may not include all of the elements shown in speech detection system 100 and remote processing system 450. Other speech detection systems may include additional components and still be within the scope of the present disclosure.

[0099]

[0209] 5A and 5B include two schematic diagrams of the optical sensing unit 116 when detecting facial skin micro-movements, consistent with some embodiments of the present disclosure. The two schematic diagrams illustrate simplified scenarios before and after muscle recruitment. As shown, the optical sensing unit 116 may include an illumination module 500, a detection module 502, and, optionally, an audio sensor 414. As described above and shown in FIG. 5 , the optical sensing unit 116 may be configured not to contact the user's skin in the facial region 108, but rather may be held a distance D away from the skin surface in the facial region 108. The distance D of the optical sensing unit 116 from the skin surface may be at least 5 mm, at least 7.5 mm, at least 10 mm, at least 15 mm, or at least 20 mm.

[0100]

[0210] In the illustrated embodiment, the illumination module 500 includes a light source 410 (e.g., an infrared laser diode) configured to generate an input light beam 504. The illumination module 500 further includes a beam-splitting element 506, such as a Dammann grating or another suitable type of diffractive optical element (DOE), configured to split the input beam 504 into multiple output beams 508, which form respective spots 106A-106E in a pattern (e.g., a matrix of locations) spread over the face region 108. In an alternative embodiment (not shown), the illumination module 500 may include multiple light sources 410 that generate respective groups of output beams 508 covering different respective subareas within the face region 108. In this alternative embodiment, the processing unit 112 may select and activate only a subset of the multiple light sources, rather than activating all of the multiple light sources. For example, to reduce power consumption of the speech detection system 100, the processing unit 112 may activate only one light source or a group of two or more light sources that illuminate a portion of the face region 108.

[0101]

[0211] The detection module 502 may include a photodetector 412, which may include an array 510 of optical sensors (e.g., an array of CMOS image sensors) having objective optics 512 for acquiring coherent light reflections 300 from the facial region 108. Due to the small dimensions and proximity of the light-sensing unit 116 to the skin surface, the detection module 502 may be configured with a wide field of view so that reflections from many spots 106 can be captured at high angles. As mentioned above, the field of view of the photodetector 412 may have an angular width of at least 60°, at least 70°, or at least 90°. Due to the roughness of the skin surface, the light patterns at the spots 106 may also be detected at these high angles.

[0102]

[0212] For example, the speech detection system 100 can analyze the light reflectance 300 to determine facial skin micro-movements resulting from the recruitment of muscle fibers 520. Determining the facial skin micro-movements can include determining the amount of skin movement, determining the direction of the skin movement, and / or determining the acceleration of the skin movement. The determined facial skin micro-movements can include voluntary and / or involuntary recruitment of muscle fibers 520. The muscle fibers 520 can be part of the zygomaticus, orbicularis oris, laughing muscle, genioglossus, or levator labii superioris alae noli muscle. The processing device 400 may be configured to perform a first speckle analysis on light reflected from a first region of the face proximate the spot 106A to determine that the first region has moved a distance d1 (i.e., first facial skin micro-movement 522A), and a second speckle analysis on light reflected from a second region of the face proximate the spot 106E to determine that the second region has moved a distance d2 (i.e., second facial skin micro-movement 522B). The processing device 400 may then use the determined movements of the first and second regions to ascertain at least one spoken word. Consistent with disclosed embodiments, the distances d1 and d2 may be less than 1000 micrometers, less than 100 micrometers, less than 10 micrometers, or less.

[0103]

[0213] FIG. 6 is a schematic diagram of a reflection image 600 related to light reflection 300 received from an area of ​​face region 108 associated with a single spot 106 (e.g., spot 106A shown in FIG. 5 ). In a disclosed embodiment, processing device 400 can receive a reflection signal indicative of coherent light reflection from face region 108. The reflection signal can be represented by reflection image 600. Processing device 400 can then determine facial skin micromotion by applying light reflection analysis. When light source 410 is a coherent light source, the light reflection analysis can include speckle analysis or any pattern-based analysis. Such analysis can be performed by processing device 400 or processing device 460 to identify speckle patterns from which movement of the corresponding area of ​​face region 108 can be derived.

[0104]

[0214] In the illustrated example, speckles 602 appear in the reflected image 600 after recruitment of muscle fibers 520. The detected speckles or any other detected pattern can then be processed to generate reflected image data. Referring to the example above, assuming that the reflected image 600 reflects spot 106A, the reflected image data can include data indicating that a first region has moved a distance d1. In some cases, the reflected image data can be processed by any image processing algorithm (e.g., CNN and RNN) to determine skin movement in at least two areas within facial region 108. Processing device 400 can then use one or more machine learning (ML) and artificial intelligence (AI) algorithms to decipher the reflected image data and extract meaning from the facial skin micro-movements.

[0105]

[0215] As shown in FIG. 7 , memory device 700 may include software modules for executing processes consistent with the present disclosure. In particular, memory device 700 may include an illumination control module 702, a sensor communication module 704, a light reflection processing module 706, an artificial neural network (ANN) training module 710, a subvocalization decoding module 708, an output determination module 712, and a database structure access module 714. The disclosed embodiments are not limited to any particular configuration of memory 700. Furthermore, processing device 400 and / or processing device 460 may execute instructions stored in any of modules 702-714 included in memory device 700. It should be understood that references to processing devices in the following description may refer individually or collectively to processing device 400 of speech detection system 100 and processing device 460 of remote processing system 450. Accordingly, steps of any of the following processes associated with modules 702-714 may be performed by one or more processors associated with speech detection system 100.

[0106]

[0216] Consistent with disclosed embodiments, the illumination control module 702, the sensor communication module 704, the light reflection processing module 706, the subvocalization decoding module 708, the ANN training module 710, the output determination module 712, and the database access module 714 can cooperate to perform various operations. For example, the illumination control module 702 can determine light characteristics for illuminating the facial region 108. The sensor communication module 704 can receive coherent light reflections from the facial region 108 and output associated reflection signals. The light reflection processing module 706 can process the reflection signals to determine facial skin micro-movements. The subvocalization decoding module 708 and the database access module 714 can cooperate to extract meaning from the facial skin micro-movements (e.g., determine silently spoken words). In some cases, the ANN training module 710 can train an artificial network using the determined silently spoken words and the determined facial skin micro-movements. The output determination module 712 can generate a presentation of the determined words.

[0107]

[0217] The illumination control module 702 can adjust the operation of the light source 410 to illuminate the face region 108. In some embodiments, the illumination control module 702 can determine values ​​for characteristics of the projected light 104, such as light intensity, pulse frequency, duty cycle, illumination pattern, luminous flux, or any other optical property. In particular embodiments, unless the user 102 is speaking, the speech detection system 100 can operate in a first illumination mode (e.g., a low frame rate) to conserve its battery power. While operating in this first illumination mode, the speech detection system 100 can process images to detect at least one trigger in the reflected signal indicative of speech (e.g., facial movement). When such a trigger is detected, the illumination control module 702 can operate the coherent light source in a second illumination mode (e.g., a high frame rate) to enable detection of changes in the coherent light pattern (e.g., speckle) resulting from unvoiced speech. The illumination control module 702 can also be configured to modify one or more characteristics of the projected light 104 based on various types of triggers. Various types of triggers can be detected by analysis of data from the sensor communications module 704.

[0108]

[0218] The sensor communications module 704 can coordinate the operation of the light detector 412, the audio sensor 414, and the additional sensors 418 to receive measurements captured from one or more sensors integrated with or connected to the speech detection system 100. In one embodiment, the sensor communications module 704 can generate sensor data related to the user 102 using signals received from the one or more sensors. In one example, the sensor communications module 704 can receive a reflected signal from the light detector 412 and generate a first data stream of reflected images that can be used to determine facial skin micromovements in the facial region. In another example, the sensor communications module 704 can receive an audio signal from the audio sensor 414 and generate a second data stream that can be used to determine words spoken aloud by the user 102. In another example, the sensor communications module 704 can receive a movement signal from a movement sensor included in the additional sensors 418 and generate a third data stream that can be used to determine activities involving the user 102. The sensor communications module 704 can communicate the sensor data to other software modules for processing.

[0109]

[0219] The light reflection processing module 706 can process sensor data received from the sensor communication module 704 in preparation for speech decoding. In one embodiment, the light reflection processing module 706 can receive from the sensor communication module 704 a reflection signal indicative of coherent light reflection from the facial region 108, originating from the light detector 412. The reflection signal can be represented by a reflection image (e.g., reflection image 600) that can be processed by at least one image processing algorithm to extract skin movement at a preselected set of locations on the face of the user 102. The number of locations to be examined can be an input to the image processing algorithm. In some cases, the locations on the skin extracted for coherent light processing can be obtained from a list of points of interest. The list of points of interest specifies anatomical locations corresponding to the zygomaticus, orbicularis oris, laughing muscle, genioglossus, or levator labio superioris alae naris. In layman's terms, the list of points of interest can include the cheeks above the mouth, the chin, the mid-chin, the cheeks below the mouth, the high points of the cheeks, and specific points behind the cheeks. Consistent with the present disclosure, the list of points of interest may be dynamically updated with more points on the face extracted during the training phase. The entire set of locations may be sorted in descending order so that any subset of the list minimizes the word error rate (WER) for a selected number of locations examined (in order). In another embodiment, the light reflection processing module 706 crops each extracted coherent light spot from the raw image frame around the coherent light spot, and the algorithm processes only the cropped image. Typically, the coherent light spot processing process involves reducing the size of the full-frame image pixel (approximately 1.5 MP) received from the sensor communications module 704 by two orders of magnitude with a very short exposure. The exposure may be dynamically set and adapted to capture only the coherent light reflection, rather than the skin segment. The cropped image of the coherent light spot can depict the coherent light pattern. In another embodiment, the light reflection processing module 706 may apply an image processing algorithm to the reflection image.For example, the light reflection processing module 706 can improve image contrast by using a threshold to remove noise and determine black pixels, and by calculating a characteristic metric of the coherent light, such as a scalar speckle energy measure, e.g., average intensity. Additionally, the light reflection processing module 706 can analyze temporal changes in the reflection pattern (e.g., average speckle intensity). Alternatively, other metrics, such as the detection of specific coherent light patterns, may be used. The light reflection processing module 706 can then assign a set of values ​​for the characteristic metric of the coherent light, which can be calculated for each frame and aggregated to generate reflection image data indicative of facial skin micromotion. The light reflection processing module 706 can communicate the reflection image data indicative of facial skin micromotion to other software modules for processing.

[0110]

[0220] The subvocalization decoding module 708 may use machine learning (ML) and artificial intelligence (AI) algorithms to decode the reflection image data indicative of facial skin micromovements received from the light reflection processing module 706. Consistent with the present disclosure, decoding the reflection image data may include extracting meaning from the detected facial skin micromovements. In one embodiment, the subvocalization decoding module 708 may use a trained ANN to correlate words with facial skin micromovements. Different types of ANNs may be used, such as a classification NN that ultimately outputs words and a sequence-to-sequence NN that outputs sentences (word sequences). In some embodiments, during the user's normal speech, the system 100 may simultaneously sample the user's 102 voice and facial movements. Automatic speech recognition (ASR) and natural language processing (NLP) algorithms may be applied by the subvocalization decoding module 708 to the actual speech, and the results of these algorithms may be used to optimize parameters of the algorithms used by the subvocalization decoding module 708. These parameters can include various neural network weights as well as the spatial distribution of the laser beam for optimal performance. Additionally, the subvocalization decoding module 708 can limit the output of the algorithm to a predefined set of words, which can significantly improve word detection accuracy in cases of ambiguity, i.e., when two different words result in similar micro-movements on the facial skin. The used word set can be personalized over time, adjusting the dictionary to the actual words used by a particular user, along with their respective frequency and context. Additionally, the subvocalization decoding module 708 can use the context of the conversation between the user 102 and the callee. This can be determined from the input of the word and sentence extraction algorithm, eliminating non-context options to improve accuracy.The context of the conversation can be understood by applying automatic speech recognition (ASR) and natural language processing (NLP) algorithms at the user 102 and callee side.

[0111]

[0221] The ANN training module 710 may be used to train an ANN to perform unvoiced speech decoding in accordance with embodiments of the present disclosure. Thousands of examples may be required to train an ANN such as that used by the subvocalization decoding module 708. To accomplish this, the ANN training module 710 may rely on a large human group (e.g., a group of reference subjects). In one example, the subvocalization decoding module 708 may perform fine-tuning on the ANN so that the ANN is customized to the user 102. In this way, within minutes of wearing the speech detection system 100, the subvocalization decoding module 708 is ready to decode facial skin micromovements. The ANN training module 710 may be used to train two different ANN types: a classification neural network that ultimately outputs words, and a sequence-to-sequence neural network that ultimately outputs sentences (word sequences). To do so, the ANN training module 710 may upload training data from memory, such as unvoiced speech data received from the light reflection processing module 706, collected from multiple reference subjects. The silent speech data can be collected from a wide variety of people (people of various ages, genders, ethnicities, physical disabilities, etc.). Note that the number of examples required for learning and generalization may be task-dependent. For word / utterance prediction (within a closed group), at least several thousand examples can be collected. The ANN training module 710 can then augment the image-processed training data to obtain more artificial data for the training process. In particular, the augmented data can include image-processed coherent light patterns using some of the image processing steps described herein.The data expansion process can include (i) a time dropout step, in which amplitudes at random time points are replaced with zero; (ii) a frequency dropout step, in which the signal is transformed into the frequency domain and random frequency chunks are filtered out; (iii) a clipping step, in which the maximum amplitude of the signal at random time points is clamped, which can add a saturation effect to the data; and (iv) a noise addition step, in which Gaussian noise is added to the signal, and a speed change step, in which the signal is resampled to achieve a slightly slower or slightly faster signal.

[0112]

[0222] The augmented data set may undergo a feature extraction process. In this process, the ANN training module 710 may calculate time-domain unvoiced speech features. To this end, for example, each signal may be divided into low- and high-frequency components, x_low and x_high, and windowed to create time frames using, for example, a frame length of 27 ms and a shift of 10 ms. For each frame, five time-domain features and nine frequency-domain features may be calculated, for a total of 14 features per signal. Specifically, the time-domain features may be expressed as follows:

number

[0113]

[0223] The ANN training module 710 may then divide the data into a training set, a validation set, and a test set. The training set may be the data used to train the model. Hyperparameter tuning may be performed using the validation set, and final evaluation may be performed using the test set. The model architecture may depend on the task. Two different examples illustrate training two networks for two conceptually different tasks. The first task may involve signal transcoding, i.e., converting silent speech to text by generating words, phonemes, or characters. This first task may be addressed by using a sequence-to-sequence model. The second task may involve predicting words or utterances, i.e., classifying utterances spoken by a user into a single category within a closed group. This second task may be addressed by using a classification model. The disclosed sequence-to-sequence model may consist of an encoder that can convert an input signal into a high-level representation (embedding) and a decoder that generates linguistic output (i.e., characters or words) from the encoded representation. The input to the encoder may be a sequence of feature vectors. In one example, the input may enter the first layer of the encoder, a temporal convolutional layer that can downsample the data to achieve good performance. A model may use hundreds of such convolutional layers.

[0114]

[0224] In some embodiments, the output from the temporal convolutional layer at each time step may be passed to three layers of a bidirectional recurrent neural network (RNN). The ANN training module 710 may employ long short-term memory (LSTM) as units in each RNN layer. Each RNN state may be a concatenation of the state of the forward RNN and the state of the backward RNN. The decoder RNN may be initialized with the final state of the encoder RNN (a concatenation of the final state of the forward encoder RNN and the first state of the backward encoder RNN). At each time step, the decoder RNN can receive as input the previous word encoded in one-hot and embedded in a 150-dimensional space with a fully connected layer. The decoder RNN output may be projected via a matrix (depending on the training data) into the space of words or phonemes. The sequence-to-sequence model can condition the prediction of the next step based on the previous prediction. During learning, the log probability is maximized as follows. [Number] where y<i is the ground truth of the previous prediction. The classification neural network may be composed of an encoder such as in the sequence-to-sequence network and an additional fully connected classification layer on top of the encoder output. The output may be projected into the space of closed words, and the scores may be converted to the probabilities of each word in the dictionary. The result of all the above procedures can include two types of trained ANNs represented by the calculated coefficients. The coefficients can be stored in a data structure related to the speech detection system 100 (e.g., data structure 422 and data structure 464). In daily use, the ANN training module 710 may receive the latest coefficients for the trained ANN. The first ANN task can include rewriting of the signal, i.e., converting silent speech to words, phonemes, characters. The second ANN task can include prediction of words or utterances, i.e., classifying the utterances spoken by the user by mouth into a single category within a closed group.

[0115]

[0225] The output determination module 712 can coordinate the operation of the output unit 114 and the operation of the network interface 420 to generate output using the speaker 404, the light indicator 406, the haptic feedback device 408, and / or to send data to a remote computing device. In some embodiments, the output generated by the output determination module 712 can include various types of output related to silent speech determined from the detected facial skin micro-movements. Specifically, the output determination module 712 can synthesize the utterance of words determined from the facial skin movements by the subvocalization decoding module 708. The synthesis may emulate the voice of the user 102 or may emulate the voice of someone other than the user 102 (e.g., the voice of a celebrity or a pre-selected template voice). The utterance of words may be presented via the speaker 404 or transmitted to the remote computing device via the network interface 420. Specifically, the output determination module 712 can generate text output from the facial skin movements by the subvocalization decoding module 708. The text output may be transmitted to a remote computing device via the network interface 420. According to another embodiment, the output generated by the output determination module 712 may relate to the operation of the speech detection system 100. In some cases, the light indicator 406 may include a light indicator that indicates the battery status of the speech detection system 100. For example, when the battery of the speech detection system 100 is low, the light indicator may begin flashing. Additional examples of the types of output that may be generated by the output determination module 712 are described throughout this disclosure.

[0116]

[0226] The database access module 714 can cooperate with the data structures 422 and 464 to retrieve stored data. The retrieved data can include, for example, correlations between multiple words and multiple facial skin movements, correlations between specific individuals and multiple facial skin micro-movements associated with the specific individuals, etc. As described above, the subvocalization decoding module 708 can perform decoding of unvoiced speech using a trained ANN. The trained ANN can extract meaning from detected facial skin micro-movements using the data stored in the data structures 422 and 464. The data structures 422 and 464 can include separate databases, including, for example, a vector database, a raster database, a tile database, a viewport database, and / or a user input database. The data stored in the data structures 422 and 464 can be received from modules 702-712 or other components of the speech detection system 100. Additionally, the data stored in the data structures 422 and 464 can be provided as input using data input, data transfer, or data upload.

[0117]

[0227] Modules 702-714 may be implemented in software, hardware, firmware, any combination thereof, etc. Processing devices of speech detection system 100 and remote processing system 450 may be configured to execute the instructions of modules 702-714. In some embodiments, aspects of modules 702-714 may be implemented in hardware, software (contained in one or more signal processing and / or application specific integrated circuits), firmware, or any combination thereof, and may be executable by one or more processors alone or in various combinations with each other. In particular, modules 702-714 may be configured to interact with each other and / or with other modules associated with speech detection system 100 to perform functions consistent with disclosed embodiments.

[0118]

[0228] Recently, image-based facial recognition technology has become a common biometric authentication method in many communication devices. It allows users to unlock their devices, make payments, and access apps or accounts using their face as a unique identifier. However, image-based facial recognition technology has limitations, such as not always being reliable and potentially being less effective in certain situations. For example, image-based facial recognition systems can be affected by factors such as poor lighting conditions, low-quality images, and occlusions such as masks or accessories. These factors can result in inaccurate or incomplete matches. Furthermore, image recognition algorithms can be biased, leading to misidentifications based on various factors such as race, gender, or age. Furthermore, false positives and false negatives are common problems in image-based facial recognition technology. As a result, individuals may be mistaken for someone else or not recognized at all. The following disclosure proposes a new and improved technical solution for providing reliable biometric authentication that can overcome the inherent shortcomings of image-based facial recognition technology.

[0119]

[0229] Some disclosed embodiments of the present disclosure may be configured to detect facial skin micro-movements of an individual, identify the individual using the detected facial skin micro-movements, and determine an action to initiate based on the identity of the individual.

[0120]

[0230] The following description illustrates exemplary implementations for identifying an individual using facial skin micro-movements, consistent with certain disclosed embodiments, with reference to Figures 8-10. Figures 8-10 are intended merely to facilitate conceptualization of exemplary implementations for performing operations for identifying an individual using facial skin micro-movements and are not intended to limit the disclosure to any particular implementation.

[0121]

[0231] Some disclosed embodiments include a head-mounted system for identifying an individual using facial skin micro-movements. Consistent with this disclosure, a head-mounted system can be understood to include any component or combination of components that can be attached to a head, as illustrated and described elsewhere in this disclosure. The term "identifying an individual" refers to a process for determining whether an individual is known to the system. Specifically, the identification process can include comparing detected characteristics of an individual with known characteristics of the individual to identify, verify, or authenticate the individual. Consistent with this disclosure, an individual can be identified based on the individual's facial skin micro-movements. The term "facial skin micro-movements" can be understood as described and exemplified elsewhere in this disclosure. In some cases, the head-mounted system can access data indicative of baseline facial skin micro-movements and use the data to determine whether the individual currently using the head-mounted system is the same individual associated with the baseline facial skin micro-movements. Depending on the implementation, the probability that the identification process described below will result in a false identification of an individual based on facial skin micro-movements may be less than 1 in 10,000, less than 1 in 100,000, or less than 1 in 1,000,000.

[0122]

[0232] Some disclosed embodiments include a wearable housing configured to be worn on an individual's head. The term "wearable housing" may be understood as described and exemplified elsewhere in this disclosure. Consistent with some disclosed embodiments, a head-mounted system includes at least one coherent light source associated with the wearable housing. The term "coherent light source" may be understood as described and exemplified elsewhere in this disclosure. The term "associated with the wearable housing" may refer to any component linked to, incorporated into, affiliated with, connected to, or associated with the wearable housing. For example, the light source may be attached to the wearable housing using screws, adhesive, clips, heat and pressure, or any other known method for adhering two elements. Alternatively, the light source may be partially or completely contained within the housing. In alternative embodiments, the light source may be associated with the housing via a wired or wireless connection. Light source 410 in FIG. 4 is an example of a coherent light source.

[0123]

[0233] Consistent with some disclosed embodiments, at least one coherent light source can be configured to project light onto a facial region of the head. Projecting coherent light may include emitting coherent light in a direction toward a portion of the face. The coherent light may be a monochromatic wave having a well-defined phase relationship across its wavefront in a defined direction, such as toward the facial region of the head. The facial region of the head refers to any anatomical part of the human body above the shoulders. The facial region may include at least some of the forehead, eyes, cheeks, ears, nose, mouth, chin, and neck. Examples of facial regions (e.g., facial region 108) are shown in FIGS. 1-3. For example, as shown in FIGS. 1 and 2, a coherent light source 410 included in the light-sensing unit 116 may be attached to the wearable housing 110 and direct light toward the facial region. The head-mounted system may also include at least one detector associated with the wearable housing. The terms "detector" and "associated with the wearable housing" may be understood as described and exemplified elsewhere in this disclosure. At least one detector may be configured to receive coherent light reflections from the facial region and output an associated reflection signal. Receiving coherent light reflections may refer to detecting, acquiring, obtaining, or measuring electromagnetic waves (e.g., in the visible or invisible spectrum) reflected from the facial region and impinging on the at least one detector. Outputting the associated reflection signal may include sending, transmitting, generating, and / or providing information representative of or corresponding to the coherent light reflections. For example, projecting coherent light onto motionless facial skin may result in a first reflection signal indicative of the coherent light reflection. However, even small micro-movements of the facial skin may cause the at least one detector to output a second reflection signal that differs from the first reflection signal. The change between the first reflection signal and the second reflection signal may be used to determine specific facial skin micro-movements. As an example, the photodetector 412 of FIG. 4 is associated with the wearable housing 110 and employed to determine facial skin micro-movements.

[0124]

[0234] Consistent with certain disclosed embodiments, the head-mounted system includes at least one processor. The term "processor" may be understood as described and exemplified elsewhere in this disclosure. The processor may be employed to provide some or all of the functionality described herein. Processing device 400 of FIG. 4 is an example of at least one processor provided for the purpose of achieving at least some of the functionality described herein.

[0125]

[0235] Some disclosed embodiments include analyzing reflected signals to determine specific facial skin micro-movements of an individual. The term "analyzing" refers to investigating, investigating, examining, and / or studying. Reflected signals may be analyzed to determine whether they are recognized or whether they correlate with other information. For example, reflected signals (or datasets derived from reflected signals) may be analyzed to determine correlations, associations, patterns, or lack thereof, for example, within a dataset or across different datasets. Specifically, reflected signals received from at least one detector may be analyzed using one or more processing techniques, such as light pattern analysis (as described and illustrated elsewhere in this disclosure). Other processing techniques may include convolution, fast Fourier transform, edge detection, pattern recognition, object detection algorithms, clustering, artificial intelligence, machine and / or deep learning, and any other processing technique for determining specific facial skin micro-movements of an individual. In some examples, a machine learning model may be trained using training examples to determine facial skin micro-movements based on reference reflectance data. One example of such a training example may include a sample reflectance data stream along with labels indicating associated facial skin micro-movements. The trained machine learning model can be used to analyze the received reflected signals against the reference reflectance data to determine facial skin micro-movements. In some examples, at least a portion of the reflected signals can be analyzed to calculate a convolution of at least a portion of the reflected signals, thereby obtaining a resultant value of the calculated convolution. Furthermore, in response to the resultant value of the calculated convolution being a first value, a first facial skin micro-movement can be determined, and in response to the resultant value of the calculated convolution being a second value, a different second facial skin micro-movement can be determined. For example, the reflected signals received by the at least one detector can be analyzed as described elsewhere in this disclosure to determine facial skin micro-movements associated with the question "what is my mom's birthday?"Additional details and examples of how the at least one processor can analyze the reflected signals to determine specific facial skin micro-movements are described herein with reference to the light reflection processing module 706.

[0126]

[0236] Consistent with some disclosed embodiments, at least some of the specific facial skin micro-movements in a facial region may include micro-movements of less than 100 microns or less than 50 microns. In other words, the output of the process for determining specific facial skin micro-movements may be accurate enough to distinguish facial skin changes in the range of 10 to 100 microns. In some embodiments, these changes may be detected over a period of 0.01 to 0.1 seconds. In some disclosed embodiments, the determined specific facial skin micro-movements may correspond to facial expressions (e.g., smiling, frowning, or worrying) or facial muscle actions corresponding to physiological events (e.g., sneezing, laughing, yawning). In other embodiments, the facial skin micro-movements may correspond to pre-uttered or spoken phonemes, syllables, words, or phrases, as described below. In still other embodiments, the facial skin micro-movements may correspond to biological processes, such as pulse or respiratory rate. In further embodiments, the facial skin micro-movements may correspond to one or more combinations of the foregoing.

[0127]

[0237] Consistent with some disclosed embodiments, certain facial skin micromovements may correspond to preparatory phonatory muscle recruitment. As described elsewhere herein, preparatory phonatory muscle recruitment refers to the effect of facial muscle movements in the absence of audible phonation or before the onset of a phonation. A facial skin micromovement corresponds to preparatory phonatory muscle recruitment when preparatory phonatory muscle recruitment is the direct or indirect cause of the facial skin micromovement. In some cases, preparatory phonatory muscle recruitment may cause a facial skin micromovement before the onset of phonation. For example, preparatory phonatory muscle recruitment may occur between 0.1 and 0.5 seconds before the actual phonation. In some cases, preparatory phonatory muscle recruitment may include voluntary muscle recruitment that occurs when an individual begins to utter a word. In other cases, preparatory phonatory muscle recruitment may include involuntary facial muscle recruitment that occurs when certain craniofacial muscles prepare to utter a word.

[0128]

[0238] Consistent with some disclosed embodiments, the specific facial skin micro-movements may correspond to muscle recruitment during the pronunciation of at least one word or a portion thereof. For example, the at least one word may correspond to a predefined expression, a password, or a secret passphrase. As described above, the actual phonation depends on the expulsion of air from the lungs to the throat. Without this airflow, no sound can be produced. Because preparatory phonation muscle recruitment occurs prior to and separately from the muscles that transmit the airflow, preparatory phonation muscle recruitment may occur with or without a subsequent phonation.

[0129]

[0239] An example of a speech detection process is shown in Figure 8. In the illustrated example, the speech detection system 100 can analyze reflected signals associated with the question "what is my mom's birthday?" to determine specific facial skin micro-movements 800 associated with an unknown individual 802.

[0130]

[0240] Some disclosed embodiments include accessing a memory correlating a plurality of facial skin micromovements with an individual. The term "accessing a memory" refers to retrieving or examining electronically stored information. This may be done, for example, by communicating with or connecting to an electronic device or component on which the data is electronically stored. Such data may be organized into a data structure for purposes of, for example, reading stored data (e.g., to obtain related information) or writing new data (e.g., to store additional information). In some cases, the accessed memory may be part of the speech detection system or part of a remote processing device (e.g., a cloud server) that can be accessed by the speech detection system. In some examples, at least one processor may access the memory, for example, upon startup, upon shutdown, at regular intervals, at selected times, in response to a query received from the at least one processor, or at any other determined time. The memory may store data correlating a plurality of facial skin micromovements with an individual. The stored data may be any electronic representation of facial skin micromovements, any electronic representation of one or more properties determined from the facial skin micromovements, or raw measurement signals detected by at least one optical detector and representing the facial skin micromovements. Correlating a plurality of facial skin micromovements with an individual may include storing in a memory or data structure relationships between the facial skin micromovements and an identifier of the individual. This may enable efficient retrieval and identification of the individual based on these relationships. For example, the memory may be associated with a built-in mechanism for linking or associating facial skin micromovements with an identifier of the individual. In one example, correlations between particular phonemes, syllables, words, or phrases and associated skin micromovements may be stored. Depending on the implementation, these correlations may be specific to an individual or to a group or subpopulation associated with the individual. (For example, micromovements associated with some particular speech may vary by individual, country, dialect, or accent of different regions.)) Correlating a plurality of facial skin micromovements with an individual may be performed through any one of the above examples. If the intention is to verify the personal identity of a particular individual, a comparison may be made with a database of correlations associated with that particular individual (e.g., based on samples previously captured from that individual). Alternatively, if the intention is to identify an individual as part of a population or subpopulation, pre-stored data associated with that population or subpopulation may be accessed.

[0131]

[0241] Consistent with the present disclosure, the fact that a plurality of facial skin micromotions correlates with an individual means that the plurality of facial skin micromotions can uniquely identify an individual or identify an individual as part of a particular group or subpopulation. In one exemplary embodiment for uniquely identifying an individual, the probability that a plurality of facial skin micromotions is the same for two different individuals may be less than 1 in 10,000, less than 1 in 100,000, less than 1 in 1 million, or less than 1 in 10 million, depending on the implementation.

[0132]

[0242] Consistent with some disclosed embodiments, the memory can correlate multiple facial skin movements with multiple individuals. Specifically, the memory can be designed to store relationships between facial skin micromovements with multiple identifiers associated with multiple individuals. For example, a specific correlation can be stored for each of many individuals so that when a current signal is received, it can be compared with various stored correlations to uniquely identify the individual associated with the stored correlation. In some disclosed embodiments, for each of the multiple individuals, the memory can store at least 10, at least 50, or at least 100 data entries related to different facial skin micromovements. In some examples, the multiple individuals may be related, e.g., they may be part of a family or the same organization. In other examples, the multiple individuals may be unrelated but may share common attributes, such as individuals belonging to the same age group or individuals associated with the same language dialect.

[0133]

[0243] Consistent with some disclosed embodiments, the at least one processor may be configured to distinguish multiple individuals from one another based on a reflected signal unique to each of the individuals. Distinguishing multiple individuals from one another means that the at least one processor can determine the individual responsible for a received reflected signal. For example, the at least one processor may identify that a particular sentence was spoken by a particular individual and not by any other individual included in the database. The at least one processor may be configured to distinguish multiple individuals from one another by detecting a reflected signal unique to each of the individuals. A unique reflected signal means that no two individuals have the same reflected signal. For example, a unique reflected signal may be associated with a unique sequence of facial micro-movements that occurs when an individual utters or pre-utters one or more phonemes, syllables, words, or phrases, such as a passphrase. In one example, the speech detection system may be used by a group of individuals, and for each individual, the speech detection system may store personal settings. In one embodiment, the at least one processor may detect a first facial skin micro-movement of a first individual during a first time period, and then detect a second facial skin micro-movement of a second individual during a subsequent second time period. Upon identifying the first individual using the first facial skin micro-movement, the at least one processor may initiate a first action (e.g., applying personalization settings associated with the first individual), and upon identifying the second individual using the second facial skin micro-movement, the at least one processor may initiate a second action (e.g., applying personalization settings associated with the second individual). Alternatively, if a correlation is identified for a particular individual, access to the application may be provided. On the other hand, if a correlation is not identified, access may be denied.

[0134]

[0244] Referring to FIG. 8 as an example, memory 804 can store multiple baseline facial skin micro-movements (e.g., 806A, 806B, 806C, and 806D) associated with user 102. While only four baseline facial skin micro-movements are shown in the figure, those skilled in the art with the benefit of this disclosure will understand that many more baseline facial skin micro-movements can be stored as reference data for identifying an individual. For example, the multiple baseline facial skin micro-movements may be for all known phonemes or for at least 1,000 words. Furthermore, memory 804 may be designed to store multiple baseline facial skin micro-movements for multiple users, thereby enabling the processor to distinguish multiple individuals from each other based on their unique reflection signals.

[0135]

[0245] Some disclosed embodiments include searching for a match between the determined specific facial skin micromovement and at least one of the multiple facial skin micromovements in memory. The term "searching for a match" can refer to finding one or more records that meet a given set of search criteria. Different types of search algorithms can be used to search for a match, such as a linear search, a binary search, a tree-based search, and various types of database searches. In addition, as described in the following paragraphs, an artificial intelligence model can be employed to search for a match within a dataset accessible by the AI ​​model. In some cases, the initiated search can be used to find which of the multiple facial skin micromovements is most likely generated by the same individual who generated the specific facial skin micromovement. A likelihood level or certainty level of the match can be determined to provide an indication of the probability or confidence in determining that the identification hypothesis is correct, i.e., that the reference facial skin micromovement stored in memory was actually generated by the same individual who generated the specific facial skin micromovement. In some disclosed embodiments, a match can be considered to have been found when the likelihood level or certainty level is, by way of example only, greater than 90%, greater than 95%, or greater than 99%.

[0136]

[0246] Consistent with the present disclosure, the at least one processor may identify a match using an artificial neural network (e.g., a deep neural network, a convolutional neural network, etc.). The artificial neural network may be configured manually, configured using machine learning methods, or configured in combination with other artificial neural networks. Other methods that the at least one processor may use to identify a match include: comparing the determined specific facial skin micro-movement with a plurality of facial skin micro-movements in memory; obtaining a difference between the determined specific facial skin micro-movement and the plurality of facial skin micro-movements in memory and comparing it to a threshold; calculating at least one statistical value (e.g., mean, variance, or standard deviation) and comparing the at least one statistical value to a threshold; calculating a distance between two vectors in a multidimensional space, where a match is identified if the distance is below a certain threshold; calculating the cosine of the angle between the two vectors in the multidimensional space, where a match is identified if the cosine value is above a certain threshold; or any other known method of identifying a match in a database.

[0137]

[0247] Referring to FIG. 8 as an example, searching for a match may result in a first result 808A indicating that a match was identified and a second result 808B indicating that a match was not identified.

[0138]

[0248] Some disclosed embodiments include initiating a first action if a match is identified and initiating a second action different from the first action if a match is not identified. The term "initiating" can refer to performing, executing, or implementing one or more operational steps. For example, at least one processor can initiate the execution of program code instructions or cause a message to be sent to another processing device to achieve a targeted (e.g., deterministic) result or goal. An action may be initiated in response to determining whether a match between a determined particular facial skin micromovement and multiple facial skin micromovements is found in memory. The term "action" can refer to the performance or execution of an activity or task. For example, performing an action can include executing at least one program code instruction to implement a function or procedure. An action may be user-defined or system-defined (e.g., software and / or hardware), or any combination thereof. The at least one processor can select which action to initiate (e.g., a first action or a second action) and can determine to initiate the selected action based on the results of the search for a match and based on various criteria. The various criteria can include user experience (e.g., preferences based on context, location, environmental conditions, usage type, user type, etc.), user requirements (e.g., context limitations, urgency or priority of the purpose behind the action), device requirements (e.g., computational capacity, computational limitations, presentation limitations, memory capacity, or memory limitations), and communication network requirements (e.g., bandwidth, latency). For example, after a match is found, a first action of sending an audio message can be initiated. The artificial voice used to generate the audio message may be selected based on the various criteria listed above.The action may be initiated by at least one processor configured in the speech detection system, a different local processing device (e.g., associated with a device proximate to the speech detection system), a remote processing device (e.g., associated with a cloud server), or any combination thereof. Thus, "initiating an action responsive to a search result" may include performing or implementing one or more operations in response to a result of a search for a match between a determined particular facial skin micro-movement and at least one of a plurality of facial skin micro-movements in memory.

[0139]

[0249] Consistent with certain disclosed embodiments, a first action is to prepare at least one predetermined setting associated with the individual. The term "predetermined setting" refers to any configuration or preference associated with the operating software of the associated computing device or any other software installed on the computing device. Examples of such predetermined settings include language settings, default actions, preferred output modes, notification types, permissions, display brightness, volume levels, default apps, network settings, and any other options selectable by the user. Consistent with the present disclosure, when a match is identified, the at least one processor can prepare (i.e., designate, establish, or set up) a specific setting associated with the identified individual. A predetermined setting associated with an individual means that data reflecting the individual's selection of the predetermined setting is stored in a database, data structure, lookup table, or linked list. In one example, the predetermined setting can govern what the speech detection system should do when it detects silent speech. Specifically, after a match is identified, the speech detection system can automatically translate silently spoken words in English into French and synthesize them with an artificial voice that sounds like the identified individual.

[0140]

[0250] Consistent with certain disclosed embodiments, the first action (i.e., when the individual is identified) includes unlocking the computing device, and the second action (i.e., when the individual is not identified) includes presenting a message indicating that the computing device remains locked. The computing device may be any electronic device with restricted access. For example, the computing device may be a laptop, PC, tablet, smartphone, wearable electronic device, electronic door lock, entrance gate, application, system, vehicle, or communication device (e.g., mobile communication device 120). In one embodiment, the computing device may be at least a portion of the speech detection system 100. The term “unlocking a computing device” generally refers to the process of gaining access to a device with a predetermined security mechanism to prevent unauthorized access. For example, upon identifying the individual, the at least one processor may send data (e.g., a passcode) to the mobile communication device 120 that causes the mobile communication device 120 to unlock. The message indicating that the computing device remains locked may be provided by the computing device or any other device in any known manner; for example, the message may be provided audibly, textually, or virtually. For example, when the individual is not identified, the speech detection system 100 may present a message that the mobile communication device 120 remains locked.

[0141]

[0251] Consistent with some disclosed embodiments, the first action (i.e., when an individual is identified) provides personal information, and the second action (i.e., when an individual is not identified) provides public information. Personal information includes data specific to an individual or information that an entity (e.g., a user, person, organization, or other data owner) may not want shared with another entity. For example, it may include any information that, if revealed to an unauthorized entity, could cause damage, loss, or injury to the associated individual or entity. Some examples of personal information (e.g., sensitive data) may include identity identification information, location information, genetic data, information regarding health, financial, business, personal, family, educational, political, religious, and / or legal matters, and / or sexual orientation or gender identity. Public information may include any information other than personal information and may be found in public databases, such as the Internet. For example, after receiving a query from an individual, the speech detection system 100 may use certain facial micro-movements to generate a response that includes personal information (when an individual is identified) or public information (when an individual is not identified).

[0142]

[0252] Consistent with certain disclosed embodiments, a first action (i.e., when the individual is identified) authorizes the transaction, and a second action (i.e., when the individual is not identified) provides information indicating that the transaction was not authorized. Authorizing a transaction refers to the process of granting approval or permission for an activity to occur. In some cases, authorizing a transaction can include verifying the legitimacy of the transaction request and confirming the identity of the individual by finding a match. Examples of transactions can include financial transactions (e.g., debiting or depositing a bank account, buying or selling goods or services using a credit card, transferring funds between accounts, paying a bill, wire transfer, or electronic funds transfer), non-financial transactions (e.g., booking a flight, booking a hotel, ordering a product online, renting a car, signing up for a subscription, updating an address or phone number), business transactions (e.g., ordering supplies, billing a customer for a provided product or service, approving a refund, or processing an invoice), and government transactions (e.g., applying for a passport or visa, paying taxes or fines, registering a vehicle, obtaining a driver's license, obtaining a permit to operate a business). If no match is found, information indicating that the transaction was not authorized may be provided. The information may be provided via the speech detection system or the mobile communication device. For example, when the speech detection system 100 is linked to a virtual wallet, upon receiving a payment request, the speech detection system 100 may prompt the individual to silently say a password. The speech detection system 100 may then determine the password using the determined specific facial skin micro-movements and compare the determined password with previously stored passwords associated with the user. If the determined password matches the stored password (i.e., the individual is identified), the speech detection system 100 may authorize the payment. Alternatively, if the determined password does not match the stored password (i.e., the individual is not identified), the speech detection system 100 may not authorize the payment.

[0143]

[0253] Consistent with certain disclosed embodiments, a first action (i.e., when the individual is identified) allows access to the application, and a second action (i.e., when the individual is not identified) prevents access to the application. Allowing access to an application can refer to the process of granting an individual authorization to use a particular software application or to use electronic hardware. The software application may be installed on the speech detection system or any computing device associated with the individual (e.g., the individual's smartphone). For example, a person's calendar application may be accessed in response to a detected query from an identified individual, such as "What was the name of the person I met with last Wednesday?" If the individual is not identified, access to the calendar application is prohibited, and therefore the query may not be answered.

[0144]

[0254] Consistent with some disclosed embodiments, the head-mounted system includes an integrated audio output, and at least one of the first actions or at least one of the second actions includes outputting audio via the audio output. The term “integrated audio output” means that the head-mounted system includes internal audio hardware configured to generate voice without the need for an external audio interface. For example, the head-mounted system may include an audio chipset capable of converting digital audio signals to analog signals and a built-in speaker or headphone jack. Additional examples of integrated audio outputs may include or be associated with loudspeakers, wireless earphones, audio headphone, hearing aid-type devices, and any other device capable of converting electrical audio signals into corresponding voices. For example, the first action may be to emit voice into the air using an audio output device such as a loudspeaker so that nearby people can hear it, and the second action may be to emit voice using an audio output device such as a wireless earphone so that the generated audio signal is only audible to the individual.

[0145]

[0255] Referring to FIG. 8 as an example, a first action 810A may be initiated when a match is found (i.e., the individual 802 is identified as a user 102), and a second action 810B may be initiated when a match is not found (i.e., the individual 802 is not identified as a user 102).

[0146]

[0256] Consistent with certain disclosed embodiments, a match may be identified upon determination of a certainty level by at least one processor. As described elsewhere herein, determining a certainty level provides an indication of confidence that the identification hypothesis is correct. In other words, with reference to FIG. 8 , the certainty level provides an indication that the unknown individual 802 is the user 102. Consistent with certain disclosed embodiments, when the certainty level is not initially reached, the at least one processor may analyze additional reflected signals to determine additional facial skin micromovements and reach a certainty level based at least in part on the analysis of the additional reflected signals. FIG. 9 (described below) shows an exemplary implementation of these embodiments.

[0147]

[0257] 9 shows a flowchart of an example process 900 for identifying an individual above a certainty level performed by a processing device (e.g., processing device 400) of speech detection system 100. For purposes of explanation, the following description refers to certain components of speech detection system 100. However, it will be understood that other implementations are possible and that other components may be used to perform example process 900. It will also be readily understood that example process 900 can be modified to modify the order of steps, eliminate steps, or further include additional steps.

[0148]

[0258] Process 900 begins when the processing device receives a reflection from a facial region (block 902), then analyzes the reflection to determine specific facial skin micro-movements (block 904) and searches for a match between the determined specific facial skin micro-movements and at least one reference facial skin micro-movement (block 906). If a match is not found (decision block 908), the processing device may initiate a second action (block 910), and the process continues by receiving additional reflection signals (block 912), analyzing them to determine additional facial skin micro-movements, and searching for a match to identify the individual 802. If a match is found (decision block 908), the processing device may determine a certainty level of the match (block 914) and compare the determined certainty level to a threshold (decision block 916). If the certainty level is greater than the threshold, the processing device may initiate a first action (block 918), and the process continues to receive (block 912), analyze (block 904), and search (block 906) additional reflected signals. However, if the certainty level is less than the threshold, the processing device may initiate a second action (block 910).

[0149]

[0259] Consistent with some disclosed embodiments, at least one processor continuously compares new facial skin micromovements with a plurality of facial skin micromovements in memory to determine an instantaneous certainty level. In this context, the term "continuously compare" means constantly or periodically comparing new facial skin micromovements with a plurality of facial skin micromovements in memory over a period of time (e.g., during a phone call). In this context, continuous comparison includes intervals between comparisons, such as multiple times per second or multiple times per minute. The term "instantaneous certainty level" refers to the confidence level of the identity of the individual associated with the new facial skin micromovements. For example, during a phone call with a banking provider, the system may periodically compare new facial skin micromovements to ensure that the same authorized individual remains on the call. Consistent with some disclosed embodiments, at least one processor is configured to initiate an associated action when the instantaneous certainty level falls below a threshold. The fact that the instantaneous certainty level is below the threshold means that there is a risk that someone other than the identified individual is responsible for the new facial skin micromovements. The associated action refers to an action associated with the fact that the instantaneous certainty level is now below a threshold and can include a second action or halting the first action. In particular, in some embodiments, after initiating the first action, when the instantaneous certainty level falls below a threshold, the at least one processor is configured to halt the first action. For example, the first action may be to authorize a transaction at a bank by speaking with a banker over the phone and providing the banker with ongoing confirmation of the individual's identity over the phone. However, if the instantaneous certainty level falls below the threshold, the transaction can be halted because it may indicate that someone other than the individual is speaking with the banker. In some cases, the second action can include halting the first action.

[0150]

[0260] 9, after a first action is initiated at block 918, additional reflections are received and the analysis step (block 904) and search step (block 906) are performed. If the determined instantaneous certainty level associated with the additional reflections falls below a threshold, the first action can be stopped by initiating a second action.

[0151]

[0261] Consistent with some disclosed embodiments, initiating the first action may be associated with an event, and the at least one processor may continuously compare new facial skin micromovements during the event. The term "event" in this context may refer to an action, activity, a change in state, or the occurrence of any other type of detectable development or stimulus. "During the event" refers to any time from the time the event is detected to the time the event ends. In one example, the event may be a point-of-sale (POS) transaction in which a user turns on a device to authorize a transaction. In another example, the event may be related to online activity (e.g., a financial transaction, a betting session, an account access session, a gaming session, an exam, a lecture, or an educational session). In another example, the event may include maintaining a secure session with access to a resource (e.g., a file, a folder, a database, a computer program, computer code, or computer settings).

[0152]

[0262] 10 shows a flowchart of an exemplary process 1000 for identifying an individual using facial skin micro-movements, consistent with some embodiments of the present disclosure. In some disclosed embodiments, process 1000 can be executed by at least one processor (e.g., processing device 400 or processing device 460) to perform the operations or functions described herein. In some embodiments, some aspects of process 1000 may be implemented as software (e.g., program code or instructions) stored in memory (e.g., memory device 402 or memory device 466) or a non-transitory computer-readable medium. In some embodiments, some aspects of process 1000 may be implemented as hardware (e.g., dedicated circuitry). In some embodiments, process 1000 may be implemented as a combination of software and hardware.

[0153]

[0263] Referring to FIG. 10 , process 1000 includes step 1002 of projecting light toward a facial region of an individual's head. For example, at least one processor may operate a wearable coherent light source (e.g., light source 410) to illuminate the facial region 108 (e.g., using multiple output beams 508). Process 1000 includes step 1004 of receiving coherent light reflections from the facial region and outputting associated reflection signals. For example, at least one processor may operate at least one detector (e.g., at least one detector 412) to receive coherent light reflections (e.g., light reflections 300) from the facial region 108. Process 1000 includes step 1006 of analyzing the reflection signals to determine specific facial skin micro-movements of the individual. For example, the light reflection processing module 706 and the subvocalization decoding module 708 are used to determine the specific facial skin micro-movements. Process 1000 includes step 1008 of accessing a memory that correlates the multiple facial skin micro-movements with the individual. Process 1000 includes step 1010 of searching for a match between the determined particular facial skin micro-movement and at least one of the plurality of facial skin micro-movements in memory. Process 1000 includes step 1012 of initiating an action based on a determination of whether a match is found. In particular, initiating a first action (e.g., first action 810A) if a match is identified, and initiating a second action (e.g., first action 810B) different from the first action if a match is not identified.

[0154]

[0264] According to one implementation, a speech detection system projects a light pattern onto a user's facial skin (e.g., cheeks). The speech detection system can then detect light reflections from various locations on the facial skin. In particular, reflections associated with certain areas may be more relevant to extracting meaning (e.g., determining communication) than other areas. The certain areas may be areas located near certain facial muscles. Identifying the certain locations may be challenging because each user has unique facial features and the position of the light source and / or detector relative to the user's face may change with each use and even during ongoing operation. The following paragraphs describe systems, methods, and computer program products for identifying the locations of those certain areas and using light reflections from the certain areas to extract meaning and ignoring light reflections from other areas to conserve processing resources.

[0155]

[0265] Some disclosed embodiments include interpreting facial skin micro-movements. The term "interpreting facial skin movements" refers to extracting meaning from detected skin movements, as described elsewhere in this disclosure. In one example, interpreting facial skin movements can include determining one or more vocalized or subvocalized words from the facial skin movements, or determining an individual's facial expression (e.g., happiness, sadness, anger, fear, surprise, disgust, contempt, or other emotion). In another example, interpreting facial skin movements can include determining an individual's identity. Facial skin movements can be detected as described elsewhere in this disclosure.

[0156]

[0266] Some disclosed embodiments include projecting light onto a plurality of facial region areas of an individual, the plurality of areas including at least a first area and a second area. The term "projecting," as discussed elsewhere in this disclosure, includes controlling a light source (e.g., a coherent light source) so that the light source emits light in a given direction (e.g., toward a portion of a face). The term "individual," as described elsewhere in this disclosure, includes a person using a speech detection system (or another person onto whom a light source is projected). As described elsewhere in this disclosure, the term "area of ​​a facial region" or simply "area" in the context of a face includes a portion of an individual's face. A facial region area is at least 1 cm 2 , at least 2 cm 2 , at least 4 cm 2 , at least 6 cm 2 , or at least 8 cm 2Consistent with some disclosed embodiments, the projected light illuminates multiple areas of the facial region. For example, the multiple areas may include 4, 8, 16, 32, or any other number of areas. In some cases, the projected light may include at least one spot, as described elsewhere in this disclosure. The at least one spot may illuminate two or more areas of the facial region; for example, as shown in FIG. 3, a single spot 106 may illuminate different portions of the facial region 108. For example, the spot 106 may include a first portion 304A associated with a first facial muscle and a second portion 304B associated with a second facial muscle. Alternatively, a single area of ​​the facial region may be illuminated by multiple light spots. Some of the multiple areas may be spaced apart, while other portions of the multiple areas may overlap. The term "spaced apart" may refer to non-overlapping or at least a distance apart. Thus, spaced apart areas may refer to two or more areas of the facial region that do not overlap each other, even having a very small gap between them. For example, a statement that a first facial region is spaced apart from a second facial region can include at least 5 mm, at least 10 mm, at least 15 mm, or any other desired distance between the first and second regions. In some embodiments, the distance may be less than 1 mm, or between 1 mm and 5 mm. In some cases, only a portion of the facial region may be illuminated by the projected light. In other cases, the entire facial region may be illuminated by the projected light. By way of example, FIGS. 11 and 12 illustrate the use of multiple spots to illuminate multiple facial regions of an individual. As shown, each of facial areas 1100A and 1100B is represented by two or more light spots.

[0157]

[0267] Some disclosed embodiments include illuminating at least a portion of a first area and at least a portion of a second area with a common light spot. As used herein, the term "at least a portion" and / or its grammatical equivalents can refer to any percentage of a total amount. For example, "at least a portion" can refer to at least about 1%, 5%, 10%, 20%, 40%, 65%, 90%, 95%, 99%, 99.9%, or 100% of the total amount or any other fraction. The term "common light spot" means that a single (common) light spot can cover some or all of the first area and the second area. The common light spot can illuminate at least a portion of the first area and the second area. In one example, the common light spot can illuminate 30% of the first area and 10% of the second area. In another example, the common light spot can illuminate 100% of the first area and 100% of the second area. Controlling the at least one coherent light source can include illuminating contiguous areas on the face, including a first area and a second area. As an example, as shown in FIG. 3, a single light spot 106 can illuminate two or more facial areas (e.g., 304A and 304B).

[0158]

[0268] Some disclosed embodiments include illuminating a first area with a first group of spots and illuminating a second area with a second group of spots separate from the first group of spots. The term "spot group" refers to two or more light spots. The number of spots in a spot group may range from two to 64 or more. For example, a spot group may include four spots, eight spots, 16 spots, 32 spots, 64 spots, or any number greater than two. As discussed elsewhere in this disclosure, there may be variation in illumination characteristics between spots or within a spot group. Illuminating an area with a spot group may refer to illuminating some or all of a facial area with two or more spots. In one example, the spot group may illuminate at least 15% of the area, at least 40% of the area, or at least 70% of the area. The first area may be illuminated by the first group of spots, and the second area may be illuminated by the second group of spots separate from the first group of spots. In this context, the term "separate" means that a first group of spots is distinguishable from a second group of spots. For example, a first group of spots may include at least one spot that is not included in a second group of spots. As an example, Figures 11 and 12 show a first area face region 1100A illuminated by a first group of spots 1108A and a second area 1100B illuminated by a second group of spots 1108B that is separate from the first group of spots.

[0159]

[0269] Some disclosed embodiments include operating a coherent light source located within the wearable housing (as described elsewhere in this disclosure) to enable illumination of multiple areas of the facial region. As used herein, enabling illumination can refer to the process of controlling a light source to generate at least one light beam and directing the at least one light beam toward multiple areas of the facial region. For example, enabling illumination can also include utilizing a beam-splitting element (as described elsewhere in this disclosure) configured to split an input beam (as described elsewhere in this disclosure) into multiple output beams that span a portion of the face. In an alternative embodiment, enabling illumination can include utilizing multiple light sources that generate each group of output beams that cover a different respective subarea within the portion of the face. FIGS. 1 and 2 show example implementations of a speech detection system (e.g., speech detection system 100) in which at least one area of ​​the facial region (e.g., facial region 108) is illuminated by multiple light spots (e.g., light spot 106). The multiple light spots may be generated by a light-sensing unit 116 that includes at least one light source 410 and at least one light detector 412 and is located within the wearable housing 110 .

[0160]

[0270] Some disclosed embodiments include operating a coherent light source remotely located from the wearable housing (as described elsewhere in this disclosure) to enable illumination of multiple facial region areas. The term “remotely located” indicates that two objects are separated from each other and have a physical distance between them such that the two objects do not appear as a physically unified entity. For example, the coherent light source may be part of a device other than the speech detection system and may be located 1 cm or more away from the wearable housing of the speech detection system. As another example, the coherent light source may be located 3 cm or more away from the wearable housing of the speech detection system. It should be understood that the distances of 1 cm and 3 cm are exemplary and non-limiting, and other distances may be used. FIG. 3 illustrates an exemplary implementation of a speech detection system in which multiple facial region areas (e.g., a first portion 304A of the facial region 108 and a second portion 304B of the facial region 108) are illuminated by a coherent light source (e.g., a non-wearable light source 302) located away from the wearable housing.

[0161]

[0271] In some disclosed embodiments, the first area is closer to at least one of the zygomaticus and the smirk muscles than the second area. The phrase "the first area is closer to the muscle than the second area" means that the distance of the first area to the specific muscle is shorter than the distance of the second area to the specific muscle. For example, the distance can be measured from the edge of the area to the edge of the specific muscle, from the center of the area to the center of the specific muscle, or any combination thereof. In this context, the center of a shape (i.e., the first area, the second area, or the specific muscle) can be the geometric center, which is the point corresponding to the average position of all points within the shape; the circumcenter, which is the center of the smallest circle that completely encloses the 2D shape; the incenter, which is the center of an inscribed circle that touches all sides of the 2D shape; or any other reference point previously defined. As described, the first area is closer to at least one of the zygomaticus and the smirk muscles than the second area. In other words, the disclosed embodiments capture two exemplary use cases, and the first exemplary use case is one in which the first area is closer to the zygomaticus than the second area. A second exemplary use case is where a first area is closer to a laugh line than a second area. By way of example, Figure 11 illustrates one implementation of the first and second exemplary use cases. Specifically, the first use case is illustrated with respect to individual 102A, and the second use case is illustrated with respect to individual 102B.

[0162]

[0272] 11 illustrates two exemplary use cases for interpreting facial skin movement. In both exemplary use cases, areas 1100 of a plurality of facial regions of an individual 102 may be illuminated by at least one light source (e.g., light source 410, not shown). The illustrated plurality of areas includes at least a first area 1100A and a second area 1100B. In the first exemplary use case involving individual 102A, the first area 1100A is closer to the zygomaticus muscle than the second area 1100B, and in the second exemplary use case involving individual 102B, the first area 1100A is closer to the laughter muscle than the second area 1100B.

[0163]

[0273] Some disclosed embodiments include receiving reflections from multiple areas. The term "receiving" can include obtaining, reclaiming, acquiring, or gaining access to data or signals. In some cases, receiving can include reading data from a memory and / or obtaining data from a computing device over a communication channel (e.g., wired and / or wireless). In other cases, receiving can include detecting electromagnetic waves (e.g., in the visible or invisible spectrum) and generating an output related to measured properties of the electromagnetic waves. In a first embodiment, at least one processor can receive data from at least one detector indicative of light reflected from multiple areas. In a second embodiment, at least one detector can receive light rays reflected from multiple areas. The term "reflection" can refer to one or more light rays bouncing off a surface (e.g., an individual's face) or data derived from one or more light rays bouncing off a surface. For example, reflection can include light detected by a light detector after being deflected off an object. The light detected by the photodetector may be generated by at least one coherent light source of the disclosed speech detection system and / or may be generated from a light source other than the disclosed speech detection system. As an example, photodetector 412 in Figures 5A and 5B is employed to receive reflection 300 resulting from light generated by light source 410.

[0164]

[0274] 11, for example, reflected image 1102A can represent a reflection received from a first area 1100A, and reflected image 1102B can represent a reflection received from a second area 1100B. As shown, in the first exemplary use case, reflected image 1102A represents a reflection received from an area near the zygomaticus muscle. In the second exemplary use case, image 1102A represents a reflection received from an area near the laughter muscle.

[0165]

[0275] Some disclosed embodiments include detecting a first facial skin movement corresponding to a reflection from a first area and a second facial skin movement corresponding to a reflection from a second area. The term "detecting" in this context refers to the process of discovering, identifying, or determining the presence of an optical reflection (or a signal associated therewith). In one example, a change in the position of the facial skin can be detected. As discussed elsewhere in this disclosure, the detection process can include using various techniques or technologies to determine the presence of a pattern or event. In some cases, the process of detecting facial skin movement can include determining whether movement has occurred and recording information representative of the detected movement. For example, at least one processor can detect facial skin movement by applying optical reflection analysis to the received reflection. In other cases, detecting facial skin movement can include determining a time when the facial skin movement occurred. In other cases, detecting facial skin movement can include determining data representative of the facial skin movement (e.g., direction, velocity, acceleration). The term "facial skin movement" broadly refers to any type of movement prompted by the recruitment of underlying facial muscles. Facial skin movements, as described elsewhere in this disclosure, include micro-facial skin movements and larger-scale skin movements (e.g., smiling, yawning, frowning) that are generally visible and detectable to the naked eye without the need for magnification. The term "facial skin movements corresponding to a reflection from a particular area" means that the detected facial skin movements occurred in a particular area of ​​the face from which the reflection was received. For example, detecting a first facial skin movement corresponding to a reflection from a first area means that the first facial skin movement can be detected by analyzing the reflection received from the first area. Detecting a second facial skin movement corresponding to a reflection from a second area means that the second facial skin movement can be detected by analyzing the reflection received from the second area.

[0166]

[0276] In some disclosed embodiments, detecting first facial skin movement includes performing a first speckle analysis on light reflected from a first area, and detecting second facial skin movement includes performing a second speckle analysis on light reflected from a second area. The term "performing" refers to the act of performing a task, activity, or function. The term "speckle analysis" can be understood as described elsewhere in this disclosure. Consistent with this disclosure, performing speckle analysis can include detecting a speckle pattern or any other pattern in a signal received from light reflected from areas of the facial region. For example, performing speckle analysis can include identifying a secondary speckle pattern resulting from coherent light reflection from each area. In other embodiments, detecting facial skin movement can include performing a pattern-based analysis or an image-based analysis in addition to or instead of performing speckle analysis.

[0167]

[0277] Consistent with some disclosed embodiments, the first speckle analysis and the second speckle analysis are performed simultaneously by at least one processor. The term "simultaneously" means that two or more events occur during contemporaneous or overlapping periods, either where one begins and ends during the duration of the other, or where the later begins before the completion of the other. In some cases, the two or more events may be speckle analyses (or any pattern-based analyses). To perform the first speckle analysis and the second speckle analysis simultaneously, the at least one processor may include multiple processors or a multi-core processor capable of performing multiple speckle analyses simultaneously.

[0168]

[0278] 11, for example, a first facial skin movement 1104A can correspond to a reflection from a first area 1100A, and a second facial skin movement 1104B can correspond to a reflection from a second area 1100B. For example, in the first exemplary use case, the first facial skin movement 1104A represents a reflection received from an area close to the zygomaticus muscle. In the second exemplary use case, the second facial skin movement 1104B represents a reflection received from an area close to the laugh muscle.

[0169]

[0279] Some disclosed embodiments include determining, based on a difference between the first and second facial skin movements, that a reflection from a first area near at least one of the zygomaticus or the laughing muscle is a stronger indicator of communication than a reflection from a second area. Determining refers to ascertaining. For example, the processor can determine, from the difference between the first and second facial skin movements, which of the first and second facial skin movements is closer to the associated muscle. The difference between the first and second facial skin movements may include any distinction, change, or difference between the first and second facial skin movements. The difference between the first and second facial skin movements may be determined using at least one of the following techniques: surface alignment, point-to-point comparison, surface registration, topological analysis, or any other technique for determining differences between two data sets. For example, the difference between the first and second facial skin movements may include differences in movement intensity, movement trajectory, movement speed, and / or various changes in facial skin topography. Based on these differences, at least one processor may determine that the reflection from the first area is a stronger indicator of communication than the reflection from the second area. The term "communication" refers to the process of conveying information through various media, such as speech, words, body language, gestures, or signals. For example, communication can include verbal cues (e.g., words, phrases, and language) and non-verbal cues (e.g., body language, facial expressions, gestures, and eye contact). The term "communication indicator" refers to a measure or sign reflecting information conveyed by an individual. For example, stating that the reflection from the first area is a stronger indicator of communication than the reflection from the second area means that the individual is attempting to convey information and that it may be easier to determine the communication the individual is attempting to convey from the first facial skin movement than from the second facial skin movement.For example, a reflection from a first area may be a stronger indicator of communication than a reflection from a second area because facial skin micro-movements determined from reflections from the first area may be associated with higher velocity, higher displacement, or other parameters indicative of the individual's attempt to communicate information and / or the content of the information the individual is attempting to communicate. Consistent with disclosed embodiments, in a first exemplary use case, when the first area is closer to the zygomaticus muscle, the first facial skin movement may reflect movement at a speed on the order of 1-10 μm / ms, and the second facial skin movement may reflect smaller movements, if any. In a second exemplary use case, when the first area is closer to the chloasma muscle, the first facial skin movement may reflect movement on the order of 0.5-2 mm, and the second facial skin movement may reflect smaller movements, if any.

[0170]

[0280] Consistent with some disclosed embodiments, the difference between the first and second facial skin movements comprises a difference of less than 100 microns. The term "difference of less than 100 microns" means that the change between a first parameter representing the first facial skin movement and a second parameter representing the second facial skin movement is less than 100 microns. In one example, the first parameter may be the magnitude of a first displacement change vector associated with the first facial skin movement, and the second parameter may be the magnitude of a second displacement change vector associated with the second facial skin movement. The displacement change is a vector that quantifies the change in distance and direction between two measurements of the facial skin. For example, the difference between the first and second facial skin movements comprises a difference of less than 50 microns, less than 10 microns, or less than 1 micron. In other embodiments, the difference between the first and second facial skin movements comprises a difference of less than 1 millimeter. Thus, a determination that the reflection from the first area is indicative of a stronger communication than the reflection from the second area is based on a difference of less than 1 millimeter, less than 100 microns, less than 50 microns, less than 10 microns, or less than 1 micron.

[0171]

[0281] Some disclosed embodiments include processing a reflection from a first area to understand communication based on a determination that the reflection from the first area is a stronger indicator of communication. The term "processing" refers to the act of performing an operation or transformation on data or information to achieve a desired result. For example, processing can include systematically manipulating, analyzing, or modifying input to generate meaningful output. The term "processing a reflection" refers to extracting information from a signal representing a received reflection. For example, processing a reflection can include actions such as filtering, amplifying, modulating, and applying optical reflection analysis, as described elsewhere in this disclosure. Based on a determination that the reflection from the first area is a stronger indicator of communication, processing a reflection from a first area to understand communication. The term "understanding communication" refers to determining speech or facial expressions associated with non-verbal communication from facial movements, as described elsewhere in this disclosure. Consistent with the present disclosure, the reflection from the first area can be processed to create an image of a speckle pattern. Even with a fast exposure time, such as 10 ms, the speed of skin movement can be sufficient to cause the speckle pattern to change during each frame, causing bright pixels to blur and wash out. The degree of speckle blur of a given spot in a given frame can indicate, for example, the instantaneous velocity of skin movement in a small area of ​​the cheek below the spot, as indicated by a loss of contrast in the image. Processing the reflection from the first area can also include extracting quantitative image features from the image of the speckle pattern. A vector of these features extracted from successive image frames can be input into a neural network to identify communication. Details of neural network architectures and training algorithms that can be used for this purpose are described elsewhere in this disclosure. Exemplary features that can be extracted for the purpose of identifying communication include speckle contrast.Any suitable measure of contrast can be used for this purpose, such as the mean square of the brightness gradient occupying the area of ​​the speckle pattern. High contrast in the speckle pattern of a given spot from the first area can indicate that the corresponding location on the cheek is stationary, while a decrease in contrast can indicate movement. Contrast decreases as the speed of movement increases. This type of contrast feature can typically be extracted from multiple spots distributed across the first area. Additionally or alternatively, other features can be extracted from the speckle image and input into the neural network. Examples of such features include, for example, the total brightness of the speckle pattern calculated by a Sobel filter and the orientation of the speckle pattern. As an example, the subvocalization decoding module 708 of FIG. 7 can be used to process reflections from the first area to understand communication.

[0172]

[0282] Consistent with certain disclosed embodiments, the communication understood from the reflections from the first area includes words spoken by the individual. "Recognizing words spoken by the individual" refers to understanding words that are vocalized or subvocalized by the individual. Signals resulting from the reflections can be processed to understand the words as described elsewhere herein. By way of example, the word "Hello" in FIG. 11 represents a word spoken by individual 102A or individual 102B that can be understood from the reflections from the first area.

[0173]

[0283] Consistent with some disclosed embodiments, the communication understood from the reflection from the first area includes an individual's nonverbal cues. The term "nonverbal cues" refers to various forms of communication that occur without the use of spoken words. Some examples of nonverbal cues may include facial expressions, body language, gestures, eye contact, tone of voice, posture, and other subtle signals that convey meaning in interpersonal interactions. For example, nonverbal cues such as facial expressions can be used to communicate basic emotions such as happiness, sadness, anger, fear, surprise, and disgust. As discussed elsewhere in this disclosure, at least one processor can determine the nonverbal cues by analyzing the reflection signals representing facial skin micromovements in the first facial area. As an example, the emoji in FIG. 11 represent nonverbal cues that may be understood from the reflection from the first area.

[0174]

[0284] Some disclosed embodiments include ignoring a reflection from a second area based on a determination that the reflection from the first area is a stronger indicator of communication. In this context, the term "ignoring a reflection" means taking fewer processing actions on a signal representing a received reflection from the second area than on a signal representing a received reflection from the first area. In one embodiment, the signal representing a received reflection from the second area may be filtered, amplified, and analyzed to determine second facial skin movement, but because communication cannot be ascertained from the signal representing a received reflection from the second area, some quantitative features may not be extracted. In another embodiment that also includes "ignoring," during a first time frame, reflections from both the first and second areas may be processed to determine which area is closer to the zygomaticus or laughter muscle. Then, during a subsequent second time frame, if it is determined that the first area is closer to the zygomaticus or laughter muscle, the reflection from the second area may be automatically discarded.

[0175]

[0285] According to some disclosed embodiments, ignoring reflections from the second area includes omitting to use reflections from the second area to understand the communication, the term "omitting to use" referring to not using information related to reflections from the second area when determining the meaning of the communication.

[0176]

[0286] 11, for example, the reflected image 1102A may be processed to grasp a communication 1106 from facial skin movements 1104A associated with the zygomaticus or mirth muscles, and the reflected image 1102B may be ignored, e.g., not used or omitted in grasping the communication. As shown, the grasped communication may include at least one word 1106A (either silently or audibly enunciated by the individual 102A or the individual 102B) and / or at least one facial expression 1106B that serves as an example of a non-verbal cue.

[0177]

[0287] Some disclosed embodiments include determining that a first area is closer to subcutaneous tissue associated with cranial nerve V or cranial nerve VII than a second area based on a difference between the first and second facial skin movements. The term "subcutaneous tissue" refers to the layer of tissue located below the skin and above the underlying muscle and bone. It is composed of fat cells, connective tissue, blood vessels, nerves, and other structures. Cranial nerve V, also known as the trigeminal nerve, is a sensory nerve in the face that controls the jaw muscles. Cranial nerve VII controls facial expression and transmits taste from the front of the tongue. Based on a difference between the first and second facial skin movements (as described above), a first area can be determined to be closer to subcutaneous tissue associated with cranial nerve V or cranial nerve VII than a second area.

[0178]

[0288] Some disclosed embodiments include operating a coherent light source to enable bi-modal illumination of multiple facial region areas. The term "coherent light source" can be understood as described elsewhere in this disclosure. In this context, operating a coherent light source refers to regulating, supervising, commanding, authorizing, and / or enabling the coherent light source to illuminate at least a portion of a face. For example, a coherent light source can be controlled to illuminate a facial region in a particular illumination mode when turned on in response to a trigger. Bi-modal illumination refers to the ability of a coherent light source to illuminate an object using at least two different illumination modes. The term "illumination mode" refers to a specific configuration or setting of the coherent light source. Each of the two modes can be associated with different values ​​of illumination parameters, such as light intensity, illumination pattern, pulse frequency, duty cycle, and luminous flux. Light source 410 in FIG. 4 is an example of either a single-mode or multi-mode (e.g., bi-mode) light source.

[0179]

[0289] In some disclosed embodiments, the first light intensity in the first illumination mode is different from the second light intensity in the second illumination mode. In some disclosed embodiments, the first illumination pattern in the first illumination mode is different from the second illumination pattern in the second illumination mode. Light intensity refers to the brightness level of illumination, and illumination pattern refers to the arrangement, distribution, or sequence of coherent or non-coherent light emitted from a light source or reflected from a surface. Light patterns can be created by the specific design, shape, or configuration of the light source to create specific visual or non-visual effects on parts of the face. Examples of illumination patterns include a grid of light spots having the same size, a grid of light spots having various sizes, a single light spot, or any other pattern.

[0180]

[0290] Some disclosed embodiments include analyzing reflections associated with a first illumination mode to identify one or more light spots associated with a first area and analyzing reflections associated with a second illumination mode to understand the communication. The term “identifying one or more light spots associated with a first area” refers to determining which of the light spots projected by the coherent light source are located in the first area. For example, identifying one or more light spots associated with the first area may be performed by comparing light intensity at a specific location with the boundary of the first area based on image analysis of the individual's face or by any other processing method. In one example, the first illumination mode may include a first illumination pattern (e.g., 64 light spots), and the second illumination mode may include a second illumination pattern (e.g., 32 light spots). For example, referring to the first exemplary use case shown in FIG. 11, the first illumination mode may be used to identify eight light spots contained within a first area 1100A associated with the zygomaticus muscle. A second illumination mode (eg, four light spots) can then be used to illuminate the first area 1100A so that communications can be captured from received reflections.

[0181]

[0291] Consistent with some disclosed embodiments, the first area is closer to the zygomaticus muscle than the second area, and the plurality of areas further includes a third area that is closer to the laugh muscles than each of the first and second areas. The terms “plurality of areas” and “close” may be understood as described elsewhere in this disclosure. Referring to FIG. 12 as an example, the plurality of facial areas 1100 includes a first area 1100A that is closer to the zygomaticus muscle than the second area 1100B, and a third area 1100C that is closer to the laugh muscles than each of the first area 1100A and the second area 1100B. In some disclosed embodiments, based on a determination that the individual 102C is engaged in unvoiced speech, a processing device of the speech detection system can process reflections from the first area 1100A to understand the communication and ignore reflections from the second area 1100B and the third area 1100C. In other embodiments, based on a determination that the individual 102C is engaged in voiced speech, the processing device of the speech detection system may process reflections from the first area 1100A to understand the communication and ignore reflections from the second area 1100B and the third area 1100C.

[0182]

[0292] Some disclosed embodiments include analyzing reflected light from a first area when speech is produced with perceptible vocalization (i.e., voiced speech) and analyzing reflected light from a third area when speech is produced in the absence of perceptible vocalization (i.e., silent speech). In other words, rather than monitoring the entire cheek and processing reflections from multiple areas, the speech detection system can process reflections received from a subset of the cheek area within these two areas (e.g., only a few square millimeters or centimeters) to detect both silent and voiced speech. Furthermore, when multiple areas are illuminated by multiple light sources (e.g., an array of laser diodes), only the light sources illuminating these two areas can be activated, thus reducing power consumption. If significant movement of the speech detection system relative to the skin is detected, a different set of light sources can be activated. In some disclosed embodiments, different processing modes can be applied to distinguish silent speech from voiced speech. For example, during silent speech, the first area closer to the zygomaticus muscle may exhibit movement at a speed on the order of 1-10 μm / ms. Therefore, the features of the speckle image itself can change rapidly, and these features can be analyzed to generate an output. However, during voiced speech, a third area near the laughter muscles can exhibit movement of approximately 0.5 to 2 mm. Therefore, the location of the spot on the cheek can shift laterally due to cheek movement. In this case, the lateral movement of the spot can indicate a change in the spot's distance from the speech detection system and thus function as a kind of depth sensor. The two processing modes (speckle sensing and depth sensing) can be used separately to detect unvoiced and voiced speech, respectively. Alternatively or additionally, these two processing modes can be used together to improve the accuracy and specificity of measurements, for example, by applying measurements of voiced speech by a given user to learn the microscopic movement patterns that would occur during unvoiced speech by the same user.

[0183]

[0293] 13 shows a flowchart of an exemplary process 1300 for identifying an individual using facial skin micro-movements, consistent with certain embodiments of the present disclosure. In some disclosed embodiments, process 1300 can be executed by at least one processor (e.g., processing device 400 or processing device 460) to perform the operations or functions described herein. In some disclosed embodiments, some aspects of process 1300 may be implemented as software (e.g., program code or instructions) stored in memory (e.g., memory device 402 or memory device 466) or a non-transitory computer-readable medium. In some disclosed embodiments, some aspects of process 1300 may be implemented as hardware (e.g., dedicated circuitry). In some disclosed embodiments, process 1300 may be implemented as a combination of software and hardware.

[0184]

[0294] Referring to FIG. 13 , process 1300 includes step 1302 of projecting light onto a plurality of facial region areas of an individual. For example, at least one processor can operate a wearable coherent light source (e.g., light source 410) to illuminate at least a first area (e.g., first area 1100A) and a second area (e.g., second area 1100A). The first area is closer to at least one of the zygomaticus muscle or the laugh muscle than the second area. Process 1300 includes step 1304 of receiving reflections from the plurality of areas. For example, at least one processor can operate at least one detector (e.g., at least one detector 412) to receive coherent light reflections (e.g., light reflections 300) from the plurality of areas 1100. Process 1300 includes step 1306 of detecting first facial skin movement corresponding to the reflection from the first area and second facial skin movement corresponding to the reflection from the second area. For example, the at least one processor may use the optical reflection processing module 706 to detect a first facial skin movement, and a second facial skin movement corresponding to a reflection from a second area. The process 1300 includes step 1308 of determining that the reflection from the first area is a stronger indicator of communication than the reflection from the second area. For example, the determination of step 1308 may be based on a difference between the first facial skin movement and the second facial skin movement. The process 1300 includes step 1310 of processing the reflection from the first area to determine communication and ignoring the reflection from the second area. For example, the determination of step 1310 may be based on a determination that the reflection from the first area is a stronger indicator of communication. The at least one word 1106A and the at least one facial expression 1106B are examples of perceived communication.

[0185]

[0295] The above-described embodiments for interpreting facial skin movement may be implemented via a non-transitory computer-readable medium, as software (e.g., as operations performed via code), as a method (e.g., process 1300 shown in FIG. 13), or as a system (e.g., speech detection system 100 shown in FIGS. 1-3), etc. When an embodiment is implemented as a system, the operations may be performed by at least one processor (e.g., processing device 400 or processing device 460 shown in FIG. 4).

[0186]

[0296] In some embodiments, an authentication or identity verification service provider uses biometrics, such as signals indicative of an individual's facial microkinesis, for authentication purposes. For example, an authentication service provider can use an individual's facial microkinesis to verify the individual's identity. The intensity and sequence of muscle activation (e.g., muscle fiber recruitment) across an individual's facial regions vary between individuals. Muscle activation or recruitment is the process of activating motor neurons to produce various levels of muscle contraction. An individual's microkinesis can be influenced by the structure of muscles, muscle fibers, skin characteristics, subcutaneous (e.g., vascular structure, fat structure, hair structure, etc.) characteristics, etc. The iris is an example of an individual's visible muscle. The iris is the colored tissue at the front of the eye that contains the pupil at its center and helps control pupil size to regulate how much light enters the eye. While every individual's iris is round, the structure of each individual's iris can be unique and may be stable throughout an individual's lifetime. The same is true for subcutaneous muscles and their activation. Facial microkinesis can generate an individual's unique biometric signature, which can be used to identify the individual. For brevity, in the following description, facial skin micro-movements may be simply referred to as facial micro-movements. Institutions requiring customer identity verification (a.k.a. authentication) can subscribe to a provider-offered authentication service to authenticate individuals (e.g., customers) before providing access to services or facilities offered by the institution. Such institutions can include financial institutions (e.g., banking and brokerage services), subscription services (e.g., providing media content, research, or other information), online gaming sites, other online platforms, government agencies, and other organizations requiring user authentication and verification, or any other entity or service desiring customer authentication. Authentication is the process of verifying or validating an individual's identity.

[0187]

[0297] Some disclosed embodiments include identity verification of an individual based on the individual's facial micro-movements. The verification may occur via a system, a computer-readable medium, or a method. The term "identity verification" is the process of determining who an individual is. It may also refer to the process of confirming or denying whether an individual is who they claim to be. For example, in some embodiments, a system of the present disclosure may determine who an individual is based on the individual's facial micro-movements. Also, in some embodiments, a system of the present disclosure may determine (e.g., confirm or deny) whether an individual is, in fact, who they claim to be, based on the individual's facial micro-movements.

[0188]

[0298] FIG. 14 is a schematic diagram of one example embodiment including a system for providing identity verification of an individual based on the individual's facial micro-movements. As shown in FIG. 14 (and FIGS. 1-4 ), a detection system 100 associated with an individual 102 can detect signals indicative of (or representative of) the individual's facial micro-movements and communicate them to a cloud server 122 using a communication network 126, e.g., directly or via a mobile communication device 120. In some embodiments, as described elsewhere in this disclosure, the server 122 can access a data structure 124 to determine, for example, correlations between words and the individual's facial micro-movements. In some embodiments, the cloud server 122 may also be configured to verify the identity of the individual based on the received signals. In some embodiments, an authentication service provider (or identity verification service provider) can use a system such as the server 122 to provide identity verification of the individual based on the individual's facial micro-movements. In some embodiments, as shown in FIG. 14 , the speech detection systems 100 associated with the institution 1400 and the individual 102 can communicate with each other and with the cloud server 122 using the communication network 126 to request and receive identity verification of the individual.

[0189]

[0299] FIGS. 15, 16A, and 16B are simplified block diagrams illustrating different aspects of an exemplary system 1500 for providing identity verification (or identity authentication) based on an individual's facial skin micromovements (or facial micromovements). Note that only elements of the authentication system 1500 relevant to the following description are shown in these figures. Embodiments within the scope of the present disclosure may include additional or fewer elements. As shown in FIG. 15, the system 1500 includes a processor 1510 and a memory 1520. While FIG. 15 shows only one processor and one memory, in some embodiments, the processor 1510 may include multiple processors and the memory 220 may include multiple devices. These multiple processors and memories may each be of similar or different construction and may be electrically connected or separated from one another. While the memory 1520 is shown separate from the processor 1510 in FIG. 15, in some embodiments, the memory 1520 may be integrated with the processor 1510. In some embodiments, memory 1520 may be located remotely from system 1500 and accessible by system 1500. Memory 1520 may include any device for storing data and / or instructions, such as, for example, random access memory (RAM), read-only memory (ROM), hard disk, optical disk, magnetic media, flash memory, other persistent, fixed, or volatile memory, etc. In some embodiments, memory 1520 may be a non-transitory computer-readable storage medium that stores instructions that, when executed by processor 1510, cause processor 1510 to perform identity verification operations based on facial micro-movements. In some embodiments, some or all of the functionality of processor 1510 and memory 1520 may be performed by a remote processing device and memory (e.g., processing device 400 and memory device 402 of remote processing system 450; see FIG. 4 ).

[0190]

[0300] Some disclosed embodiments include receiving a reference signal to reliably verify a correspondence between a particular individual and an account at an institution. The term "receiving" can include, for example, retrieving, obtaining, or gaining access to data. Receiving can include reading data from memory and / or receiving data from a computing device via a communication channel (e.g., wired and / or wireless). The at least one processor can receive data via synchronous and / or asynchronous communication protocols, for example, by polling a memory buffer for data and / or by receiving data as an interrupt event. The term "signals" or "signal" can refer to information encoded for transmission over a physical medium or wirelessly. Examples of signals may include signals in the electromagnetic radiation spectrum (e.g., AM or FM radio, Wi-Fi, Bluetooth, radar, visible light, lidar, IR, Zigbee, Z-wave, and / or GPS signals), voice or ultrasonic signals, electrical signals (e.g., voltage, current, or charge signals), electronic signals (e.g., as digital data), tactile signals (e.g., touch), and / or any other type of information encoded for transmission over a physical medium between two entities, either via a physical medium or wirelessly (e.g., via a communications network). In some embodiments, the signals may include or represent "speckle," reflected image data, or optical reflectance analysis data (e.g., speckle analysis, pattern-based analysis, etc.) as described elsewhere in this disclosure.

[0191]

[0301] Receiving a signal in a "trusted" manner refers to receiving a trustworthy signal. For example, receiving a signal in such a way that the authenticity and / or validity of the signal can be trusted. In some embodiments, when receiving a signal in a trustworthy manner, there may be some level of assurance that the signals are valid or that they are what is expected. In some embodiments, receiving a signal in a trustworthy manner may indicate that these signals are transmitted in a secure manner so that they cannot be easily intercepted and / or decrypted by a third party. Generally, signals can be sent and received in a trustworthy manner using any known secure transmission method. In some embodiments, receiving a signal in a trustworthy manner may refer to receiving an encrypted signal. The signals may be encrypted using any now known or hereafter developed encryption technology (e.g., Wired Equivalent Privacy (WEP), Wi-Fi Protected Access (WPA), Wi-Fi Protected Access Version 2 (WPA2), Wi-Fi Protected Access Version 3 (WPA3), etc.). In some embodiments, the encrypted signals may include key(s) that can be used to decrypt the encrypted signals by methods known in the art.

[0192]

[0302] As used herein, the term “reference signal” refers to a signal used as a basis for understanding something. For example, the reference signal may be a baseline signal used for comparison purposes, e.g., to determine if a characteristic of the signal has changed. In some embodiments, the reference signal may represent one or more properties or characteristics of an individual. In some embodiments, the reference signal may represent one or more properties / characteristics of an individual. In some embodiments, the reference signal may be (or may be a representation of) a speckle pattern (e.g., reflectance image 600 in FIG. 6 ) or another light reflection pattern output by speech detection system 100 associated with the individual. In some embodiments, the reference signal may include or be a representation of one or more characteristics of the individual. In some embodiments, the reference signal may be (or may include) characteristics or features extracted from the individual's light reflection pattern. In some embodiments, one or more algorithms may be used to extract these characteristics or features of the individual's facial micro-movements that are incorporated into the reference signal. These extracted features may include reliable and / or non-reliable features. Reliable features may include measurable characteristics of an individual's facial micro-movements (e.g., temporal or amplitude onset, peak (minimum or maximum), offset, interval, time difference between peaks, and other measurable characteristics). On the other hand, non-reliable feature extraction may apply time and / or frequency analysis to obtain statistical characteristics of an individual's facial micro-movements. In some embodiments, the reference signal may represent multiple biometric signals of the individual (e.g., a combination of facial micro-movements and one or more of pulse, cardiac signal, ECG, body temperature, pressure, or other biometric signals). In some embodiments, it is also contemplated that the detected facial micro-movement signal or light reflex pattern output by speech detection system 100 may itself be used as the individual's reference signal.

[0193]

[0303] The reference signal can be configured to enable verification of a correspondence between a particular individual and an account at an institution. The term “correspondence” refers to a degree of similarity, connection, equivalence, match, or connection. For example, in some embodiments, a particular individual's reference signal can be used to determine equivalence, similarity, match, or connection between that individual and an account at an institution (e.g., a customer). An institution may associate and maintain biometric or other data for customers, and that data or related data may be included in the reference signal. The term “institution” refers to any facility or organization, without limitation. In some embodiments, an institution may be, for example, an organization that provides some type of service to multiple individuals, each of whom may have an account with the institution. In some embodiments, an institution may be a financial institution (e.g., a bank, stockbroker, mutual fund broker, etc.) where multiple customers may have accounts (e.g., cash accounts, money market accounts, stock accounts, online accounts, safe deposit boxes, etc.). In some embodiments, the institution may be a company associated with online activities (e.g., gaming activities, betting activities, exam / test providers, education / class providers, etc.), or a university or educational institution where multiple students have accounts (to access classes, bills, etc.). In some embodiments, the institution may be a healthcare provider (e.g., hospital, clinic, laboratory, etc.) or insurance provider (e.g., insurance company) where multiple patients or customers have accounts, a company where multiple employees have accounts, etc. In other embodiments, the institution may be a government agency or institution. The reference signal may be received from any source (e.g., an individual, an institution, etc.).

[0194]

[0304] In some embodiments, an institution may engage an authentication service provider and / or subscribe to an authentication service to verify the identity of an individual (or customer) in connection with providing services to the individual (e.g., before granting access to an account). The authentication service provider may use a system (e.g., system 1500 of FIGS. 15, 16A, and 16B) to verify the identity of an individual using a reference signal. In some embodiments, the system may have access to the reference signals of all customers of the institution (e.g., all account holders at a bank, all students enrolled in a university class, etc.). For example, in some embodiments, reference signals 1502 of all customers (e.g., account holders) of institution 1400 (e.g., a bank) may be sent to system 1500 (e.g., during registration) as shown in FIG. 16. System 1500 may securely store correlations 1504 between reference signals 1502 and the identities of different customers in a secure data structure (e.g., data structure 124) accessible by system 1500. In some embodiments, the customer's name and / or other identifying information (such as an account number or other information identifying the individual associated with the reference signal) may also be stored and associated with the reference data in the stored correlations 1504. As described in more detail below, the system 1500 may authenticate individuals using the stored reference signals and correlations. For example, as shown in FIG. 16B , when an individual engages in a transaction with the institution 1400 (e.g., attempts to access the customer's account), the institution 1400 may request 1506 an authentication service provider (or system 1500) to authenticate the individual (e.g., confirm the individual's identity, confirm the individual is the customer associated with the account, etc.). The system 1500 may receive the individual's real-time facial micro-motor signal 1508 as the individual engages in the transaction, and the system 1500 may compare 1512 the received real-time signal 1508 to the stored reference signals 1502 or correlations 1504 to determine whether the individual is a customer.For example, the system 1500 may determine whether the received signal is associated with a customer authorized to access the account by comparing the two signals to determine whether one or more characteristics of the received signal correspond to or sufficiently match characteristics of a stored reference signal.

[0195]

[0305] Consistent with some disclosed embodiments, a reference signal may be derived based on reference facial micro-movements detected using a first coherent light reflected from a particular individual's face. The term "reference" in "reference facial micro-movements" indicates that these facial micro-movements are used to generate the reference signal. As described elsewhere in this disclosure, the term "coherent light" includes light that is highly ordered and exhibits a high degree of spatial and temporal coherence. As described in detail elsewhere in this disclosure, when coherent light strikes an individual's facial skin, some of it is absorbed, some is transmitted, and some is reflected. The amount and type of light reflected depends on the properties of the skin and the angle at which the light strikes. For example, coherent light illuminating a rough, uneven, or textured skin surface may be reflected or scattered in a variety of different directions, resulting in a pattern of bright and dark areas known as "speckle." In some embodiments, when coherent light is reflected from an individual's face, optical reflectance analysis performed on the reflected light can include speckle analysis or any pattern-based analysis to derive information about the skin (e.g., facial skin micromotion) represented in the reflected signal. In some embodiments, the speckle pattern can result from overlapping interference of coherent light waves, resulting in intensity variations. In some embodiments, the detected speckle pattern (or any other detected pattern) can be processed to generate reflected image data from which a reference signal can be generated.

[0196]

[0306] As described elsewhere in this disclosure with reference to Figures 1-6, a speech detection system 100 associated with an individual may detect facial micro-movements of the individual. For example, with particular reference to Figures 5-7, in some embodiments, the speech detection system 100 may analyze coherent light reflections 300 from the individual's facial region 108 to determine facial micro-movements (e.g., amount of skin movement, direction of skin movement, acceleration of skin movement, speckle pattern, etc.) resulting from recruitment of muscle fibers 520 and output signals representative of the detected facial micro-movements. In some embodiments, the determined facial micro-movements may correspond to muscle activation.

[0197]

[0307] Consistent with certain disclosed embodiments, the reference signal for authentication can correspond to muscle activation during the pronunciation of at least one word. The term “authentication” (and other configurations of this term, such as “authenticate”) refers to determining the identity of an individual or determining whether an individual is, in fact, who they claim to be. In some embodiments, authentication is a security process that relies on unique characteristics of an individual to identify who they are or verify whether they are who they claim to be. For example, authentication can be a security measure that matches biometric characteristics of an individual attempting to access a resource (e.g., a device, a system, a service). As used herein, the term “pronunciation” (or other configurations, such as pronunciations, pronouncing, etc.) refers to when an individual actually utters (or vocalizes) at least one word (or syllable, etc.) or before an individual actually utters a word(s) (e.g., during silent speech or pre-vocalization). As described elsewhere in this disclosure, speech-related muscle activity occurs before speech (e.g., when airflow from the lungs is absent but facial muscles produce the desired sound; when some air flows from the lungs but the word is produced in a manner that is imperceptible using an audio sensor; etc.). For example, with reference to Figures 15, 16A, and 16B, a reference signal 1502 that may be used to verify correspondence between a particular individual and an account at an institution may correspond to a signal caused by muscle activation that occurs during or before the utterance of at least one word (e.g., during silent speech). Note that real-time signal 1508 (described below) may be generated in a similar manner.

[0198]

[0308] In some disclosed embodiments, the at least one specific muscle includes muscle activation associated with at least one specific muscle selected from the group consisting of the zygomaticus, orbicularis oris, mirthus, genioglossus, and levator labii superioris nasalis. "Muscle activation" refers to muscle tension, force, and / or movement. Such activation can occur when the brain recruits a muscle. In some embodiments, as described elsewhere in this disclosure, muscle activation or recruitment is the process of activating motor neurons to produce muscle contraction. As described elsewhere in this disclosure, facial skin micromotors include various types of voluntary and involuntary movements (e.g., movements ranging from a few micrometers to a few millimeters and lasting a fraction of a second to a few seconds) caused by muscle recruitment or activation. Some muscles, such as the quadriceps (a powerful muscle group responsible for exerting force very quickly), have a high ratio of muscle fibers to motor neurons. Other muscles, such as the eye muscles, have a much lower ratio because they use more precise, fine movements and produce smaller-scale skin deformations. As described elsewhere in this disclosure, the zygomaticus, orbicularis oris, laughing muscle, genioglossus, and levator labio naris superioris muscles can articulate specific points on an individual's cheeks above the mouth, chin, chin center, cheeks below the mouth, the high points of the cheeks, and the back of the cheeks. In some embodiments, a reference signal for authentication can be based on facial micro-movements detected from the individual's face (e.g., based on coherent light reflection) while the individual is engaged in a normal activity (e.g., speaking normally, reading quietly, etc.). In some embodiments, a reference signal can be generated based on facial micro-movements when the individual speaks or silently speaks (enunciates, articulates, enunciates, etc.) a selected word(s), syllable(s), or phrase.

[0199]

[0309] Consistent with some disclosed embodiments, the identity verification operation may further include presenting at least one word to the particular individual to pronounce. As used herein, the term "present" generally refers to making something known. For example, in some embodiments, words may be presented to the individual by visually displaying the words to the individual, and the individual may attempt to pronounce the displayed words. In some embodiments, one or more words may also be presented to the individual audibly, and the individual may repeat or attempt to repeat the words, and a signal may be generated when or before the individual utters the presented word(s). In some embodiments, one or more shapes representing one or more words (e.g., dog, cat) may be presented to the individual to pronounce.

[0200]

[0310] For example, the individual may be presented with one or more words (e.g., words, sentences, etc.) to pronounce, and the reference signal 1502 (and / or real-time signal 1508) may be generated based on facial micro-movements resulting from the individual pronouncing one or more of the presented word(s) or one or more syllables within the word(s). The one or more words may be presented to the individual to pronounce in any manner and with any device. For example, with reference to FIG. 14 , in some embodiments, the word(s) used to generate the reference signal 1502 (and / or real-time signal 1508) may be displayed to the individual in text on the display screen 1402 of the mobile communication device 120, and the reference signal 1502 (and / or real-time signal 1508) may be generated as the user pronounces the displayed word(s). In some embodiments, at least one word may be presented to the user graphically. For example, an image (e.g., a photo, a cartoon, etc.) representing a word (e.g., dog, cat, etc.) may be displayed to the individual, and the reference signal 1502 (and / or real-time signal 1508) may be generated when the individual pronounces the word represented by the image. Generally, any word (e.g., a random word) or multiple words may be presented to the individual to pronounce.

[0201]

[0311] Consistent with some disclosed embodiments, presenting at least one word to the individual to pronounce includes presenting the at least one word textually to the particular individual. For example, presenting the word “dog” may be presented by displaying the word “dog” textually. In some embodiments, presenting the word “dog” may be done by graphically showing an image (a picture, cartoon, line drawing, or another similar pictorial representation) of a dog. For example, the individual may be presented with one or more words (words, sentences, etc.) to pronounce, and the reference signal 1502 (and / or real-time signal 1508) may be generated based on facial micromovements resulting from the individual pronouncing one or more of the presented words or one or more syllables within the word(s). The one or more words may be presented to the individual to pronounce in any manner and with any device. For example, in some embodiments, a word(s) may be displayed to the individual textually on the display screen 1402 of the mobile communication device 120, and the reference signal 1502 (and / or the real-time signal 1508) may be generated when the user pronounces the displayed word(s). In some embodiments, at least one word may be presented to the user graphically. For example, an image (e.g., a picture, a cartoon, etc.) representing a word (e.g., dog, cat, etc.) may be displayed to the individual, and the reference signal 1502 (and / or the real-time signal 1508) may be generated when the individual pronounces the word represented by the image. In general, any word (e.g., a random word) or multiple words may be presented to the individual to pronounce.

[0202]

[0312] Consistent with certain disclosed embodiments, presenting at least one word to an individual to pronounce includes audibly presenting at least one word to a particular individual. For example, one or more words may be presented to an individual by, for example, audibly pronouncing the word(s) over a speaker. For example, with reference to FIG. 16 , when an individual is setting up an account with an institution, one or more words can be presented to the individual to pronounce, and a reference signal 1502 can be generated based on the resulting facial micromovements. As another example, when using a mobile communication device 120 to conduct a transaction with institution 1400 (e.g., setting up an account or attempting to access an account), the word(s) used to generate reference signal 1502 (and / or real-time signal 1508) may be audibly presented to the individual using a speaker of device 120, output unit 114 of speech detection system 100, or another speaker. The speech detection system 100 associated with the individual can then generate a baseline signal 1502 (and / or a real-time signal 1508) based on muscle activation when the user pronounces a word(s) or one or more syllables within a word(s).

[0203]

[0313] It should be noted that while the mobile communication device 120 is described as being used to audibly, textually, and / or graphically display to the individual the word(s) used to generate the reference signal 1502 and / or the real-time signal 1508, this is merely exemplary. In general, the word(s) may be presented to the individual on any device. For example, in some embodiments, the words may be presented visually (e.g., textually, graphically, etc.) on a screen 1600 (see FIG. 16B ) of any device accessible to the individual (e.g., a visual display of a smartphone, tablet, smartwatch, personal digital assistant, desktop computer, laptop computer, Internet of Things (IoT) device, dedicated terminal, wearable communication device, VR / XR glasses, etc.). Similarly, the words may be presented audibly to the individual on any device (e.g., a speaker on any one of the above devices, etc.). In some embodiments, instead of presenting the user with the word(s) used to generate the reference signal 1502 and / or real-time signal 1508, it is contemplated that a question or prompt that generates the word(s) may be presented to the user (e.g., audibly, textually, graphically, etc.). For example, a query such as "what is your password?", "what is the city of your birth?", etc. may be presented to the individual, and the reference signal 1502 (and / or real-time signal 1508) may be generated from the response. In some embodiments, both the reference signal 1502 and the real-time signal 1508 may be generated by presenting the individual with the same word(s) or syllable(s) to pronounce.

[0204]

[0314] Consistent with certain disclosed embodiments, the at least one word presented may be a password. In general, a "password" may be any word or character string. In some embodiments, a password may be a character string, one or more words, or a phrase that must be used to gain entry to something. For example, with reference to FIG. 16 , when an individual is setting up an account with an institution, the individual may be asked to pronounce (e.g., vocalize or pre-vocalize) the account's password, and reference signal 1502 may be generated based on the resulting facial micro-movements. As another example, in an embodiment in which an individual is attempting to access a customer's account at a financial institution, the individual may be asked to pronounce the password associated with the account, for example, by being presented with a query (e.g., "what is your password?"). Reference signal 1502 and / or real-time signal 1508 may then be generated based on coherent light reflections from the individual's face as the individual pronounces the password.

[0205]

[0315] In some embodiments, the reference signal for authentication may correspond to muscle activation during pronunciation of one or more syllables. For example, the reference signal may be generated when an individual pronounces (speaks or pre-speaks) a syllable, such as, for example, a vowel or any other syllable. Although not required, in some embodiments, one or more syllables (e.g., a vowel or any other character), or one or more words containing syllables, may be presented to the individual, and the reference signal 1502 for authentication (and / or real-time signal 1508) may be generated by system 1500 based on facial micro-movements as the individual pronounces the one or more syllables.

[0206]

[0316] Some disclosed embodiments include storing correlations between the identity of a particular individual and reference signals reflecting facial micromovements in a secure data structure. A "secure data structure" is a location where data or information can be securely stored without being susceptible to unauthorized access. Unauthorized access may include access by members of an organization (e.g., an agency, an authentication service provider, etc.) who are not authorized to access the stored data, or access by members outside the organization. Data structures consistent with this disclosure can include any collection of data values ​​and relationships between them. Data may be stored linearly, horizontally, hierarchically, relationally, non-relationally, unidimensionally, multidimensionally, operationally, ordered, unordered, object-oriented, centralized, decentralized, distributed, custom, or in any manner that allows data access. By way of non-limiting example, data structures may include arrays, associative arrays, linked lists, binary trees, balanced trees, heaps, stacks, queues, sets, hash tables, records, tagged unions, ER models, and graphs. For example, a data structure may include an XML database, an RDBMS database, an SQL database, or NoSQL alternatives for data storage / retrieval, such as MongoDB, Redis, Couchbase, Datastax Enterprise Graph, Elastic Search, Splunk, Solr, Cassandra, Amazon DynamoDB, Scylla, HBase, and Neo4J. The data structure may be a component of a disclosed system or a component of remote computing (e.g., a cloud-based data structure). The data in the data structure may be stored in contiguous or non-contiguous memory. Furthermore, a data structure as used herein does not require that the information be co-located. It may, for example, be distributed across multiple servers that may be owned or operated by the same or different entities. Thus, the term "data structure" as used herein in the singular includes multiple data structures.

[0207]

[0317] In some embodiments, the secure data structure may be a secure database. The stored information may be encrypted within the secure data structure. As described elsewhere in this disclosure, the term "database" refers to a collection of data, which may be distributed or non-distributed. In some embodiments, the secure data structure may be a secure enclave (also known as a trusted execution environment). A secure enclave is a computing environment that provides code and data isolation from the operating system using hardware-based isolation or by isolating entire virtual machines by placing a hypervisor within a Trusted Computing Base (TCB). A Trusted Computing Base (TCB) may be a computing system that provides a secure environment for operation. This includes its hardware, firmware, software, operating system, physical location, built-in security controls, and prescribed security and safety procedures. A hypervisor, also known as a virtual machine monitor or VMM, is software that creates and runs virtual machines (VMs). A hypervisor allows a single host computer to support multiple guest VMs by virtually sharing resources such as memory and processing. Even users with physical or root access to the machine and operating system may not be able to access the contents of a secure enclave or tamper with the execution of code within the enclave. Secure enclaves provide CPU hardware-level isolation and memory encryption on servers by isolating application code and data and encrypting memory. Secure enclaves are at the core of confidential computing. In some embodiments, a set of security-related instruction codes can be built into the processor to protect stored data.Data within a security enclave can be protected because the enclave is decrypted on-the-fly only within the processor and then only for code and data executing within the enclave itself. Using appropriate software, a secure enclave can enable encryption of stored data and provide full-stack security for the stored data. In some embodiments, secure enclave support may be built into one or more processors of system 1500 (such as processor 1510). In some embodiments, a secure data structure may include encrypted key / value storage. The secure data structure may, in some embodiments, be on a dedicated chip, in a separate IC circuit, or part of processor 1510. In some embodiments, the secure data structure may include remote authentication. For example, corresponding authentication keys may be stored locally on system 1500 and on a remote server, and access may be provided to a stored database based on a successful comparison of the two authentication keys.

[0208]

[0318] Consistent with some disclosed embodiments, a correlation between a particular individual's identity and a reference signal (reflecting that individual's facial micro-movements) can be stored in a secure data structure. A "correlation" refers to a relationship or connection between an individual's identity and that individual's reference signal. For example, a correlation is a measure of the degree to which the two are related. In some embodiments, a representation (or signature) of the individual's received reference signal can be stored as the correlation. While not required, in some embodiments, the stored signature may be a reduced-size version of the received reference signal. In some embodiments, an encrypted version of the signature can be stored in the secure data structure. In some embodiments, a "hash" of the received reference signal can be stored as the correlation. As will be appreciated by those skilled in the art, a hash is a unique digital signature generated from an input signal (e.g., a received reference signal) using, for example, a commercially available algorithm. To reduce the likelihood of unauthorized access to data, for example, a hashed / encrypted signature of the individual can be stored as the correlation in a secure data structure. In some embodiments, the correlation can be or include a feature or characteristic of the reference signal extracted using, for example, a feature extraction algorithm. In some embodiments, the correlation may include significant information or landmarks in the reference signal (e.g., the location and orientation of peaks and / or valleys, spatial and / or temporal gaps between peaks and / or valleys). In some embodiments, the encrypted reference signal itself may be stored as a correlation. Because the stored correlation is a representation of an individual's facial micro-movements as influenced by that individual's human traits (e.g., muscle fiber structure, vascular structure, tissue structure, etc.), the stored correlation may uniquely identify the individual to whom the reference signal corresponds. In some embodiments, the correlation may include the identity (e.g., name, account number, or other identifying information) of the individual to whom the reference signal corresponds or is associated.15, system 1500 stores correlations 1504 of reference signals 1502 for an individual in a secure data structure in memory 1520. In another exemplary embodiment, as shown in Figures 16A and 16B, system 1500 stores correlations 1504 of reference signals 1502 for different individuals (e.g., Tom, Amy, Ron, etc.) in a secure data structure (e.g., data structure 124) in a remote database.

[0209]

[0319] Some disclosed embodiments include receiving, via the authority, a request to authenticate the particular individual following the storing. As previously mentioned, the term “authenticating” refers to determining an individual's identity or whether an individual is, in fact, who they (implicitly or explicitly) claim to be. In some embodiments, authentication is a security process that relies on unique characteristics of an individual to identify who they are or verify that they are who they claim to be. For example, authentication is a security measure that matches, for example, biometric characteristics of an individual attempting to access a resource (e.g., a device, system, service). In some embodiments, access to the resource is granted only if the individual's biometric characteristics match those stored in a secure data structure for that particular individual. Consistent with its common usage, the term “request” means requesting something. In some embodiments, the request may be an electronic or digital signal. For example, in some embodiments, as shown in FIGS. 15, 16A, and 16B, a system 1500 can receive a request 1506 for authentication of an individual. In some embodiments, the request 1506 may originate from an institution with which the individual is involved in a transaction (e.g., institution 1400). In some embodiments, the individual can send the request 1506 to the institution (e.g., as part of a transaction), and the institution can forward the request to the system 1500.

[0210]

[0320] In some embodiments, upon receiving (or in response to) a request for a transaction from an individual, the institution 1400 can send a request 1506 to an authentication service provider to authenticate the individual. Without limitation, a transaction can include any type of interaction between two parties (e.g., the individual and the institution 1400). In some embodiments, a transaction between an individual and the institution 1400 can include a request from the individual to the institution 1400 to take some action (e.g., a request for information, a request for access to an account, a request to transfer funds, etc.).

[0211]

[0321] Consistent with certain disclosed embodiments, authentication relates to financial transactions at an institution. As described elsewhere in this disclosure, the term “transaction” refers to any type of interaction between two parties (e.g., an individual and an institution). For example, an individual may request access to a customer's account at a financial institution (e.g., a bank, a stockbroker, etc.), and in response to the request, the institution may request authentication of the individual from an authentication service to authenticate the individual (e.g., to verify that the individual requesting access is the customer associated with the account) before allowing the individual to access the account and conduct another transaction. The institution may require authentication when the individual attempts to conduct any type of transaction. Consistent with certain embodiments, a financial transaction includes at least one of transferring funds, purchasing stock, selling stock, accessing financial data, or accessing a specific individual's account. For example, an individual may attempt to trade stocks from an account at a stockbroker, transfer funds from an account, or view financial statements, and the broker may send a request for authentication of the individual to system 1500.

[0212]

[0322] Any type of institution can use the disclosed system and authentication service. Consistent with some embodiments, the institution is associated with a certain online activity, and once authenticated, a particular individual is provided access to perform the online activity. The term “online activity” can refer to any activity performed using the Internet or other computer network. For example, when an individual wants to log in to a customer account at an online stock brokerage (or other financial institution) and / or trade stocks, if the system (in response to an authentication request) indicates that the individual is a customer or an individual authorized to operate the account (in some embodiments only), the individual may be allowed to continue trading. An institution can be involved in providing any type of online activity to an individual. Consistent with some embodiments, the online activity is at least one of a financial transaction, a betting session, an account access session, a gaming session, an exam, a lecture, or an educational sessio...

Claims

1. 1. A head-mounted system for identifying an individual using facial skin micro-movements, comprising: a wearable housing configured to be worn on the individual's head; at least one coherent light source associated with the wearable housing and configured to project light onto a facial region of the head; at least one detector associated with the wearable housing and configured to receive coherent light reflections from the facial region and output an associated reflection signal; analyzing the reflected signals to determine specific facial skin micro-movements of the individual; accessing a memory correlating a plurality of facial skin micro-movements with said individual; searching for a match between the determined particular facial skin micro-movement and at least one of the plurality of facial skin micro-movements in the memory; If a match is identified, initiating a first action; at least one processor configured to initiate a second action different from the first action if a match is not identified; a head-mounted system including:

2. The head-mounted system of claim 1 , wherein the first action establishes at least one predetermined setting associated with the individual.

3. The head-mounted system of claim 1 , wherein the first action unlocks the computing device and the second action includes presenting a message indicating that the computing device remains locked.

4. The head-mounted system of claim 1 , wherein the first action provides personal information and the second action provides public information.

5. The head-mounted system of claim 1 , wherein the first action authorizes the transaction and the second action provides information indicating that the transaction is not authorized.

6. The head-mounted system of claim 1 , wherein the first action allows access to an application and the second action prevents access to the application.

7. The head-mounted system of claim 1 , wherein at least some of the specific facial skin micro-movements in the facial region are micro-movements of less than 100 microns.

8. The head-mounted system of claim 1 , wherein the specific facial skin micro-movements correspond to muscle recruitment of pre-phonation.

9. The head-mounted system of claim 1 , wherein the specific facial skin micro-movements correspond to muscle recruitment during pronunciation of at least one word.

10. The head-mounted system of claim 9 , wherein the at least one word corresponds to a password.

11. 2. The head-mounted system of claim 1, wherein the memory is configured to correlate a plurality of facial skin movements with a plurality of individuals, and the at least one processor is configured to distinguish the plurality of individuals from one another based on a reflected signal unique to each of the plurality of individuals.

12. 10. The head-mounted system of claim 1, further comprising an integrated audio output, wherein at least one of the first actions or at least one of the second actions comprises outputting audio via the audio output.

13. The head-mounted system of claim 1 , wherein the match is identified upon a confidence level determination by the at least one processor.

14. 14. The head-mounted system of claim 13, wherein when the confidence level is not initially reached, the at least one processor is configured to analyze additional reflected signals to determine additional facial skin micro-movements and reach the confidence level based at least in part on analysis of the additional reflected signals.

15. 14. The head-mounted system of claim 13, wherein the at least one processor is further configured to continually compare new facial skin micro-movements with the plurality of facial skin micro-movements in the memory to determine an instantaneous confidence level.

16. 16. The head-mounted system of claim 15, wherein after initiating the first action, when the momentary confidence level falls below a threshold, the at least one processor is configured to stop the first action.

17. The head-mounted system of claim 15 , wherein the at least one processor is configured to initiate an associated action when the instantaneous confidence level falls below a threshold.

18. 16. The head-mounted system of claim 15, wherein initiating the first action is associated with an event, and the at least one processor is configured to continuously compare the new facial skin micro-movements during the event.

19. 1. A method for identifying an individual using facial skin micro-movements, comprising: operating a wearable coherent light source configured to project light toward a facial region of the individual's head; operating at least one detector configured to receive coherent light reflections from the facial region and output an associated reflection signal; analyzing the reflected signals to determine specific facial skin micro-movements of the individual; accessing a memory correlating a plurality of facial skin micro-movements with an individual; searching for a match between the determined particular facial skin micro-movement and at least one of the plurality of facial skin micro-movements in a memory; If a match is identified, initiating a first action; if no match is identified, initiating a second action different from the first action; A method comprising:

20. 1. A non-transitory computer-readable medium comprising instructions that, when executed by at least one processor, cause the at least one processor to perform operations for identifying an individual using facial skin micro-movements, the operations comprising: operating a wearable coherent light source configured to project light toward a facial region of the individual's head; operating at least one detector configured to receive coherent light reflections from the facial region and output an associated reflection signal; analyzing the reflected signals to determine specific facial skin micro-movements of the individual; accessing a memory correlating a plurality of facial skin micro-movements with an individual; searching for a match between the determined particular facial skin micro-movement and at least one of the plurality of facial skin micro-movements in a memory; If a match is identified, initiating a first action; and if no match is identified, initiating a second action different from the first action. Non-transitory computer-readable medium.

21. 1. A non-transitory computer-readable medium comprising instructions that, when executed by at least one processor, cause the at least one processor to perform operations for interpreting facial skin movement, the operations comprising: projecting light onto a plurality of facial region areas of the individual, the plurality of facial region areas including at least a first area and a second area, the first area being closer to at least one of the zygomaticus muscle or the laughing muscle than the second area; receiving reflections from the plurality of areas; Detecting first facial skin movements corresponding to a reflection from the first area and second facial skin movements corresponding to a reflection from the second area; determining, based on a difference between the first facial skin movement and the second facial skin movement, that the reflection from the first area near at least one of the zygomaticus or the laughing muscle is a stronger indicator of communication than the reflection from the second area; and processing the reflections from the first area to confirm the communication and ignoring the reflections from the second area based on the determination that the reflections from the first area are indicative of a stronger communication. Non-transitory computer-readable medium.

22. 22. The non-transitory computer-readable medium of claim 21, wherein the first area and the second area are spaced apart.

23. 22. The non-transitory computer-readable medium of claim 21, wherein the communication ascertained from the reflection from the first area includes words spoken by the individual.

24. 22. The non-transitory computer-readable medium of claim 21, wherein the communication ascertained from the reflection from the first area includes non-verbal cues of the individual.

25. 22. The non-transitory computer-readable medium of claim 21, wherein the operating further comprises operating a coherent light source located within the wearable housing to enable illumination of the plurality of facial region areas.

26. 22. The non-transitory computer-readable medium of claim 21, wherein the operating further comprises operating a coherent light source located remotely from the wearable housing to enable illumination of the plurality of facial region areas.

27. 22. The non-transitory computer-readable medium of claim 21, wherein the operations further comprise illuminating at least a portion of the first area and at least a portion of the second area with a common spot of light.

28. 22. The non-transitory computer-readable medium of claim 21 , wherein the operations further include illuminating the first area with a first set of spots and illuminating the second area with a second set of spots different from the first set of spots.

29. 22. The non-transitory computer-readable medium of claim 21, wherein the operations further include operating a coherent light source to enable bi-modal illumination of the plurality of facial region areas; analyzing reflections associated with a first illumination mode to identify one or more light spots associated with the first area; and analyzing reflections associated with a second illumination mode to confirm the communication.

30. 30. The non-transitory computer-readable medium of claim 29, wherein a first light intensity of the first illumination mode is different from a second light intensity of the second illumination mode.

31. 30. The non-transitory computer-readable medium of claim 29, wherein a first illumination pattern of the first illumination mode is different from a second illumination pattern of the second illumination mode.

32. 22. The non-transitory computer-readable medium of claim 21 , wherein the operations further include determining that the first area is closer to subcutaneous tissue associated with cranial nerve V or cranial nerve VII than the second area based on a difference between the first facial skin movement and the second facial skin movement.

33. 22. The non-transitory computer-readable medium of claim 21, wherein the first area is closer to the zygomaticus muscle than the second area, and the plurality of light areas further includes a third area closer to the chloasma muscle than each of the first area and the second area.

34. 34. The non-transitory computer-readable medium of claim 33, wherein the operations further include analyzing reflected light from the first area when speech is generated with perceptible vocalization, and analyzing reflected light from the third area when speech is generated in the absence of perceptible vocalization.

35. 22. The non-transitory computer-readable medium of claim 21, wherein the difference between the first facial skin movement and the second facial skin movement comprises a difference of less than 100 microns, and wherein the determination that the reflection from the first area is a stronger indication of communication than the reflection from the second area is based on the difference of less than 100 microns.

36. 22. The non-transitory computer-readable medium of claim 21, wherein ignoring the reflection from the second area comprises omitting use of the reflection from the second area to verify the communication.

37. 22. The non-transitory computer-readable medium of claim 21, wherein detecting the first facial skin movement comprises performing a first speckle analysis on light reflected from the first area, and detecting the second facial skin movement comprises performing a second speckle analysis on light reflected from the second area.

38. 38. The non-transitory computer-readable medium of claim 37, wherein the first speckle analysis and the second speckle analysis occur simultaneously by the at least one processor.

39. 1. A method for interpreting facial skin movement, comprising: projecting light onto a plurality of facial region areas of the individual, the plurality of facial region areas including at least a first area and a second area, the first area being closer to at least one of the zygomaticus muscle or the laughing muscle than the second area; receiving reflections from the plurality of areas; Detecting first facial skin movements corresponding to a reflection from the first area and second facial skin movements corresponding to a reflection from the second area; determining, based on a difference between the first facial skin movement and the second facial skin movement, that the reflection from the first area near at least one of the zygomaticus or the laughing muscle is a stronger indicator of communication than the reflection from the second area; and processing the reflections from the first area to confirm the communication and ignoring the reflections from the second area based on the determination that the reflections from the first area are a stronger indicator of communication.

40. 1. A system for interpreting facial skin movement, comprising: projecting light onto a plurality of facial region areas of the individual, the plurality of facial region areas including at least a first area and a second area, the first area being closer to at least one of the zygomaticus muscle or the laughing muscle than the second area; receiving reflections from the plurality of areas; detecting a first facial skin movement corresponding to a reflection from the first area and a second facial skin movement corresponding to a reflection from the second area; determining, based on a difference between the first facial skin movement and the second facial skin movement, that the reflection from the first area near at least one of the zygomaticus muscle or the laughing muscle is a stronger indicator of communication than the reflection from the second area; at least one processor configured to process the reflections from the first area to confirm the communication and ignore the reflections from the second area based on the determination that the reflections from the first area are indicative of a stronger communication; A system including:

41. 1. A non-transitory computer-readable medium comprising instructions that, when executed by at least one processor, cause the at least one processor to perform an identity verification operation based on facial micro-movements, the operation comprising: receiving a reference signal that reliably verifies a correspondence between a particular individual and an account at the institution, the reference signal being derived based on reference facial micro-movements detected using a first coherent light reflected from the face of the particular individual; storing in a secure data structure the correlation between the identity of the particular individual and the reference signals reflecting facial micro-movements; receiving, via said authority, a request to authenticate said particular individual after storing said request; receiving a real-time signal indicative of a second coherent light reflex derived from a second facial micro-movement of the particular individual; comparing the real-time signal to the reference signal stored in the secure data structure, thereby authenticating the particular individual; and upon authentication, notifying the institution that the particular individual has been authenticated. Non-transitory computer-readable medium.

42. 42. The non-transitory computer-readable medium of claim 41, wherein the authentication is associated with a financial transaction at the institution.

43. 43. The non-transitory computer-readable medium of claim 42, wherein the financial transaction comprises at least one of a transfer of funds, a purchase of stock, a sale of stock, access to financial data, or access to the particular individual's account.

44. 42. The non-transitory computer-readable medium of claim 41, wherein receiving the real-time signal and comparing the real-time signal occurs multiple times during a transaction, and the operations further include reporting a discrepancy if a subsequent difference is detected following the notifying.

45. 45. The non-transitory computer-readable medium of claim 44, wherein the operations further comprise determining a confidence level that an individual associated with the real-time signal is the particular individual.

46. 46. ​​The non-transitory computer-readable medium of claim 45, wherein the operations further comprise terminating the transaction when the confidence level falls below a threshold.

47. 46. ​​The non-transitory computer-readable medium of claim 45, wherein the transaction is a financial transaction including providing access to an account of the particular individual, and when the confidence level is below a threshold, the operation further includes blocking the individual associated with the real-time signal from the particular individual's account.

48. 42. The non-transitory computer-readable medium of claim 41, wherein the reference signal for authentication corresponds to muscle activation during pronunciation of at least one word.

49. 49. The non-transitory computer-readable medium of claim 48, wherein the muscle activation is associated with at least one specific muscle including the zygomaticus, orbicularis oris, laughing muscle, genioglossus, or levator labii superioris alae naris.

50. 49. The non-transitory computer-readable medium of claim 48, wherein the at least one word is a password.

51. 49. The non-transitory computer-readable medium of claim 48, wherein the operations further comprise presenting the at least one word to the particular individual for pronunciation.

52. 52. The non-transitory computer-readable medium of claim 51, wherein presenting the at least one word for pronunciation to the particular individual comprises presenting the at least one word aloud.

53. 52. The non-transitory computer-readable medium of claim 51, wherein presenting the at least one word for pronunciation to a particular individual comprises presenting the at least one word textually.

54. 42. The non-transitory computer-readable medium of claim 41, wherein the reference signal for authentication corresponds to muscle activation during pronunciation of one or more syllables.

55. 42. The non-transitory computer-readable medium of claim 41, wherein the authority is associated with an online activity and, upon authentication, provides the particular individual with access to perform the online activity.

56. 56. The non-transitory computer-readable medium of claim 55, wherein the online activity is at least one of a financial transaction, a betting session, an account access session, a gaming session, an exam, a lecture, or an educational session.

57. 42. The non-transitory computer-readable medium of claim 41, wherein the authority is associated with a resource, and upon authentication, the particular individual is provided with access to the resource.

58. 58. The non-transitory computer-readable medium of claim 57, wherein the resource is at least one of a file, a folder, a data structure, a computer program, computer code, or a computer configuration.

59. 1. A method for providing identity verification based on facial micro-movements, comprising: receiving a reference signal that reliably verifies a correspondence between a particular individual and an account at the institution, the reference signal being derived based on reference facial micro-movements detected using a first coherent light reflected from the face of the particular individual; storing in a secure data structure the correlation between the identity of the particular individual and the reference signals reflecting facial micro-movements; receiving, via said authority, a request to authenticate said particular individual after storing said request; receiving a real-time signal indicative of a second coherent light reflex derived from a second facial micro-movement of the particular individual; comparing the real-time signal to the reference signal stored in the secure data structure, thereby authenticating the particular individual; and upon authentication, notifying the institution that the particular individual has been authenticated.

60. 1. A system for providing identity verification based on facial micro-movements, comprising: receiving a reference signal that reliably verifies a correspondence between a particular individual and an account at the institution, the reference signal being derived based on reference facial micro-movements detected using a first coherent light reflected from the face of the particular individual; storing the correlation between the identity of the particular individual and the reference signals reflecting facial micro-movements in a secure data structure; receiving, via the authority, a request to authenticate the particular individual after storing the request; receiving a real-time signal indicative of a second coherent light reflex derived from a second facial micro-movement of the particular individual; comparing the real-time signal to the reference signal stored in the secure data structure, thereby authenticating the particular individual; at least one processor configured, upon authentication, to notify the institution that the particular individual has been authenticated; A system including:

61. 1. A non-transitory computer-readable medium comprising instructions that, when executed by at least one processor, cause the at least one processor to perform operations for continuous authentication based on facial skin micro-movements, the operations comprising: receiving, during an ongoing electronic transaction, a first signal representative of a coherent light reflex associated with a first facial skin micro-movement during a first period of time; Using the first signal to determine the identity of a particular individual associated with the first facial skin micro-movement; receiving, during the ongoing electronic transaction, a second signal representative of a coherent light reflex associated with a second facial skin micro-movement, the second signal being received during a second time period subsequent to the first time period; using the second signal to determine that the particular individual is also associated with the second facial skin micro-movement; receiving, during the ongoing electronic transaction, a third signal representative of a coherent light reflex associated with a third facial skin micro-movement, the third signal being received during a third time period subsequent to the second time period; using the third signal to determine that the third facial skin micromovement is not associated with the particular individual; and and initiating an action based on the determination that the third facial skin micromovement is not associated with the particular individual. Non-transitory computer-readable medium.

62. 62. The non-transitory computer-readable medium of claim 61, wherein the ongoing electronic transaction is by telephone.

63. 62. The non-transitory computer-readable medium of claim 61 , wherein during the second period of time, the operations further include continuously outputting data confirming that the particular individual is associated with the second facial skin micromovement.

64. 62. The non-transitory computer-readable medium of claim 61, wherein the action includes providing an indication that the particular individual is not responsible for the detected third facial skin micromovement.

65. 62. The non-transitory computer-readable medium of claim 61, wherein the action includes performing a process to identify another individual responsible for the third facial skin micromovement.

66. 62. The non-transitory computer-readable medium of claim 61, wherein the first time period, the second time period, and the third time period are part of a single online activity associated with the ongoing electronic transaction.

67. 67. The non-transitory computer-readable medium of claim 66, wherein the online activity is at least one of a financial transaction, a betting session, an account access session, a gaming session, an exam, a lecture, or an educational session.

68. 67. The non-transitory computer-readable medium of claim 66, wherein the online activity includes a plurality of sessions, and the operations further include using received signals associated with facial skin micro-movements to determine that the particular individual will participate in each of the plurality of sessions.

69. 67. The non-transitory computer-readable medium of claim 66, wherein the action includes notifying an entity associated with the online activity that an individual other than the particular individual is currently participating in the online activity.

70. 77. The non-transitory computer-readable medium of claim 76, wherein the action includes preventing the particular individual from participating in the online activity until the individual's identity is verified.

71. 62. The non-transitory computer-readable medium of claim 61, wherein the first period of time, the second period of time, and the third period of time are part of a secure session involving access to a resource.

72. 72. The non-transitory computer-readable medium of claim 71, wherein the resource is at least one of a file, a folder, a database, a computer program, computer code, or a computer configuration.

73. 72. The non-transitory computer-readable medium of claim 71, wherein the action includes notifying an entity associated with the resource that an individual other than the particular individual has gained access to the resource.

74. 72. The non-transitory computer-readable medium of claim 71, wherein the action includes terminating the access to the resource.

75. 62. The non-transitory computer-readable medium of claim 61 , wherein the first period of time, the second period of time, and the third period of time are part of a single communication session, the communication session being at least one of a telephone call, a teleconference, a videoconference, or a real-time virtual communication.

76. 76. The non-transitory computer-readable medium of claim 75, wherein the action includes notifying an entity associated with the communication session that an individual other than the particular individual has joined the communication session.

77. 62. The non-transitory computer-readable medium of claim 61, wherein determining the identity of the particular individual comprises accessing a memory that correlates a plurality of reference facial skin micro-movements with individuals, and determining a match between the first facial skin micro-movement and at least one of the plurality of reference facial skin micro-movements.

78. 62. The non-transitory computer-readable medium of claim 61, wherein the operations further include determining the first facial skin micro-movement, the second facial skin micro-movement, and the third facial skin micro-movement by analyzing a received signal indicative of coherent light reflection to identify time and intensity changes in speckle.

79. 1. A method for continuous authentication based on facial skin micro-motion, comprising: receiving, during an ongoing electronic transaction, a first signal representative of a coherent light reflex associated with a first facial skin micro-movement during a first period of time; Using the first signal to determine the identity of a particular individual associated with the first facial skin micro-movement; receiving, during the ongoing electronic transaction, a second signal representative of a coherent light reflex associated with a second facial skin micro-movement, the second signal being received during a second time period subsequent to the first time period; using the second signal to determine that the particular individual is also associated with the second facial skin micro-movement; receiving, during the ongoing electronic transaction, a third signal representative of a coherent light reflex associated with a third facial skin micro-movement, the third signal being received during a third time period subsequent to the second time period; using the third signal to determine that the third facial skin micromovement is not associated with the particular individual; and Initiating an action based on the determination that the third facial skin micromovement is not associated with the particular individual; and A method comprising:

80. 1. A system for providing facial micro-movement based identity verification, comprising: receiving a first signal representative of a coherent light reflex associated with a first facial skin micro-movement during a first time period during an ongoing electronic transaction; Using the first signal, determine the identity of a particular individual associated with the first facial skin micro-movement; receiving, during the ongoing electronic transaction, a second signal representative of a coherent light reflex associated with a second facial skin micro-movement, the second signal being received during a second time period subsequent to the first time period; using the second signal to determine that the particular individual is also associated with the second facial skin micro-movement; receiving, during the ongoing electronic transaction, a third signal representative of a coherent light reflex associated with a third facial skin micro-movement, the third signal being received during a third time period subsequent to the second time period; Using the third signal, determine that the third facial skin micromovement is not associated with the particular individual; at least one processor configured to initiate an action based on the determination that the third facial skin micromovement is not associated with the particular individual; A system including:

81. 1. A non-transitory computer-readable medium comprising instructions that, when executed by at least one processor, cause the at least one processor to perform a thresholding operation for facial skin micro-movement interpretation, the operation comprising: Detecting facial micro-movements in the absence of perceptible vocalizations associated with the facial micro-movements; determining an intensity level of said facial micro-movements; comparing the determined intensity level to a threshold; interpreting the facial micro-movement when the intensity level is above the threshold; and and ignoring the facial micro-movement when the intensity level is below the threshold. Non-transitory computer-readable medium.

82. 82. The non-transitory computer-readable medium of claim 81, wherein the operations further comprise enabling adjustment of the threshold.

83. 82. The non-transitory computer-readable medium of claim 81, wherein the threshold is variable depending on environmental conditions.

84. 84. The non-transitory computer-readable medium of claim 83, wherein the environmental conditions include a background noise level.

85. 85. The non-transitory computer-readable medium of claim 84, wherein the operations further include receiving data indicative of the background noise level and determining a value for the threshold based on the received data.

86. 82. The non-transitory computer-readable medium of claim 81, wherein the threshold is variable depending on at least one physical activity involving an individual associated with the facial micro-movement.

87. 87. The non-transitory computer-readable medium of claim 86, wherein the at least one physical activity includes walking, running, or breathing.

88. 90. The non-transitory computer-readable medium of claim 87, wherein the operations further include receiving data indicative of the at least one physical activity involving the individual and determining a value for the threshold based on the received data.

89. 82. The non-transitory computer-readable medium of claim 81, wherein the threshold is customized for a user.

90. 90. The non-transitory computer-readable medium of claim 89, further comprising receiving a personalized threshold value for a particular individual and storing the personalized threshold value in a setting associated with the particular individual.

91. 90. The non-transitory computer-readable medium of claim 89, further comprising receiving a plurality of thresholds for a particular individual, each threshold being associated with a different situation.

92. 92. The non-transitory computer-readable medium of claim 91, wherein at least one of the different conditions comprises a physical condition of the particular individual, an emotional condition of the particular individual, or a location of the particular individual.

93. 93. The non-transitory computer-readable medium of claim 92, wherein the operations further include receiving data indicative of a current situation of the particular individual, and selecting one of the plurality of thresholds based on the received data.

94. 92. The non-transitory computer-readable medium of claim 91, wherein interpreting the facial micro-movements includes synthesizing speech associated with the facial micro-movements.

95. 82. The non-transitory computer-readable medium of claim 81, wherein interpreting the facial micro-movements comprises understanding and executing commands based on the facial micro-movements.

96. 96. The non-transitory computer-readable medium of claim 95, wherein executing the command comprises generating a signal that triggers an action.

97. 82. The non-transitory computer-readable medium of claim 81, wherein determining the intensity level comprises determining a value associated with a series of fine movements within a period of time.

98. 82. The non-transitory computer-readable medium of claim 81, wherein facial micro-movements having intensity levels below the threshold are interpretable but are nevertheless ignored.

99. 1. A method of thresholding for facial skin micromotion interpretation, comprising: Detecting facial micro-movements in the absence of perceptible vocalizations associated with the facial micro-movements; determining an intensity level of said facial micro-movements; comparing the determined intensity level to a threshold; interpreting the facial micro-movement when the intensity level is above the threshold; and and ignoring the facial micro-movement when the intensity level is below the threshold.

100. 1. A system for thresholding for facial skin micromotion interpretation, comprising: Detecting facial micro-movements in the absence of perceptible vocalizations associated with the facial micro-movements; determining an intensity level of the facial micro-movement; comparing the determined intensity level to a threshold; interpreting the facial micro-movement when the intensity level is above the threshold; at least one processor configured to ignore the facial micro-movement when the intensity level is below the threshold; A system including:

101. 1. A non-transitory computer-readable medium comprising instructions that, when executed by at least one processor, cause the at least one processor to perform operations for establishing an unvoiced conversation, the operations comprising: establishing a wireless communication channel for enabling non-vocalized conversation via a first wearable device and a second wearable device, each of which includes a coherent light source and a photodetector configured to detect facial skin micro-movements from coherent light reflections; Detecting, by the first wearable device, a first facial skin micro-movement occurring in the absence of perceptible vocalization; transmitting a first communication from the first wearable device to the second wearable device via the wireless communication channel, the first communication being derived from the first facial skin micromovements and transmitted for presentation via the second wearable device; receiving a second communication from the second wearable device via the wireless communication channel, the second communication derived from second facial skin micromovements detected by the second wearable device; presenting the second communication to a wearer of the first wearable device. Non-transitory computer-readable medium.

102. 102. The non-transitory computer-readable medium of claim 101, wherein the first communication includes a signal reflecting the first facial skin micromovement.

103. 102. The non-transitory computer-readable medium of claim 101, wherein the operation further comprises interpreting the first facial skin micromovement as a word, and the first communication comprises transmitting the word.

104. 102. The non-transitory computer-readable medium of claim 101, wherein presenting the second communication to the wearer of the first wearable device includes synthesizing words derived from the second facial skin micromovements.

105. 102. The non-transitory computer-readable medium of claim 101, wherein presenting the second communication to the wearer of the first wearable device includes providing a text output reflecting words derived from the second facial skin micromovements.

106. 102. The non-transitory computer-readable medium of claim 101, wherein presenting the second communication to the wearer of the first wearable device includes providing a graphical output reflecting at least one facial expression derived from the second facial skin micromovements.

107. 107. The non-transitory computer-readable medium of claim 106, wherein the graphical output includes at least one emoji.

108. 102. The non-transitory computer-readable medium of claim 101, wherein the operation further comprises determining that the second wearable device is located in proximity to the first wearable device.

109. 109. The non-transitory computer-readable medium of claim 108, wherein the operations further include automatically establishing the wireless communication channel between the first wearable device and the second wearable device.

110. 109. The non-transitory computer-readable medium of claim 108, wherein the operations further include presenting, via the first wearable device, a suggestion to establish a non-vocalized conversation with the second wearable device.

111. 102. The non-transitory computer-readable medium of claim 101, wherein the operations further include determining an intent of the wearer of the first wearable device to initiate a non-voiced conversation with the wearer of the second wearable device and automatically establishing the wireless communication channel between the first wearable device and the second wearable device.

112. 112. The non-transitory computer-readable medium of claim 111, wherein the intent is determined from the first facial skin micro-movement.

113. 102. The non-transitory computer-readable medium of claim 101, wherein the wireless communication channel is established directly between the first wearable device and the second wearable device.

114. 102. The non-transitory computer-readable medium of claim 101, wherein the wireless communication channel is established from the first wearable device to the second wearable device through at least one intermediate communication device.

115. 115. The non-transitory computer-readable medium of claim 114, wherein the at least one communication device includes at least one of a first smartphone associated with the wearer of the first wearable device, a second smartphone associated with the wearer of the second wearable device, a router, or a server.

116. 102. The non-transitory computer-readable medium of claim 101, wherein the first communication includes a signal reflecting a first word spoken in a first language, the second communication includes a signal reflecting a second word spoken in a second language, and presenting the second communication to the wearer of the first wearable device includes translating the second word into the first language.

117. 102. The non-transitory computer-readable medium of claim 101, wherein the first communication includes details identifying the wearer of the first wearable device and the second communication includes a signal identifying the wearer of the second wearable device.

118. 102. The non-transitory computer-readable medium of claim 101, wherein the first communication includes a timestamp indicating when the first facial skin micromovement was detected.

119. 1. A method for establishing a non-voiced conversation, comprising: establishing a wireless communication channel for enabling non-vocalized conversation via a first wearable device and a second wearable device, each of which includes a coherent light source and a photodetector configured to detect facial skin micro-movements from coherent light reflections; Detecting, by the first wearable device, a first facial skin micro-movement occurring in the absence of perceptible vocalization; transmitting a first communication from the first wearable device to the second wearable device over the wireless communication channel, the first communication being derived from the first facial skin micromovements and transmitted for presentation to a wearer of the second wearable device; receiving a second communication from the second wearable device via the wireless communication channel, the second communication derived from second facial skin micromovements detected by the second wearable device; presenting the second communication to a wearer of the first wearable device.

120. 1. A system for establishing a non-voiced conversation, comprising: establishing a wireless communication channel for enabling non-vocalized conversation via a first wearable device and a second wearable device, each of which includes a coherent light source and a photodetector configured to detect facial skin micro-movements from coherent light reflections; detecting, by the first wearable device, a first facial skin micro-movement occurring in the absence of perceptible vocalization; transmitting a first communication from the first wearable device to the second wearable device over the wireless communication channel, the first communication being derived from the first facial skin micromovements and transmitted for presentation to a wearer of the second wearable device; receiving a second communication from the second wearable device via the wireless communication channel, the second communication derived from second facial skin micromovements detected by the second wearable device; at least one processor configured to present the second communication to a wearer of the first wearable device; A system including:

121. 1. A non-transitory computer-readable medium comprising instructions that, when executed by at least one processor, cause the at least one processor to initiate a content interpretation operation prior to vocalization of content to be interpreted, the operation comprising: receiving signals representative of facial skin micro-movements; determining at least one word to be spoken from the signal prior to utterance of the at least one word in a source language; providing an interpretation of the at least one word prior to the utterance of the at least one word; and causing the interpretation of the at least one word to be presented when the at least one word is spoken. Non-transitory computer-readable medium.

122. 122. The non-transitory computer-readable medium of claim 121, wherein the interpretation is a translation of the at least one word from a source language to at least one target language other than the source language.

123. 123. The non-transitory computer-readable medium of claim 122, wherein the interpretation of the at least one word includes transcribing the at least one word into text in the at least one target language.

124. 123. The non-transitory computer-readable medium of claim 122, wherein the interpretation of the at least one word comprises speech synthesis of the at least one word in the at least one target language.

125. 123. The non-transitory computer-readable medium of claim 122, further comprising receiving a selection of the at least one destination language.

126. 126. The non-transitory computer-readable medium of claim 125, wherein the selecting the at least one target language comprises selecting multiple target languages, and causing the presentation of the interpretation of the at least one word comprises causing presentation in the multiple languages ​​simultaneously.

127. 122. The non-transitory computer-readable medium of claim 121, wherein the interpretation of the at least one word includes transcribing the at least one word into text in the source language.

128. 128. The non-transitory computer-readable medium of claim 127, wherein presenting the interpretation of the at least one word includes outputting a textual representation of the rewrite along with a video of an individual associated with the facial skin micro-movements.

129. 122. The non-transitory computer-readable medium of claim 121, wherein receiving a signal occurs via at least one detector of coherent light reflections from a facial region of a human uttering the at least one word.

130. 130. The non-transitory computer-readable medium of claim 129, wherein causing the presentation of the interpretation of the at least one word occurs contemporaneously with the human vocalizing the at least one word.

131. 122. The non-transitory computer-readable medium of claim 121, wherein causing the interpretation of the at least one word to be presented comprises using a wearable speaker to output an audible presentation of the at least one word.

132. 122. The non-transitory computer-readable medium of claim 121, wherein causing the interpretation of the at least one word to be presented includes transmitting an audio signal over a network.

133. 122. The non-transitory computer-readable medium of claim 121, further comprising: determining at least one expected word to be spoken following the at least one word being spoken; providing an interpretation of the at least one expected word prior to utterance of the at least one word; and causing presentation of the interpretation of the at least one expected word following presentation of the at least one word when the at least one word is spoken.

134. 122. The non-transitory computer-readable medium of claim 121, wherein causing the interpretation of the at least one word to be presented includes transmitting a text translation of the at least one word over a network.

135. 122. The non-transitory computer-readable medium of claim 121, wherein the operations further comprise determining at least one non-verbal interjection from the signal and outputting a representation of the non-verbal interjection.

136. 122. The non-transitory computer-readable medium of claim 121, wherein determining at least one word from the signal comprises interpreting the facial skin micro-movements using speckle analysis.

137. 122. The non-transitory computer-readable medium of claim 121, wherein the signal representing facial skin micro-movements corresponds to muscle activation before the utterance of the at least one word.

138. 138. The non-transitory computer-readable medium of claim 137, wherein the muscle activation is associated with at least one specific muscle including the zygomaticus, orbicularis oris, laughing muscle, genioglossus, or levator labii superioris alae naris.

139. 1. A method of initiating content interpretation prior to vocalization of the content to be interpreted, comprising: receiving signals representative of facial skin micro-movements; determining at least one word to be spoken from the signal prior to utterance of the at least one word in a source language; providing an interpretation of the at least one word prior to the utterance of the at least one word; causing the interpretation of the at least one word to be presented as the at least one word is spoken.

140. 1. A system for initiating content interpretation prior to vocalization of the content to be interpreted, comprising: Receive signals representing facial skin micro-movements, determining at least one word to be spoken from the signal prior to uttering the at least one word in a source language; providing an interpretation of the at least one word prior to the utterance of the at least one word; at least one processor configured to cause the interpretation of the at least one word to be presented when the at least one word is spoken; A system including:

141. A non-transitory computer-readable medium comprising instructions that, when executed by at least one processor, cause the at least one processor to perform a private voice assistant operation, The operation is receiving a signal indicative of a specific facial skin micro-movement reflecting a private request to an assistant, wherein responding to the private request requires identification of a specific individual associated with the specific facial skin micro-movement; accessing a data structure that maintains correlations between the particular individual and a plurality of facial skin micro-movements associated with the particular individual; searching said data structure for a match indicating a correlation between said particular individual's stored identity and said particular facial skin micro-movement; initiating a first action in response to the request in response to determining the existence of the match in the data structure, the first action including enabling access to information specific to the particular individual; and if the match is not identified in the data structure, initiating a second action different from the first action. Non-transitory computer-readable medium.

142. 142. The non-transitory computer-readable medium of claim 141, wherein the second action includes providing non-personal information.

143. 142. The non-transitory computer-readable medium of claim 141, wherein the second action includes notification that access to information specific to the particular individual has been denied.

144. 142. The non-transitory computer-readable medium of claim 141, wherein the second action includes blocking access to the information specific to the particular individual.

145. 142. The non-transitory computer-readable medium of claim 141, wherein the second action includes attempting to authenticate the particular individual using additional data.

146. 146. The non-transitory computer-readable medium of claim 145, wherein the additional data comprises additional detected facial skin micro-movements.

147. 146. The non-transitory computer-readable medium of claim 145, wherein the additional data includes data other than facial skin micro-motion.

148. 142. The non-transitory computer-readable medium of claim 141, wherein the operations further include initiating additional action to identify another individual other than the particular individual when the match is not identified.

149. 149. The non-transitory computer-readable medium of claim 148, wherein in response to identification of another individual other than the particular individual, the operations further include initiating a third action responsive to the request.

150. 150. The non-transitory computer-readable medium of claim 149, wherein the third action includes enabling access to information specific to the other individual.

151. 142. The non-transitory computer-readable medium of claim 141, wherein the private request is to activate software code, the first action activates the software code, and the second action prevents activation of the software code.

152. 142. The non-transitory computer-readable medium of claim 141, wherein the private request is for sensitive information, and the operation further comprises determining that the particular individual has permission to access the sensitive information.

153. 142. The non-transitory computer-readable medium of claim 141, wherein the receiving, accessing, and retrieving occur repeatedly during an ongoing session.

154. 154. The non-transitory computer-readable medium of claim 153, wherein during a first period during the ongoing session, the specific individual is identified and the first action is initiated, and during a second period during the ongoing session, the specific individual is not identified and any remaining first actions are terminated in favor of the second action.

155. 142. The non-transitory computer-readable medium of claim 141, wherein the operation further includes operating at least one coherent light source to enable illumination of a portion of the face of the individual making the private request other than the lips, and wherein receiving the signal occurs via at least one detector of coherent light reflection from the portion of the face other than the lips.

156. 156. The non-transitory computer-readable medium of claim 155, wherein the at least one processor, the at least one coherent light source, and the at least one detector are integrated within a wearable housing configured to be supported by an ear of the individual.

157. 156. The non-transitory computer-readable medium of claim 155, wherein the operations further include analyzing the received signal to determine preparatory phonation muscle recruitment and determining the private request based on the determined preparatory phonation muscle recruitment.

158. 156. The non-transitory computer-readable medium of claim 155, wherein the operations further include determining the private request when a perceptible vocalization of the private request is lacking.

159. 1. A method for operating a private voice assistant, comprising: receiving a signal indicative of a specific facial skin micro-movement reflecting a private request to an assistant, wherein responding to the private request requires identification of a specific individual associated with the specific facial skin micro-movement; accessing a data structure that maintains correlations between the particular individual and a plurality of facial skin micro-movements associated with the particular individual; searching said data structure for a match indicating a correlation between said particular individual's stored identity and said particular facial skin micro-movement; initiating a first action in response to the request in response to determining the existence of the match in the data structure, the first action including enabling access to information specific to the particular individual; and if the match is not identified in the data structure, initiating a second action different from the first action.

160. 1. A system for operating a private voice assistant, comprising: receiving a signal indicative of a specific facial micro-movement reflecting a private request to the assistant, and responding to the private request requiring identification of a specific individual associated with the specific facial micro-movement; accessing a data structure that maintains correlations between the particular individual and a plurality of facial skin micro-movements associated with the particular individual; searching said data structure for a match indicating a correlation between said particular individual's stored identity and said particular facial skin micro-movement; initiating a first action in response to the request in response to determining the existence of the match in the data structure, the first action enabling access to information specific to the particular individual; at least one processor configured to initiate a second action different from the first action if the match is not identified in the data structure; A system including:

161. 1. A non-transitory computer-readable medium comprising instructions that, when executed by at least one processor, cause the at least one processor to perform operations for determining subvocalized phonemes from facial skin micro-movements, the operations comprising: controlling at least one coherent light source to enable illumination of a first region of a face and a second region of the face; performing a first pattern analysis on light reflected from the first region of the face to determine a first micro-movement of facial skin in the first region of the face; performing a second pattern analysis on the light reflected from the second region of the face to determine second micro-movements of facial skin in the second region of the face; identifying at least one subvocalized phoneme using the first micro-movement of the facial skin in the first region of the face and the second micro-movement of the facial skin in the second region of the face; Non-transitory computer-readable medium.

162. 162. The non-transitory computer-readable medium of claim 161, wherein performing the second pattern analysis occurs after performing the first pattern analysis.

163. 162. The non-transitory computer-readable medium of claim 161, wherein performing the second pattern analysis occurs concurrently with performing the first pattern analysis.

164. 162. The non-transitory computer-readable medium of claim 161, wherein the first region is spaced apart from the second region.

165. 162. The non-transitory computer-readable medium of claim 161, wherein ascertaining the at least one subvocalized phoneme comprises ascertaining a sequence of phonemes, and the operation further comprises extracting meaning from the sequence of phonemes.

166. 166. The non-transitory computer-readable medium of claim 165, wherein each phoneme in the sequence of phonemes is derived from the first pattern analysis and the second pattern analysis.

167. 166. The non-transitory computer-readable medium of claim 165, wherein the operations further include identifying at least one phoneme in the sequence of phonemes as private and omitting to generate audio output reflecting the at least one private phoneme.

168. 162. The non-transitory computer-readable medium of claim 161, wherein the operations further comprise determining both the first fine movement and the second fine movement during a common period of time.

169. 162. The non-transitory computer-readable medium of claim 161, wherein the operations further include receiving the first light reflection and the second light reflection via at least one detector, wherein the at least one detector and the at least one coherent light source are integrated within a wearable housing.

170. 162. The non-transitory computer-readable medium of claim 161, wherein controlling the at least one coherent light source comprises projecting different light patterns onto the first region and the second region.

171. 171. The non-transitory computer-readable medium of claim 170, wherein the different light patterns include a plurality of light spots such that the first region of the face is illuminated by at least a first light spot and the second region of the face is illuminated by at least a second light spot different from the first light spot.

172. 162. The non-transitory computer-readable medium of claim 161, wherein controlling the at least one coherent light source comprises illuminating the first region and the second region with a common spot of light.

173. 162. The non-transitory computer-readable medium of claim 161, wherein the first micro-movement of the facial skin and the second micro-movement of the facial skin correspond to simultaneous muscle recruitment, the determined first micro-movement of the facial skin in the first region of the face corresponds to recruitment of a first muscle selected from the zygomaticus, orbicularis oris, laughing muscle, and levator labii superioris alae naris, and the determined second micro-movement of the facial skin in the second region of the face corresponds to recruitment of a second muscle, different from the first muscle, selected from the zygomaticus, orbicularis oris, laughing muscle, and levator labii superioris alae naris.

174. 162. The non-transitory computer-readable medium of claim 161, wherein the operations further include accessing a default language of an individual associated with the facial skin micro-movement and extracting meaning from the at least one subvocalized phoneme using the default language.

175. 162. The non-transitory computer-readable medium of claim 161, wherein the operations further comprise generating an audio output reflective of the at least one sub-vocalized phoneme using a synthesized voice.

176. 162. The non-transitory computer-readable medium of claim 161, wherein the at least one phoneme comprises a sequence of phonemes, and the operations further comprise determining a prosody associated with the sequence of phonemes and extracting meaning based on the determined prosody.

177. 162. The non-transitory computer-readable medium of claim 161, wherein the operations further include determining an emotional state of an individual associated with the facial skin micro-movements and extracting meaning from the at least one subvocalized phoneme and the determined emotional state.

178. 162. The non-transitory computer-readable medium of claim 161, wherein the operations further include identifying at least one irrelevant phoneme as part of the filler and omitting to generate audio output reflecting the irrelevant phoneme.

179. 1. A method for determining subvocalized phonemes from facial skin micro-movements, comprising: controlling at least one coherent light source to enable illumination of a first region of a face and a second region of the face; performing a first pattern analysis on light reflected from the first region of the face to determine a first micro-movement of facial skin in the first region of the face; performing a second pattern analysis on the light reflected from the second region of the face to determine second micro-movements of facial skin in the second region of the face; and identifying at least one subvocalized phoneme using the first micro-movement of the facial skin in the first region of the face and the second micro-movement of the facial skin in the second region of the face.

180. 1. A system for determining subvocalized phonemes from facial micro-movements, comprising: controlling at least one coherent light source to enable illumination of a first region of a face and a second region of the face; performing a first pattern analysis on light reflected from the first region of the face to determine a first micro-movement of facial skin in the first region of the face; performing a second pattern analysis on the light reflected from the second region of the face to determine second micro-movements of facial skin in the second region of the face; at least one processor configured to identify at least one subvocalized phoneme using the first micro-movement of the facial skin in the first region of the face and the second micro-movement of the facial skin in the second region of the face; A system including:

181. 1. A non-transitory computer-readable medium comprising instructions that, when executed by at least one processor, cause the at least one processor to perform operations for generating a synthetic representation of a facial expression, the operations comprising: controlling at least one coherent light source to enable illumination of a portion of a face; receiving an output signal from the photodetector corresponding to a reflection of coherent light from the portion of the face; applying speckle analysis to the output signal to determine speckle analysis-based facial skin micromotion; using the determined speckle analysis-based facial skin micro-movements to identify at least one word that has been pre-uttered or vocalized during a period of time; using the determined speckle analysis-based facial skin micro-motions to identify at least one change in facial expression during the period of time; outputting data for causing the virtual representation of the face to mimic the at least one change in facial expression in coordination with the audio presentation of the at least one word during the period of time. Non-transitory computer-readable medium.

182. 182. The non-transitory computer-readable medium of claim 181, wherein controlling the at least one coherent light source to enable illumination of the portion of the face comprises projecting a light pattern onto the portion of the face.

183. 183. The non-transitory computer-readable medium of claim 182, wherein the light pattern comprises a plurality of spots.

184. 183. The non-transitory computer-readable medium of claim 182, wherein the portion of the face includes cheek skin.

185. 183. The non-transitory computer-readable medium of claim 182, wherein the portion of the face excludes lips.

186. 182. The non-transitory computer-readable medium of claim 181, wherein the output signal from the photodetector emanate from a wearable device.

187. 182. The non-transitory computer-readable medium of claim 181, wherein the output signal from the photodetector emanate from a non-wearable device.

188. 182. The non-transitory computer-readable medium of claim 181, wherein the determined speckle analysis-based facial skin micromotion is associated with recruitment of at least one of the zygomaticus, orbicularis oris, genioglossus, laugher, or levator labii superioris alae naris muscles.

189. 182. The non-transitory computer-readable medium of claim 181, wherein the at least one change in facial expression over a period of time includes a speech-related facial expression and a non-speech-related facial expression.

190. 190. The non-transitory computer-readable medium of claim 189, wherein the virtual representation of the face is associated with an avatar of the individual from whom the output signal is derived, and wherein mimicking the at least one change in facial expression comprises causing a visual change in the avatar that reflects at least one of the speech-related facial expression and the non-speech-related facial expression.

191. 191. The non-transitory computer-readable medium of claim 190, wherein the visual change to the avatar comprises changing a color of at least a portion of the avatar.

192. 182. The non-transitory computer-readable medium of claim 181, wherein the audio presentation of the at least one word is based on a personal recording.

193. 182. The non-transitory computer-readable medium of claim 181, wherein the audio presentation of the at least one word is based on synthetic speech.

194. 200. The non-transitory computer-readable medium of claim 193, wherein the synthesized voice corresponds to a voice of an individual from whom the output signal is derived.

195. 200. The non-transitory computer-readable medium of claim 193, wherein the synthesized voice corresponds to a template voice selected by the individual from whom the output signal is derived.

196. 182. The non-transitory computer-readable medium of claim 181, wherein the operations further include determining an emotional state of the individual from whom the output signal is derived based at least in part on the facial skin micro-movements, and augmenting the virtual representation of the face to reflect the determined emotional state.

197. 182. The non-transitory computer-readable medium of claim 181, wherein the operations further include receiving a selection of a desired emotional state and augmenting the virtual representation of the face to reflect the selected emotional state.

198. 182. The non-transitory computer-readable medium of claim 181, wherein the operations further include identifying an undesirable facial expression, and wherein the output data that causes the virtual representation omits data that causes the undesirable facial expression.

199. 1. A method for generating a synthetic representation of a facial expression, comprising: controlling at least one coherent light source to enable illumination of a portion of a face; receiving an output signal from the photodetector corresponding to a reflection of coherent light from the portion of the face; applying speckle analysis to the output signal to determine speckle analysis-based facial skin micromotion; using the determined speckle analysis-based facial skin micro-movements to identify at least one word that has been pre-uttered or vocalized during a period of time; using the determined speckle analysis-based facial skin micro-motions to identify at least one change in facial expression during the period of time; and outputting data during the period of time in coordination with the audio presentation of the at least one word to cause the virtual representation of the face to mimic the at least one change in facial expression.

200. 1. A system for generating a synthetic representation of a facial expression, comprising: controlling at least one coherent light source to enable illumination of a portion of the face; receiving an output signal from the photodetector corresponding to a reflection of coherent light from the portion of the face; applying speckle analysis to the output signal to determine speckle analysis-based facial skin micromotion; using the determined speckle analysis-based facial skin micro-movements to identify at least one word that was pre-uttered or vocalized during a period of time; using the determined speckle analysis-based facial skin micro-motions to identify at least one change in facial expression during the period of time; at least one processor configured to output data for causing the virtual representation of the face to mimic the at least one change in facial expression in coordination with the audio presentation of the at least one word during the period of time; A system including:

201. 1. A non-transitory computer-readable medium comprising instructions that, when executed by at least one processor, cause the at least one processor to perform operations for attention-related interaction based on facial skin micro-movements, the operations comprising: determining facial skin micromotion of the individual based on reflection of coherent light from a facial region of the individual; using the facial skin micromotor movements to determine a particular engagement level of the individual; receiving data associated with an anticipated interaction with the individual; accessing a data structure that correlates information reflecting alternative engagement levels with different presentation methods; determining a particular presentation for the anticipated interaction based on the particular engagement level and the correlation information; and associating the particular presentation with the expected interaction for subsequent engagement with the individual. Non-transitory computer-readable medium.

202. 202. The non-transitory computer-readable medium of claim 201, wherein the operations further include generating an output reflecting the expected interaction according to the determined particular presentation method.

203. 202. The non-transitory computer-readable medium of claim 201, wherein the operations further include operating at least one coherent light source to enable illumination of portions of the individual's face other than the lips, and receiving a signal indicative of reflection of the coherent light from portions of the face other than the lips.

204. 204. The non-transitory computer-readable medium of claim 203, wherein the operations further include performing speckle analysis on the coherent light reflections from portions of the face other than the lips to determine the facial skin micromotion.

205. 202. The non-transitory computer-readable medium of claim 201, wherein the particular level of engagement is a category of engagement.

206. 202. The non-transitory computer-readable medium of claim 201, wherein the particular level of engagement comprises a magnitude of engagement.

207. 202. The non-transitory computer-readable medium of claim 201, wherein the particular engagement level reflects the degree to which the individual is involved in an activity including at least one of talking, thinking, or resting.

208. 208. The non-transitory computer-readable medium of claim 207, wherein the operations further include determining the degree to which the individual is engaged in the activity based on facial skin micro-movements corresponding to recruitment of at least one muscle from a muscle group including the zygomaticus, orbicularis oris, laughing muscle, or levator labii superioris alae nosi.

209. 202. The non-transitory computer-readable medium of claim 201, wherein the received data associated with the anticipated interaction includes an incoming call, and the associated different presentation methods include notifying the individual of the incoming call and sending the incoming call to voicemail.

210. 202. The non-transitory computer-readable medium of claim 201, wherein the received data associated with the anticipated interaction includes an incoming text message, and the associated different presentation methods include presenting the text message to the individual in real time and deferring presentation of the text message until a later time.

211. 202. The non-transitory computer-readable medium of claim 201, wherein determining the particular presentation method for the anticipated interaction includes determining how to notify the individual of the anticipated interaction.

212. 212. The non-transitory computer-readable medium of claim 211, wherein determining how to notify the individual of the expected interaction is based at least in part on an identification of multiple electronic devices currently being used by the individual.

213. 202. The non-transitory computer-readable medium of claim 201, wherein the received data associated with the expected interaction indicates an importance of the expected interaction, and the particular presentation method is determined based at least in part on the importance.

214. 202. The non-transitory computer-readable medium of claim 201, wherein the received data associated with the anticipated interaction indicates an urgency level of the anticipated interaction, and the particular presentation manner is determined based at least in part on the particular urgency level.

215. 202. The non-transitory computer-readable medium of claim 201, wherein the particular presentation method includes withholding presentation of content until a detected period of low engagement, and the operations further include detecting low engagement at a subsequent time and presenting the content at the subsequent time.

216. 202. The non-transitory computer-readable medium of claim 201, wherein the operations further include using the facial skin micro-movements to determine that the individual is engaged in a conversation with another individual and determining whether the expected interaction is relevant to the conversation, and wherein the particular presentation method is determined based at least in part on the relevance of the expected interaction to the conversation.

217. 217. The non-transitory computer-readable medium of claim 216, wherein the operations further include determining a topic of the conversation using the facial skin micro-movements, and determining that the expected interaction is relevant to the conversation is based on the received data associated with the expected interaction and the topic of the conversation.

218. 217. The non-transitory computer-readable medium of claim 216, wherein a first presentation method is used for the expected interaction when the expected interaction is determined to be relevant to the conversation, and a second presentation method is used for the expected interaction when the expected interaction is determined to be not relevant to the conversation.

219. 1. A method for facial skin micro-movement based attention-related interaction, comprising: determining facial skin micromotion of the individual based on reflection of coherent light from a facial region of the individual; using the facial skin micromotor movements to determine a particular engagement level of the individual; receiving data associated with an anticipated interaction with the individual; accessing a data structure that correlates information reflecting alternative engagement levels with different presentation methods; determining a particular presentation for the anticipated interaction based on the particular engagement level and the correlation information; and associating the particular presentation method with the expected interaction for subsequent engagement with the individual.

220. 1. A system for facial skin micro-movement based attention-related interaction, comprising: determining facial skin micromotion of the individual based on reflection of coherent light from a facial region of the individual; using said facial skin micro-motor movements to determine a particular engagement level of said individual; receiving data associated with an anticipated interaction with the individual; accessing a data structure that correlates information reflecting alternative engagement levels with different presentation methods; determining a particular presentation for the anticipated interaction based on the particular engagement level and the correlation information; at least one processor configured to associate the particular presentation with the expected interaction for subsequent engagement with the individual; A system including:

221. 1. A non-transitory computer-readable medium comprising instructions that, when executed by at least one processor, cause the at least one processor to perform a speech synthesis operation from detected facial skin micro-movements, the operation comprising: determining specific facial skin micro-movements of a first individual speaking with a second individual based on light reflection from a facial region of the first individual; accessing a data structure correlating facial micro-movements with words; performing a lookup in the data structure for a particular word associated with the particular facial skin micro-movement; obtaining input related to preferred speech consumption characteristics of the second individual; adopting the preferred speech consumption characteristics; synthesizing an audible output of the particular word using the adopted preferred speech consumption characteristics. Non-transitory computer-readable medium.

222. 222. The non-transitory computer-readable medium of claim 221, further comprising presenting a user interface to at least one of the first individual and the second individual for changing the preferred speech consumption characteristics.

223. 222. The non-transitory computer-readable medium of claim 221, wherein obtaining the input associated with the preferred speech consumption characteristic of the second individual comprises receiving the input from the first individual.

224. 222. The non-transitory computer-readable medium of claim 221, wherein obtaining the input associated with the preferred speech consumption characteristic of the second individual comprises receiving the input from the second individual.

225. 222. The non-transitory computer-readable medium of claim 221, wherein obtaining the input associated with the preferred speech consumption characteristic of the second individual includes searching for information about the second individual.

226. 226. The non-transitory computer-readable medium of claim 225, wherein obtaining the input associated with the preferred speech consumption characteristics of the second individual includes determining the information based on image data captured by an image sensor worn by the first individual.

227. 222. The non-transitory computer-readable medium of claim 221, wherein the input associated with the preferred speech consumption characteristic of the second individual indicates an age of the second individual.

228. 222. The non-transitory computer-readable medium of claim 221, wherein the input associated with the preferred speech consumption characteristic of the second individual is indicative of an environmental condition associated with the second individual.

229. 222. The non-transitory computer-readable medium of claim 221, wherein the input associated with the preferred speech consumption characteristic of the second individual is indicative of a hearing impairment of the second individual.

230. 222. The non-transitory computer-readable medium of claim 221, wherein the second individual is one of a plurality of individuals, and the operations further include obtaining additional input from the plurality of individuals and classifying the plurality of individuals based on the additional input.

231. 222. The non-transitory computer-readable medium of claim 221, wherein adopting the preferred speech consumption characteristics includes presetting predicted facial micro-movement speech synthesis controls.

232. 222. The non-transitory computer-readable medium of claim 221, wherein the input associated with the preferred speech consumption characteristic includes a preferred pace of speech, and wherein the synthesized audible output of the particular word occurs at the preferred pace of speech.

233. 222. The non-transitory computer-readable medium of claim 221, wherein the input associated with the preferred speech consumption characteristic includes a speech volume, and the synthesized audible output of the particular word occurs at the preferred speech volume.

234. 222. The non-transitory computer-readable medium of claim 221, wherein the input associated with the preferred speech consumption characteristic includes a target language of speech other than a language associated with the particular facial skin micro-movement, and the synthesized audible output of the particular word occurs in the target language of the utterance.

235. 222. The non-transitory computer-readable medium of claim 221, wherein the input associated with the preferred speech consumption characteristic includes a preferred voice, and the synthesized audible output of the particular word occurs in the preferred voice.

236. 236. The non-transitory computer-readable medium of claim 235, wherein the preferred voice is at least one of a celebrity voice, an accented voice, or a gender-based voice.

237. 222. The non-transitory computer-readable medium of claim 221, wherein the operations further include presenting a first synthesized version of the intended speech based on the facial micro-movements and presenting a second synthesized version of the speech based on the facial micro-movements in combination with the preferred speech consumption characteristics.

238. 238. The non-transitory computer-readable medium of claim 237, wherein presenting the first composite version and the second composite version to the first individual occurs sequentially.

239. 1. A method for performing speech synthesis from detected facial micro-movements, comprising: determining specific facial skin micro-movements of a first individual speaking with a second individual based on light reflection from a facial region of the first individual; accessing a data structure correlating facial micro-movements with words; performing a lookup in the data structure for a particular word associated with the particular facial skin micro-movement; obtaining input related to preferred speech consumption characteristics of the second individual; adopting the preferred speech consumption characteristics; and synthesizing an audible output of the particular word using the adopted preferred speech consumption characteristics.

240. 1. A system for performing speech synthesis from detected facial micro-movements, comprising: determining specific facial skin micro-movements of a first individual speaking with a second individual based on light reflections from a facial region of the first individual; accessing a data structure that correlates facial micro-movements with words; performing a lookup in the data structure for a particular word associated with the particular facial skin micro-movement; obtaining input related to preferred speech consumption characteristics of the second individual; adopting said preferred speech consumption characteristics; at least one processor configured to synthesize an audible output of the particular word using the adopted preferred speech consumption characteristics; A system including:

241. 1. A non-transitory computer-readable medium comprising instructions that, when executed by at least one processor, cause the at least one processor to perform operations for personalized presentation of a preliminary utterance, the operations comprising: receiving a reflection signal corresponding to light reflected from a facial region of the individual; using the received reflected signals to determine specific facial skin micro-movements of an individual in the absence of perceptible vocalizations associated with the specific facial skin micro-movements; accessing a data structure correlating facial skin micro-movements with words; performing a lookup in said data structure of a particular unuttered word associated with said particular facial skin micro-movement; and audibly presenting the particular unspoken word to the individual prior to the individual's utterance of the particular word. Non-transitory computer-readable medium.

242. 242. The non-transitory computer-readable medium of claim 241, wherein the operations further include recording data associated with the particular unspoken word for future use.

243. 243. The non-transitory computer-readable medium of claim 242, wherein the data includes at least one of the audible representation of the particular unspoken word or a textual representation of the particular unspoken word.

244. 242. The non-transitory computer-readable medium of claim 241, wherein the light reflected from the facial region of the individual comprises coherent light reflection.

245. 244. The non-transitory computer-readable medium of claim 243, wherein the operations further include adding punctuation to the textual representation.

246. 242. The non-transitory computer-readable medium of claim 241, wherein the operations further comprise adjusting a rate of the audible presentation of the particular unspoken word based on input from the individual.

247. 242. The non-transitory computer-readable medium of claim 241, wherein the operations further include adjusting a volume of the audible presentation of the particular unspoken word based on input from the individual.

248. 242. The non-transitory computer-readable medium of claim 241, wherein the causing an audible presentation comprises outputting an audio signal to a personal hearing device configured to be worn by the individual.

249. 249. The non-transitory computer-readable medium of claim 248, wherein the operation further includes operating at least one coherent light source to enable illumination of the facial region of the individual, the at least one coherent light source being integrated with the personal hearing device.

250. 242. The non-transitory computer-readable medium of claim 241, wherein the audible presentation of the particular unspoken word is a synthesis of a selected voice.

251. 251. The non-transitory computer-readable medium of claim 250, wherein the selected voice is a synthesis of the individual's voice.

252. 251. The non-transitory computer-readable medium of claim 250, wherein the selected voice is a synthesis of a voice of another individual other than the individual associated with the facial skin micromovement.

253. 242. The non-transitory computer-readable medium of claim 241, wherein the particular unspoken word corresponds to a speakable word in a first language, and the audible presentation includes a synthesis of the speakable word in a second language different from the first language.

254. 254. The non-transitory computer-readable medium of claim 253, wherein the operations further include associating the particular facial skin micromovement with a plurality of utterable words in the second language and selecting a most suitable utterable word from the plurality of utterable words, and wherein the audible presentation includes the most suitable utterable word in the second language.

255. 242. The non-transitory computer-readable medium of claim 241, wherein the operation further comprises determining that an intensity of a portion of the particular facial skin micro-movement is below a threshold and providing associated feedback to the individual.

256. 242. The non-transitory computer-readable medium of claim 241, wherein the audible presentation of the particular unuttered word is provided to the individual at least 20 milliseconds before the individual utters the particular word.

257. 242. The non-transitory computer-readable medium of claim 241, wherein the operations further include ceasing the audible presentation of the particular unspoken word in response to a detected trigger.

258. 258. The non-transitory computer-readable medium of claim 257, wherein the operations further include detecting the trigger from determined facial skin micro-movements of the individual.

259. 1. A method for personalized presentation of preliminary utterances, comprising: receiving a reflection signal corresponding to light reflected from a facial region of the individual; using the received reflected signals to determine specific facial skin micro-movements of an individual in the absence of perceptible vocalizations associated with the specific facial skin micro-movements; accessing a data structure correlating facial skin micro-movements with words; performing a lookup in said data structure of a particular unuttered word associated with said particular facial skin micro-movement; and audibly presenting the particular unspoken word to the individual prior to utterance of the particular word by the individual.

260. 1. A system for personal presentation of preliminary utterances, comprising: receiving a reflection signal corresponding to light reflected from a facial region of the individual; using the received reflected signals to determine specific facial skin micro-movements of an individual in the absence of perceptible vocalizations associated with the specific facial skin micro-movements; accessing a data structure that correlates facial micromotor movements with words; performing a lookup in said data structure of a particular unuttered word associated with said particular facial skin micro-movement; at least one processor configured to cause the particular unspoken word to be audibly presented to the individual prior to utterance of the particular word by the individual; A system including:

261. 1. A non-transitory computer-readable medium comprising instructions that, when executed by at least one processor, cause the at least one processor to perform operations for determining facial skin micro-movements, the operations comprising: controlling at least one coherent light source to project a plurality of light spots onto a facial area of ​​the individual, the plurality of light spots including at least a first light spot and a second light spot spaced apart from the first light spot; analyzing reflected light from the first light spot to determine a change in first spot reflection; analyzing reflected light from the second light spot to determine a change in second spot reflectance; determining the facial skin micro-movements based on the determined changes in the first spot reflex and the second spot reflex; interpreting the facial skin micro-movements derived from analyzing the first spot reflection and analyzing the second spot reflection; generating an output of the interpretation. Non-transitory computer-readable medium.

262. 262. The non-transitory computer-readable medium of claim 261, wherein the plurality of light spots further includes a third light spot and a fourth light spot, each of the third light spot and the fourth light spot being spaced apart from each other and from the first light spot and the second light spot.

263. 263. The non-transitory computer-readable medium of claim 262, wherein the facial skin micromovements are determined based on the determined changes in the first spot reflex and the second spot reflex, and changes in the third spot reflex and the fourth spot reflex.

264. 262. The non-transitory computer-readable medium of claim 261, wherein the plurality of light spots comprises at least 16 spaced apart light spots.

265. 262. The non-transitory computer-readable medium of claim 261, wherein the plurality of light spots are projected onto an area other than the lips of the individual.

266. 262. The non-transitory computer-readable medium of claim 261, wherein the change in the first spot reflex and the change in the second spot reflex correspond to simultaneous muscle recruitment.

267. 267. The non-transitory computer-readable medium of claim 266, wherein the first spot reflex and the second spot reflex each correspond to recruitment of a single muscle selected from the zygomaticus, orbicularis oris, genioglossus, laughing muscle, or levator labii superioris alae naris.

268. 267. The non-transitory computer-readable medium of claim 266, wherein the first spot reflex corresponds to recruitment of a muscle selected from the zygomaticus, orbicularis oris, laughing muscle, genioglossus, or levator labii superioris nasalis, and the second spot reflex corresponds to recruitment of another muscle selected from the zygomaticus, orbicularis oris, laughing muscle, genioglossus, or the levator labii superioris nasalis.

269. 262. The non-transitory computer-readable medium of claim 261, wherein the at least one coherent light source is associated with a detector, and the at least one coherent light source and the detector are integrated within a wearable housing.

270. 262. The non-transitory computer-readable medium of claim 261, wherein determining the facial skin micro-movements comprises analyzing the change in the first spot reflection relative to the change in the second spot reflection.

271. 262. The non-transitory computer-readable medium of claim 261, wherein the determined facial skin micro-motions within the facial region include micro-motions of less than 100 microns.

272. 262. The non-transitory computer-readable medium of claim 261, wherein the interpretation includes an emotional state of the individual.

273. 262. The non-transitory computer-readable medium of claim 261, wherein the interpretation includes at least one of the individual's heart rate or respiratory rate.

274. 262. The non-transitory computer-readable medium of claim 261, wherein the interpretation includes an identification of the individual.

275. 262. The non-transitory computer-readable medium of claim 261, wherein the interpretation comprises a word.

276. 276. The non-transitory computer-readable medium of claim 275, wherein the output comprises a textual representation of the word.

277. 276. The non-transitory computer-readable medium of claim 275, wherein the output comprises an audible presentation of the word.

278. 276. The non-transitory computer-readable medium of claim 275, wherein the output includes metadata indicative of facial expressions or prosody associated with words.

279. 1. A method for determining facial skin micro-movements, comprising: controlling at least one coherent light source to project a plurality of light spots onto a facial area of ​​the individual, the plurality of light spots including at least a first light spot and a second light spot spaced apart from the first light spot; analyzing reflected light from the first light spot to determine a change in first spot reflection; analyzing reflected light from the second light spot to determine a change in second spot reflectance; determining the facial skin micro-movements based on the determined changes in the first spot reflex and the second spot reflex; interpreting the facial skin micro-movements derived from analyzing the first spot reflection and analyzing the second spot reflection; generating an output of the interpretation. method.

280. A system for determining facial skin micro-movements, comprising: controlling at least one coherent light source to project a plurality of light spots onto a facial area of ​​the individual, the plurality of light spots including at least a first light spot and a second light spot spaced apart from the first light spot; analyzing the reflected light from the first light spot to determine a change in first spot reflection; analyzing the reflected light from the second light spot to determine a change in second spot reflectance; determining the facial skin micro-movements based on the determined changes in the first spot reflex and the second spot reflex; interpreting the facial skin micro-movements derived from analyzing the first spot reflection and analyzing the second spot reflection; at least one processor configured to generate an output of the interpretation; A system including:

281. 1. A non-transitory computer-readable medium comprising instructions that, when executed by at least one processor, cause the at least one processor to perform operations for interpreting impaired speech based on facial movements, the instructions comprising: The operation is receiving signals associated with specific facial skin movements of an individual with a speech disorder that affects how the individual pronounces words; accessing a data structure containing correlations between the plurality of words and a plurality of facial skin movements corresponding to the manner in which the individual pronounces the plurality of words; identifying a particular word associated with the particular facial skin movement based on the received signals and the correlation; and generating an output of the particular word for presentation, the output being different from the pronunciation of the particular word by the individual. Non-transitory computer-readable medium.

282. 282. The non-transitory computer-readable medium of claim 281, wherein the facial skin movement is facial skin micro-movement.

283. 283. The non-transitory computer-readable medium of claim 282, wherein the signal is received from a sensor that detects light reflection from a portion of the individual's face other than the lips.

284. 284. The non-transitory computer-readable medium of claim 283, wherein the facial skin micro-movement corresponds to recruitment of at least one muscle from a muscle group including the zygomaticus, genioglossus, orbicularis oris, lolis, or levator labii superioris alae naris.

285. 282. The non-transitory computer-readable medium of claim 281, wherein the signal is received from an image sensor configured to measure incoherent light reflection.

286. 282. The non-transitory computer-readable medium of claim 281, wherein the data structure is personalized to the individual's unique facial skin movements.

287. 282. The non-transitory computer-readable medium of claim 281, wherein the operations further include employing a training model to populate the data structure.

288. 282. The non-transitory computer-readable medium of claim 281, wherein the particular facial skin movement is associated with the utterance of the particular word, and the utterance of the particular word is in a non-canonical manner.

289. 282. The non-transitory computer-readable medium of claim 281, wherein the output of the particular word is audible and is used to correct the speech disorder of the individual.

290. 290. The non-transitory computer-readable medium of claim 289, wherein the speech disorder is stuttering and the modification includes outputting the particular word spoken in a non-stuttering manner.

291. 290. The non-transitory computer-readable medium of claim 289, wherein the speech disorder is hoarseness and the modification includes outputting the particular word in a non-hoarse form.

292. 290. The non-transitory computer-readable medium of claim 289, wherein the speech impairment is low volume and the modification includes outputting the particular word at a louder volume than the particular word was spoken.

293. 282. The non-transitory computer-readable medium of claim 281, wherein the output of the particular word is text.

294. 300. The non-transitory computer-readable medium of claim 293, wherein the operations further include adding punctuation to the text output of the particular word.

295. 282. The non-transitory computer-readable medium of claim 281, wherein the data structure includes data associated with at least one record of the individual previously pronouncing the particular word.

296. 282. The non-transitory computer-readable medium of claim 281, wherein the identified specific words associated with the specific facial skin movements are muted.

297. 282. The non-transitory computer-readable medium of claim 281, wherein the particular facial skin movement is associated with a subvocalization of the particular word, and the generated output comprises a private audible presentation of the subvocalized word to the individual.

298. 282. The non-transitory computer-readable medium of claim 281, wherein the particular facial skin movement is associated with a subvocalization of the particular word, and the generated output comprises a non-private audible presentation of the subvocalized word.

299. 1. A method for interpreting impaired speech based on facial movements, comprising: receiving signals associated with specific facial skin movements of an individual with a speech disorder that affects how the individual pronounces words; accessing a data structure containing correlations between the plurality of words and a plurality of facial skin movements corresponding to the manner in which the individual pronounces the plurality of words; identifying a particular word associated with the particular facial skin movement based on the received signals and the correlation; and generating for presentation an output of the particular word, the output being different from a pronunciation of the particular word by the individual.

300. 1. A system for interpreting impaired speech based on facial movements, comprising: receiving signals associated with specific facial skin movements of an individual having a speech disorder that affects how the individual pronounces words; accessing a data structure containing correlations between the plurality of words and a plurality of facial skin movements corresponding to the manner in which the individual pronounces the plurality of words; identifying a particular word associated with the particular facial skin movement based on the received signals and the correlation; at least one processor configured to generate an output of the particular word for presentation, the output being different from a pronunciation of the particular word by the individual; A system including:

301. 1. A non-transitory computer-readable medium comprising instructions that, when executed by at least one processor, cause the at least one processor to perform operations for ongoing verification of authenticity of communications based on light reflection from facial skin, the operations comprising: generating a first data stream representing a communication by a subject, the communication having a duration; generating a second data stream from facial skin light reflectance captured during the duration of the communication to verify the subject's identity; transmitting the first data stream to a destination; transmitting the second data stream to the destination; the second data stream, when received at the destination, is correlated with the first data stream such that the second data stream can be used to repeatedly check that the communication originated from the subject during the duration of the communication. Non-transitory computer-readable medium.

302. 302. The non-transitory computer-readable medium of claim 301, wherein checking that the communication originated from the subject comprises verifying that every word in the communication originated from the subject.

303. 302. The non-transitory computer-readable medium of claim 301, wherein checking that the communication originated from the subject comprises verifying that utterances captured at regular time intervals originated from the subject at the regular time intervals during the duration of the conversation.

304. 302. The non-transitory computer-readable medium of claim 301, wherein the first data stream and the second data stream are intermixed within a common omnibus data stream.

305. 302. The non-transitory computer-readable medium of claim 301, wherein the destination is a social networking service, and the second data stream enables the social networking service to publish the communication with a trustworthiness index value.

306. 302. The non-transitory computer-readable medium of claim 301, wherein the destination is an entity involved in a real-time transaction with the subject, and the second data stream enables the entity to verify the identity of the subject in real time during the duration of the communication.

307. 307. The non-transitory computer-readable medium of claim 306, wherein verifying the identity includes verifying the subject's name.

308. 307. The non-transitory computer-readable medium of claim 306, wherein verifying the identity includes verifying that the subject spoke the words presented in the communication at least at periodic intervals throughout the communication.

309. 302. The non-transitory computer-readable medium of claim 301, wherein the operations further include determining a biometric signature of the subject from light reflexes associated with facial skin captured prior to the communication, and wherein the identity of the subject is determined using the verifying facial skin light reflexes and the biometric signature.

310. 310. The non-transitory computer-readable medium of claim 309, wherein the biometric signature is determined based on a microvein pattern of the facial skin.

311. 310. The non-transitory computer-readable medium of claim 309, wherein the biometric signature is determined based on a sequence of facial skin micro-movements associated with phonemes spoken by the subject.

312. 302. The non-transitory computer-readable medium of claim 301, wherein the second data stream indicates a liveliness state of the subject, and transmitting the second data stream enables verification of authenticity of the communication based on the liveliness state of the subject.

313. 302. The non-transitory computer-readable medium of claim 301, wherein the first data stream is indicative of an expression of the subject, and the second data stream allows for verification of the expression.

314. 302. The non-transitory computer-readable medium of claim 301, wherein the operations further include storing in a data structure identifying the subject's facial skin micro-movements of an utterance or preparatory utterance of a passphrase, and identifying the subject based on the utterance or preparatory utterance of the passphrase.

315. 302. The non-transitory computer-readable medium of claim 301, wherein the operations further include storing a profile of the subject based on a pattern of facial skin micro-movements in a data structure, and identifying the subject based on the pattern.

316. 302. The non-transitory computer-readable medium of claim 301, wherein the first data stream is based on signals associated with sounds captured by a microphone during the duration of the communication.

317. 302. The non-transitory computer-readable medium of claim 301, wherein the first data stream and the second data stream are determined based on signals from a same photodetector.

318. 318. The non-transitory computer-readable medium of claim 317, wherein generating the first data stream representing the communication by the subject includes reproducing speech based on the demonstrating facial skin light reflexes.

319. 1. A method for ongoing verification of reliability of communications based on light reflection from facial skin, comprising: generating a first data stream representing a communication by a subject, the communication having a duration; generating a second data stream from facial skin light reflectance captured during the duration of the communication to verify the subject's identity; transmitting the first data stream to a destination; transmitting the second data stream to the destination; wherein the second data stream, when received at the destination, is correlated with the first data stream such that the second data stream can be used to repeatedly check that the communication originated from the subject during the duration of the communication.

320. 1. A system for determining facial skin micro-movements, comprising: generating a first data stream representing a communication by a subject, the communication having a duration; generating a second data stream from facial skin light reflectance captured during the duration of the communication to verify the subject's identity; transmitting the first data stream to a destination; transmitting the second data stream to a destination; at least one processor configured to correlate the second data stream with the first data stream such that, upon receipt at the destination, the second data stream can be used to repeatedly check during the communication that the communication originated from the subject; A system including:

321. 1. A head-mounted system for noise suppression, comprising: a wearable housing configured to be worn on the head of a wearer; at least one coherent light source associated with the wearable housing and configured to project light toward a facial region of the head; at least one detector associated with the wearable housing and configured to receive coherent light reflections from the facial region associated with facial skin micro-movements and output an associated reflection signal; analyzing the reflected signal to determine speech timing based on the facial skin micro-movements within the face region; receiving an audio signal from at least one microphone, the audio signal including sounds of words spoken by the wearer along with ambient sounds; correlating the reflected signal with the received audio signal based on the speech timing to determine portions of the audio signal associated with the words spoken by the wearer; and at least one processor configured to output the determined portions of the audio signal associated with the words spoken by the wearer, while omitting output of other portions of the audio signal that do not include the words spoken by the wearer.

322. 322. The head-mounted system of claim 321, wherein the at least one processor is further configured to record the determined portion of the audio signal.

323. 322. The head-mounted system of claim 321, wherein the at least one processor is further configured to determine that the other portions of the audio signal are not associated with the words spoken by the wearer.

324. 322. The head-mounted system of claim 321, wherein the other portion of the audio signal comprises ambient noise.

325. 322. The head-mounted system of claim 321, wherein the at least one processor is further configured to determine that the other portion of the audio signal includes speech from at least one person other than the wearer.

326. 326. The head-mounted system of claim 325, wherein the at least one processor is further configured to record the speech of the at least one person.

327. 326. The head-mounted system of claim 325, wherein the at least one processor is further configured to receive input indicating a wearer's desire to output the speech of the at least one person, and to output a portion of the audio signal associated with the speech of the at least one person.

328. 326. The head-mounted system of claim 325, wherein the at least one processor is further configured to identify the at least one person, determine a relationship between the at least one person and the wearer, and automatically output a portion of the audio signal associated with speech of the at least one person based on the determined relationship.

329. 322. The head-mounted system of claim 321, wherein the at least one processor is further configured to analyze the audio signal and the reflected signal to identify non-verbal interjections of the wearer and omit the non-verbal interjections from the output.

330. 322. The head-mounted system of claim 321, wherein outputting the determined portion of the audio signal includes synthesizing a vocalization of the word spoken by the wearer.

331. 331. The head-mounted system of claim 330, wherein the synthetic utterances emulate the wearer's voice.

332. 331. The head-mounted system of claim 330, wherein the synthetic vocalizations emulate the voice of a particular individual other than the wearer.

333. 331. The head-mounted system of claim 330, wherein the synthetic utterances include translations of the words spoken by the wearer.

334. 322. The head-mounted system of claim 321, wherein the at least one processor is further configured to analyze the reflected signal to identify an intention to speak and activate at least one microphone in response to the identified intention.

335. 322. The head-mounted system of claim 321, wherein the at least one processor is further configured to analyze the reflected signal to identify pauses in the words spoken by the wearer and disable at least one microphone during the identified pauses.

336. 322. The head-mounted system of claim 321, wherein the at least one microphone is part of a communication device configured to be wirelessly paired with the head-mounted system.

337. 322. The head-mounted system of claim 321, wherein the at least one microphone is integrated with the wearable housing, and the wearable housing is configured, when worn, to aim the at least one coherent light source to illuminate at least a portion of the wearer's cheek.

338. 338. The head-mounted system of claim 337, wherein a first portion of the wearable housing is configured to be positioned in the wearer's ear canal, a second portion is configured to be positioned outside the ear canal, and the at least one microphone is included in the second portion.

339. 1. A method for noise suppression using facial skin micro-movements, comprising: operating a wearable coherent light source configured to project light toward a facial region of a head of a wearer; operating at least one detector configured to receive coherent light reflections from said facial regions associated with facial skin micro-movements and output associated reflection signals; analyzing the reflected signal to determine speech timing based on the facial skin micro-movements within the face region; receiving an audio signal from at least one microphone, the audio signal including sounds of words spoken by the wearer along with ambient sounds; correlating the reflected signal with the received audio signal based on the speech timing to determine portions of the audio signal associated with the words spoken by the wearer; outputting the determined portions of the audio signal associated with the words spoken by the wearer, while omitting to output other portions of the audio signal that do not include the words spoken by the wearer.

340. 1. A non-transitory computer-readable medium comprising instructions that, when executed by at least one processor, cause the at least one processor to perform operations for noise suppression using facial skin micro-movements, the operations comprising: operating a wearable coherent light source configured to project light toward a facial region of a head of a wearer; operating at least one detector configured to receive coherent light reflections from said facial regions associated with facial skin micro-movements and output associated reflection signals; analyzing the reflected signal to determine speech timing based on the facial skin micro-movements within the face region; receiving an audio signal from at least one microphone, the audio signal including sounds of words spoken by the wearer along with ambient sounds; correlating the reflected signal with the received audio signal based on the speech timing to determine portions of the audio signal associated with the words spoken by the wearer; outputting the determined portions of the audio signal associated with the words spoken by the wearer while omitting output of other portions of the audio signal that do not include the words spoken by the wearer. Non-transitory computer-readable medium.

341. 1. A non-transitory computer-readable medium comprising instructions that, when executed by at least one processor, cause the at least one processor to perform operations for providing private answers to unvoiced questions, the instructions comprising: The operation is receiving signals indicative of specific facial micro-movements in the absence of perceptible vocalization; accessing a data structure correlating facial micro-movements with words; using the received signal to perform a lookup in the data structure for a particular word associated with the particular facial micro-movement; determining a query from the particular words; accessing at least one data structure to perform a lookup for an answer to said query; generating a discrete output comprising the answer to the query. Non-transitory computer-readable medium.

342. 342. The non-transitory computer-readable medium of claim 341, wherein the received signal is obtained via a head-mounted optical detector and is derived from skin micro-movements of a portion of the face other than the mouth.

343. 343. The non-transitory computer-readable medium of claim 342, wherein the head-mounted optical detector is configured to detect incoherent light reflections from portions of the face.

344. 343. The non-transitory computer-readable medium of claim 342, wherein the operations further include controlling at least one coherent light source to enable illumination of the facial portion, and wherein the head-mounted light detector is configured to detect coherent light reflections from the facial portion.

345. 343. The non-transitory computer-readable medium of claim 342, wherein the discrete output comprises an audible output delivered to a wearer of the head-mounted optical detector via at least one earphone.

346. 343. The non-transitory computer-readable medium of claim 342, wherein the discrete output comprises a text output delivered to a wearer of the head-mounted optical detector.

347. 343. The non-transitory computer-readable medium of claim 342, wherein the discrete output comprises a tactile output delivered to a wearer of the head-mounted optical detector.

348. 342. The non-transitory computer-readable medium of claim 341, wherein the facial micro-movement corresponds to muscle activation of at least one of the zygomaticus, orbicularis oris, laughing muscle, genioglossus, or levator labii superioris alae naris.

349. 342. The non-transitory computer-readable medium of claim 341, wherein the operations further include receiving image data, and wherein the query is determined based on the unspoken articulation of the particular word and the image data.

350. 350. The non-transitory computer-readable medium of claim 349, wherein the image data is obtained from a wearable image sensor.

351. 350. The non-transitory computer-readable medium of claim 349, wherein the image data reflects an identity of a person, the query is for a name of the person, and the discrete output includes the name of the person.

352. 350. The non-transitory computer-readable medium of claim 349, wherein the image data reflects the identity of an edible product, the query is for a list of allergens contained in the edible product, and the discrete output includes the list of allergens.

353. 350. The non-transitory computer-readable medium of claim 349, wherein the image data reflects the identity of an inanimate object, the query is for details of the inanimate object, and the discrete output includes the requested details of the inanimate object.

354. 342. The non-transitory computer-readable medium of claim 341, wherein the operations further include using the particular facial micro-movement to attempt to authenticate an individual associated with the particular facial micro-movement.

355. When the individual is authenticated, the operations further include providing a first response to the query, the first response including personal information; If the individual is not authenticated, the operations further include providing a second response to the query, the second response omitting the personal information.

355. The non-transitory computer readable medium of claim 354.

356. 355. The non-transitory computer-readable medium of claim 354, wherein the operations further include accessing personal data associated with the individual and using the personal data to generate the discrete output that includes the answer to the query.

357. 357. The non-transitory computer-readable medium of claim 356, wherein the personal data includes at least one of the individual's age, the individual's gender, the individual's current location, the individual's occupation, the individual's home address, the individual's education level, or the individual's health status.

358. 342. The non-transitory computer-readable medium of claim 341, wherein the operations further include using the facial micro-movements to determine an emotional state of an individual associated with the facial micro-movements, and wherein the answer to the query is determined based in part on the determined emotional state.

359. 1. A method for providing private answers to silent questions, comprising: receiving signals indicative of specific facial micro-movements in the absence of perceptible vocalization; accessing a data structure correlating facial micro-movements with words; using the received signal to perform a lookup in the data structure for a particular word associated with the particular facial micro-movement; determining a query from the particular words; accessing at least one data structure to perform a lookup for an answer to said query; generating a discrete output comprising the answer to the query.

360. 1. A system for providing private answers to unvoiced questions, comprising: receiving signals indicative of specific facial micro-movements in the absence of perceptible vocalization; accessing a data structure that correlates facial micro-movements with words; using the received signal to perform a lookup in the data structure for a particular word associated with the particular facial micro-movement; determining a query from the particular word; accessing at least one data structure to perform a lookup for an answer to said query; at least one processor configured to generate a discrete output comprising the answer to the query; A system including:

361. 1. A non-transitory computer-readable medium comprising instructions that, when executed by at least one processor, cause the at least one processor to execute control commands based on facial skin micro-movements, the operations comprising: operating at least one coherent light source to enable illumination of portions of the face other than the lips; receiving a particular signal representing a coherent light reflex associated with a particular non-lip facial skin micro-movement; accessing a data structure associating a plurality of non-lip facial skin micro-movements with control commands; identifying in the data structure a particular control command associated with the particular signal associated with the particular non-lip facial skin micro-movement; and executing the specific control command. Non-transitory computer-readable medium.

362. 362. The non-transitory computer-readable medium of claim 361, wherein the facial skin micro-movements correspond to unspoken articulation of at least one word associated with the particular control command.

363. 362. The non-transitory computer-readable medium of claim 361, wherein the facial skin micro-movement corresponds to the recruitment of at least one specific muscle.

364. 364. The non-transitory computer-readable medium of claim 363, wherein the at least one specific muscle includes the zygomaticus, orbicularis oris, lolis, or levator labii superioris alae naris.

365. 362. The non-transitory computer-readable medium of claim 361, wherein the facial skin micro-movements include a sequence of facial skin micro-movements from which the specific control commands are derived.

366. The non-transitory computer-readable medium of claim 361, wherein the facial skin micro-movements include involuntary micro-movements.

367. 367. The non-transitory computer-readable medium of claim 366, wherein the involuntary fine movement is triggered by an individual thinking about speaking the specific control command.

368. 367. The non-transitory computer-readable medium of claim 366, wherein the involuntary fine movements are not noticeable to the human eye.

369. 362. The non-transitory computer-readable medium of claim 361, wherein manipulating the at least one coherent light source includes determining an intensity or light pattern for illuminating portions of the face other than the lips.

370. 362. The non-transitory computer-readable medium of claim 361, wherein the particular signal is received at a rate between 50 Hz and 200 Hz.

371. 362. The non-transitory computer-readable medium of claim 361, wherein the operation further comprises analyzing the particular signal to identify time and intensity variations in speckle caused by light reflection from parts of the face other than the lips.

372. 362. The non-transitory computer-readable medium of claim 361, wherein the operations further include processing data from at least one sensor to determine a context of the particular non-lip facial skin micro-movement, and determining an action to be initiated based on the particular control command and the determined context.

373. 362. The non-transitory computer-readable medium of claim 361, wherein the specific control commands are configured to cause an audible translation of words from a source language into at least one target language other than the source language.

374. 362. The non-transitory computer-readable medium of claim 361, wherein the particular control command is configured to cause an action in a media player application.

375. 362. The non-transitory computer-readable medium of claim 361, wherein the specific control command is configured to cause an action associated with an incoming call.

376. 362. The non-transitory computer-readable medium of claim 361, wherein the specific control command is configured to cause an action associated with an ongoing call.

377. 362. The non-transitory computer-readable medium of claim 361, wherein the specific control command is configured to cause an action associated with a text message.

378. 362. The non-transitory computer-readable medium of claim 361, wherein the specific control command is configured to cause activation of a virtual personal assistant.

379. 1. A method for executing control commands based on facial skin micro-movements, comprising: operating at least one coherent light source to enable illumination of portions of the face other than the lips; receiving a particular signal representing a coherent light reflex associated with a particular non-lip facial skin micro-movement; accessing a data structure associating a plurality of non-lip facial skin micro-movements with control commands; identifying in the data structure a particular control command associated with the particular signal associated with the particular non-lip facial skin micro-movement; and executing the specific control command.

380. 1. A head-mounted system for executing control commands based on facial skin micro-movements, comprising: a wearable housing configured to be worn on the individual's head; at least one coherent light source associated with the wearable housing and configured to illuminate portions of the individual's face other than the lips; At least one detector associated with the wearable housing and configured to receive specific signals representative of coherent light reflections associated with specific non-lip facial skin micro-movements; accessing a data structure associating a plurality of non-lip facial skin micro-movements with control commands; identifying in the data structure a particular control command associated with the particular signal associated with the particular non-lip facial skin micro-movement; at least one processor configured to execute the specific control commands; a head-mounted system including:

381. 1. A non-transitory computer-readable medium comprising instructions that, when executed by at least one processor, cause the at least one processor to initiate operations for detecting changes in neuromuscular activity over time, the operations comprising: Establishing a baseline of neuromuscular activity from coherent light reflexes associated with past skin micromovements; receiving a current signal representing a coherent light reflection associated with the individual's current skin micromotion; identifying a deviation of the current skin micromotor activity from a baseline of the neuromuscular activity; and outputting an index value of the deviation. Non-transitory computer-readable medium.

382. 382. The non-transitory computer-readable medium of claim 381, wherein the operations further include establishing the baseline from historical signals representing previous coherent light reflections associated with persons other than the individual.

383. 382. The non-transitory computer-readable medium of claim 381, wherein the operation further comprises establishing the baseline from historical signals representing previous coherent light reflections associated with the individual.

384. 384. The non-transitory computer-readable medium of claim 383, wherein the history signal is based on skin micromovements occurring over a period of more than one day.

385. 384. The non-transitory computer-readable medium of claim 383, wherein the historical signal is based on skin micromovements that occurred at least one year prior to receipt of the current signal.

386. 382. The non-transitory computer-readable medium of claim 381, wherein the operations further include receiving the current signal from the wearable optical detector while the wearable optical detector is worn by the individual.

387. 387. The non-transitory computer-readable medium of claim 386, wherein the operation further includes controlling at least one wearable coherent light source to enable illumination of a portion of the individual's face, and wherein the current signal is associated with coherent light reflection from the portion of the face illuminated by the at least one wearable coherent light source.

388. 382. The non-transitory computer-readable medium of claim 381, wherein the current skin micromovement corresponds to recruitment of at least one of the zygomaticus, orbicularis oris, genioglossus, laughing muscle, or levator labii superioris alae naris muscles.

389. 382. The non-transitory computer-readable medium of claim 381, wherein the operations further include receiving the current signal from a non-wearable photodetector.

390. 390. The non-transitory computer-readable medium of claim 389, wherein the coherent light reflection associated with current skin micromotion is received from skin other than facial skin.

391. 391. The non-transitory computer-readable medium of claim 390, wherein the non-facial skin originates from the individual's neck, wrist, or chest.

392. The non-transitory computer-readable medium of claim 381, wherein the operations further include receiving an additional signal associated with the individual's skin micromovement during a period prior to the current skin micromovement, and determining a trend in change in the individual's neuromuscular activity based on the current signal and the additional signal, wherein the index value indicates the trend in the change.

393. 382. The non-transitory computer-readable medium of claim 381, wherein the operation further comprises determining a probable cause of the deviation of the current skin micromovement from a baseline of the neuromuscular activity, and the index value is indicative of the probable cause.

394. 394. The non-transitory computer-readable medium of claim 393, wherein the operations further comprise outputting an additional indicator value of the probable cause of the deviation.

395. 394. The non-transitory computer-readable medium of claim 393, wherein the operations further include receiving data indicative of at least one environmental condition, and wherein determining the probable cause of the deviation is based on the at least one environmental condition and the identified deviation.

396. 394. The non-transitory computer-readable medium of claim 393, wherein the operations further include receiving data indicative of at least one physical condition of the individual, and determining the probable cause of the deviation is based on the at least one physical condition and the identified deviation.

397. 394. The non-transitory computer-readable medium of claim 393, wherein the probable cause corresponds to at least one physical condition including being under influence, fatigue, or stress.

398. 394. The non-transitory computer-readable medium of claim 393, wherein the probable cause corresponds to at least one health condition including a heart attack, multiple sclerosis (MS), Parkinson's disease, epilepsy, or stroke.

399. 1. A method for detecting changes in neuromuscular activity over time, comprising: Establishing a baseline of neuromuscular activity from coherent light reflexes associated with an individual's past skin micromovements; and receiving a signal representative of a coherent light reflection associated with the individual's current skin micromotion; identifying a deviation of the current skin micromotor activity from a baseline of the neuromuscular activity; and outputting an indication of said deviation.

400. 1. A system for detecting changes in neuromuscular activity over time, comprising: Establish a baseline of neuromuscular activity from coherent light reflexes associated with an individual's past skin micromovements, receiving a signal representative of a coherent light reflection associated with the individual's current skin micromotion; Identifying deviations of the current skin micromotor activity from a baseline of the neuromuscular activity; at least one processor configured to output an indication of said deviation; A system including:

401. 1. A dual-use head-mounted system for projecting graphical content and interpreting non-verbal speech, comprising: a wearable housing configured to be worn on the individual's head; at least one light source associated with the wearable housing and configured to project light onto a facial area of ​​the individual in a graphical pattern, the graphical pattern configured to visually convey information; a sensor that detects a portion of the light reflected from the facial area; receiving an output signal from the sensor; determining facial skin micro-movements associated with non-verbalizations from the output signal; at least one processor configured to process the output signals to interpret the facial skin micro-movements; Dual-use head-mounted systems, including:

402. 402. The head-mounted system of claim 401, wherein the at least one processor is further configured to receive a selection of the graphical pattern and control the at least one light source to project the selected graphical pattern.

403. 402. The head-mounted system of claim 401, wherein the graphical pattern is comprised of a plurality of spots for use in determining the facial skin micro-movements via speckle analysis.

404. 402. The head-mounted system of claim 401, wherein the projected light is configured to be visible to individuals other than the individual via the human eye.

405. 402. The head-mounted system of claim 401, wherein the projected light is visible via an infrared sensor.

406. 402. The head-mounted system of claim 401, wherein the projected light source includes a laser.

407. 402. The head-mounted system of claim 401, wherein the at least one processor is configured to vary the graphical pattern over time.

408. 402. The head-mounted system of claim 401, wherein the at least one processor is configured to receive position information and modify the graphical pattern based on the received position information.

409. 402. The head-mounted system of claim 401, wherein the graphical pattern includes a scrolling message, and the at least one processor is configured to scroll the message.

410. 402. The head-mounted system of claim 401, wherein the at least one processor is further configured to detect a trigger and cause the graphical pattern to be displayed in response to the trigger.

411. 402. The head-mounted system of claim 401, wherein processing the output signal to interpret the facial skin micro-movements includes determining non-verbal speech from the facial skin micro-movements.

412. 412. The head-mounted system of claim 411, wherein the at least one processor is configured to determine the graphical pattern from the non-verbal speech.

413. 402. The head-mounted system of claim 401, wherein processing the output signal to interpret the facial skin micro-movements includes determining an emotional state from the facial skin micro-movements.

414. 414. The head-mounted system of claim 413, wherein the at least one processor is configured to determine the graphical pattern from the determined emotional state.

415. 402. The head-mounted system of claim 401, further comprising an integrated audio output, wherein the at least one processor is configured to initiate an action including outputting audio via the audio output.

416. 402. The head-mounted system of claim 401, wherein the at least one processor is configured to identify a trigger and modify the pattern based on the trigger.

417. 417. The head-mounted system of claim 416, wherein the at least one processor is configured to analyze the facial skin micro-movements to identify the trigger.

418. 417. The head-mounted system of claim 416, wherein modifying the pattern includes ceasing the projection of the graphical pattern.

419. 1. A non-transitory computer-readable medium comprising instructions that, when executed by at least one processor, cause the at least one processor to perform operations for projecting graphical content and interpreting non-verbal utterances, the instructions comprising: The operation is operating a wearable light source configured to project light in a graphical pattern onto a facial area of ​​the individual, the graphical pattern configured to visually convey information; receiving an output signal from a sensor corresponding to a portion of the light reflected from the facial region; determining facial skin micro-movements associated with non-verbalizations from the output signal; and and processing the output signal to interpret the facial skin micro-movements. Non-transitory computer-readable medium.

420. 1. A method for projecting graphical content and interpreting non-verbal speech, comprising: projecting light in a graphical pattern onto a facial area of ​​the individual, the graphical pattern configured to visually convey information; receiving the light reflected from the facial area; determining micro-movements of skin associated with non-verbal speech from the reflected light; and processing the output signal to interpret the facial skin micro-movements.

421. 1. A head-mounted system for interpreting facial skin micro-movements, comprising: a housing configured to be worn on the head of a wearer; at least one detector integrated with the housing and configured to receive light reflections from a facial region of the head and output an associated reflection signal; at least one microphone associated with the housing and configured to capture sounds produced by the wearer and output an associated audio signal; at least one processor in the housing configured to use both the reflected signal and the audio signal to generate an output corresponding to a word spoken by the wearer; a head-mounted system including:

422. 422. The head-mounted system of claim 421, further comprising at least one light source integrated with the housing and configured to project coherent light toward the facial region of the head.

423. 422. The head-mounted system of claim 421, wherein the at least one processor is configured to receive a spoken form of the words and determine at least one of the words prior to utterance of the at least one word.

424. 422. The head-mounted system of claim 421, wherein the words spoken by the wearer include at least one word spoken in a non-vocalized manner, and the at least one processor is configured to determine the at least one word without using the audio signal.

425. 422. The head-mounted system of claim 421, wherein the at least one processor is configured to use the reflected signals to identify one or more spoken words in the absence of perceptible speech.

426. 426. The head-mounted system of claim 425, wherein the at least one processor is configured to use the reflected signals to determine specific facial skin micro-movements and to correlate the specific facial skin micro-movements with reference skin micro-movements corresponding to the words.

427. 427. The head-mounted system of claim 426, wherein the at least one processor is configured to use the audio signal to determine the baseline skin micro-movements.

428. 422. The head-mounted system of claim 421, further comprising a speaker integrated with the housing and configured to generate an audio output.

429. 422. The head-mounted system of claim 421, wherein the output includes an audible presentation of the word spoken by the wearer.

430. 430. The head-mounted system of claim 429, wherein the audible presentation includes a synthesis of a voice of an individual other than the wearer.

431. 430. The head-mounted system of claim 429, wherein the audible presentation comprises a synthesis of the wearer's voice.

432. 432. The head-mounted system of claim 431, wherein the words spoken by the wearer are in a first language and the generated output includes words spoken in a second language.

433. 432. The head-mounted system of claim 431, wherein the at least one processor is configured to use audio signals to determine the voice of the individual to synthesize spoken words in the absence of perceptible vocalization.

434. 422. The head-mounted system of claim 421, wherein the output includes a textual presentation of the word spoken by the wearer.

435. 435. The head-mounted system of claim 434, wherein the at least one processor is configured to cause the textual representation of the word to be transmitted over a wireless communication channel to a remote computing device.

436. 422. The head-mounted system of claim 421, wherein the at least one processor is configured to cause the generated output to be transmitted to a remote computing device that executes a control command corresponding to the word spoken by the wearer.

437. 422. The head-mounted system of claim 421, wherein the at least one processor is further configured to analyze the reflected signals to determine facial skin micro-movements corresponding to recruitment of at least one specific muscle.

438. 438. The head-mounted system of claim 437, wherein the at least one specific muscle includes the zygomaticus, orbicularis oris, lolis, or levator labii superioris alae naris.

439. 1. A method for interpreting facial skin micro-movements, comprising: receiving coherent light reflections from facial regions associated with facial skin micromovements of the individual; outputting a reflection signal associated with the light reflection; capturing sounds made by the individual; and outputting an audio signal associated with the captured sound; and using both the reflected signal and the audio signal to generate an output corresponding to a word spoken by the individual.

440. 1. A non-transitory computer-readable medium comprising instructions that, when executed by at least one processor, cause the at least one processor to perform operations for interpreting facial skin micro-movements, the instructions comprising: The operation is receiving coherent light reflections from facial regions associated with facial skin micromovements of the individual; outputting a reflection signal associated with the light reflection; capturing sounds made by the individual; and outputting an audio signal associated with the captured sound; and using both the reflected signal and the audio signal to generate an output corresponding to a word produced by the individual. Non-transitory computer-readable medium.

441. 1. A non-transitory computer-readable medium comprising instructions that, when executed by at least one processor, cause the at least one processor to initiate training of operations for interpreting facial skin micro-movements, the operations comprising: receiving a first signal representative of a pre-phonation facial skin micro-movement during a first time period; receiving a second signal representative of a sound during a second period following the first period; analyzing the sounds to identify words spoken during the second period of time; correlating words spoken during the second time period with facial skin micromovements of the preliminary utterance received during the first time period; storing the correlation; and receiving a third signal representative of facial skin micromovements received in the absence of vocalization during a third time period; and using the stored correlation to identify a language associated with the third signal; and and outputting the language. Non-transitory computer-readable medium.

442. 442. The non-transitory computer-readable medium of claim 441, wherein the operations further include identifying additional correlations between additional words spoken over an additional extended period of time and additional pre-vocal facial skin micromovements detected during the additional extended period of time, and training a neural network using the additional correlations.

443. 442. The non-transitory computer-readable medium of claim 441, wherein the outputted language includes an indication of the words spoken during the second period of time.

444. 442. The non-transitory computer-readable medium of claim 441, wherein the outputted language includes an indication of at least one word that is different from the words spoken during the second period of time.

445. 445. The non-transitory computer-readable medium of claim 444, wherein the at least one word includes a similar phoneme sequence to the at least one word spoken during the second period of time.

446. 442. The non-transitory computer-readable medium of claim 441, wherein the first signal is associated with a first individual and the third signal is associated with a second individual.

447. 442. The non-transitory computer-readable medium of claim 441, wherein the first signal and the third signal are associated with the same individual.

448. 448. The non-transitory computer-readable medium of claim 447, wherein the operations further include using the correlation to continually update a user profile associated with the individual.

449. 442. The non-transitory computer-readable medium of claim 441, wherein the correlation is stored in a cloud-based data structure.

450. 442. The non-transitory computer-readable medium of claim 441, wherein the operations further include accessing a voice signature of the individual associated with facial skin micro-movements, and analyzing the sounds to identify words spoken during the second period of time is based on the voice signature.

451. 442. The non-transitory computer-readable medium of claim 441, wherein the second period of time begins less than 350 milliseconds after the first period of time.

452. 452. The non-transitory computer-readable medium of claim 451, wherein the third time period begins at least one day after the second time period.

453. 442. The non-transitory computer-readable medium of claim 441, wherein the first signal is based on a coherent light reflection, and the operation further includes controlling at least one coherent light source to project coherent light onto a facial area of ​​the individual from which the light reflection is received.

454. 454. The non-transitory computer-readable medium of claim 453, wherein the first signal is received from a photodetector, the photodetector and the coherent light source being part of a wearable assembly.

455. 455. The non-transitory computer-readable medium of claim 454, wherein the second signal representing sound is received from a microphone that is part of the wearable assembly.

456. 442. The non-transitory computer-readable medium of claim 441, wherein outputting the language includes presenting the words associated with the third signal in text.

457. 442. The non-transitory computer-readable medium of claim 441, wherein the operations further include, when a confidence level for identifying the language associated with the third signal falls below a threshold, processing additional signals captured during a fourth time period following the third time period to increase the confidence level.

458. 442. The non-transitory computer-readable medium of claim 441, wherein the operations further include receiving a fourth signal representing facial skin micromovements of an additional preliminary utterance during a fourth period of time, receiving a fifth signal representing a sound during a fifth period of time following the fourth period of time, and identifying a word spoken during the fifth period of time using the fourth signal.

459. 1. A method for interpreting facial skin micro-movements, comprising: receiving a first signal representative of a pre-phonation facial skin micro-movement during a first time period; receiving a second signal representative of a sound during a second period following the first period; analyzing the sounds to identify words spoken during the second period of time; correlating words spoken during the second time period with facial skin micromovements of the preliminary utterance received during the first time period; storing the correlation; and receiving a third signal representative of facial skin micromovements received in the absence of vocalization during a third time period; and using the stored correlation to identify a language associated with the third signal; and and outputting said language.

460. A system for interpreting facial skin micro-movements, comprising: receiving a first signal representative of a pre-phonation facial skin micro-movement during a first time period; receiving a second signal representative of a sound during a second period following the first period; analyzing the sounds to identify words spoken during the second period of time; correlating words uttered during the second time period with the facial skin micromovements of the preliminary utterances received during the first time period; storing the correlation; receiving a third signal representative of facial skin micromovements received in the absence of vocalization during a third time period; using the stored correlation to identify a language associated with the third signal; at least one processor configured to output said language; A system including:

461. A multi-function earpiece, an ear-worn housing; a speaker integrated with the ear-worn housing for presenting sound; a light source integrated with the ear-worn housing for projecting light toward the wearer's facial skin; a photodetector integrated with the ear-worn housing configured to receive reflections from the skin corresponding to facial skin micro-movements indicative of a pre-uttered word of the wearer; The multi-function earpiece is configured to simultaneously present the sound from the speaker, project the light toward the skin, and detect the received reflection indicative of the pre-spoken word.

462. 462. The multi-function earpiece of claim 461, wherein at least a portion of the ear-worn housing is configured to be placed within the ear canal.

463. 462. The multi-function earpiece of claim 461, wherein at least a portion of the ear-worn housing is configured to be positioned on or behind the ear.

464. 462. The multi-function earpiece of claim 461, further comprising at least one processor configured to output via the speaker an audible simulation of the pre-uttered word derived from the reflections.

465. 465. The multi-function earpiece of claim 464, wherein the audible simulation of the pre-spoken words includes a synthesis of the voice of an individual other than the wearer.

466. 465. The multi-function earpiece of claim 464, wherein the audible simulation of the pre-spoken words includes a synthesis of the pre-spoken words in a first language other than a second language of the pre-spoken words.

467. 462. The multi-function earpiece of claim 461, further comprising a microphone integrated with the ear-worn housing for receiving audio indicative of the wearer's speech.

468. 462. The multi-function earpiece of claim 461, wherein the light source is configured to project a pattern of coherent light, the pattern including a plurality of spots, toward the wearer's facial skin.

469. 462. The multi-function earpiece of claim 461, wherein the photodetector is configured to output an associated reflected signal indicative of muscle fiber recruitment.

470. 470. The multi-functional earpiece of claim 469, wherein the recruited muscle fibers include at least one of zygomaticus muscle fibers, orbicularis oris muscle fibers, lolis muscle fibers, or levator labii superioris alaris muscle fibers.

471. 462. The multi-function earpiece of claim 461, further comprising at least one processor configured to analyze the light reflections to determine the facial skin micro-movements.

472. The multi-function earpiece of claim 471, wherein the analysis includes speckle analysis.

473. 472. The multi-function earpiece of claim 471, further comprising a microphone integrated with the ear-worn housing for receiving audio indicative of the wearer's speech, wherein the at least one processor is configured to use the audio received via the microphone and the reflections received via the photodetector to correlate facial skin micro-movements with spoken words and to train a neural network to determine subsequent pre-spoken words from subsequent facial skin micro-movements.

474. 472. The multi-function earpiece of claim 471, wherein the at least one processor is configured to identify a trigger in the determined facial skin micro-movements for activating the microphone.

475. 472. The multi-function earpiece of claim 471, further comprising a pairing interface for pairing with a communication device, wherein the at least one processor is configured to transmit an audible simulation of the pre-spoken word to the communication device.

476. 472. The multi-function earpiece of claim 471, further comprising a pairing interface for pairing with a communication device, wherein the at least one processor is configured to send a text representation of the pre-spoken words to the communication device.

477. 462. The multi-function earpiece of claim 461, wherein the light source is configured to project coherent light toward the wearer's facial skin.

478. 462. The multi-function earpiece of claim 461, wherein the light source is configured to project incoherent light toward the wearer's facial skin.

479. 1. A method of operating a multi-function earpiece, comprising: activating a speaker integrated with an ear-worn housing associated with the multi-function earpiece to present sound; operating a light source integrated with the ear-worn housing to project light toward the wearer's facial skin; operating a photodetector integrated with the ear-worn housing and configured to receive reflections from the wearer's facial skin corresponding to facial skin micro-movements indicative of a pre-uttered word; simultaneously presenting the sound from the speaker, projecting the light toward the skin, and detecting the received reflection indicative of the pre-uttered word.

480. 1. A non-transitory computer-readable medium comprising instructions that, when executed by at least one processor, cause the at least one processor to perform operations for operating a multi-function earpiece, the operations comprising: activating a speaker integrated with an ear-worn housing associated with the multi-function earpiece to present sound; operating a light source integrated with the ear-worn housing to project light toward the wearer's facial skin; operating a photodetector integrated with the ear-worn housing and configured to receive reflections from the wearer's facial skin corresponding to facial skin micro-movements indicative of a pre-uttered word; simultaneously presenting the sound from the speaker, projecting the light toward the skin, and detecting the received reflection indicative of the pre-uttered word. Non-transitory computer-readable medium.

481. a driver for integration with a software program and for enabling a neuromuscular sensing device to interface with said software program, an input handler for receiving inaudible muscle activation signals from the neuromuscular detection device; a lookup component for mapping particular ones of the non-audible activation signals to corresponding commands within the software program; a signal processing module for receiving the non-audible muscle activation signals from the input handler, providing the particular one of the non-audible muscle activation signals to the lookup component, and receiving an output as the corresponding command; a communication module that communicates the corresponding commands to the software program, thereby enabling control within the software program based on non-audible muscle activity detected by the neuromuscular detection device.

482. 482. The driver of claim 481, wherein the input handler, the lookup component, the signal processing module, and the control code are incorporated into the software program.

483. 482. The driver of claim 481, wherein the input handler, the lookup component, the signal processing module, and the control code are incorporated into the neuromuscular detection device.

484. 482. The driver of claim 481, wherein the input handler, the lookup component, the signal processing module, and the control code are incorporated into an application programming interface (API).

485. 484. The driver of claim 483, wherein the neuromuscular detection device includes: a light source configured to project light toward the skin; a light detector configured to sense reflection of the light from the skin; and at least one processor configured to generate the non-audible muscle activation signal based on the sensed light reflection.

486. 486. A driver as described in claim 485, wherein the sensed reflection of the light from the skin corresponds to micro-movements of the skin.

487. 482. The driver of claim 481, wherein the lookup component is pre-populated based on training data correlating the non-audible muscle activation signals with the corresponding commands.

488. 482. The driver of claim 481, including a training module for determining a correlation between the corresponding command and the non-audible muscle activation signal and populating the lookup component.

489. 482. The driver of claim 481, wherein the lookup component comprises a lookup table.

490. 482. The driver of claim 481, wherein the lookup component includes an artificial intelligence data structure.

491. 482. The driver of claim 481, wherein the neuromuscular detection device includes a light source for projecting light toward the skin, a light detector configured to sense reflection of the light from the skin, and at least one processor configured to generate the non-audible muscle activation signal based on the sensed light reflection.

492. 492. A driver as described in claim 491, wherein the light source is configured to output coherent light.

493. 493. A driver as described in claim 492, wherein the at least one processor is configured to generate the non-audible muscle activation signal based on speckle analysis of received reflections of the coherent light.

494. 482. The driver of claim 481, wherein the lookup component is further configured to map a portion of the particular one of the non-audible activation signals to text.

495. 495. The driver of claim 494, wherein the text corresponds to subvocalization appearing in the non-audible muscle activation signal.

496. 495. The driver of claim 494, wherein the lookup component is further configured to map the particular portion of the non-audible muscle activation signal to a command for causing at least one of a visual output of the text or an audible synthesis of the text.

497. 482. The driver of claim 481, further comprising a return output for transmitting data to the neuromuscular detection device.

498. 498. The driver of claim 497, wherein the data is configured to cause at least one of an audio, tactile, or textual output via the neuromuscular detection device.

499. 482. The driver of claim 481, further comprising a detection and correction routine that detects and corrects errors that occur during data transmission.

500. 482. The driver of claim 481, further comprising a configuration management routine that allows the driver to be configured by applications other than the software program.

501. 1. A non-transitory computer-readable medium comprising instructions that, when executed by at least one processor, cause the at least one processor to perform context-driven facial fine motor operations, the operations comprising: receiving a first signal representative of a first coherent light reflection associated with a first facial skin micro-movement during a first time period; analyzing the first coherent light reflection to determine a first plurality of words associated with the first facial skin micro-movement; receiving first information indicative of a first contextual condition under which the first facial skin micromovement occurred; receiving a second signal representative of a second coherent light reflection associated with a second facial skin micro-movement during a second time period; analyzing the second coherent light reflection to determine a second plurality of words associated with the second facial skin micro-movement; receiving second information indicative of a second contextual condition under which the second facial skin micromovement occurred; accessing a plurality of control rules correlating a plurality of actions with a plurality of contextual conditions, a first control rule defining a manner of private presentation based on the first contextual condition and a second control rule defining a manner of non-private presentation based on the second contextual condition; upon receiving the first information, implementing the first control rule to privately output the first plurality of words; upon receiving the second information, implementing the second control rule to non-privately output the second plurality of words. Non-transitory computer-readable medium.

502. 502. The non-transitory computer-readable medium of claim 501, wherein the first information indicative of the first contextual condition includes an indication that the first facial skin micromovement is associated with private thoughts.

503. 502. The non-transitory computer-readable medium of claim 501, wherein the first information indicative of the first contextual condition includes an indication that the first facial skin micromovement occurs in a private setting.

504. 502. The non-transitory computer-readable medium of claim 501, wherein the first information indicative of the first contextual condition includes an indication that the individual generating the facial micromovement is looking down.

505. 502. The non-transitory computer-readable medium of claim 501, wherein the second information indicative of the second contextual condition includes an indication that the second facial skin micromovement occurs during a phone call.

506. 502. The non-transitory computer-readable medium of claim 501, wherein the second information indicative of the second contextual condition includes an indication that the second facial skin micromovement occurs during a videoconference.

507. 502. The non-transitory computer-readable medium of claim 501, wherein the second information indicative of the second contextual condition includes an indication that the second facial skin micromovement is performed during a social interaction.

508. 502. The non-transitory computer-readable medium of claim 501, wherein at least one of the first information and the second information indicates an activity of an individual that generates the facial micro-movement, and wherein the operation further includes implementing either the first control rule or the second control rule based on the activity.

509. 502. The non-transitory computer-readable medium of claim 501, wherein at least one of the first information and the second information indicates a location of an individual generating the facial micro-movement, and wherein the operation further includes implementing either the first control rule or the second control rule based on the location.

510. 502. The non-transitory computer-readable medium of claim 501, wherein at least one of the first information and the second information indicates a type of engagement of an individual generating the facial micro-movement on a computing device, and wherein the operation further includes implementing either the first control rule or the second control rule based on the type of engagement.

511. 502. The non-transitory computer-readable medium of claim 501, wherein privately outputting the first plurality of words includes generating an audio output to a personal sound generating device.

512. 502. The non-transitory computer-readable medium of claim 501, wherein privately outputting the first plurality of words includes generating a text output to a personal text generation device.

513. 502. The non-transitory computer-readable medium of claim 501, wherein non-privately outputting the second plurality of words includes transmitting an audio output to a mobile communication device.

514. 502. The non-transitory computer-readable medium of claim 501, wherein non-privately outputting the second plurality of words includes causing text output to be presented on a shared display.

515. 502. The non-transitory computer-readable medium of claim 501, wherein the operations further include determining a trigger for switching between a private output mode and a non-private output mode.

516. 516. The non-transitory computer-readable medium of claim 515, wherein the operations further include receiving third information indicative of a change in a contextual condition, and wherein the trigger is determined from the third information.

517. 516. The non-transitory computer-readable medium of claim 515, wherein the operations further include determining the trigger based on the first plurality of words or the second plurality of words.

518. 516. The non-transitory computer-readable medium of claim 515, wherein the operations further include receiving an output mode selection from an associated user interface and determining the trigger based on the output mode selection.

519. 1. A method for generating context-driven facial fine motor output, comprising: receiving a first signal representative of a first coherent light reflection associated with a first facial skin micro-movement during a first time period; analyzing the first coherent light reflection to determine a first plurality of words associated with the first facial skin micro-movement; receiving first information indicative of a first contextual condition under which the first facial skin micromovement occurred; receiving a second signal representative of a second coherent light reflection associated with a second facial skin micro-movement during a second time period; analyzing the second coherent light reflection to determine a second plurality of words associated with the second facial skin micro-movement; receiving second information indicative of a second contextual condition under which the second facial skin micromovement occurred; accessing a plurality of control rules correlating a plurality of actions with a plurality of contextual conditions, a first control rule defining a manner of private presentation based on the first contextual condition and a second control rule defining a manner of non-private presentation based on the second contextual condition; upon receiving the first information, implementing the first control rule to privately output the first plurality of words; Upon receiving the second information, implementing the second control rule to non-privately output the second plurality of words.

520. 1. A system for generating context-driven facial fine motor output, comprising: receiving a first signal representative of a first coherent light reflection associated with a first facial skin micro-movement during a first time period; analyzing the first coherent light reflection to determine a first plurality of words associated with the first facial skin micromovement; receiving first information indicative of a first contextual condition under which the first facial skin micromovement occurred; receiving a second signal representative of a second coherent light reflection associated with a second facial skin micro-movement during a second time period; analyzing the second coherent light reflection to determine a second plurality of words associated with the second facial skin micromovement; receiving second information indicative of a second contextual condition under which the second facial skin micromovement occurred; accessing a plurality of control rules correlating a plurality of actions with a plurality of contextual conditions, a first control rule defining a manner of private presentation based on the first contextual condition and a second control rule defining a manner of non-private presentation based on the second contextual condition; Upon receiving the first information, implementing the first control rule to privately output the first plurality of words; at least one processor configured to implement the second control rules to non-privately output the second plurality of words upon receiving the second information; system.

521. 1. A non-transitory computer-readable medium comprising instructions that, when executed by at least one processor, cause the at least one processor to perform operations for extracting a response to content based on facial skin micro-movements, the operations comprising: determining the facial skin micro-movements of the individual based on reflection of coherent light from a facial region of the individual while the individual is consuming content; determining at least one specific micro-expression from the facial skin micro-movements; accessing at least one data structure containing correlations between a plurality of micro-expressions and a plurality of non-verbal perceptions; determining a particular non-verbal perception of the content consumed by the individual based on the at least one particular micro-expression and the correlations in the data structure; and and initiating an action associated with the particular non-verbalized perception. Non-transitory computer-readable medium.

522. 522. The non-transitory computer-readable medium of claim 521, wherein the at least one particular micro-expression is imperceptible to the human eye.

523. 522. The non-transitory computer-readable medium of claim 521, wherein the facial skin micro-movements used to determine the at least one particular micro-expression correspond to recruitment of at least one muscle from a muscle group including the zygomaticus, genioglossus, orbicularis oris, laughing muscle, or levator labii superioris alae naris.

524. 522. The non-transitory computer-readable medium of claim 521, wherein the at least one particular micro-expression comprises a sequence of micro-expressions associated with the particular non-verbalized perception.

525. 525. The non-transitory computer-readable medium of claim 524, wherein the operations further include determining a degree of the particular non-verbalized perception based on the sequence of micro-expressions, and determining an action to initiate based on the degree of the particular non-verbalized perception.

526. 522. The non-transitory computer-readable medium of claim 521, wherein the at least one data structure includes past non-verbalized perceptions of previously consumed content, and the operations further include determining a degree of the particular non-verbalized perception relative to the past non-verbalized perceptions, and determining an action to initiate based on the degree of the particular non-verbalized perception.

527. 522. The non-transitory computer-readable medium of claim 521, wherein the non-verbalized perception comprises an emotional state of the individual.

528. 522. The non-transitory computer-readable medium of claim 521, wherein the operations further include determining an action to initiate based on the consumed content and the particular non-verbalized perception.

529. 522. The non-transitory computer-readable medium of claim 521, wherein the initiated action includes causing transmission of a message reflecting a correlation between the particular non-verbalized perception and the consumed content.

530. 522. The non-transitory computer-readable medium of claim 521, wherein the initiated action includes storing in memory a correlation between the particular non-verbalized perception and the consumed content.

531. 522. The non-transitory computer-readable medium of claim 521, wherein the action includes determining additional content to present to the individual based on the particular non-verbalized perception and the consumed content.

532. 532. The non-transitory computer-readable medium of claim 531, wherein the consumed content is of a first type and the additional content is of a second type different from the first type.

533. 522. The non-transitory computer-readable medium of claim 521, wherein the consumed content is part of a chat with at least one other individual, and the action includes generating a visual representation of the particular non-verbalized perception within the chat.

534. 522. The non-transitory computer-readable medium of claim 521, wherein the action includes selecting an alternative manner for presenting the consumed content.

535. 522. The non-transitory computer-readable medium of claim 521, wherein the action varies based on the type of content consumed.

536. 522. The non-transitory computer-readable medium of claim 521, wherein the operations further include operating at least one wearable coherent light source to enable illumination of portions of the individual's face other than the lips, and receiving a signal indicative of coherent light reflections from the portions of the face other than the lips.

537. 537. The non-transitory computer-readable medium of claim 536, wherein the facial skin micromotion is determined based on speckle analysis of the coherent light reflection.

538. 522. The non-transitory computer-readable medium of claim 521, wherein the reflection of the coherent light is received by a wearable photodetector.

539. 1. A method for extracting a response to content based on facial skin micro-movements, comprising: determining the facial skin micro-movements of the individual based on reflection of coherent light from a facial region of the individual while the individual is consuming content; determining at least one specific micro-expression from the facial skin micro-movements; accessing a data structure comprising correlations between a plurality of micro-expressions and a plurality of non-verbal perceptions; determining a particular non-verbal perception of the content consumed by the individual based on the at least one particular micro-expression and the correlations in the data structure; and and initiating an action associated with the particular non-verbalized perception.

540. 1. A system for extracting a response to content based on facial skin micro-movements, comprising: determining the facial skin micro-movements of the individual based on reflection of coherent light from a facial region of the individual while the individual is consuming content; determining at least one specific micro-expression from the facial skin micro-movements; accessing a data structure that includes correlations between a plurality of micro-expressions and a plurality of non-verbal percepts; determining a particular non-verbal perception of the content consumed by the individual based on the at least one particular micro-expression and the correlations in the data structure; at least one processor configured to initiate an action associated with the particular non-verbal perception; system.

541. 1. A non-transitory computer-readable medium comprising instructions that, when executed by at least one processor, cause the at least one processor to perform operations of removing noise from facial skin micromotor signals, the operations comprising: operating a light source to enable illumination of a facial skin area of ​​the individual during a period in which the individual is engaged in at least one non-speech related physical activity; receiving a signal representative of light reflection from the facial skin area; analyzing the received signals to identify a first reflection component indicative of pre-speech facial skin micro-movements and a second reflection component associated with the at least one non-speech related physical activity; filtering out the second reflection component to enable interpretation of words from the first reflection component indicative of facial skin micro-movements of the preliminary utterance. Non-transitory computer-readable medium.

542. 542. The non-transitory computer-readable medium of claim 541, wherein the light source is a coherent light source.

543. 542. The non-transitory computer-readable medium of claim 541, wherein the second reflected component is a result of walking.

544. 542. The non-transitory computer-readable medium of claim 541, wherein the second reflected component is a result of running.

545. 542. The non-transitory computer-readable medium of claim 541, wherein the second reflected component is a result of respiration.

546. The non-transitory computer-readable medium of claim 541, wherein the second reflex component is the result of a blink and is based on neural activation of at least one orbicularis oculi muscle.

547. 542. The non-transitory computer-readable medium of claim 541, wherein when the individual is simultaneously engaged in a first physical activity and a second physical activity, the operations further include identifying a first portion of the second reflectance component associated with the first physical activity and a second portion of the second reflectance component associated with the second physical activity, and filtering out the first portion of the second component and the second portion of the second component from the first component to enable interpretation of a word from facial skin micro-movements of the preliminary utterance associated with the first component.

548. 542. The non-transitory computer-readable medium of claim 541, wherein the operations further include receiving data from a mobile communication device, the data indicative of the at least one non-speech related physical activity.

549. 549. The non-transitory computer-readable medium of claim 548, wherein the mobile communication device lacks a light sensor that detects the light reflection.

550. 549. The non-transitory computer-readable medium of claim 548, wherein the data received from the mobile communication device includes at least one of data indicative of the individual's heart rate, data indicative of the individual's blood pressure, or data indicative of the individual's exercise.

551. 542. The non-transitory computer-readable medium of claim 541, wherein the operations further include presenting the words in a synthesized voice.

552. 542. The non-transitory computer-readable medium of claim 541, wherein the signal is received from a sensor associated with a wearable housing, and the instructions further include analyzing the signal to determine the at least one non-speech-related physical activity.

553. 553. The non-transitory computer-readable medium of claim 552, wherein the sensor is an image sensor configured to capture at least one event in the individual's environment, and the at least one processor is configured to determine that the event is associated with the at least one non-speech-related physical activity.

554. 542. The non-transitory computer-readable medium of claim 541, wherein the operations further include using a neural network to identify the second reflex component associated with the at least one non-speech-related physical activity.

555. 542. The non-transitory computer-readable medium of claim 541, wherein the preparatory vocalization facial skin micro-movements correspond to the recruitment of one or more involuntary muscle fibers.

556. 556. The non-transitory computer-readable medium of claim 555, wherein the involuntary recruitment of muscle fibers is a result of the individual thinking about saying the word.

557. 556. The non-transitory computer readable medium of claim 555, wherein the recruitment of one or more muscle fibers includes recruitment of at least one of zygomaticus muscle fibers, orbicularis oris muscle fibers, genioglossus muscle fibers, lolis muscle fibers, or levator labii superioris alae noli muscle fibers.

558. 542. The non-transitory computer-readable medium of claim 541, wherein the signal is received at a rate between 50 Hz and 200 Hz.

559. A method for removing noise from facial skin micro-movements, comprising: operating a light source to enable illumination of a facial skin area of ​​the individual during a period in which the individual is engaged in at least one non-speech related physical activity; receiving a signal representative of light reflection from the facial skin area; analyzing the received signals to identify a first reflection component indicative of pre-speech facial skin micro-movements and a second reflection component associated with the at least one non-speech related physical activity; filtering out the second reflection component to enable interpretation of words from the first reflection component indicative of facial skin micro-movements of the preliminary utterance.

560. A system for determining facial skin micro-movements, comprising: operating a light source to enable illumination of a facial skin area of ​​the individual during a period in which the individual is engaged in at least one non-speech related physical activity; receiving a signal representative of light reflection from the facial skin area; analyzing the received signals to identify a first reflection component indicative of pre-speech facial skin micro-movements and a second reflection component associated with the at least one non-speech related physical activity; at least one processor configured to filter out the second reflection component to enable interpretation of words from the first reflection component indicative of facial skin micro-movements of the preliminary utterance; system.

Citation Information

Cited By

  • Information processing system, information processing method, and program

    JP7904655B1