Detecting and utilizing facial micro-motion
By analyzing facial reflected signals using wearable coherent light sources and processors, the problem of facial micromotion detection and interpretation during silent reading is solved, and effective detection and interpretation of silent communication information is achieved, and a variety of application scenarios are supported.
Patent Information
- Application Number
- CN202380066666.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-02-28
- Filing Date
- 2023-07-19
- Publication Date
- 2025-07-11
AI Technical Summary
The prior art is difficult to effectively detect and interpret facial skin micromovement occurring during silent reading, especially in the case of silent speech and pre-vocalization.
By using a wearable coherent light source to project light to the face area, a light detector is used to receive reflected signals, and analyze it through the processor to determine the facial skin micromovement, access the data structures related to the movement in the memory for matching and interpretation, so as to realize the detection and interpretation of facial micromovement.
It realizes facial communication information detection and interpretation in silent speech and pre-voice situations, supports identity verification, continuous authentication, wireless communication, voice assistance, content interpretation and other interactive functions, and improves communication efficiency and accuracy.
Smart Images

Figure CN120303605A_ABST
Abstract
Description
[0001] Cross - reference to related applications
[0002] This application claims the priority of U.S. Provisional Patent Application No. 63 / 390,653, filed on July 20, 2022; U.S. Provisional Patent Application No. 63 / 394,329, filed on August 2, 2022; U.S. Provisional Patent Application No. 63 / 438,061, filed on January 10, 2023; U.S. Provisional Patent Application No. 63 / 441,183, filed on January 26, 2023; and U.S. Provisional Patent Application No. 63 / 487,299, filed on February 28, 2023, the entire contents of all of which are incorporated herein by reference. Technical field
[0003] The present disclosure generally relates to the field of discerning information from neuromuscular activity. One example is discerning communication by detecting facial skin movements that occur during subvocalization. Other examples include enabling control - based neuromuscular activity and discerning changes in neuromuscular activity over time. Background art
[0004] Human brain and nerve activities are complex and involve many subsystems. One of these subsystems is the facial region that humans use to communicate with others. From birth, humans are trained to activate craniofacial muscles to produce sounds. Even before full language capabilities are fully developed, infants use facial expressions (including micro - expressions) to convey deeper information about themselves. However, after learning language capabilities, speech is the primary technology humans use for communication.
[0005] The normal process of vocalized speech uses multiple groups of muscles and nerves, from the chest and abdomen, through the throat, up through the mouth and face. To produce a given phoneme, motor neurons activate muscle groups in the face, larynx, and mouth to prepare to push an airstream out of the lungs, and these muscles continue to move during speech to produce words and sentences. Without such an airstream, no sound is emitted from the mouth. When there is no airflow from the lungs, silent speech occurs, while the muscles in the face, larynx, and mouth clearly produce the desired sounds or move in an interpretable manner.
[0006] Some disclosed embodiments aim to provide a new method for extracting meaning from neuromuscular activity that detects facial skin micromovements that occur during subvocalization, such as silent speech. Summary of the invention
[0007] Embodiments according to the present disclosure provide systems, methods, and devices for detecting and using facial movements.
[0008] Some disclosed embodiments may include systems, methods, and non-transitory computer-readable media for using facial skin micromovements to identify an individual. These embodiments may include operating a wearable coherent light source configured to project light toward a facial region of an individual's head; operating at least one detector configured to receive coherent light reflections from the facial region and output an associated reflection signal; analyzing the reflection signal to determine a specific facial skin micromovement of the individual; accessing a memory that associates a plurality of facial skin micromovements with the individual; searching the memory for a match between the determined specific facial skin micromovement and at least one of the plurality of facial skin micromovements; if a match is identified, initiating a first action; and if no match is identified, initiating a second action different from the first action.
[0009] Some disclosed embodiments may include systems, methods, and non-transitory computer-readable media for interpreting facial skin movements. These embodiments may include projecting light onto a plurality of facial regions of an individual, where the plurality of regions includes at least a first region and a second region, and the first region is closer to at least one of the zygomaticus or risorius muscles than the second region; receiving reflections from the plurality of regions; detecting a first facial skin movement corresponding to the reflection from the first region and a second facial skin movement corresponding to the reflection from the second region; based on a difference between the first facial skin movement and the second facial skin movement, determining that the reflection from the first region, which is closer to at least one of the zygomaticus or risorius muscles, is a stronger communication indicator than the reflection from the second region; and based on determining that the reflection from the first region is a stronger communication indicator, processing the reflection from the first region to determine communication and ignoring the reflection from the second region.
[0010] Some disclosed embodiments may include systems, methods, and non-transitory computer-readable media for performing authentication operations based on facial micromovements. These embodiments may include receiving, in a trusted manner, a reference signal for verifying a correspondence between a specific individual and an account at an institution, the reference signal being based on reference facial micromovements detected using first coherent light reflected from the face of the specific individual; storing a correlation between the identity of the specific individual and the reference signal reflecting the facial micromovements in a secure data structure; after storing, receiving, via the institution, a request to authenticate the specific individual; receiving a real-time signal indicative of a second coherent light reflection from second facial micromovements of the specific individual; comparing the real-time signal with the reference signal stored in the secure data structure to authenticate the specific individual; and after authentication, notifying the institution that the specific individual is authenticated.
[0011] Some disclosed embodiments may include systems, methods, and non-transitory computer-readable media for continuous authentication based on facial skin micromovements. These embodiments may include receiving a first signal during an ongoing electronic transaction, the first signal representing coherent light reflection associated with a first facial skin micromovement during a first time period; using the first signal to determine the identity of a particular individual associated with the first facial skin micromovement; receiving a second signal during the ongoing electronic transaction, the second signal representing coherent light reflection associated with a second facial skin micromovement, the second signal received during a second time period after the first time period; using the second signal to determine that the particular individual is also associated with the second facial skin micromovement; receiving a third signal representing coherent light reflection associated with a third facial skin micromovement during the ongoing electronic transaction, the third signal received during a third time period after the second time period; using the third signal to determine that the third facial skin micromovement is not associated with the particular individual; and initiating an action based on the determination that the third facial skin micromovement is not associated with the particular individual.
[0012] Some disclosed embodiments may include systems, methods, and non-transitory computer-readable media for performing threshold operations for interpreting facial skin micromovements. These embodiments may include detecting a facial skin micromovement in the absence of a perceivable vocalization associated with the facial micromovement; determining an intensity level of the facial skin micromovement; comparing the determined intensity level to a threshold; interpreting the facial skin micromovement when the intensity level is above the threshold; and ignoring the facial skin micromovement when the intensity level is below the threshold.
[0013] Some disclosed embodiments may include systems, methods, and non-transitory computer-readable media for establishing a silent session. These embodiments may include establishing a wireless communication channel for enabling a silent conversation via a first wearable device and a second wearable device, wherein the first wearable device and the second wearable device each include a coherent light source and a light detector configured to detect facial skin micromovements based on coherent light reflection; detecting, by the first wearable device, a first facial skin micromovement occurring in the absence of a perceivable vocalization; transmitting, via the wireless communication channel, a first communication from the first wearable device to the second wearable device, wherein the first communication is derived from the first facial skin micromovement and transmitted by the second wearable device for presentation; receiving, via the wireless communication channel, a second communication from the second wearable device, wherein the second communication is derived from a second facial skin micromovement detected by the second wearable device; and presenting the second communication to a wearer of the first wearable device.
[0014] Some disclosed embodiments may include systems, methods, and non-transitory computer-readable media for initiating a content interpretation operation prior to the utterance of content to be interpreted. These embodiments may include receiving a signal representative of facial skin micromovements; determining, based on the signal, at least one word to be spoken prior to uttering at least one word in a source language; beginning an interpretation of the at least one word prior to the utterance of the at least one word; and presenting an interpretation of the at least one word when the at least one word is spoken.
[0015] Some disclosed embodiments may include systems, methods, and non-transitory computer-readable media for performing a private voice assistance operation. These embodiments may include receiving a signal indicative of a particular facial skin micromovement that reflects a private request for an assistant, where responding to the private request requires identifying a particular individual associated with the particular facial skin micromovement; accessing a data structure that maintains a correlation between the particular individual and a plurality of facial skin micromovements associated with the particular individual; searching the data structure for a match indicative of a correlation between a stored identity of the particular individual and the particular facial skin micromovement; in response to a determination that the match exists in the data structure, initiating a first action in response to the request, where the first action involves enabling access to unique information of the particular individual; and if the match is not identified in the data structure, initiating a second action different from the first action.
[0016] Some disclosed embodiments may include systems, methods, and non-transitory computer-readable media for determining a silent phoneme based on facial skin micromovements. These embodiments may include controlling at least one coherent light source in a manner that enables illumination of a first region of a face and a second region of the face; performing a first pattern analysis on light reflected from the first region of the face to determine a first micromovement of facial skin in the first region of the face; performing a second pattern analysis on light reflected from the second region of the face to determine a second micromovement of facial skin in the second region of the face; and using the first micromovement of facial skin in the first region of the face and the second micromovement of facial skin in the second region of the face to determine at least one silent phoneme.
[0017] Some disclosed embodiments may include systems, methods, and non-transitory computer-readable media for generating a synthetic representation of a facial expression. These embodiments may include controlling at least one coherent light source in a manner that enables illumination of a portion of a face; receiving an output signal from a light detector, where the output signal corresponds to the reflection of coherent light from a portion of the face; applying speckle analysis to the output signal to determine facial skin micromovements based on the speckle analysis; using the determined facial skin micromovements based on the speckle analysis to identify at least one word pre-spoken or spoken during a period of time; using the determined facial skin micromovements based on the speckle analysis to identify at least one change in facial expression during the period of time; and during the period of time, outputting data for causing a virtual representation of the face to mimic the at least one change in facial expression in combination with an audio rendering of the at least one word.
[0018] Some disclosed embodiments may include systems, methods, and non-transitory computer-readable media for performing attention-related interactions based on facial skin micromovements. These embodiments may include determining an individual's facial skin micromovements based on the reflection of coherent light from a facial region of the individual; using the facial skin micromovements to determine a particular level of focus of the individual; receiving data associated with an expected interaction with the individual; accessing a data structure that associates information reflecting alternative levels of focus with different presentation modalities; determining a particular presentation modality for the expected interaction based on the particular level of focus and the associated information; and associating the particular presentation modality with the expected interaction for subsequent interaction with the individual.
[0019] Some disclosed embodiments may include systems, methods, and non-transitory computer-readable media for performing speech synthesis operations based on detected facial skin micromovements. These embodiments may include determining particular facial skin micromovements of a first individual talking to a second individual based on light reflection from a facial region of the first individual; accessing a data structure that associates facial micromovements with words; performing a lookup of a particular word associated with the particular facial skin micromovements in the data structure; obtaining an input associated with preference speech consumption characteristics of the second individual; adopting the preference speech consumption characteristics; and using the adopted preference speech consumption characteristics to synthesize an audible output of the particular word.
[0020] Some disclosed embodiments may include systems, methods, and non-transitory computer-readable media for performing operations for a pre-vocal personal presentation. These embodiments may include receiving a reflected signal corresponding to light reflected from an individual's facial region; using the received reflected signal to determine a particular facial skin micro-movement of the individual in the absence of a perceivable vocalization associated with a particular facial skin micro-movement; accessing a data structure that associates facial skin micro-movements with words; performing a lookup of a particular non-vocal word associated with the particular facial skin micro-movement in the data structure; and causing the particular non-vocal word to be audibly presented to the individual before the individual vocalizes the particular word.
[0021] Some disclosed embodiments may include systems, methods, and non-transitory computer-readable media for interpreting speech disorders based on facial movements. These embodiments may include receiving a signal associated with a particular facial skin movement of an individual having a speech disorder that affects the way the individual pronounces a plurality of words; accessing a data structure that includes a correlation between the plurality of words and a plurality of facial skin movements corresponding to the way the individual pronounces the plurality of words; identifying a particular word associated with the plurality of particular facial skin movements based on the received signal and the correlation; and generating an output of the particular word for presentation, wherein the output is different from how the individual pronounces the particular word.
[0022] Some disclosed embodiments may include systems, methods, and non-transitory computer-readable media for continuously verifying the authenticity of a communication based on light reflection from facial skin. These embodiments may include generating a first data stream representing a communication of an object, the communication having a duration; generating a second data stream for authenticating the identity of the object based on facial skin light reflections captured during the duration of the communication; transmitting the first data stream to a destination; transmitting the second data stream to the destination; and wherein the second data stream is related to the first data stream in such a way that once the second data stream is received at the destination, the second data stream is enabled for repeatedly checking during the duration of the communication that the communication originated from the object.
[0023] Some disclosed embodiments may include systems, methods, and non-transitory computer-readable media for noise suppression using facial skin micromovements. These embodiments may include operating a wearable coherent light source configured to project light toward a facial region of a wearer's head; operating at least one detector configured to receive coherent light reflections from the facial region associated with facial skin micromovements and output an associated reflection signal; analyzing the reflection signal to determine speech timing based on the facial skin micromovements in the facial region; receiving an audio signal from at least one microphone, the audio signal including sounds of words spoken by the wearer and ambient sounds; correlating the reflection signal with the received audio signal based on the speech timing to determine a portion of the audio signal associated with the words spoken by the wearer; and outputting the determined portion of the audio signal associated with the words spoken by the wearer while omitting output of other portions of the audio signal that do not include words spoken by the wearer.
[0024] Some disclosed embodiments may include systems, methods, and non-transitory computer-readable media for providing private responses to silent questions. These embodiments may include receiving a signal indicating a specific facial micromovement without a perceivable vocalization; accessing a data structure that associates facial micromovements with words; using the received signal to perform a lookup of a specific word associated with the specific facial micromovement in the data structure; determining a query based on the specific word; accessing at least one data structure to perform a lookup of a response to the query; and generating a covert output including the response to the query.
[0025] Some disclosed embodiments may include systems, methods, and non-transitory computer-readable media for performing control commands based on facial skin micromovements. These embodiments may include operating at least one coherent light source in a manner that enables illumination of non-lip portions of the face; receiving a specific signal representing coherent light reflections associated with a specific non-lip facial skin micromovement; accessing a data structure that associates multiple non-lip facial skin micromovements with control commands; identifying a specific control command associated with the specific signal in the data structure, the specific signal being associated with the specific non-lip facial skin micromovement; and performing the specific control command.
[0026] Some disclosed embodiments may include systems, methods, and non-transitory computer-readable media for detecting changes in neuromuscular activity over time. These embodiments may include establishing a baseline of neuromuscular activity from coherent light reflections associated with historical skin micromovements; receiving a current signal representing coherent light reflections associated with an individual's current skin micromovements; identifying a deviation of the current skin micromovements from the baseline of neuromuscular activity; and outputting an indicator of the deviation.
[0027] Some disclosed embodiments may include systems, methods, and non-transitory computer-readable media for projecting graphical content and for interpreting non-verbal expressions. These embodiments may include operating a wearable light source configured to project light in a graphical pattern onto a facial region of an individual, wherein the graphical pattern is configured to visually convey information; receiving an output signal from a sensor corresponding to a portion of the light reflected from the facial region; determining facial skin micromovements associated with the non-verbal expression based on the output signal; and processing the output signal to interpret the facial skin micromovements.
[0028] Some disclosed embodiments may include systems, methods, and non-transitory computer-readable media for interpreting facial skin micromovements. These embodiments may include receiving coherent light reflections from a facial region associated with an individual's facial skin micromovements; outputting a reflection signal associated with the light reflections; capturing sound produced by the individual; outputting an audio signal associated with the captured sound; and using both the reflection signal and the audio signal to generate an output corresponding to a word expressed by the individual.
[0029] Some disclosed embodiments may include systems, methods, and non-transitory computer-readable media for interpreting facial skin micromovements. These embodiments may include receiving a first signal representing pre-articulatory facial skin micromovements during a first time period; receiving a second signal representing sound during a second time period after the first time period; analyzing the sound to identify a word spoken during the second time period; associating the word spoken during the second time period with the pre-articulatory facial skin micromovements received during the first time period; storing the association; receiving a third signal representing facial skin micromovements received without articulation during a third time period; using the stored association to identify a language associated with the third signal; and outputting the language.
[0030] Some disclosed embodiments may include systems, methods, and non-transitory computer-readable media for operating a multi-functional headset. These embodiments may include operating a speaker integrated with an ear-mountable housing associated with the multi-functional headset for presenting sound; operating a light source integrated with the ear-mountable housing to project light toward the skin of the wearer's face; operating a light detector integrated with the ear-mountable housing, the light detector configured to receive a reflection from the skin, the reflection corresponding to a facial skin micromovement indicative of a pre-articulated word of the wearer; and simultaneously presenting sound through the speaker, projecting light onto the skin, and detecting the received reflection indicative of the pre-articulated word.
[0031] Some disclosed embodiments may include a driver for integrating with a software program and enabling a neuromuscular detection device to connect to the software program. The driver includes: an input processing program for receiving an inaudible muscle activation signal from the neuromuscular detection device; a lookup component for mapping a specific inaudible activation signal in the inaudible activation signal to a corresponding command in the software program; a signal processing module for receiving the inaudible muscle activation signal from the input processing program, providing a specific signal in the inaudible muscle activation signal to the lookup component, and receiving an output as the corresponding command; and a communication module for transmitting the corresponding command to the software program, thereby enabling control within the software program based on inaudible muscle activity detected by the neuromuscular detection device.
[0032] Some disclosed embodiments may include systems, methods, and non-transitory computer-readable media for performing context-driven facial micromotion operations. These embodiments may include receiving, during a first time period, a first signal representing a first coherent light reflection associated with a first facial skin micromotion; analyzing the first coherent light reflection to determine a first plurality of words associated with the first facial skin micromotion; receiving first information indicating a first context condition in which the first facial skin micromotion occurs; receiving, during a second time period, a second signal representing a second coherent light reflection associated with a second facial skin micromotion; analyzing the second coherent light reflection to determine a second plurality of words associated with the second facial skin micromotion; receiving second information indicating a second context condition in which the second facial skin micromotion occurs; accessing a plurality of control rules associating a plurality of actions with a plurality of context conditions, wherein a first control rule specifies a form of private presentation based on the first context condition, and a second control rule specifies a form of non-private presentation based on the second context condition; executing the first control rule to privately output the first plurality of words upon receiving the first information; and executing the second control rule to non-privately output the second plurality of words upon receiving the second information.
[0033] Some disclosed embodiments may include systems, methods, and non-transitory computer-readable media for extracting reactions to content based on facial skin micromotions. These embodiments may include determining, during a time period in which an individual consumes content, the individual's facial skin micromotions based on coherent light reflections from the individual's facial region; determining at least one specific micro-expression based on the facial skin micromotions; accessing at least one data structure containing correlations between a plurality of micro-expressions and a plurality of non-verbalized perceptions; determining a specific non-verbalized perception of the content consumed by the individual based on the at least one specific micro-expression and the correlations in the data structure; and initiating an action associated with the specific non-verbalized perception.
[0034] Some disclosed embodiments may include systems, methods, and non-transitory computer-readable media for removing noise from facial skin micromotion signals. These embodiments may include operating a light source in a manner that illuminates an individual's facial skin region during a time period in which the individual participates in at least one non-verbal-related physical activity; receiving a signal representing light reflections from the facial skin region; analyzing the received signal to identify a first reflection component indicating pre-articulatory facial skin micromotions and a second reflection component associated with the at least one non-verbal-related physical activity; and filtering out the second reflection component so that words can be interpreted based on the first reflection component indicating pre-articulatory facial skin micromotions.
[0035] Consistent with other disclosed embodiments, a non-transitory computer-readable storage medium can store program instructions that are executed by at least one processing device and perform any of the methods described herein.
[0036] The foregoing general description and the following detailed description are merely exemplary and explanatory and are not restrictive of the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] The drawings incorporated in and constituting a part of this disclosure illustrate various disclosed embodiments. In the drawings:
[0038] Figure 1 is a schematic diagram of a user using a first example speech detection system according to some embodiments of the present disclosure.
[0039] Figure 2A is a schematic diagram of a user using a second example speech detection system according to some embodiments of the present disclosure.
[0040] Figure 2B is a perspective view of a user using a third example speech detection system according to some embodiments of the present disclosure.
[0041] Figure 3 is a schematic diagram of a user using a fourth example speech detection system according to some embodiments of the present disclosure.
[0042] Figure 4 is a block diagram showing some components of a speech detection system and a remote processing system according to some embodiments of the present disclosure.
[0043] Figure 5A and Figure 5B is a schematic diagram of a part of a speech detection system detecting facial skin micromovements according to some embodiments of the present disclosure.
[0044] Figure 6 is a schematic diagram of a reflected image associated with light reflections received from a region of a facial area associated with a single speckle according to some embodiments of the present disclosure.
[0045] Figure 7 is a block diagram of a memory according to disclosed embodiments.
[0046] Figure 8 is a diagram of an exemplary alternative action speech detection process according to some embodiments of the present disclosure.
[0047] Figure 9 is a flowchart of an example process for identifying an individual according to some embodiments of the present disclosure.
[0048] Figure 10is a flow chart of an example process for identifying individuals using facial skin micro-movements according to some embodiments of the present disclosure.
[0049] Figure 11 are illustrations of two example use cases for interpreting facial skin motion based on light reflections, according to some embodiments of the present disclosure.
[0050] Figure 12 is an illustration of another example use case for interpreting facial skin motion based on light reflections according to some embodiments of the present disclosure.
[0051] Figure 13 is a flow chart of an example process for interpreting facial skin movements according to some embodiments of the present disclosure.
[0052] Figure 14 is a schematic diagram of the operation of an exemplary authentication service according to some embodiments of the present disclosure, the authentication service being configured to provide identity verification of an individual based on facial micro-movements.
[0053] Figure 15 , Figure 16A and Figure 16B is a simplified illustration of an exemplary system for individual identity verification using facial micro-movements according to some embodiments of the present application.
[0054] Figure 17A is a flowchart of an exemplary process for using facial micro-movements for individual identity verification according to some embodiments of the present application.
[0055] Figure 17B is a flowchart of an exemplary process for generating a reference signal for individual identity verification according to some embodiments of the present application.
[0056] Figure 18 is a schematic diagram of an exemplary authentication system and service according to some embodiments of the present application, which is configured to provide continuous authentication of an individual based on facial skin micro-movements.
[0057] Figure 19 is a simplified illustration of an exemplary system configured to provide continuous authentication of an individual using facial micro-movements in accordance with some embodiments of the present disclosure.
[0058] Figure 20 is a flow chart of an exemplary process for continuously authenticating an individual using facial micro-movements according to some embodiments of the present application.
[0059] Figure 21 is a flowchart of another exemplary process for continuously authenticating an individual using facial micro-movements according to some embodiments of the present application.
[0060] Figure 22 It is a flowchart of another exemplary process for continuously authenticating an individual using facial micro - movements according to some embodiments of the present application.
[0061] Figure 23 It is a flowchart of another exemplary process for continuously authenticating an individual using facial micro - movements according to some embodiments of the present application.
[0062] Figure 24 It includes a series of displacement - relative - to - time graphs according to some embodiments of the present disclosure, the time graph including threshold levels associated with multiple facial positions.
[0063] Figure 25A and Figure 25B It is a schematic diagram of exemplary displacement levels of facial micro - movements according to some embodiments of the present application, where a threshold - trigger mechanism can be employed.
[0064] Figure 26 It is a block diagram of an exemplary speech detection system using thresholds and threshold adjustments as a trigger mechanism according to some embodiments of the present disclosure.
[0065] Figure 27 It is a graph of displacement relative to time including background noise according to some embodiments of the present disclosure.
[0066] Figure 28A and Figure 28B It shows an example of measuring skin potential difference to determine facial micro - movements according to some embodiments of the present disclosure.
[0067] Figure 29 It is a flowchart of an exemplary method for using a threshold to interpret or ignore facial micro - movements according to some embodiments of the present application.
[0068] Figure 30 It is a schematic diagram of a system configured to enable silent reading conversations between individuals according to some embodiments of the present disclosure.
[0069] Figure 31 It is a schematic diagram of an exemplary process of detected facial skin micro - movements of an individual according to some embodiments of the present application.
[0070] Figure 32 It is a schematic diagram of another system configured to enable silent reading conversations between individuals according to some embodiments of the present disclosure.
[0071] Figure 33 It is a flowchart of an exemplary process for establishing a silent reading conversation according to some embodiments of the present disclosure.
[0072] Figure 34It is a schematic diagram of an exemplary content interpretation process initiated before the utterance of the content to be interpreted according to some embodiments of the present disclosure.
[0073] Figure 35 It is a flowchart of an exemplary process for initiating content interpretation before the utterance of the content to be interpreted according to an embodiment of the present disclosure.
[0074] Figure 36 It illustrates an exemplary protocol for performing a private voice-assisted operation with different facial skin micromovements according to an embodiment of the present disclosure.
[0075] Figure 37 It illustrates an example of a second action initiated if no match is recognized in an exemplary data structure according to an embodiment of the present disclosure.
[0076] Figure 38 It is a flowchart of an exemplary process for performing a private voice-assisted operation according to an embodiment of the present disclosure.
[0077] Figure 39 It is an exemplary diagram showing how different regions of facial skin are used to detect silent phonemes according to some embodiments of the present application.
[0078] Figure 40 It shows three curves depicting exemplary alternative timings for completing a process involving detecting silent phonemes according to an embodiment of the present disclosure.
[0079] Figure 41 It is a flowchart of an exemplary process for determining silent phonemes from facial skin micromovements according to an embodiment of the present disclosure.
[0080] Figure 42A It is a perspective view of a user wearing an exemplary head-mounted device and a resulting virtual representation of one facial expression of the user according to some embodiments of the present disclosure.
[0081] Figure 42B It is another perspective view of a user wearing an example head-mounted device and a virtual representation of another facial expression of the resulting user according to some embodiments of the present disclosure.
[0082] Figure 43 It is a block diagram of an exemplary operating environment for generating a synthetic representation of a facial expression according to some embodiments of the present disclosure.
[0083] Figure 44 It is a block diagram of an exemplary system for generating a synthetic representation of a facial expression and / or for determining voiced phonemes from facial skin micromovements according to some embodiments of the present application.
[0084] Figure 45 is a flowchart showing an exemplary method for generating a synthetic representation of a facial expression and / or for determining a phoneme being pronounced from facial skin micromovements according to some embodiments of the present application.
[0085] Figure 46 is a flowchart showing another exemplary method for generating a synthetic representation of a facial expression and / or for determining a phoneme being pronounced from facial skin micromovements according to some embodiments of the present application.
[0086] Figure 47 is a schematic diagram of an example process for determining a presentation manner based on facial skin micromovements according to some embodiments of the present application.
[0087] Figure 48 is a schematic diagram of an exemplary system for a user to use an attention-related interaction based on facial skin micromovements according to some embodiments of the present application.
[0088] Figure 49 is a schematic diagram of receiving an expected interaction via a smart phone according to some embodiments of the present disclosure.
[0089] Figure 50 is a flowchart of an example process for determining a presentation manner based on facial skin micromovements according to some embodiments of the present disclosure.
[0090] Figure 51 shows a first individual wearing a speech detection system when communicating with at least one second individual according to some embodiments of the present disclosure.
[0091] Figure 52 shows a flowchart of an example process for initiating content interpretation before the content to be interpreted is pronounced according to embodiments of the present disclosure.
[0092] Figure 53A and Figure 53B is a schematic diagram of an audible presentation of silently reading a word before pronunciation according to some embodiments of the present disclosure.
[0093] Figure 54 is a block diagram of an exemplary speech detection system for determining a silently read word from facial micromovements that cause an audible presentation using received reflections according to some embodiments of the present disclosure.
[0094] Figure 55 shows an exemplary schematic diagram of a synthetic translation between languages according to some embodiments of the present disclosure.
[0095] Figure 56 shows an exemplary additional function of a pre-pronunciation personal presentation according to some embodiments of the present disclosure.
[0096] Figure 57 is a flowchart illustrating an exemplary method for determining a subvocalized word from facial micromovements using received reflections for audible rendering, according to some embodiments of the present disclosure.
[0097] Figure 58 is a perspective view of an individual using a first example speech detection system, according to some embodiments of the present disclosure.
[0098] Figure 59A and Figure 59B is a schematic diagram of a portion of a speech detection system detecting facial skin micromovements, according to some embodiments of the present application.
[0099] Figure 60 is a block diagram illustrating exemplary components of a first example of a speech detection system, according to some embodiments of the present disclosure.
[0100] Figure 61 is a flowchart illustrating an exemplary method for determining facial skin micromovements, according to some embodiments of the present application.
[0101] Figure 62 is an illustration of an example system for correcting speech disorders based on facial movements, according to some embodiments of the present disclosure.
[0102] Figure 63 is a flowchart illustrating an example process for correcting speech disorders based on facial movements, according to some embodiments of the present disclosure.
[0103] Figure 64 is a schematic diagram of an exemplary speech detection system that sends two data streams to a destination to verify communication authenticity, according to some embodiments of the present disclosure.
[0104] Figure 65 is a schematic diagram of an example function for authenticating a communication at a destination, according to some embodiments of the present disclosure.
[0105] Figure 66 is a flowchart illustrating an exemplary method for using received reflections to verify communication authenticity, according to some embodiments of the present disclosure.
[0106] Figure 67 illustrates an exemplary headset system for noise suppression, according to some embodiments of the present disclosure.
[0107] Figure 68 illustrates an example of audio signal processing for noise suppression, according to some embodiments of the present disclosure.
[0108] Figure 69 is a flowchart illustrating an example process for noise suppression, according to some embodiments of the present disclosure.
[0109] Figure 70 Illustrates an exemplary system for providing private responses to silent questions according to an embodiment of the present disclosure.
[0110] Figure 71 Illustrates an example of image data application that can be used to provide private responses to silent questions according to an embodiment of the present disclosure.
[0111] Figure 72 Illustrates a flowchart of an exemplary process for providing private responses to silent questions according to an embodiment of the present disclosure.
[0112] Figure 73 Is a schematic diagram of an individual using a first example speech detection system according to some embodiments of the present disclosure.
[0113] Figure 74 Is a schematic diagram of two individuals each using an example speech detection system according to some embodiments of the present disclosure.
[0114] Figure 75 Is a flowchart of an exemplary method for performing silent speech control according to some embodiments of the present disclosure.
[0115] Figure 76 Is a schematic diagram of an exemplary timeline of the progression of a medical condition that can be detected by measuring skin micromovements over time, consistent with some embodiments of the present disclosure.
[0116] Figure 77 Is a block diagram of an exemplary system capable of detecting changes in neuromuscular activity over time according to some embodiments of the present disclosure.
[0117] Figure 78 Is a block diagram of an example function for detecting deviations in a medical condition according to some embodiments of the present disclosure.
[0118] Figure 79 Is a flowchart showing an exemplary method for using received light reflection to detect changes in neuromuscular activity over time according to some embodiments of the present disclosure.
[0119] Figure 80 Is a schematic diagram of using a projected graphic pattern to detect non-verbal information from an individual according to some embodiments of the present disclosure.
[0120] Figure 81 Is a schematic diagram of changing a projected graphic pattern according to some embodiments of the present disclosure.
[0121] Figure 82A flowchart of an exemplary process for detecting non-verbal information using a projected graphical pattern according to some embodiments of the present disclosure.
[0122] Figure 83 An exemplary embodiment of a user wearing a head-mounted system for interpreting facial skin micromovements is shown.
[0123] Figure 84 A flowchart of an exemplary method for interpreting facial skin micromovements is shown.
[0124] Figures 85A to 85C An exemplary embodiment of a training operation for interpreting facial skin micromovements in first to third time periods according to some disclosed embodiments is shown.
[0125] Figure 86 Is according to some disclosed embodiments with an example additional extended time period Figures 85A to 85C A flowchart of an example of the first to third time periods shown in
[0126] Figure 87 A flowchart of an exemplary method for interpreting facial skin micromovements according to some disclosed embodiments is shown.
[0127] Figure 88 A schematic diagram of a user wearing an exemplary head-mounted device with added facial micromovement detection capabilities according to some embodiments of the present disclosure.
[0128] Figure 89 A schematic diagram of an exemplary facial micromovement detection process according to some embodiments of the present application.
[0129] Figure 90 A flowchart of an example process for operating a multifunctional earphone according to some embodiments of the present disclosure.
[0130] Figure 91 A schematic diagram of a user wearing an exemplary head-mounted device with an alternative form factor according to some embodiments of the present disclosure.
[0131] Figure 92 A block diagram of an exemplary driver for docking with a software program and a device according to disclosed embodiments is shown.
[0132] Figure 93 A schematic diagram of an exemplary driver for integration with a software program and a neuromuscular detection device according to disclosed embodiments is shown.
[0133] Figure 94 A schematic diagram of an exemplary system for integration with a software program and for enabling a device to interface with a software program according to embodiments of the present disclosure is illustrated.
[0134] Figure 95 is a block diagram showing an exemplary operating environment for generating context - driven facial micro - motion outputs according to some embodiments of the present disclosure.
[0135] Figure 96 is a block diagram showing an exemplary system for generating context - driven facial micro - motion outputs according to some embodiments of the present application.
[0136] Figure 97 is a flowchart showing an exemplary method for generating context - driven facial micro - motion outputs according to some embodiments of the present application.
[0137] Figure 98 is a flowchart showing another exemplary method for generating context - driven facial micro - motion outputs according to some embodiments of the present application.
[0138] Figure 99 is a schematic diagram of a user wearing an example head - mounted device and context - driven outputs generated based on facial micro - motions according to some embodiments of the present disclosure.
[0139] Figure 100 is a schematic diagram of an example system for extracting reactions to content based on facial skin micro - motions according to some embodiments of the present disclosure.
[0140] Figure 101 includes a block diagram of two example use cases for initiating actions based on reactions to content according to some embodiments of the present disclosure.
[0141] Figure 102 is a flowchart of an example process for extracting reactions to content based on facial skin micro - motions according to some embodiments of the present disclosure.
[0142] Figure 103 shows an individual performing a first non - verbal related activity (e.g., walking) and a second non - verbal related activity (e.g., sitting) while wearing a speech recognition system according to embodiments of the present disclosure.
[0143] Figure 104 shows according to embodiments of the present disclosure Figure 103 an exemplary close - up view of a speech detection system.
[0144] Figure 105 shows an exemplary comparison between a first signal of speech - related facial skin motion when an individual is walking and a second signal of speech - related facial skin motion when the individual is sitting according to embodiments of the present disclosure.
[0145] Figure 106Illustrated is an exemplary decomposition and classification of an electronic representation of an optical signal into a first reflected component indicative of pre-articulatory facial skin micro-movements and a second reflected component associated with at least one non-verbal related body activity, according to an embodiment of the present disclosure.
[0146] Figure 107 Illustrated is an exemplary second reflected component of an optical signal reflected from a facial region of an individual simultaneously participating in a first body activity and a second body activity, according to an embodiment of the present disclosure.
[0147] Figure 108 Illustrated is a flowchart of an example process for removing noise from a facial skin micro-movement signal, according to an embodiment of the present disclosure.
[0148] Figure 109 Illustrated is another exemplary decomposition and classification of a representation of an optical signal to identify a first reflected component indicative of pre-articulatory facial skin micro-movements, according to an embodiment of the present disclosure. Detailed Description
[0149] The following detailed description includes references to the accompanying drawings. Wherever possible, the same reference numbers are used in the drawings and the specification to refer to the same or like parts. Although several illustrative embodiments are described herein, modifications, adaptations, and other embodiments are possible. For example, components shown in the drawings may be replaced, added, or modified, and the illustrative methods described herein may be modified by replacing, reordering, removing, or adding steps to the disclosed methods. Accordingly, the following detailed description is not limited to the disclosed embodiments and examples. Instead, the appropriate scope is defined by the appended claims.
[0150] When discussed in the context of different disclosed embodiments, various terms used in the specification and claims may be defined or construed differently. It should be understood that the definitions, construals, and explanations of terms in each example apply to all examples, even if not repeated, unless a transitional definition, explanation, or construal would render an embodiment inoperable. It should also be understood that once a term is defined herein, in the absence of an inherent inconsistency, that definition applies to all other uses of that term herein. Additionally, the exemplary embodiments of the drawings and their descriptions should not be considered as definitions of the claim terms, but rather as non-limiting examples for illustrating particular embodiments.
[0151] Throughout this document, the present disclosure refers to "embodiments" and "disclosed embodiments," which refer to examples of the inventive concepts, ideas, and / or manifestations described herein. Many related and unrelated embodiments are described throughout the present disclosure. The fact that some "disclosed embodiments" are described as exhibiting a feature or characteristic does not mean that other disclosed embodiments necessarily share that feature or characteristic.
[0152] The present disclosure uses open license language, indicating that, for example, some embodiments "may" adopt, relate to, or include a particular feature. The use of the term "may" and other open terms is intended to indicate that although not every embodiment may adopt the specifically disclosed feature, at least one embodiment does adopt the specifically disclosed feature.
[0153] The different embodiments of the present disclosure may relate to systems, methods, and / or computer-readable media containing instructions. A system refers to at least two interconnected or interrelated components or parts that work together to achieve a common goal, function, or sub-function. A method refers to at least two steps, actions, or techniques to be followed in order to complete a task or sub-task, achieve a goal, or reach the next step. A computer-readable media containing instructions refers to any storage mechanism that contains, for example, program code instructions to be executed by a computer processor. Examples of computer-readable media are further described elsewhere in the present disclosure. The instructions may be written in any type of computer programming language, such as an interpreted language (e.g., a scripting language such as HTML and JavaScript), a procedural or functional language (e.g., C or Pascal that can be compiled to convert to executable code), an object-oriented programming language (e.g., Java or Python), a logic programming language (e.g., Prolog or answer set programming), and / or any other programming language. The instructions executed by at least one processor may include one or more program code instructions implemented in hardware, in software (including in one or more signal processing and / or application specific integrated circuits), in firmware, or in any combination thereof, as previously described. Causing the processor to perform an operation may involve causing the processor to compute, execute, or otherwise implement one or more arithmetic, mathematical, logical, reasoning, or inferential steps.
[0154] Some disclosed embodiments may relate to detecting facial skin micromovements. The phrase "facial skin micromovement" broadly refers to skin movements on the face that can be detected by a sensor but may not be easily detectable by the naked eye. Facial skin micromovements include various types of movements, including involuntary movements caused by muscle recruitment and other types of small-scale skin deformations, ranging from micrometers to millimeters in scale and lasting from fractions of a second to several seconds. In some cases, facial skin micromovements are part of a larger-scale skin movement that is visible to the naked eye (e.g., smiling may involve many facial skin micromovements). In other cases, facial skin micromovements are not part of any larger-scale skin movement that is visible to the naked eye. Although such micromovements can occur over a facial area of several square millimeters, they can occur in surface areas of the facial skin that are less than 1 square centimeter, less than 1 square millimeter, less than 0.1 square millimeter, less than 0.01 square millimeter, or even smaller areas. In some embodiments, facial skin micromovements correspond to the recruitment of one or more muscles in the facial region of an individual's head. The facial region can include specific anatomical regions such as a portion of the cheek above the mouth, a portion of the cheek below the mouth, a portion of the middle of the jaw, a portion of the cheek below the eyes, the neck, the chin, and other regions associated with specific muscle recruitments that may cause facial skin micromovements. In some embodiments, a specific muscle may be attached to skin tissue rather than to any bone. In particular, the specific muscle may be located in the subcutaneous tissue associated with cranial nerve V or cranial nerve VII. As discussed in more detail herein, Figure 5A The first facial skin micromovement 522A and the second facial skin micromovement 522B in
[0155] When a particular muscle contracts, the muscle pulls on the facial skin and causes movement of the facial skin. Some of the movements that occur when a particular muscle contracts may be micromovements. For example, the particular muscles that may cause micromovements of the facial skin in the context of the present disclosure can be generally divided into four groups: orbital, nasal, oral, and lingual. The orbital group of facial muscles contains two muscles associated with the eye socket. These muscles control the movement of the eyelids, which is important for protecting the cornea from damage. They are both innervated by cranial nerve VII. The nasal group of facial muscles is associated with the movement of the nose and the skin around it. There are three muscles in this group, and they are also both innervated by cranial nerve VII. The oral group is the most important group of facial expression muscles: responsible for the movement of the mouth and lips. Such movements are required in singing and whistling and emphasize vocal communication. The oral muscle group consists of the orbicularis oris muscle, the buccinator muscle, and various smaller muscles. In a particular embodiment, the disclosed system can monitor micromovements of the facial skin corresponding to the recruitment of the buccinator muscle. The buccinator muscle is located between the mandible and the maxilla and is relatively deep compared to other muscles of the face. The lingual muscle group consists of the following muscles: four intrinsic muscles for changing the shape of the tongue (e.g., the superior longitudinal muscle, the inferior longitudinal muscle, the vertical muscle, and the transverse muscle); and four extrinsic muscles for changing the position of the tongue (e.g., the genioglossus muscle, the hyoglossus muscle, the styloglossus muscle, and the palatoglossus muscle). Any of the lingual muscles listed above may cause movement of the tongue, which can be detected by analyzing the detected micromovements of the facial skin. As discussed in more detail herein, Figure 5A and Figure 5B the muscle fibers 520 in are non-limiting examples of facial muscles that cause micromovements of the facial skin according to the present disclosure.
[0156] According to the present disclosure, facial skin micromovements can be detected during silent reading. The phrase "during silent reading" refers to any speech-related activity that occurs without vocalization, before vocalization, or before imperceptible vocalization. In one embodiment, the speech-related activity may include silent speech (i.e., when there is no airflow from the lungs but the facial muscles clearly express the desired sound). In another embodiment, the speech-related activity may include speaking silently (i.e., when some air flows out of the lungs, but the words are pronounced in a way that is imperceptible to an audio sensor). In yet another embodiment, the speech-related activity may include prevocalization muscle recruitment (i.e., silent reading that occurs before the onset of vocalization is sometimes referred to as prevocalization in this document). In some cases, prevocalization facial skin micromovements may be triggered by voluntary muscle recruitment that occurs when certain craniofacial muscles begin to speak a word. In other cases, prevocalization facial skin micromovements may be triggered by involuntary facial muscle recruitment that an individual performs when certain craniofacial muscles are preparing to speak a word. For example, involuntary facial muscle recruitment may occur between 0.1 second and 0.5 seconds before actual vocalization. In some cases, the proposed system may use the detected facial skin micromovements that occur during silent reading to identify the upcoming spoken word. Determining the word that a user wants to say before actual vocalization can have many benefits, since the system does not have to wait for the user to vocalize the word to begin processing the word. In one example, the disclosed system may generate captions for a live broadcast without delay. In another example, the disclosed system may translate in real time what the user is saying into a different language. Additionally, since the disclosed system can detect words before they are vocalized, the actual vocalization of these words is not necessary. Thus, facial skin micromovements that occur during silent reading can be detected in the absence of perceptible vocalization. Movements of the facial skin or muscles in the absence of vocalization but still conveying speech-related information are referred to herein as silent speech. Detecting silent speech can have various uses, including but not limited to enabling silent communication with other users, initiating commands, or enabling interaction with a virtual personal assistant. As discussed in more detail herein, Figure 7 the silent reading decryption module 708 in
[0157] In some embodiments, a speech detection system is used to detect facial skin micro-movements. Although the shorthand "speech detection system" is used, it should be understood that the system can alternatively or additionally be configured to detect non-verbal commands, expressions, or emotions. The system can also be used for user authentication. The speech detection system can include any device in a set of devices operably coupled together. As used herein, the term "system" includes any device or set of devices that are operably connected together and configured to perform a function. In some embodiments, the system can include a computer (e.g., a desktop computer, a laptop computer, a server, a smart phone, a portable digital assistant (PDA), or a similar device) or multiple computers or servers operably connected together (e.g., using wired or wireless means) to share information and / or data. The computer can include a dedicated computer (e.g., hardwired and coded to perform the desired function) or can include a general-purpose computer (e.g., using software to perform any desired function). In some embodiments, the system can include a cloud server. As described elsewhere in this disclosure, a cloud server can be a computer platform that provides services via a network such as the Internet. In one embodiment, the speech detection system can include a wearable housing, a coherent light source or an incoherent light source, a light detector, and a processor. However, the specific list of the above components is not intended to limit the systems covered by this disclosure. As will be understood by those skilled in the art who benefit from this disclosure, many variations and / or modifications can be made to the exemplary speech detection system. For example, in all cases, not all components are essential for the detection of facial skin micro-movements. Additionally, the components can be rearranged into various configurations while providing the functions of the various disclosed embodiments. In some cases, the speech detection system according to some embodiments of this disclosure does not have to be wearable, but can be aimed at the skin from a position not connected to the human body. A wearable or non-wearable system can project coherent light towards the user's facial area, analyze the reflected light, and determine facial skin micro-movements. Alternatively, in other cases, the speech detection system according to some embodiments of this disclosure does not have to include a coherent light source. Specifically, the light detector can be an ultra-high-resolution image sensor (e.g., over 120 megapixels) or any other sensor capable of performing facial micro-movement detection, and one or more image processing algorithms can be used to complete the detection of facial skin micro-movements. As discussed in more detail herein, Figures 1 to 3 The speech detection system 100 in
[0158] Some disclosed embodiments include a wearable housing configured to be worn on an individual's head. The term "wearable housing" broadly includes any structure or housing designed to be attached to the human head, such as in a manner configured to be worn by a user. Such a wearable housing can be configured to contain or support one or more electronic components or sensors. In one example, the wearable housing is configured to be associated with a pair of glasses. In another example, the wearable housing is associated with in-ear headphones. The wearable housing can have a cross-section that is button-shaped, P-shaped, square, rectangular, rounded rectangular, or any other regular or irregular shape that can be worn by a user. Such a structure can allow the wearable housing to be worn on, in, or around a body part associated with the user's head (e.g., on the ear, in the ear, around the neck). The wearable housing can be made of plastic, metal, composite material, a combination of two or more of plastic, metal, and composite material, or other suitable materials. According to embodiments of the present disclosure, the housing can be worn on the ear. There are several ways to attach the housing to the ear: 1. In-the-ear (ITE): The housing can be directly inserted into the ear canal and held in place by the shape of the ear. Examples include earbuds and earplugs. In some cases, the housing can be customized to fit the specific shape of an individual's ear and placed in an ear bowl. 2. Behind-the-ear (BTE): The housing can be placed behind the ear and have a small tube that extends into the ear canal. Examples include hearing aids and Bluetooth headsets. 3. On-the-ear (OTE): The housing can be placed on top of the ear and held in place by a headband or other support. Examples include structures such as headphones, ear cups, etc. 4. Over-the-head (OTH): The housing can be held in place by a headband that goes over the top of the head. In other embodiments, the wearable housing can be attached to an auxiliary device, such as glasses (sunglasses or corrective vision glasses), a hat, a helmet, goggles, or any other type of head-mounted device. In some cases, the wearable housing can be attached to the auxiliary device using at least one adapter. Specifically, at least one adapter can be configured to enable an individual to wear the speech detection system in two or more different ways. For example, a single adapter can enable the wearable housing to be attached to both glasses and in-ear headphones. As discussed in more detail herein, Figure 1 and Figure 2A the wearable housing 110 in
[0159] Some embodiments include a coherent light source configured to project light toward a user's facial region. Other embodiments include an incoherent light source configured to project light toward a user's facial region. As used herein, the term "light source" broadly refers to any device configured to emit light. The term "coherent light" includes light that is highly ordered and exhibits a high degree of spatial and temporal coherence. For example, this may occur when light waves are in phase with each other and have a uniform frequency and wavelength, resulting in a highly directional light beam with limited outward spread as it propagates. Alternatively, coherent light can include cases where the light waves have a constant phase difference. In some examples, coherent light can be generated by a coherent light source, such as a laser and other types of light sources with a narrow spectral range and a high degree of monochromaticity (i.e., the light consists of a single wavelength). In contrast, incoherent light can be generated by incoherent light sources such as incandescent bulbs and natural sunlight, which have a wide spectral range and a low degree of monochromaticity.
[0160] For example, coherent light can include many waves with the same frequency, having different phases and amplitudes, not necessarily at the same time and location. To control interference, it may be necessary to pre-identify the light phase information. In one embodiment, the coherent light source can be a laser, such as a solid-state laser, a laser diode, a high-power laser, a quantum cascade laser (QCL), or an alternative light source, such as an LED-based light source. Additionally, the coherent light source can emit light in different formats, such as light pulses, continuous wave (CW), quasi-CW, etc. For example, one type of light source that can be used is a vertical cavity surface emitting laser (VCSEL). Another type of light source that can be used is an external cavity diode laser (ECDL). In some examples, the light source can include a laser diode configured to emit light with a wavelength between about 650 nm and 1150 nm. Alternatively, the coherent light source can include a laser diode configured to emit light with a wavelength between about 800 nm and about 1020 nm, between about 850 nm and about 950 nm, or between about 1300 nm and about 1700 nm. Unless otherwise specified, the terms "about" and "substantially the same" with respect to a numerical value can include a difference of up to 5% relative to the stated value. As discussed in more detail herein, Figure 4 and Figure 5A and Figure 5BThe light source 410 therein is a non-limiting example of a light source according to the present disclosure. In the context of the present disclosure, it should be recognized that the use of a coherent light source is intended as a non-limiting example implementation in the context of speech detection systems, methods, and computer-readable media. Many of the embodiments described herein can be practiced with coherent light or incoherent light, and the reference to either by way of example herein is not intended to be limiting. For example, even if not explicitly stated, the speech detection systems, methods, and computer program products described and claimed can be configured to measure incoherent light reflection to detect facial skin micro-movements.
[0161] Some embodiments include at least one detector configured to receive light reflection from a facial region of a user. The phrase "light detector" or simply "detector" broadly refers to any device, element, or system capable of measuring one or more properties of an electromagnetic wave (e.g., power, frequency, phase, pulse timing, pulse duration, or other characteristics) and generating an output related to the measured property. Examples of detectors according to the present disclosure can include: photosensitive sensors, imaging sensors, phase detectors, MEMS sensors, wavemeters, spectrometers, spectrophotometers, homodyne detectors, or heterodyne detectors. In some embodiments, at least one detector can be configured to detect coherent light reflection. Additionally or alternatively, at least one detector can be configured to detect incoherent light reflection. The at least one detector can include a plurality of detectors formed by a plurality of detection elements. The at least one detector can include different types of light detectors. The at least one detector can include a plurality of detectors of the same type, which can differ in other characteristics (e.g., sensitivity, size). Combinations of several types of detectors can be used for different reasons. According to some embodiments, the at least one detector can measure any form of light reflection and scattering, including secondary speckle patterns, different types of specular reflection, diffuse reflection, speckle interferometry, and any other form of light scattering. In some embodiments, the at least one detector is configured to output an associated reflection signal based on the detected coherent light reflection. In the context of the present disclosure, the phrase "reflection signal" broadly refers to any form of data obtained from the at least one light detector in response to light reflection from a facial region. The reflection signal can be any electronic representation having properties determined based on the light reflection, or the raw measurement signal detected by the at least one light detector. As discussed in more detail herein, Figure 4 and Figure 5A and Figure 5B the light detector 412 therein is a non-limiting example of a light detector according to the present disclosure.
[0162] Some embodiments include at least one processor configured to use reflected signals from a detector and determine facial skin micromovements. The phrase "at least one processor" can refer to any physical device or group of devices having circuitry that performs logical operations on inputs. For example, the at least one processor can include one or more integrated circuits (ICs), including application-specific integrated circuits (ASICs), microchips, microcontrollers, microprocessors, all or part of a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a field-programmable gate array (FPGA), a server, a virtual server, or other circuitry suitable for executing instructions or performing logical operations. The instructions executed by the at least one processor can, for example, be pre-loaded into a memory integrated with or embedded in the controller, or can be stored in a separate memory. The memory can include random access memory (RAM), read-only memory (ROM), a hard disk, an optical disk, magnetic media, flash memory, other permanent, fixed, or volatile memory, or any other mechanism capable of storing instructions. In some embodiments, the at least one processor can include more than one processor. Each processor can have a similar construction, or the processors can have different constructions that are electrically connected or disconnected from each other. For example, the processors can be separate circuits or integrated in a single circuit. When more than one processor is used, the processors can be configured to work independently or collaboratively and can be co-located or located remotely from each other. The processors can be coupled by electrical, magnetic, optical, acoustic, mechanical, or other means that allow them to interact. As discussed in more detail herein, Figure 1 processing unit 112 in Figure 4 and processing device 400 in
[0163] In some embodiments, the at least one processor may determine facial skin micromovements by applying optical reflection analysis. The term "optical reflection analysis" refers to evaluating the properties of a surface by analyzing the pattern of light scattered from the surface. When light strikes a surface (e.g., facial skin), some of it is absorbed, some is transmitted, and some is reflected. The amount and type of light that is reflected depends on the properties of the surface and the angle at which the light strikes the surface. In one example, when using an incoherent light source, optical reflection analysis may include scatter analysis that involves measuring the scattering of light from a surface (e.g., facial skin). In another example, when using a coherent light source, optical reflection analysis may include speckle analysis or any pattern-based analysis. For example, coherent light that strikes a rough, undulating, or textured surface may be reflected or scattered in many different directions, creating a pattern of bright and dark regions known as "speckle". A computer (e.g., including a processor) may be used to perform such analysis to identify the speckle pattern and obtain information about the surface (e.g., facial skin) represented in the reflected signal received from at least one light detector. The speckle pattern may occur because the interfering coherent light waves add together to give a composite wave with intensity variations. The detected speckle pattern or any other detected pattern may then be processed to generate reflected image data. As discussed in more detail herein, Figure 7 The optical reflection processing module 706 depicted in
[0164] According to the present disclosure, the reflected image data can be processed by any image processing algorithm, including classical and / or artificial neural network (ANN)-based algorithms such as convolutional neural network (CNN), recurrent neural network (RNN). In some examples, the reflected image data can be preprocessed by transforming the image data using a transformation function to obtain a transformed speckle image. For example, the transformed reflected image data can include one or more convolutions of the speckle image. The transformation function can include one or more image filters such as low-pass filter, high-pass filter, band-pass filter, all-pass filter, etc. In some examples, the transformation function can include a non-linear function. In some examples, at least part of the reflected image data can be smoothed, for example, by using Gaussian convolution, using a median filter, etc., so as to preprocess the reflected image data. In some examples, the reflected image data can be preprocessed to obtain different representations of the reflected image data. For example, the reflected image data can include: a representation of at least part of the reflected image data in the frequency domain; a discrete Fourier transform of at least part of the reflected image data; a discrete wavelet transform of at least part of the reflected image data; a time / frequency representation of at least part of the reflected image data; a representation of at least part of the reflected image data in a lower dimension; a lossy representation of at least part of the reflected image data; a lossless representation of at least part of the reflected image data; any of the above time series; any combination of the above. In some examples, the reflected image data can be preprocessed to extract edges, and the preprocessed reflected image data can include information based on the extracted edges and / or information related to the extracted edges. In some examples, the reflected image data can be preprocessed to extract features from the reflected image data. Some examples of these features can include information related to the following: edges, corners, spots, ridges, scale-invariant feature transform (SIFT) features, temporal features, etc.
[0165] In some embodiments, performing the optical reflection analysis can include using one or more rules, functions, procedures, artificial neural networks, object detection algorithms, visual event detection algorithms, action detection algorithms, motion detection algorithms, background subtraction algorithms, inference models, etc. to evaluate the reflected image data and / or the preprocessed reflected image data. Some non-limiting examples of such inference models can include: manually pre-programmed inference models; classification models; regression models; results of training algorithms (such as machine learning algorithms and / or deep learning algorithms) on training examples, where the training examples can include examples of data instances, and in some cases, the data instances can be labeled with corresponding desired labels and / or results; and so on. In some embodiments, performing the speckle analysis can include analyzing pixels, voxels, point clouds, range data, etc. included in the reflected image data.
[0166] Some embodiments may involve analyzing reflected image data to decipher speech. The process of decrypting speech from reflected image data may involve identifying patterns or discerning features in the reflected image data. For example, known data, patterns, or signatures may be associated with certain phonemes, combinations of phonemes, words, combinations of words, or any other speech-related component. By discerning such information in the reflected image data, the speech can be decrypted. Such identification and / or decryption may be assisted by machine learning. For example, machine learning models or algorithms may be employed to identify and / or understand speech or commands. Some non-limiting examples of machine learning algorithms that may be used include classification algorithms, data regression algorithms, image segmentation algorithms, visual detection algorithms (such as object detectors, motion detectors, edge detectors, etc.), visual recognition algorithms (such as object recognition, etc.), speech recognition algorithms, mathematical embedding algorithms, natural language processing algorithms, support vector machines, random forests, nearest neighbor algorithms, deep learning algorithms, artificial neural network algorithms, convolutional neural network algorithms, recurrent neural network algorithms, linear machine learning models, non-linear machine learning models, ensemble algorithms, etc. For example, a trained machine learning algorithm may include an inference model, such as a prediction model, classification model, regression model, clustering model, segmentation model, artificial neural network (such as a deep neural network, convolutional neural network, recurrent neural network, etc.), random forest, support vector machine, etc. In some examples, training examples may include example inputs and corresponding desired outputs. Additionally, in some examples, using the training examples to train a machine learning algorithm may generate a trained machine learning algorithm, and the trained machine learning algorithm may be used to estimate an output for an input not included in the training examples. In some examples, the engineer, scientist, process, and machine that train the machine learning algorithm may further use validation examples and / or test examples. For example, validation examples and / or test examples may include example inputs and corresponding desired outputs, the trained machine learning algorithm and / or an intermediate-trained machine learning algorithm may be used to estimate the output for the example inputs of the validation examples and / or test examples, the estimated output may be compared with the corresponding desired output, and the trained machine learning algorithm and / or the intermediate-trained machine learning algorithm may be evaluated based on the results of the comparison. In some examples, a machine learning algorithm may have parameters and hyperparameters, where the hyperparameters are set manually by a person or automatically by a process external to the machine learning algorithm (such as a hyperparameter search algorithm), and the parameters of the machine learning algorithm are set by the machine learning algorithm based on the training examples. In some embodiments, the hyperparameters are set according to the training examples and validation examples, and the parameters are set according to the training examples and the selected hyperparameters.
[0167] In some examples, decrypting speech based on reflected image data can involve a trained machine learning algorithm that serves as an inference model to generate an inferred output when provided with an input. For example, the trained machine learning algorithm can include a classification algorithm, the input can include a sample, and the inferred output can include a classification of the sample. In another example, the trained machine learning algorithm can include a regression model, the input can include a sample, and the inferred output can include an inferred value of the sample. In yet another example, the trained machine learning algorithm can include a clustering model, the input can include a sample, and the inferred output can include an assignment of the sample to at least one cluster. In additional examples, the trained machine learning algorithm can include a classification algorithm, the input can include an image, and the inferred output can include a classification of an item depicted in the image. In yet another example, the trained machine learning algorithm can include a regression model, the input can include an image, and the inferred output can include an inferred value of an item depicted in the image (such as an estimated facial skin movement, etc.). In further examples, the trained machine learning algorithm can include an image segmentation model, the input can include an image, and the inferred output can include a segmentation of the image. In yet another example, the trained machine learning algorithm can include an object detector, the input can include an image, and the inferred output can include one or more detected objects in the image and / or one or more locations of the objects within the image. In some examples, the trained machine learning algorithm can include one or more formulas and / or one or more functions and / or one or more rules and / or one or more procedures, the input can be used as an input to the formula and / or function and / or rule and / or procedure, and the inferred output can be based on the output of the formula and / or function and / or rule and / or procedure (e.g., selecting one of the outputs of the formula and / or function and / or rule and / or procedure, using a statistical measure of the output of the formula and / or function and / or rule and / or procedure, etc.). As discussed in more detail herein, Figure 6 the reflected image 600 in
[0168] In some embodiments, an artificial neural network can be configured to analyze an input and generate a corresponding output. Some non-limiting examples of such artificial neural networks can include shallow artificial neural networks, deep artificial neural networks, feedback artificial neural networks, feedforward artificial neural networks, autoencoder artificial neural networks, probabilistic artificial neural networks, time-delay artificial neural networks, convolutional artificial neural networks, recurrent artificial neural networks, long / short-term memory artificial neural networks, and the like. In some examples, an artificial neural network can be configured manually. For example, the structure of the artificial neural network can be manually selected, the type of artificial neurons of the artificial neural network can be manually selected, the parameters of the artificial neural network (such as the parameters of the artificial neurons of the artificial neural network) can be manually selected, and so on. In some examples, machine learning algorithms can be used to configure the artificial neural network. For example, a user can select hyperparameters for the artificial neural network and / or the machine learning algorithm, and the machine learning algorithm can use the hyperparameters and training examples to determine the parameters of the artificial neural network, such as using backpropagation, using gradient descent, using stochastic gradient descent, using mini-batch gradient descent, and the like. In some examples, an artificial neural network can be created from two or more other artificial neural networks by combining the two or more other artificial neural networks into a single artificial neural network.
[0169] The disclosed embodiments may include and / or access a data structure or data. A data structure according to the present disclosure may include any collection of data values and the relationships between them. For example, a data structure may contain the correlation of facial micro-movements with words or phonemes, and at least one processor may perform a lookup of a particular word or phoneme associated with detected facial skin micro-movements in the data structure. Data may be stored linearly, horizontally, hierarchically, relationally, non-relationally, one-dimensionally, multi-dimensionally, operably, in an ordered manner, in an unordered manner, in an object-oriented manner, in a centralized manner, in a decentralized manner, in a distributed manner, in a customized manner, or in any manner enabling data access. As a non-limiting example, a data structure may include an array, an associative array, a linked list, a binary tree, a balanced tree, a heap, a stack, a queue, a set, a hash table, a record, a tagged union, an ER model, and a graph. For example, a data structure may include an XML database, an RDBMS database, a SQL database, or a NoSQL alternative for data storage / search, such as MongoDB, Redis, Couchbase, DatastaxEnterprise Graph, ElasticSearch, Splunk, Solr, Cassandra, Amazon DynamoDB, Scylla, HBase, and Neo4J. A data structure may be a component of the disclosed system or a remote computing component (e.g., a cloud-based data structure). The data in the data structure may be stored in contiguous or non-contiguous memory. Additionally, as used herein, a data structure does not require the information to be located in the same place. The information may be distributed across multiple servers, e.g., servers owned or operated by the same or different entities. Thus, the phrase "data structure" used herein in the singular includes plural data structures. As discussed in more detail herein, Figure 1 the data structure 124 in Figure 4 and the data structures 422 and 464 in
[0170] Consistent with the present application, at least one processor may generate an output associated with the determined facial skin micromovements. The phrase "generate an output" broadly refers to issuing a command, issuing data, and / or causing any type of electronic device to initiate an action. In some embodiments, the output may be sound (e.g., delivered via a speaker configured to fit in a user's ear), and the sound may be an audible rendering of words associated with silent speech or pre-articulatory speech. In one example, the audible rendering of words may include a response to a question that the user silently asks a virtual personal assistant. In another example, the audible rendering of words may include synthetic speech (e.g., the artificial generation of human speech). According to other disclosed embodiments, the output may be directed to a display (e.g., a visual display such as a computer monitor, a television, a mobile communication device, VR or XR glasses, or any other device that enables visual perception), and the generated output may include a graphical, image, or text rendering of words (e.g., subtitles) associated with pre-articulatory or articulatory speech. The text rendering of words may be presented while the words are being articulated. In other embodiments, the output may be directed to a communication device associated with the user, and the generated output may be any data exchanged with the communication device. The phrase "communication device" is intended to include all possible types of devices capable of exchanging data using a network configured to transmit data. In some examples, the communication device may include a smart phone, a tablet device, a smart watch, a personal digital assistant, a desktop computer, a laptop computer, an Internet of Things (IoT) device, a dedicated terminal, a wearable communication device, and any other device that enables data communication. As discussed in more detail herein, Figure 7 the output determination module 712 in Figure 7 is a non-limiting example of a software module for generating an output associated with the determined facial skin micromovements.
[0171] The disclosed embodiments may relate to using a network to exchange data (e.g., text data). The phrase "communication network" or simply "network" may include any type of physical or wireless computer network arrangement for exchanging data. For example, the network may be the Internet, a private data network, a virtual private network using a public network, a Wi-Fi network, a LAN or WAN network, a combination of one or more of the foregoing, and / or any other suitable connection that can enable the exchange of information between the various components of the system. In some embodiments, the network may include one or more physical links for exchanging data, such as Ethernet, coaxial cable, twisted pair cable, optical fiber, or any other suitable physical medium for exchanging data. The network may also include a public switched telephone network ("PSTN") and / or a wireless cellular network. The network may be a secure network or an insecure network. In other embodiments, one or more components of the system may communicate directly via a dedicated communication network. Direct communication may use any suitable technology, including for example BLUETOOTH TM , BLUETOOTH LE TM (BLE), Wi-Fi, near field communication (NFC), or any other suitable communication method that provides a medium for exchanging data and / or information between separate entities. As discussed in more detail herein, Figure 1 the communication network 126 shown is a non-limiting example of a communication network according to the present disclosure.
[0172] As used herein, a non-transitory computer-readable storage medium (or a similar construct such as a non-transitory computer-readable medium) refers to any type of physical memory that can store information or data readable by at least one processor. Examples include random access memory (RAM), read-only memory (ROM), volatile memory, non-volatile memory, a hard disk drive, a CD ROM, a DVD, a flash drive, a magnetic disk, any other optical data storage medium, any physical medium having a hole pattern, a mark, or other readable elements, a PROM, an EPROM, a FLASH-EPROM, or any other flash memory, NVRAM, a cache, a register, any other memory chip or cartridge, and network versions thereof. The phrases "memory" and "computer-readable storage medium" may refer to multiple structures, such as multiple memories or computer-readable storage media located within a wearable device or at a remote location. Additionally, one or more computer-readable storage media may be used to implement a computer-implemented method. Thus, the term computer-readable storage medium should be understood to include tangible articles and to exclude carrier waves and transient signals.
[0173] Now refer to Figure 1 , which shows an individual 102 using a speech detection system according to some embodiments of the present disclosure. Figure 1is a single exemplary representation, and it should be understood that within the scope of the present disclosure, some of the illustrated elements may be omitted and other elements may be added. In the illustrated exemplary embodiment, the speech detection system 100 may be mounted on the head of the user 102. Specifically, the speech detection system 100 (also referred to herein simply as the "system") may have the form and appearance of an earcup over-ear headset. Alternatively, in one of many other ways within the scope of the present disclosure, the system may be wearable, including in-ear earphones, integrated into or attachable to the temple of glasses, a headband, or any other mechanism capable of securing the system or a portion thereof to a human head. The speech detection system 100 may be configured to direct projected light 104 (e.g., coherent light) toward corresponding locations on the face of the user 102, thereby creating an array of light spots 106 that extends over the facial region 108 of the face. The facial region 108 may have an area of at least 1 cm2, at least 2 cm 2 , at least 4 cm 2 , at least 6 cm 2 or at least 8 cm 2 . In some embodiments, the size of the facial region 108 may be determined such that the movement of different parts of the facial muscles can be sensed. In the depicted example, only one beam of the projected light 104 is shown; however, it is contemplated that each light spot projected toward the facial region 108 may be associated with a corresponding beam or with one or more beams. In other embodiments, the light source may project light in a manner other than an array of light spots. For example, the region of the face may be illuminated uniformly or non-uniformly.
[0174] For a wearable embodiment, the speech detection system 100 may include a wearable housing 110 configured to be worn on the head of the user 102. The wearable housing 110 may include a processing unit 112 or be associated with a processing unit 112 configured to interpret facial skin micromovements; an output unit 114 configured to fit into the user's ear and present an auditory and / or vibration output; and an optical sensing unit 116 configured to project light toward a non-lip portion of the face of the user 102 and detect the reflection of the projected light. In the illustrated example, the optical sensing unit 116 may be connected to the output unit 114 by an arm 118 and may thus be held in a position proximate to and / or facing the user's face. According to some disclosed embodiments, the optical sensing unit 116 does not contact the skin of the user at the facial region 108; instead, the optical sensing unit 116 may be held at a distance from the skin surface of the facial region 108. The distance of the optical sensing unit 116 from the skin surface may be at least 5 mm, at least 7.5 mm, at least 10 mm, at least 15 mm, or at least 20 mm.
[0175] The optical sensing unit 116 can be configured to receive the reflection of light 104 from the facial region 108 and output an associated reflection signal. Specifically, the reflection signal can indicate a light pattern (e.g., a secondary speckle pattern) that may occur due to the reflection of coherent light from each light spot 106 within the field of view of the speech detection system 100. To cover a sufficiently large facial region 108, the detector of the speech detection system 100 can have a wide field of view. For example, the field of view can have an angular width of at least 60°, at least 70°, or at least 90°. Within this field of view, the speech detection system 100 can sense and process signals that reflect light patterns from all of the light spots 106 or only a subset of the light spots 106. For example, the processing unit 112 can select a subset of the light spots 106 that is determined to give the maximum amount of useful and reliable information about the relevant movement of the skin surface of the user 102 and can avoid processing data from other light spots 106. Additional details of the structure and operation of the optical sensing unit 116 are described below with reference to FIG. 5.
[0176] According to the present disclosure, even if the user 102 does not vocalize speech or make any other sound, the speech detection system 100 can be capable of detecting the facial skin micromovements of the user 102 and extracting meaning from the detected movements. The extracted meaning can be an identification of the user 102 wearing the speech detection system 100, an identification of the user's silent reading (such as words silently spoken by the user 102), an identification of the words spoken aloud by the user 102, an identification of the phonemes silently spoken by the user 102, or an identification of the phonemes spoken aloud by the user 102. Similarly, the extracted meaning can include an identification of the heart rate of the user 102, an identification of the breathing rate of the user 102, and / or other characteristics associated with the verbal or non-verbal communication of the user 102. In one example, the speech detection system 100 can generate an output signal that includes data associated with identification information, UI commands, a synthesized audio signal, a text transcription, or any combination thereof. In one example, the synthesized audio signal can be played back to the user 102 via the speaker in the output unit 114. This playback can be useful in giving the user 102 feedback regarding the speech output.
[0177] According to the present disclosure, the speech detection system 100 may exchange data (e.g., output signals) with various communication devices associated with a user (e.g., the mobile communication device 120 or the server 122). The term "communication device" is intended to include all possible types of devices capable of exchanging data using a digital communication network, an analog communication network, or any other communication network configured to transmit data. In some examples, the communication device may include wearable communication devices such as smart phones, tablet computers, smart watches, personal digital assistants, laptop computers, IoT devices, dedicated terminals, industrial machinery, vehicles, smart homes, appliances, or any other electronic device capable of exchanging information or data with another electronic device. In other examples, the communication device may include non-wearable communication devices such as desktop computers, smart home hubs, routers, servers, or any other network-connected equipment. In some cases, the processing device of the mobile communication device 120 or the server 122 may supplement or replace some of the functions of the processing unit 112 of the speech detection system 100. In some embodiments, the output signal generated by the speech detection system 100 may be transmitted via a communication link to the mobile communication device 120 or the cloud server. The term "cloud server" refers to a computer platform that provides services via a network such as the Internet. In Figure 1 In the example embodiment shown, the server 122 may use one or more virtual machines that may not correspond to individual hardware. For example, the computing and / or storage capabilities may be achieved by allocating an appropriate portion of the desired computing / storage capabilities from a scalable repository such as a data center or a distributed computing environment. In one example configuration, the server 122 may be a cloud server that determines the neural activity of the user 102 based on facial skin micromovements. In one example, the server 122 may implement the methods described herein using custom hardwired logic, one or more application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), firmware, and / or program logic that, in combination with a computer system, cause the server 122 to be a special purpose machine.
[0178] In some embodiments, server 122 may access data structure 124 to determine, for example, the correlation between words and multiple facial movements. Data structure 124 may utilize volatile or non-volatile, magnetic, semiconductor, tape, optical, removable, non-removable, other types of storage devices, or tangible or non-transitory computer-readable media, or any medium or mechanism for storing information. Data structure 124 may be part of server 122 or separate from server 122, as shown. When data structure 124 is not part of server 122, server 122 may exchange data with data structure 124 via a communication link. Data structure 124 may include one or more memory devices that store data and instructions for performing one or more features of the disclosed methods. In one embodiment, data structure 124 may include any of a plurality of suitable data structures, ranging from small data structures hosted on a workstation to large data structures distributed among multiple data centers. Data structure 124 may also include any combination of one or more data structures controlled by a memory controller device (e.g., a server) or software. According to the present disclosure, speech detection system 100 may communicate with mobile communication device 120 or server 122 using communication network 126 as defined above.
[0179] Now referring to Figure 2A , which shows another example embodiment of speech detection system 100 according to the present disclosure. In this example, wearable housing 110 may be integrated with or otherwise attached to a pair of glasses 200 having a frame 202. In this example embodiment, glasses 200 may include a nasal electrode 204 and a temporal electrode 206 attached to frame 202 and contacting the user's skin surface. Electrodes 204 and 206 may receive surface electromyogram (sEMG) signals, which provide additional information about the activation of the user's facial muscles. Speech detection system 100 may use the electrical activity sensed by electrodes 204 and 206 together with the output of optical sensing unit 116 to generate, for example, a synthetic audio signal. Additionally or alternatively, speech detection system 100 may include one or more additional optical sensing units 208 similar to optical sensing unit 116 for sensing skin movements in other regions of the user's face, such as eye movements. These additional optical sensing units may be used together with or instead of optical sensing unit 116. In the example shown, optical sensing unit 116 may illuminate a first facial region 108A, and optical sensing unit 208 may illuminate a second facial region 108B. First facial region 108A and second facial region 108B may be non-overlapping.
[0180] In some disclosed embodiments, a speech detection system may be incorporated into, integrated with, or otherwise attached to an extended reality device. As used herein, the phrase "extended reality device" may include any type of device or system that enables a user to perceive and / or interact with an extended reality environment. The phrase "extended reality environment" refers to all types of real and virtual combined environments and human-machine interactions that are at least partially generated by computer technology. A non-limiting example of an extended reality environment may be a virtual reality (VR) environment. A virtual reality environment may be an immersive simulated non-physical environment that provides a user with the perception of being present in the virtual environment. Another non-limiting example of an extended reality environment may be an augmented reality (AR) environment. An augmented reality environment may involve a real-time direct or indirect view of a physical reality world environment enhanced with virtual computer-generated perceptual information (e.g., virtual objects with which a user may interact). Another non-limiting example of an extended reality environment is a mixed reality (MR) environment. A mixed reality environment may be a mixture of a physical real world and a virtual environment, where physical and virtual objects may coexist and interact in real time. Examples of extended reality devices may include VR headsets, AR headsets, MR headsets, smart glasses, and wearable projection devices.
[0181] Now refer to Figure 2B , which shows another example embodiment of a speech detection system 100 according to some embodiments of the present disclosure. In the depicted example, the speech detection system 100 may be part of an extended reality device 250. The extended reality device 250 may include all the sensors and the like discussed above with reference to the glasses 200. For example, the extended reality device 250 may include a gyroscope, an accelerometer, a magnetometer, an image sensor, a depth sensor, an infrared sensor, a proximity sensor, and / or any other one or more sensors configured to measure one or more characteristics associated with an individual wearing the extended reality device 250 and generate an output related to the one or more measured characteristics. In some cases, the speech detection system 100 may use the input from any of the sensors of the extended reality device 250 to determine the spoken or unspoken words expressed by the individual 102. For example, the speech detection system 100 may use the input from the image sensor of the extended reality device 250 and the data from the optical sensing unit 116 (see Figure 1 ) to extract the meaning of facial movements. In other cases, the extended reality device 250 may generate an output including a visual and / or audible presentation associated with the words detected by the speech detection system 100. For example, the individual 102 may interact with the extended reality device 250 using silent commands.
[0182] Now refer to Figure 3 , Figure 3Shows another exemplary implementation of the speech detection system 100 according to the present disclosure. In Figure 3 In the illustrated implementation, the speech detection system 100 can be integrated with the mobile communication device 120. Specifically, the mobile communication device 120 can include a light detector configured to detect the reflection 300 of light from the facial region 108. In this example, the light projected onto the facial region 108 is from the non-wearable light source 302, which can be a coherent light source or an incoherent light source. In some configurations, the non-wearable light source 302 can be included in the mobile communication device 120. Alternatively, the non-wearable light source 302 can be separate from the mobile communication device 120.
[0183] According to the present disclosure, and as Figure 3 depicted, the pattern of light projected onto the facial region 108 can be a single light spot 106 large enough to illuminate different parts of the facial region 108. For example, the light spot 106 can include a first part 304A associated with the first facial muscle and a second part 304B associated with the second facial muscle. Thereafter, the processing device of the mobile communication device 120 can apply optical reflection analysis to the received reflection 300 to determine facial skin micromovements. Specifically, the processing device of the mobile communication device 120 can determine the first facial skin micromovement of the first part 304A and the second facial skin micromovement of the second part 304B. The processing device can use both the first facial skin micromovement and the second facial skin micromovement to extract meaning (e.g., determine speech or commands, or authenticate the user 102) and generate an output. When the extracted meaning includes continuous authentication of the user 102, the exemplary implementation of the speech detection system 100 shown in Figure 3 can be used. Specifically, the speech detection system 100 can provide an authentication service that uses the biometrics of facial micromovements for continuous authentication during the use of the mobile communication device 120.
[0184] Figure 4 is a block diagram of an exemplary configuration of the speech detection system 100 and an exemplary configuration of the remote processing system 450. It should be noted that Figure 4is merely a representation of one implementation, and it should be understood that within the scope of the present disclosure, some of the illustrated elements may be omitted and other elements may be added. In the depicted implementation, the speech detection system 100 includes: a processing unit 112, which includes a processing device 400 and a memory device 402; an output unit 114, which includes a speaker 404, a light indicator 406, and a haptic feedback device 408; an optical sensing unit 116, which includes at least one light source 410 and at least one light detector 412; an audio sensor 414; a power source 416; one or more additional sensors 418; a network interface 420; and a data structure 422. The speech detection system 100 can directly or indirectly access a bus 424 (or any other communication mechanism) that interconnects the above subsystems and components to transfer information and commands within the speech detection system 100. Some of the above-listed subsystems and components are mentioned herein in the singular, but may be plural in alternative configurations. For example, in some configurations, the speech detection system 100 may include multiple light sources 410 or multiple light detectors 412.
[0185] Figure 4 The illustrated processing device 400 may comprise any physical device or group of devices having circuitry that performs logical operations on one or more inputs. Instructions executed by at least one processor may, for example, be pre-loaded into a memory integrated with or embedded in the processing device 400, or may be stored in a separate memory (e.g., the memory device 402 or the data structure 422). As described above, the processing device may include more than one processor. Each processor may have a similar construction, or the processors may have different constructions that are electrically connected or disconnected from each other. For example, the processors may be separate circuits or integrated in a single circuit. When more than one processor is used, the processors may be configured to work independently or collaboratively, and may be co-located or located remotely from each other. The processors may be coupled by electrical, magnetic, optical, acoustic, mechanical, or other means that allow them to interact. According to the present disclosure, at least some of the functions described below with respect to the processing device 400 may be performed by a processing device of a remote processing system 450.
[0186] Figure 4The memory device 402 shown in Figure 7 may include high-speed random access memory and / or non-volatile memory, such as one or more disk storage devices, one or more optical storage devices, and / or flash memory (e.g., NAND, NOR). According to the present disclosure, the components of the memory device 402 may be distributed across more than one unit of the speech detection system 100 and / or more than one memory device. In particular, the memory device 402 may be used to store software products and / or data stored on a non-transitory computer-readable medium. As described above, the terms "memory" and "computer-readable storage medium" may refer to multiple structures, such as multiple memories or computer-readable storage media located within the speech detection system 100 or at a remote location (e.g., at the remote processing system 450). Additionally, one or more computer-readable storage media may be used to implement a computer-implemented method. The following refers to
[0187] Figure 4 The output unit 114 shown in
[0188] Figure 4 may cause output from various output devices, such as the speaker 404, the light indicator 406, and the haptic feedback device 408. Examples of the speaker 404 may include loudspeakers, in-ear headphones, audio headsets, hearing aid-type devices, bone conduction headphones, and any other device capable of converting an electrical audio signal into a corresponding sound, or may be incorporated therewith. In some embodiments, the speaker 404 may be configured to allow only the user 102 to listen to the generated audio signal. Alternatively, the speaker 404 may be configured to emit sound into the open air for anyone nearby to hear. The light indicator 406 may include one or more light sources, e.g., an array of LEDs associated with different colors. The light indicator 406 may be used to indicate the battery status of the speech detection system 100 or to indicate its operating mode. The haptic feedback device 408 may include a vibration motor, a linear actuator, a vibration transducer, or any other force feedback device that provides a haptic or tactile cue or is capable of converting an electrical signal into a corresponding vibration or application of force.
[0188] Figure 4The optical sensing unit 116 shown may include a light source 410 and a light detector 412. The light source 410 may project coherent or incoherent light onto the facial region 108. As described above, the light source 410 may be a laser such as a solid-state laser, a laser diode, a high-power laser, or an alternative light source such as a light-emitting diode (LED)-based light source. Additionally, the light source 410 may emit light in different formats, such as light pulses, continuous wave (CW), quasi-CW, etc. In one embodiment, the light source 410 may be an infrared laser diode configured to emit an input beam of coherent radiation. The light source 410 may be associated with a beam-splitting element, such as a Dammann grating or another suitable type of diffractive optical element (DOE), for splitting the input beam into a plurality of output beams that form corresponding light spots 106 at a position matrix extending over the facial region 108. In another embodiment (not shown in the figures), the light source 410 may include a plurality of laser diodes or other emitters that generate corresponding sets of output beams that cover different respective sub-regions within the facial region 108. In one embodiment, the processing unit 112 may select and activate only a subset of the emitters and not activate all of the emitters. For example, to reduce the power consumption of the speech detection system 100, the processing unit 112 may activate only one emitter or a subset of two or more emitters that illuminate a specific region on the user's face that has been found to provide the most useful information for generating the desired speech output.
[0189] Figure 4 The light detector 412 shown therein may be used to detect reflections from the facial region 108 indicative of facial skin movement. As described above, the light detector may be capable of measuring properties of coherent or incoherent light, such as power, frequency, phase, pulse timing, pulse duration, and other properties. In some embodiments, the light detector 412 may include an array of detection elements, e.g., a set of charge-coupled device (CCD) sensors and / or a set of complementary metal-oxide semiconductor (CMOS) sensors, and have an objective lens optics for imaging the facial region 108 onto the array. Due to the small size of the optical sensing unit 116 and its proximity to the skin surface, the light detector 412 may have a wide enough field of view to detect many light spots 106 at high angles of at least 60°, at least 70°, or at least 90°. The light detector 412 may be configured to generate an output related to the measured properties of the detected light. According to the present disclosure, the output of the light detector 412 may include any form of data determined in response to light reflections received from the facial region 108. In some embodiments, the output may include a reflected signal that includes an electronic representation of one or more characteristics determined based on coherent or incoherent light reflections. In other embodiments, the output may include the raw measurements detected by at least one light detector 412.
[0190] In some embodiments, the light detector 412 can measure one or more optical properties associated with skin changes. The phrase "skin changes" refers to any detectable movement, alteration, or modification that occurs in the skin. Such skin changes can include changes in the epidermis (i.e., the outermost layer of the skin), the dermis (i.e., the middle layer of the skin), the subcutaneous tissue (i.e., the deepest layer of the skin), and deeper muscle tissue. The optical properties can be measured without contacting the skin of the individual 102. Examples of one or more optical properties of the reflected light that can be measured by the light detector 412 can include intensity, frequency, reflection, angle, sharpness, bidirectional reflectance distribution function, color, brightness, glossiness, transparency, opacity, surface texture, surface undulation, surface movement, and other optical properties that can be obtained from the analysis of light reflection. The output of the light detector 412 can be used to determine information associated with the skin changes. In some embodiments, the information associated with those skin changes can be obtained from changes in the distance from the skin to the detector when the skin moves, and in other embodiments, the changes can be obtained without relying on changes in the distance between the skin and the light detector 412. For example, by detecting changes in non-distance measurements (e.g., image sharpness) over time, the determination speed or angular velocity of facial skin changes can be determined. Thus, in one non-limiting example, optical properties can be detected based on the random intensity changes observed when coherent light interacts with a rough or scattering surface such as human skin. In another non-limiting example, optical properties can be detected based on, for example, the interference of light waves when an interference pattern is used to measure the phase difference or amplitude change between two or more optical paths.
[0191] In some embodiments, the optical sensing unit 116 may not require the parameters of a reference light source (such as the wavelength, intensity, or coherence of the light source), and may not require a reference beam (commonly used with a beam splitter) to measure one or more optical properties of the reflected light. For example, the optical sensing unit 116 can use a single beam to illuminate the skin and then process the light reflection that returns to the light detector 412. While some speech detection systems may include a single pixel sensor (e.g., a photodiode), in other embodiments, the light detector 412 can include one or more multi-pixel sensors (e.g., each pixel sensor includes more than 4 million pixels, more than 10 million pixels, or more than 10 million pixels), which enables the generation of an image that provides spatial information beyond a single point. For example, a reflected image as depicted in Figure 6 can be generated based on the output of the light detector 412. As described throughout this disclosure, image processing can be used to analyze the output of the light detector 412 to determine the pattern of light scattered from the surface. For example, the characteristics of secondary speckles can be determined.
[0192] In some non - limiting examples, the optical sensing unit 116 can split the outgoing light beam into multiple light beams using a diffraction element and can cause interference without relying on the superposition of coherent light waves. In some non - limiting examples, the optical sensing unit 116 can be arranged such that the light detector 412 can be positioned along a different optical axis from the light source 410. In other non - limiting examples, aligning the light source and the sensor along the same optical axis can be used to maintain coherence, achieve path - length matching, ensure spatial overlap, and maintain the sensitivity and accuracy of the interference pattern. However, since some implementations of the light detector 412 detect reflected images rather than the distance to a point, the optical sensing unit 116 can include a first optical axis for the outgoing light and a second optical axis for the incoming light that is not aligned with the first optical axis. In some embodiments, the light detector 412 is configured to measure sub - microbic speed and depth changes in the range of 5 - 500 micrometers. In alternative embodiments, the light detector 412 is configured to measure changes less than one micrometer. All examples provided in this paragraph are alternative and can be implemented in many alternative embodiments provided herein, depending on the implementation details.
[0193] Figure 4The audio sensor 414 shown in FIG. may include one or more audio sensors configured to capture audio by converting sound into digital information. Some examples of audio sensors may include microphones, unidirectional microphones, bidirectional microphones, cardioid microphones, omnidirectional microphones, on-board microphones, wired microphones, wireless microphones, or any combination of the foregoing. The audio sensor 414 may be configured to capture sound emitted by the user 102 such that the user 102 can use the speech detection system 100 as a conventional headset when needed. Additionally or alternatively, the audio sensor 414 may be used in conjunction with the silent speech sensing capabilities of the speech detection system 100. In one embodiment, the audio signal output by the audio sensor 414 may be used to change the operating state of the speech detection system 100. For example, the processing unit 112 may generate a voice output only when the audio sensor 414 does not detect the user 102's vocalization of a word. In another embodiment, the audio sensor 414 may be used in a calibration process, wherein the optical sensing unit 116 detects micro-movements of the skin when the user 102 emits certain phonemes or words. The processing unit 112 may compare the reflected signal output by the light detector 412 with the sound sensed by the audio sensor 414 to calibrate the optical sensing unit 116. The calibration may include prompting the user 102 to move the position of the optical sensing unit 116 to align the optical components at a desired position relative to the facial region 108. In yet another embodiment, the audio sensor 414 enables the immediate training of the neural network of the speech detection system 100. For example, the speech detection system 100 may be configured to associate facial skin micro-movements with words using audio signals captured simultaneously with the micro-movements. After identifying the recorded words, the speech detection system 100 may perform a review to identify the facial micro-movements prior to the expression of those words, thereby training the speech detection system 100. In a similar manner, the speech detection system may be trained for expressions, commands, user identification, and emotions.
[0194] Figure 4 The power supply 416 shown may provide electrical energy to power the speech detection system 100. The power supply may include any device or system that can store, distribute, or transfer electrical power, including but not limited to one or more batteries (e.g., lead-acid batteries, lithium-ion batteries, nickel-metal hydride batteries, nickel-cadmium batteries), one or more capacitors, one or more connections to an external power source, one or more power converters, or any combination of the foregoing. Referring Figure 4 to the example shown, the power supply 416 may be removable, which means that the speech detection system 100 may be wearable. The mobility of the power supply enables the user 102 to use the speech detection system 100 in various situations. In other embodiments, the power supply 416 may be associated with a connection to an external power source (such as the power grid), which may be used to charge the power supply 416.
[0195] Figure 4 The additional sensor 418 shown in FIG. may include various sensors, such as an image sensor, a motion sensor, an environmental sensor, an electromyography (EMG) sensor, a resistance sensor, an ultrasonic sensor, a proximity sensor, a biometric sensor, or other sensing devices configured to facilitate related functions. For example, the speech detection system 100 may include one or more image sensors configured to capture visual information from the environment of the user 102 by converting light (not emitted from the light source 410) into image data. According to the present disclosure, the image sensor may be included in any device or system capable of detecting light signals in the near-infrared, infrared, visible, and / or ultraviolet spectra and converting them into electrical signals. Examples of image sensors may include digital cameras, semiconductor charge-coupled devices (CCDs), active pixel sensors in complementary metal-oxide semiconductors (CMOS), or N-type metal-oxide semiconductors (NMOS, Live MOS). The electrical signals may be used to generate image data. According to the present disclosure, the image data may include pixel data streams, digital images, digital video streams, data obtained from captured images, and data that can be used to construct one or more 3D images, 3D image sequences, 3D videos, or virtual 3D representations. The image data acquired by the one or more image sensors may be sent to the processing unit 112 or the remote processing system 450 via wired or wireless transmission.
[0196] The speech detection system 100 may also include one or more motion sensors configured to measure the motion of the user 102. Specifically, the motion sensor may perform at least one of the following: detecting the motion of the user 102; measuring the speed of the user 102; measuring the acceleration of the user 102; or measuring any other action related to motion. In some embodiments, the motion sensor may include one or more accelerometers configured to detect changes in acceleration (e.g., intrinsic acceleration) and / or measure the acceleration of the speech detection system 100. In some embodiments, the motion sensor may include one or more gyroscopes configured to detect changes in the orientation of the speech detection system 100 and / or measure information related to the orientation of the speech detection system 100. In some embodiments, the motion sensor may include one or more of an image sensor, a lidar sensor, a radar sensor, or a proximity sensor. For example, by analyzing the captured images, the processing device 400 may determine the motion of the speech detection system 100, for example, using a self-motion algorithm. In addition, the processing device may determine the motion of an object in the environment of the speech detection system 100, for example, by object tracking.
[0197] The speech detection system 100 may also include one or more different types of environmental sensors configured to capture data reflecting the environment of the user 102. In some embodiments, the environmental sensors may include one or more chemical sensors configured to perform at least one of the following: measuring chemical properties in the environment of the user 102; measuring changes in chemical properties in the environment of the user 102; detecting the presence of chemical substances in the environment of the user 102; and / or measuring the concentration of chemical substances in the environment of the user 102. Examples of measurable chemical properties include: pH level; toxicity; and temperature. Examples of measurable chemical substances or phenomena include: electrolytes; specific enzymes; specific hormones; specific proteins; smoke; carbon dioxide; carbon monoxide; oxygen; ozone; hydrogen; and hydrogen sulfide. In other embodiments, the environmental sensors may include one or more temperature sensors configured to detect changes in the temperature of the environment of the user 102 and / or measure the temperature of the environment of the user 102. In other embodiments, the environmental sensors may include one or more barometers configured to detect changes in the atmospheric pressure in the environment of the user 102 and / or measure the atmospheric pressure in the environment of the user 102. In other embodiments, the environmental sensors may include one or more light sensors configured to detect changes in the ambient light in the environment of the user 102.
[0198] Figure 4 The illustrated network interface 420 may provide two-way data communication to a network such as communication network 126. In one embodiment, network interface 420 may include an Integrated Services Digital Network (ISDN) card, a cellular modem, a satellite modem, or a modem that provides a data communication connection over the Internet. As another example, network interface 420 may include a Wireless Local Area Network (WLAN) card. In another embodiment, network interface 420 may include an Ethernet port connected to a radio frequency receiver and transmitter and / or an optical (e.g., infrared) receiver and transmitter. The specific design and implementation of network interface 420 may depend on the one or more communication networks through which the speech detection system 100 is intended to operate. For example, in some embodiments, the speech detection system 100 may include a network interface 420 designed to operate over a GSM network, a GPRS network, an EDGE network, a Wi-Fi or WiMax network, and a Bluetooth network. In any such embodiment, network interface 420 may be configured to send and receive electrical, electromagnetic, or optical signals carrying digital data streams or digital signals representing various types of information.
[0199] Figure 4The data structure 422 shown in [figure] can include any hardware, software, firmware, or combination thereof for storing and facilitating retrieval of information from a database. The term "database" can be understood to include a collection of data that may or may not be distributed. The database can include a database management system that controls the organization, storage, and retrieval of data contained within the database. As described above, the data included in the database can be stored linearly, horizontally, hierarchically, relationally, non-relationally, one-dimensionally, multi-dimensionally, operationally, in an ordered manner, in an unordered manner, in an object-oriented manner, in a centralized manner, in a decentralized manner, in a distributed manner, in a customized manner, or in any manner that enables data access. In the disclosed embodiments, the data structure 422 can include correlations of facial micromovements with words, commands, emotions, expressions, and / or biological conditions. At least one processor can perform lookups in the data structure to interpret detected facial skin micromovements. According to one embodiment, at least some of the data stored in the data structure 422 can alternatively or additionally be stored in the remote processing system 450.
[0200] According to the present disclosure, the speech detection system 100 can be configured to communicate with a remote processing system 450 (e.g., a mobile communication device 120 or a server 122). The remote processing system 450 can directly or indirectly access a bus 452 (or other communication mechanism) that interconnects subsystems and components for transferring information within the remote processing system 450. For example, the bus 452 can interconnect a memory interface 454, a network interface 456, a power supply 458, a processing device 460, one or more additional sensors 462, a data structure 464, and a memory device 466.
[0201] Figure 4 The memory interface 454 shown in [figure] can be used to access software products and / or data stored on a non-transitory computer-readable medium or other memory device (such as the memory devices 402, 466, the data structure 422, or the data structure 464). The memory device 466 can contain software modules for performing the processes according to the present disclosure. In a particular embodiment, the memory device 466 can include a shared memory module 472, a node registration module 473, a load balancing module 474, one or more computing nodes 475, an internal communication module 476, an external communication module 477, and a database access module (not shown). The modules 472 to 477 can contain software instructions for execution by at least one processor (e.g., the processing device 460) associated with the remote processing system 450. The shared memory module 472, the node registration module 473, the load balancing module 474, the computing module 475, and the external communication module 477 can cooperate to perform various operations.
[0202] The shared memory module 472 may allow information sharing between the remote processing system 450 and other devices associated with one or more speech detection systems 100. In some embodiments, the shared memory module 472 may be configured to enable the processing device 460 to access, retrieve, and store data. For example, using the shared memory module 472, the processing device 460 may perform at least one of the following: execute a software program stored on the memory devices 402, 466, data structures 422, or data structures 464; store information in the memory devices 402, 466, data structures 422, or data structures 464; or retrieve information from the memory devices 402, 466, data structures 422, or data structures 464.
[0203] The node registration module 473 may be configured to track the availability of one or more computing nodes 475. In some examples, the node registration module 473 may be implemented as: a software program (such as a software program executed by one or more computing nodes 475); a hardware solution; or a combined software and hardware solution. In some embodiments, the node registration module 473 may communicate with one or more computing nodes 475, for example, using the internal communication module 476. In some examples, one or more computing nodes 475 may notify the node registration module 473 of their status, for example, by sending a message: at startup, at shutdown, at a constant interval, at a selected time, in response to a query received from the node registration module 473, or at any other determined time. In some examples, the node registration module 473 may query the status of one or more computing nodes 475, for example, by sending a message: at startup, at a constant interval, at a selected time, or at any other determined time.
[0204] The load balancing module 474 can be configured to divide the workload among one or more computing nodes 475. In some examples, the load balancing module 474 can be implemented as a software program (such as a software program executed by one or more of the computing nodes 475), a hardware solution, or a combined software and hardware solution. In some embodiments, the load balancing module 474 can interact with the node registration module 473 to obtain information about the availability of one or more computing nodes 475. In some embodiments, the load balancing module 474 can communicate with one or more computing nodes 475, for example, using the internal communication module 476. In some examples, one or more computing nodes 475 can notify the load balancing module 474 of their status, for example, by sending a message: at startup, at shutdown, at a constant interval, at a selected time, in response to a query received from the load balancing module 474, or at any other determined time. In some examples, the load balancing module 474 can query the status of one or more computing nodes 475, for example, by sending a message: at startup, at a constant interval, at a preselected time, or at any other determined time.
[0205] The internal communication module 476 can be configured to receive information from and / or send information to one or more components of the remote processing system 450. For example, control signals and / or synchronization signals can be sent and / or received through the internal communication module 476. In one embodiment, input information for a computer program, output information of a computer program, and / or intermediate information of a computer program can be sent and / or received through the internal communication module 476. In another embodiment, the information received through the internal communication module 476 can be stored in the memory device 466 or the data structure 464. For example, the internal communication module 476 can be used to transmit information retrieved from the data structure 464. In another example, a reference signal reflecting the facial micro-movements of the user 102 can be stored in the data structure 464 and accessed using the internal communication module 476.
[0206] The external communication module 477 can be configured to receive and / or transmit information from and / or to one or more speech detection systems 100. For example, control signals can be transmitted and / or received via the external communication module 477. In one embodiment, the information received via the external communication module 477 can be stored in the memory device 466, the data structure 464, and / or any memory device in one or more speech detection systems 100. In another embodiment, the information retrieved from the data structure 464 can be transmitted to the speech detection system 100 or to any entity communicating with the user 102 using the external communication module 477. For example, when the user 102 communicates with a financial institution (e.g., a bank), the information retrieved from the data structure 464 can be transmitted to effectuate the authentication of the user 102. In another embodiment, sensor data can be transmitted and / or received using the external communication module 477. Examples of such input data can include data received from the speech detection system 100, information captured from the environment of the user 102 using one or more sensors such as the additional sensor 418 and the additional sensor 462.
[0207] In some embodiments, aspects of the modules 472-477 can be implemented by hardware, software (including in one or more signal processing and / or application specific integrated circuits), firmware, or any combination thereof, and can be executed by one or more processors individually or in various combinations with each other. Specifically, the modules 472-477 can be configured to interact with each other and / or with other modules of the speech detection system 100 to perform functions in accordance with the disclosed embodiments. The memory device 466 can include additional modules and instructions or fewer modules and instructions.
[0208] Figure 4 The network interface 456, the power supply 458, the processing device 460, the additional sensor 462, and the data structure 464 shown can share similar functions with the corresponding elements in the speech detection system 100, as described above. The specific design and implementation of the above components can vary based on the implementation of the remote processing system 450. Additionally, the remote processing system 450 can include more or fewer components. For example, when the remote processing system 450 is a mobile communication device associated with the user 102 (e.g., the mobile communication device 120), it can include a speaker, a microphone, and additional sensors.
[0209] As Figure 4The components and arrangements of the speech detection system 100 and the remote processing system 450 shown are not intended to limit the disclosed embodiments. As will be understood by those skilled in the art benefiting from this disclosure, many variations and / or modifications can be made to the depicted configurations of the speech detection system 100 and the remote processing system 450. For example, in all cases, not all components are essential for the operation of the input unit. Any component can be located in any suitable part of the speech detection system 100 or the remote processing system 450. Additionally, the components can be rearranged into various configurations while providing the functionality of the disclosed embodiments. For example, some speech detection systems may not include all of the elements shown in the speech detection system 100 and the remote processing system 450. Other speech detection systems can include additional components and still fall within the scope of this disclosure.
[0210] Figure 5A and Figure 5B include two schematic diagrams of the optical sensing unit 116 in detecting facial skin micromovements according to some embodiments of the present disclosure. These two schematic diagrams illustrate simplified scenarios before and after muscle recruitment. As depicted, the optical sensing unit 116 can include an illumination module 500, a detection module 502, and an optional audio sensor 414. As discussed above and shown in FIG. 5, the optical sensing unit 116 can be configured to not touch the skin at the facial region 108 of the user, but can be held at a distance D from the skin surface of the facial region 108. The distance D of the optical sensing unit 116 from the skin surface can be at least 5 mm, at least 7.5 mm, at least 10 mm, at least 15 mm, or at least 20 mm.
[0211] In the depicted embodiment, the illumination module 500 includes a light source 410 (e.g., an infrared laser diode) configured to generate an input beam 504. The illumination module 500 also includes a beam splitting element 506 (e.g., a Dammann grating or another suitable type of diffractive optical element (DOE)) configured to split the input beam 504 into a plurality of output beams 508 that form corresponding light spots 106A - 106E on a pattern (e.g., a position matrix) extending over the facial region 108. In an alternative embodiment (not shown in the figures), the illumination module 500 can include a plurality of light sources 410 that generate corresponding sets of output beams 508 covering different corresponding sub-regions within the facial region 108. In this alternative embodiment, the processing unit 112 can selectively activate only a subset of the plurality of light sources without activating all of them. For example, to reduce the power consumption of the speech detection system 100, the processing unit 112 can activate only one light source or a group of two or more light sources that illuminate only a portion of the facial region 108.
[0212] The detection module 502 may include a light detector 412, and the light detector 412 may include an optical sensor array 510 (e.g., a CMOS image sensor array) having an objective lens optics 512 for obtaining the reflection 300 of the coherent light from the facial region 108. Due to the small size of the optical sensing unit 116 and its proximity to the skin surface, the detection module 502 may be configured to have a wide field of view so as to acquire the reflections of many light spots 106 at large angles. As described above, the field of view of the light detector 412 may have an angular width of at least 60°, at least 70°, or at least 90°. Due to the roughness of the skin surface, the light patterns at the light spots 106 can also be detected at these large angles.
[0213] The speech detection system 100 may analyze the light reflection 300 to determine the facial skin micromovements caused by the recruitment of the muscle fibers 520. Determining the facial skin micromovements may include determining the amount of skin movement, determining the direction of skin movement, and / or determining the acceleration of skin movement. The determined facial skin micromovements may include the voluntary and / or involuntary recruitment of the muscle fibers 520. The muscle fibers 520 may be a part of: the zygomaticus muscle; the orbicularis oris muscle; the risorius muscle; the genioglossus muscle; or the levator labii alaeque nasi muscle. The processing device 400 may be configured to perform a first light spot analysis on the light reflected from a first region of the face near the light spot 106A to determine that the first region has moved a distance d1, i.e., the first facial skin micromovement 522A; and perform a second light spot analysis on the light reflected from a second region of the face near the light spot 106E to determine that the second region has moved a distance d2, i.e., the second facial skin micromovement 522B. Thereafter, the processing device 400 may use the determined movements of the first and second regions to identify at least one spoken word. According to the disclosed embodiments, the distances d1 and d2 may be less than 1000 micrometers, less than 100 micrometers, less than 10 micrometers, or smaller.
[0214] Figure 6 is a schematic diagram of a reflection image 600 associated with the light reflection 300, which is received from a region of the facial region 108 associated with a single light spot 106 (e.g., the light spot 106A depicted in FIG. 5). In the disclosed embodiments, the processing device 400 may receive a reflection signal indicating the coherent light reflection from the facial region 108. The reflection signal may be represented by the reflection image 600. Thereafter, the processing device 400 may determine the facial skin micromovements by applying light reflection analysis. When the light source 410 is a coherent light source, the light reflection analysis may include speckle analysis or any pattern-based analysis. Such analysis may be performed by the processing device 400 or the processing device 460 to identify the speckle pattern and obtain the movement of the corresponding region of the facial region 108.
[0215] In the depicted example, speckles 602 appear in the reflected image 600 after the recruitment of muscle fibers 520. The detected speckles or any other detected pattern can then be processed to generate reflected image data. Referring to the example discussed above, assuming that the reflected image 600 reflects the light spot 106A, the reflected image data can include data indicating the movement distance d1 of the first region. In some cases, the reflected image data can be processed by any image processing algorithm (e.g., CNN and RNN) to determine the skin movement of at least two regions within the facial region 108. Thereafter, the processing device 400 can use one or more machine learning (ML) algorithms and artificial intelligence (AI) algorithms to decrypt the reflected image data and extract meaning from the facial skin micromovements.
[0216] As Figure 7 shown, the memory device 700 can contain software modules to perform the processes according to the present disclosure. In particular, the memory device 700 can include an illumination control module 702, a sensor communication module 704, a light reflection processing module 706, an artificial neural network (ANN) training module 710, a lip-reading decryption module 708, an output determination module 712, and a database structure access module 714. The disclosed embodiments are not limited to any specific configuration of the memory 700. Additionally, the processing device 400 and / or the processing device 460 can execute instructions in any of the modules 702-714 included in the memory device 700. It should be understood that references to the processing device in the following discussion can refer to the processing device 400 of the speech detection system 100 and / or the processing device 460 of the remote processing system 450, either individually or jointly. Thus, steps of any of the following processes associated with the modules 702-714 can be performed by one or more processors associated with the speech detection system 100.
[0217] According to the disclosed embodiments, the illumination control module 702, the sensor communication module 704, the light reflection processing module 706, the lip-reading decryption module 708, the ANN training module 710, the output determination module 712, and the database access module 714 can cooperate to perform various operations. For example, the illumination control module 702 can determine the light characteristics for illuminating the facial region 108. The sensor communication module 704 can receive the coherent light reflection from the facial region 108 and output an associated reflection signal. The light reflection processing module 706 can process the reflection signal to determine the facial skin micromovements. The lip-reading decryption module 708 and the database access module 714 can cooperate to extract meaning from the facial skin micromovements (e.g., determine the silently spoken word). In some cases, the ANN training module 710 can use the determined silently spoken word and the determined facial skin micromovements to train the artificial network. The output determination module 712 can generate a presentation of the determined word.
[0218] The illumination control module 702 can adjust the operation of the light source 410 to illuminate the facial region 108. In some embodiments, the illumination control module 702 can determine values of the characteristics of the projected light 104, such as light intensity, pulse frequency, duty cycle, illumination pattern, luminous flux, or any other optical characteristic. In a particular embodiment, as long as the user 102 is not speaking, the speech detection system 100 can operate in a first illumination mode (e.g., low frame rate) to conserve the power of its battery. While the speech detection system 100 is operating in this first illumination mode, it can process images to detect at least one trigger (e.g., movement of the face) in the reflected signal that indicates speech. When such a trigger is detected, the illumination control module 702 can cause the coherent light source to operate in a second illumination mode (e.g., high frame rate) to enable the detection of changes in the coherent light pattern (e.g., speckles) that occur due to silent speech. The illumination control module 702 can also be configured to change one or more characteristics of the projected light 104 based on various types of triggers. Various types of triggers can be detected by analyzing data from the sensor communication module 704.
[0219] The sensor communication module 704 can adjust the operation of the light detector 412, the audio sensor 414, and the additional sensors 418 to receive captured measurements from one or more sensors integrated with or connected to the speech detection system 100. In one embodiment, the sensor communication module 704 can use the signals received from one or more sensors to generate sensor data associated with the user 102. In one example, the sensor communication module 704 can receive a reflected signal from the light detector 412 and can generate a first data stream of the reflected image, based on which facial skin micromovements in the facial region can be determined. In another example, the sensor communication module 704 can receive an audio signal from the audio sensor 414 and can generate a second data stream, based on which the words spoken by the voice of the user 102 can be determined. In another example, the sensor communication module 704 can receive a motion signal from a motion sensor included in the additional sensors 418 and generate a third data stream, based on which the activities participated in by the user 102 can be determined. The sensor communication module 704 can transmit the sensor data to other software modules for processing.
[0220] The light reflection processing module 706 can process the sensor data received from the sensor communication module 704 to prepare for speech decryption. In one embodiment, the light reflection processing module 706 can receive a reflection signal from the sensor communication module 704, the reflection signal indicating the coherent light reflection from the facial area 108 originating from the light detector 412. The reflection signal can be represented by a reflection image (e.g., reflection image 600), which can be processed by at least one image processing algorithm to extract skin movement at a set of pre-selected locations on the face of the user 102. The number of locations to be examined can be an input to the image processing algorithm. In some cases, the location on the skin extracted for coherent light processing can be taken from a list of points of interest. The list of points of interest specifies anatomical locations corresponding to the zygomaticus, orbicularis oris, risorius, genioglossus, or levator labii superioris nasalis. In layman's terms, the list of points of interest can include specific points at the cheek above the mouth, at the chin, at the mandible, at the cheek below the mouth, at the high cheek, and at the back of the cheek. According to the present disclosure, the list of points of interest can be dynamically updated with more points on the face extracted during the training phase. The entire set of positions can be sorted in descending order so that any subset of the list (in order) minimizes the word error rate (WER) relative to the selected number of examined positions. In another embodiment, the light reflection processing module 706 can crop each coherent light spot extracted from the original image frame around the coherent light spot, and the algorithm processes only the cropped image. Generally, the process of coherent light spot processing involves reducing the size of the full frame image pixels (about 1.5MP) received from the sensor communication module 704 by two orders of magnitude with a very short exposure. The exposure can be dynamically set and adjusted to be able to capture only the coherent light reflections instead of the skin segments. The cropped image of the coherent light spot can depict the coherent light pattern. In other embodiments, the light reflection processing module 706 can apply an image processing algorithm to the reflection image. For example, the light reflection processing module 706 can improve the contrast of the image by removing noise using a threshold to determine black pixels and calculating a characteristic metric of the coherent light (such as a scalar speckle energy metric, such as average intensity). In addition, the light reflection processing module 706 can analyze the temporal changes in the reflection pattern (e.g., average speckle intensity). Alternatively, other metrics may be used, such as detection of specific coherent light patterns. Thereafter, the light reflection processing module 706 may assign a sequence of values of characteristic metrics of coherent light, which may be calculated frame by frame and aggregated to generate reflection image data indicating facial skin micro-movements. The light reflection processing module 706 may transmit the reflection image data indicating facial skin micro-movements to other software modules for processing.
[0221] The silent decryption module 708 can use machine learning (ML) algorithms and artificial intelligence (AI) algorithms to decrypt the reflected image data indicating the facial skin micromovements received from the light reflection processing module 706. According to the present disclosure, decrypting the reflected image data can include extracting meaning from the detected facial skin micromovements. In one embodiment, the silent decryption module 708 can use a trained ANN to associate words with the facial skin micromovements. Different types of ANNs can be used, such as a classification NN that outputs a final word and a sequence-to-sequence NN that outputs a sentence (a sequence of words). In some embodiments, during the normal speech of the user, the system 100 can sample the speech and facial movements of the user 102 simultaneously. Automatic speech recognition (ASR) and natural language processing (NLP) algorithms can be applied by the silent decryption module 708 to the actual speech, and the results of these algorithms can be used to optimize the parameters of the algorithms used by the silent decryption module 708. These parameters can include the weights of various neural networks and the spatial distribution of the laser beams for optimal performance. In addition, the silent decryption module 708 can limit the output of the algorithms to a predetermined set of words, which can significantly improve the accuracy of word detection in the case of ambiguity (i.e., when two different words result in similar micromovements on the facial skin). The set of words used can be personalized over time, adjusting the lexicon to the actual words used by a specific user and their respective frequencies and contexts. In addition, the silent decryption module 708 can use the context of the conversation between the user 102 and the callee. The context can be determined based on the input of the word and sentence extraction algorithms to improve the accuracy by eliminating out-of-context options. The context of the conversation can be understood by applying automatic speech recognition (ASR) and natural language processing (NLP) algorithms on both the user 102 side and the callee side.
[0222] According to an embodiment of the present disclosure, the ANN training module 710 can be used to train an ANN to perform silent speech decryption. To train the ANN (e.g., the ANN that can be used by the silent reading decryption module 708), thousands of examples may be required. To achieve this, the ANN training module 710 can rely on a large group of people (e.g., a group of reference human subjects). In one example, the silent reading decryption module 708 can perform fine-tuning on the ANN such that it is customized for the user 102. In this way, within a few minutes or less of wearing the speech detection system 100, the silent reading decryption module 708 can be ready to decrypt facial skin micro-movements. The ANN training module 710 can be used to train two different types of ANNs: a classification neural network that outputs words finally and a sequence-to-sequence neural network that outputs sentences (sequences of words). For this purpose, the ANN training module 710 can upload training data from the memory, such as the silent speech data collected from multiple reference human subjects received from the optical reflection processing module 706. The silent speech data can be collected from a variety of people (people of different ages, genders, races, physical disabilities, etc.). It should be noted that the number of examples required for learning and generalization can be task-related. For word / utterance prediction (within a closed group), at least thousands of examples can be collected. Thereafter, the ANN training module 710 can enhance the image-processed training data to obtain more artificial data for the training process. In particular, the enhanced data can include the image-processed coherent light patterns, and some image processing steps are described herein. The data enhancement process can include the following steps: (i) time dropout, where the amplitude at random time points is replaced by zero; (ii) frequency dropout, where the signal is transformed into the frequency domain and random frequency blocks are filtered out; (iii) clipping, where the maximum amplitude of the signal at random time points is limited. This clipping can add a saturation effect to the data; (iv) adding noise, where Gaussian noise is added to the signal, and variable speed, where the signal is resampled to achieve a slightly lower or slightly faster signal.
[0223] The enhanced dataset can undergo a feature extraction process. In this process, the ANN training module 710 can calculate the time-domain silent speech features. For this purpose, for example, each signal can be split into a low-frequency component x_low and a high-frequency component x_high, and windowed to create time frames, e.g., using a frame length of 27 ms and a shift of 10 ms. For each frame, five time-domain features and nine frequency-domain features can be calculated, a total of 14 features for each signal. Specifically, the time-domain features can be represented as follows:
[0224]
[0225] Among them, ZCR is the zero crossing rate. Additionally, in this example, the amplitude values used are from a 16-point short-time Fourier transform, i.e., the frequency domain features and all features are normalized to zero mean and unit variance.
[0226] Thereafter, the ANN training module 710 can divide the data into a training set, a validation set, and a test set. The training set can be the data used to train the model. The validation set can be used to complete hyperparameter tuning, and the test set can be used to complete the final evaluation. The model architecture can be task-related. Two different examples describe training two networks for two conceptually different tasks. The first task can include signal transcription, i.e., converting silent speech into text by generating words, phonemes, or letters. This first task can be solved by using a sequence-to-sequence model. The second task can include predicting words or utterances, i.e., classifying the utterances spoken by a user into a single category within a closed group. This second task can be solved by using a classification model. The disclosed sequence-to-sequence model can consist of an encoder and a decoder. The encoder can transform the input signal into a high-level representation (embedding), while the decoder generates a language output (i.e., characters or words) from the encoded representation. The input into the encoder can be a sequence of feature vectors. In one example, the input can enter the first layer (temporal convolutional layer) of the encoder, which can downsample the data to achieve good performance. The model can use one hundred such convolutional layers.
[0227] In some embodiments, at each time step, the output from the temporal convolutional layer can be passed to three layers of a bidirectional recurrent neural network (RNN). The ANN training module 710 can employ long short-term memory (LTSM) as the unit in each RNN layer. Each RNN state can be a concatenation of the state of the forward RNN and the state of the backward RNN. The decoder RNN can be initialized with the final state of the encoder RNN (a concatenation of the final state of the forward encoder RNN and the first state of the backward encoder RNN). At each time step, the decoder RNN can receive the previous word as input, which is one-hot encoded and embedded in a 150-dimensional space through a fully connected layer. The decoder RNN output can be projected into the space of words or phonemes (depending on the training data) through a matrix. The sequence-to-sequence model can adjust the next-step prediction based on the previous prediction. During learning, the log probability can be maximized:
[0228]
[0229] Among them, y is the previously predicted true value. The classification neural network can be composed of an encoder such as in a sequence-to-sequence network and an additional fully connected classification layer on top of the encoder output. The output can be projected into the space of closed words, and the scores can be converted into probabilities for each word in the lexicon. The result of the above entire process can include two types of trained ANNs (represented by the calculated coefficients). The coefficients can be stored in data structures associated with the speech detection system 100 (e.g., data structure 422 and data structure 464). In daily use, the ANN training module 710 can receive the latest coefficients for the ANN to be trained. The first ANN task can be signal transcription, that is, converting silent speech into text through word / phoneme / letter generation. The second ANN task can be word / utterance prediction, that is, classifying the utterances spoken by the user into a single category within a closed group.
[0230] The output determination module 712 can adjust the operations of the output unit 114 and the network interface 420 to generate outputs using the speaker 404, the light indicator 406, the haptic feedback device 408, and / or send data to a remote computing device. In some embodiments, the outputs generated by the output determination module 712 can include various types of outputs associated with the silent speech determined based on the detected facial skin micromovements. Specifically, the output determination module 712 can synthesize the pronunciation of the words determined by the silent reading decryption module 708 based on the facial skin movements. The synthesis can simulate the voice of the user 102 or simulate the voice of someone other than the user 102 (e.g., the voice of a celebrity or a preselected template voice). The pronunciation of the words can be presented via the speaker 404 or transmitted to the remote computing device via the network interface 420. Alternatively, the output determination module 712 can generate a text output based on the facial skin movements by the silent reading decryption module 708. The text output can be transmitted to the remote computing device via the network interface 420. According to another embodiment, the outputs generated by the output determination module 712 can relate to the operation of the speech detection system 100. In some cases, the light indicator 406 can include a light indicator showing the battery status of the speech detection system 100. For example, when the speech detection system 100 has a low battery, the light indicator can start to blink. Additional examples of the types of outputs that can be generated by the output determination module 712 are described throughout this disclosure.
[0231] The database access module 714 can cooperate with data structures 422 and 464 to retrieve stored data. The retrieved data can include, for example, the correlations between multiple words and multiple facial skin movements, the correlations between a specific individual and multiple facial skin micromovements associated with the specific individual, and so on. As described above, the silent speech decryption module 708 can use the trained ANN to perform silent speech decryption. The trained ANN can extract meanings from the detected facial skin micromovements using the data stored in data structures 422 and 464. Data structures 422 and 464 can include separate databases, including, for example, vector databases, raster databases, tile databases, viewport databases, and / or user input databases. The data stored in data structures 422 and 464 can be received from modules 702-712 or other components of the speech detection system 100. In addition, the data stored in data structures 422 and 464 can be provided using data input, data transmission, or data upload as inputs.
[0232] Modules 702-714 can be implemented in software, hardware, firmware, a hybrid of any of these, and so on. The processing devices of the speech detection system 100 and the remote processing system 450 can be configured to execute the instructions of modules 702-714. In some embodiments, aspects of modules 702-714 can be implemented by hardware, software (including in one or more signal processing and / or application specific integrated circuits), firmware, or any combination thereof, capable of being executed by one or more processors individually or in various combinations with each other. Specifically, modules 702-714 can be configured to interact with each other and / or with other modules associated with the speech detection system 100 to perform the functions according to the disclosed embodiments.
[0233] Today, image-based facial recognition technology is commonly used as a biometric authentication method in many communication devices. It allows users to use their faces as a unique identifier to unlock their devices, make payments, and access applications or accounts. However, image-based facial recognition technology is not always reliable and has limitations that may make it less effective in certain situations. For example, image-based facial recognition systems may be affected by factors such as poor lighting conditions, low-quality images, and occlusions such as masks or accessories. These factors may result in inaccurate or incomplete matches. In addition, image recognition algorithms may exhibit biases, leading to misidentifications based on various factors such as race, gender, or age. Moreover, false positives and false negatives are prevalent problems in image-based face recognition technology; thus, individuals may be misidentified as others or not identified at all. The following disclosure presents a new and improved technical solution for providing reliable biometric authentication, which can overcome the inherent defects of image-based facial recognition technology.
[0234] Some disclosed embodiments of the present disclosure may be configured to detect facial skin micromovements of an individual, identify the individual using the detected facial skin micromovements, and determine an action to initiate based on the identification of the individual.
[0235] The following description refers to Figures 8 to 10 exemplary embodiments for using facial skin micromovements to identify an individual according to some disclosed embodiments. Figures 8 to 10 It is only intended to facilitate the conceptualization of exemplary embodiments for performing the operation of using facial skin micromovements to identify an individual, and does not limit the present disclosure to any specific embodiment.
[0236] Some disclosed embodiments include a head-mounted system for using facial skin micromovements to identify an individual. According to the present disclosure, a head-mounted system can be understood to include any component or combination of components attachable to the head, as illustrated and described elsewhere in the present disclosure. The phrase "identifying an individual" refers to the process of determining whether an individual is known to the system. Specifically, the identification process may involve comparing the detected features of the individual with the known features of the individual to identify, verify, or authenticate the individual. According to the present disclosure, an individual can be identified based on the facial skin micromovements of the individual. The phrase "facial skin micromovements" can be understood as described and illustrated elsewhere in the present disclosure. In some cases, the head-mounted system can access data indicating reference facial skin micromovements and use the data to determine whether the individual currently using the head-mounted system is the same individual associated with the reference facial skin micromovements. Depending on the embodiment, the probability of misidentifying an individual based on his / her facial skin micromovements described below can be less than one in ten thousand, less than one in one hundred thousand, or less than one in one million.
[0237] Some disclosed embodiments include a wearable housing configured to be worn on an individual's head. The phrase "wearable housing" can be understood as described and illustrated elsewhere in the present disclosure. According to some disclosed embodiments, the head-mounted system includes at least one coherent light source associated with the wearable housing. The phrase "coherent light source" can be understood as described and illustrated elsewhere in the present disclosure. The phrase "associated with the wearable housing" can refer to any component linked, incorporated, attached, connected, or related to the wearable housing. For example, the light source can be mounted to the wearable housing with screws, adhesives, clips, heat and pressure, or any other known means of attaching two elements. Alternatively, the light source can be partially or completely contained within the housing. In alternative embodiments, the light source can be associated with the housing by a wired or wireless connection. Figure 4 The light source 410 in
[0238] According to some disclosed embodiments, at least one coherent light source can be configured to project light towards a facial region of the head. Projecting coherent light can include radiating coherent light in a direction towards a portion of the face. The coherent light can be a monochromatic wave that has a well-defined phase relationship on its wavefront in a defined direction (such as towards the facial region of the head). The facial region of the head refers to any anatomical portion above the shoulders of a human body. The facial region can include at least some of the following: forehead, eyes, cheeks, ears, nose, mouth, chin, and neck. Examples of the facial region are shown in Figures 1 to 3 (e.g., facial region 108). For example, as shown in Figure 1 and FIG. 2, the coherent light source 410 included in the optical sensing unit 116 is attached to the wearable housing 110 and can direct light towards the facial region. The head-mounted system can also include at least one detector associated with the wearable housing. The terms "detector" and "associated with the wearable housing" can be understood as described and illustrated elsewhere in this disclosure. The at least one detector can be configured to receive the coherent light reflection from the facial region and output a related reflection signal. Receiving the coherent light reflection can refer to detecting, acquiring, obtaining, or otherwise measuring the electromagnetic wave (e.g., in the visible or invisible spectrum) reflected from the facial region and impinging on the at least one detector. Outputting the related reflection signal can include sending, transmitting, generating, and / or providing information representing or corresponding to the coherent light reflection. For example, projecting coherent light onto non-moving facial skin can result in a first reflection signal indicating the coherent light reflection. However, even a micro-movement of the facial skin may cause the at least one detector to output a second reflection signal different from the first reflection signal. The change between the first reflection signal and the second reflection signal can be used to determine a specific facial skin micro-movement. As an example, Figure 4 the light detector 412 in is associated with the wearable housing 110 and is used to determine facial skin micro-movements.
[0239] According to some disclosed embodiments, the head-mounted system includes at least one processor. The term "processor" can be understood as described and illustrated elsewhere in this disclosure. A processor can be employed to provide some or all of the functions described herein. Figure 4 The processing device 400 in is an example of at least one processor provided for the purpose of implementing at least some of the functions described herein.
[0240] Some disclosed embodiments include analyzing reflected signals to determine specific facial skin micromovements of an individual. The term "analyze" refers to examining, investigating, scrutinizing, and / or studying. The reflected signals can be analyzed to determine if they are recognized or if they are related to other information. For example, the reflected signals (or a data set derived from the reflected signals) can be analyzed, such as to determine correlations, associations, patterns within the data set or relative to different data sets, or the lack thereof. Specifically, one or more processing techniques (such as light pattern analysis (as described and illustrated elsewhere in this disclosure)) can be used, for example, to analyze the reflected signals received from at least one detector. Other processing techniques can include convolution, fast Fourier transform, edge detection, pattern recognition, object detection algorithms, clustering, artificial intelligence, machine and / or deep learning, and any other processing techniques for determining specific facial skin micromovements of an individual. In some examples, training examples can be used to train a machine learning model to determine facial skin micromovements based on reference reflection data. Examples of such training examples can include sample reflection data streams, and labels indicating the associated facial skin micromovements. The trained machine learning model can be used to analyze received reflected signals relative to the reference reflection data to determine facial skin micromovements. In some examples, at least a portion of the reflected signals can be analyzed to compute a convolution of at least a portion of the reflected signals, thereby obtaining a result value of the computed convolution. Additionally, in response to the result value of the computed convolution being a first value, a first facial skin micromovement can be determined, and in response to the result value of the computed convolution being a second value, a different second facial skin micromovement can be determined. For example, the reflected signals received from at least one detector can be analyzed as described elsewhere in this disclosure, and facial skin micromovements associated with the question "what is my mom’s birthday?" can be determined. Additional details and examples of how at least one processor analyzes reflected signals to determine specific facial skin micromovements are described herein with reference to the light reflection processing module 706.
[0241] According to some disclosed embodiments, at least some specific facial skin micromovements in the facial region can include micromovements less than 100 micrometers or less than 50 micrometers. In other words, the output of the process for determining specific facial skin micromovements can be accurate enough to distinguish changes in facial skin in the range of 10 to 100 micrometers. In some embodiments, these changes can be detected within a time period of 0.01 to 0.1 seconds. In some disclosed embodiments, the determined specific facial skin micromovements can correspond to facial expressions (e.g., smiling, frowning, worried) or facial muscle movements corresponding to physiological events (e.g., sneezing, laughing, yawning). In other embodiments, the facial skin micromovements can correspond to pre-articulatory or articulatory phonemes, syllables, words, or phrases, as described below. In other embodiments, the facial skin micromovements can correspond to biological processes, such as pulse or respiration rate. In further embodiments, the facial skin micromovements can correspond to a combination of one or more of the foregoing.
[0242] According to some disclosed embodiments, specific facial skin micromovements can correspond to pre-articulatory muscle recruitment. As described elsewhere herein, pre-articulation or subvocalization refers to the influence of facial muscle movements in the absence of audible articulation or prior to the occurrence of articulation. The facial skin micromovements correspond to pre-articulatory muscle recruitment, which is the direct or indirect cause of the facial skin micromovements. In some cases, pre-articulatory muscle recruitment may cause facial skin micromovements prior to the onset of articulation. For example, pre-articulatory muscle recruitment can occur between 0.1 second and 0.5 seconds prior to actual articulation. In some cases, pre-articulatory muscle recruitment can include voluntary muscle recruitment that occurs when an individual begins to articulate a word. In other cases, pre-articulatory muscle recruitment can include involuntary facial muscle recruitment that occurs when certain craniofacial muscles prepare for articulation.
[0243] According to some disclosed embodiments, specific facial skin micromovements can correspond to muscle recruitment during the articulation of at least one word or a portion thereof. For example, at least one word can correspond to a predefined expression, password, or secret passphrase. As described above, actual articulation depends on whether air is expelled from the lungs and into the throat. Without such an air stream, no sound is produced. Since pre-articulatory muscle recruitment occurs before and separately from the muscles that convey the air stream, pre-articulatory muscle recruitment can occur with or without a subsequent articulation.
[0244] Figure 8 An exemplary speech detection process is shown. In the example shown, the speech detection system 100 can analyze the reflected signal associated with the question "what is my mom’s birthday?" to determine the specific facial skin micromovements 800 associated with the unknown individual 802.
[0245] Some disclosed embodiments include accessing a memory that associates multiple facial skin micromovements with an individual. The phrase "accessing a memory" refers to retrieving or examining electronically stored information. This can occur, for example, by communicating with or connecting to an electronic device or component in which data is electronically stored. For the purpose of reading the stored data (e.g., obtaining relevant information) or for the purpose of writing new data (e.g., storing additional information), such data can be organized, for example, in a data structure. In some cases, the accessed memory can be part of a speech detection system or part of a remote processing device (e.g., a cloud server) accessible by the speech detection system. In some examples, at least one processor can access the memory, for example, at startup, at shutdown, at a constant interval, at a selected time, in response to a query received from at least one processor, or at any other determined time. The memory can store data associating multiple facial skin micromovements with an individual. The stored data can be any electronic representation of the facial skin micromovements, any electronic representation of one or more characteristics determined from the facial skin micromovements, or an original measurement signal detected by at least one light detector and representing the facial skin micromovements. Associating multiple facial skin micromovements with an individual can include storing in the memory or data structure a relationship between the facial skin micromovements and an identifier of the individual. This can allow for the efficient retrieval and identification of individuals based on these relationships. For example, the memory can be associated with a built-in mechanism for linking or associating facial skin micromovements with an identifier of the individual. In one example, correlations between specific phonemes, syllables, words, or phrases and associated skin micromovements can be stored. Depending on the embodiment, these correlations can be unique to an individual or specific to a group or subgroup associated with the individual. (For example, micromovements associated with certain parts of speech can vary between individuals, countries, dialects, or based on different regional accents). Associating multiple facial skin micromovements with an individual can occur by any of the above examples. If the intention is to verify the personal identity of a specific individual, a comparison can be made to a database of correlations associated with that specific individual (e.g., based on a sample previously captured from that individual). Or, if the intention is to identify an individual as part of a group or subgroup, pre-stored data associated with that group or subgroup can be accessed.
[0246] According to the present disclosure, the fact that multiple facial skin micromovements are associated with an individual means that the multiple facial skin micromovements can uniquely identify the individual or identify the individual as part of a particular group or subgroup. In one exemplary implementation for uniquely identifying an individual, the probability that the multiple facial skin micromovements are the same for two different individuals can be less than one in ten thousand, less than one in one hundred thousand, less than one in one million, or less than one in ten million, depending on the implementation.
[0247] According to some disclosed implementations, a memory can associate multiple facial skin movements with multiple individuals. Specifically, the memory can be designed to store the relationship between facial skin micromovements and multiple identifiers associated with multiple individuals. For example, a specific correlation for each of many individuals can be stored such that when a current signal is received, it can be compared with the various stored correlations to uniquely identify the individual associated with the stored correlation. In some disclosed implementations, for each of the multiple individuals, the memory can store at least 10, at least 50, or at least 100 data entries associated with different facial skin micromovements. In some examples, the multiple individuals can be related; for example, the multiple individuals can be family members or part of the same organization. In other examples, the multiple individuals can be unrelated but include a common attribute; for example, individuals from the same age group, or individuals associated with the same language dialect.
[0248] According to some disclosed embodiments, at least one processor may be configured to distinguish a plurality of individuals from one another based on reflected signals unique to each of the plurality of individuals. Distinguishing a plurality of individuals from one another means that at least one processor may be able to determine which individual caused the received reflected signal. For example, at least one processor may identify that a certain sentence was spoken by a particular individual and not by any other individual included in a database. The at least one processor may be configured to distinguish the plurality of individuals from one another by detecting reflected signals unique to each individual. A unique reflected signal means that no two individuals have the same reflected signal. For example, a unique reflected signal may be associated with a unique sequence of facial skin micromovements that occur when an individual vocalizes or pre-vocalizes one or more phonemes, syllables, words, or phrases, such as a passphrase. In one example, a speech detection system may be used by a group of individuals, and for each individual, the speech detection system may store personal settings. In one embodiment, at least one processor may detect first facial skin micromovements of a first individual during a first time period and second facial skin micromovements of a second individual during a subsequent second time period. When identifying the first individual using the first facial skin micromovements, at least one processor may initiate a first action (e.g., apply personal settings associated with the first individual), and when identifying the second individual using the second facial skin micromovements, at least one processor may initiate a second action (e.g., apply personal settings associated with the second individual). Alternatively, access to an application may be provided if a particular individual's relevance is recognized, while access may be denied if no relevance is recognized.
[0249] By reference Figure 8 to an example, the memory 804 may store a plurality of reference facial skin micromovements (e.g., 806A, 806B, 806C, and 806D) associated with the user 102. In the drawings, only four reference facial skin micromovements are shown, but as will be understood by those skilled in the art who benefit from the present disclosure, a greater number of reference facial skin micromovements may be stored as reference data to identify individuals. For example, the plurality of reference facial skin micromovements may be used for all known phonemes, or for at least 1,000 words. Additionally, the memory 804 may be designed to store the plurality of reference facial skin micromovements of a plurality of users, enabling the processor to distinguish a plurality of individuals from one another based on reflected signals unique to each of the plurality of individuals.
[0250] Some disclosed embodiments include searching the memory for a match between the determined specific facial skin micro - movement and at least one of a plurality of facial skin micro - movements. The phrase "searching for a match" can refer to finding one or more records that satisfy a given set of search criteria. Different types of search algorithms can be used to search for a match, such as linear search, binary search, tree - based search, and various types of database searches. Additionally, an artificial intelligence model can be employed and used to search for matches in a dataset accessible to the AI model, as described in the following paragraphs. In some cases, the initiated search can be used to discover which one of the plurality of facial skin micro - movements is most likely generated by the same individual who generated the specific facial skin micro - movement. A likelihood level or degree of certainty of the match can be determined to provide an indication of the probability or confidence that the recognition hypothesis is correct, i.e., that the reference facial skin micro - movement stored in the memory was indeed generated by the same individual who generated the specific facial skin micro - movement. In some disclosed embodiments, a match can be considered found when the degree of likelihood or certainty (by way of example only) is greater than 90%, greater than 95%, or greater than 99%.
[0251] According to the present disclosure, at least one processor can use an artificial neural network (such as a deep neural network, a convolutional neural network) to identify a match. The artificial neural network can be configured manually, using machine - learning methods, or by combining other artificial neural networks. Other ways in which at least one processor can be used to identify a match include comparing the determined specific facial skin micro - movement with the plurality of facial skin micro - movements in the memory; obtaining the difference between the determined specific facial skin micro - movement and the plurality of facial skin micro - movements in the memory and comparing the difference with a threshold; calculating at least one statistical value (e.g., mean, variance, or standard deviation) and comparing the at least one statistical value with a threshold; calculating the distance between two vectors in a multi - dimensional space, where if the distance is below a certain threshold, a match is identified; calculating the cosine of the angle between two vectors in a multi - dimensional space, where if the cosine value is above a certain threshold, a match is identified; and any other known ways of identifying a match in a database.
[0252] For reference Figure 8 As an example, searching for a match can result in a first result 808A indicating that a match has been identified and a second result 808B indicating that no match has been identified.
[0253] Some disclosed embodiments include initiating a first action if a match is identified and initiating a second action different from the first action if no match is identified. The term "initiate" can refer to performing, executing, or implementing one or more operational steps. For example, at least one processor can initiate the execution of program code instructions or cause a message to be sent to another processing device to achieve a targeted (e.g., deterministic) result or goal. The action can be an initiated response to whether a match is found between the determined specific facial skin micro-movement and multiple facial skin micro-movements in a memory. The term "action" can refer to the implementation or execution of an activity or task. For example, performing an action can include executing at least one program code instruction to implement a function or process. An action can be user-defined or system-defined (e.g., software and / or hardware) or any combination thereof. At least one processor can select which action to initiate (e.g., the first action or the second action) and can determine to initiate the selected action based on the result of the search for a match and based on various criteria. The various criteria can include user experience (e.g., preferences, such as based on context, location, environmental conditions, type of use, type of user), user requirements (e.g., context limitations, urgency or priority of the purpose behind the action), device requirements (e.g., computing power, computing limitations, rendering limitations, memory capacity or memory limitations), communication network requirements (e.g., bandwidth, latency). For example, after a match is found, a first action of sending an audio message can be initiated. An artificial voice for generating the audio message can be selected based on the various criteria listed above. The action can be initiated by at least one processor configured with a speech detection system, a different local processing device (e.g., associated with a device near the speech detection system), and / or by a remote processing device (e.g., associated with a cloud server) or any combination thereof. Thus, "initiating an action in response to a search result" can include performing or implementing one or more operations in response to a search result of a match between the determined specific facial skin micro-movement and at least one of multiple facial skin micro-movements in a memory.
[0254] According to some disclosed embodiments, a first action establishes at least one predetermined setting associated with an individual. The phrase "predetermined setting" refers to any configuration or preference associated with the operating software of the associated computing device or any other software installed on the computing device. Examples of such predetermined settings can include language settings, default actions, preferred output modes, notification types, permissions, display brightness, volume levels, default applications, network settings, and any other options that can be selected by the user. According to the present disclosure, when a match is recognized, at least one processor can formulate (i.e., specify, establish, or set) a specific setting associated with the recognized individual. The statement that a predetermined setting is associated with an individual means that data reflecting the individual's choice of the predetermined setting is stored in a database, data structure, lookup table, or linked list. In one example, the predetermined setting can manage what the speech detection system should do when it detects silent speech. Specifically, after a match is recognized, the speech detection system can automatically translate words spoken silently in English into French and synthesize them with artificial speech that sounds like the recognized individual.
[0255] According to some disclosed embodiments, a first action (i.e., when an individual is recognized) includes unlocking the computing device, and a second action (i.e., when an individual is not recognized) includes presenting a message indicating that the computing device remains locked. The computing device can be any electronic device with restricted access. For example, the computing device can be a laptop computer, PC, tablet computer, smart phone, wearable electronic device, electronic door lock, entrance gate, application, system, vehicle, communication device (e.g., mobile communication device 120). In one embodiment, the computing device can be at least a part of the speech detection system 100. The phrase "unlock the computing device" generally refers to the process of obtaining access to a device that has appropriate security mechanisms to prevent unauthorized access. For example, when an individual is recognized, at least one processor can send data (e.g., a password) to the mobile communication device 120 to unlock the mobile communication device 120. The message indicating that the computing device remains locked can be provided by the computing device or by any other device in any known manner. For example, the message can be provided audibly, textually, or virtually. For example, when an individual is not recognized, the speech detection system 100 can present a message that the mobile communication device 120 remains locked.
[0256] According to some disclosed embodiments, a first action (i.e., when an individual is recognized) provides personal information, and a second action (i.e., when an individual is not recognized) provides public information. Personal information includes data or entities specific to an individual (e.g., a user, a person, an organization, or other data owner) that may not be desired to be shared with another entity. For example, the information may include any information that could cause harm, loss, or injury to the individual or entity associated with the information if disclosed to an unauthorized entity. Some examples of personal information (e.g., sensitive data) can include identification information, location information, genetic data, information related to health, finance, business, personal, family, education, politics, religion, and / or legal matters, and / or sexual orientation or gender identity. Public information can include any information other than personal information and can be found in a public database (such as the Internet). For example, after receiving a query from an individual, the speech detection system 100 can use specific facial skin micromovements to generate a response that includes personal information (when the individual is recognized) or includes public information (when the individual is not recognized).
[0257] According to some disclosed embodiments, a first action (i.e., when an individual is recognized) authorizes a transaction, and a second action (i.e., when an individual is not recognized) provides information indicating that the transaction is not authorized. Authorizing a transaction refers to the process of granting approval or permission for an activity to occur. In some cases, authorizing a transaction may involve verifying the legitimacy of a transaction request and confirming the identity of an individual by finding a match. Examples of transactions can include financial transactions (e.g., withdrawals or deposits from a bank account, purchases or sales of goods or services using a credit card, fund transfers between accounts, bill payments, wire transfers, or electronic fund transfers), non-financial transactions (e.g., booking a flight, booking a hotel, ordering a product online, renting a car, registering for a subscription, updating an address or phone number), business transactions (e.g., ordering supplies, billing a customer for products or services provided, approving a refund, or processing an invoice), and government transactions (e.g., applying for a passport or visa, paying taxes or fines, registering a vehicle, obtaining a driver's license, obtaining a business license). When no match is found, information can be provided to indicate that the transaction is not authorized. The information can be provided via the speech detection system or via a mobile communication device. For example, when the speech detection system 100 is linked to a virtual wallet, upon receiving a payment request, the speech detection system 100 can prompt the individual to silently speak a password. Thereafter, the speech detection system 100 can use the determined specific facial skin micromovements to determine the password and compare the determined password with a previously stored password (stored in association with the user). When the determined password matches the stored password, the speech detection system 100 can authorize the payment (i.e., when the individual is recognized). Alternatively, when the determined password does not match the stored password, the speech detection system 100 can not authorize the payment (i.e., when the individual is not recognized).
[0258] Consistent with some disclosed embodiments, the first action (i.e., when the individual is recognized) permits access to the application, and the second action (i.e., when the individual is not recognized) blocks access to the application. Permitting access to the application can refer to the process of granting an individual permission to use a specific software application or use electronic hardware. The software application can be installed in the speech detection system or in any computing device associated with the individual (e.g., the individual's smart phone). For example, access to the individual's calendar application can be made in response to a detected query from the recognized individual (such as: "What was the name of the person I met with last Wednesday"). If the individual is not recognized, access to the calendar application will be prohibited, and thus the query may not be answered.
[0259] According to some disclosed embodiments, the head-mounted system includes an integrated audio output, wherein at least one of the first action or at least one of the second action includes outputting audio via the audio output. The term integrated audio output means that the head-mounted system includes internal audio hardware that is configured to generate sound without the need for an external audio interface. For example, the head-mounted system can include an audio chipset that can convert digital audio signals to analog signals and a built-in speaker or headphone jack. Additional examples of integrated audio output can include speakers, in-ear headphones, audio headsets, hearing aid type devices, and any other device capable of converting an electrical audio signal to a corresponding sound, or can be associated with speakers, in-ear headphones, audio headsets, hearing aid type devices, and any other device capable of converting an electrical audio signal to a corresponding sound. For example, the first action can be to emit sound to the open air using an audio output device (such as a speaker) for anyone nearby to hear, and the second action can be to emit sound using an audio output device (such as in-ear headphones) to allow only the individual to listen to the generated audio signal.
[0260] For reference Figure 8 As an example, when a match is found (i.e., individual 802 is recognized as user 102), the first action 810A can be initiated, and when no match is found (i.e., individual 802 is not recognized as user 102), the second action 810B can be initiated.
[0261] According to some disclosed embodiments, a match can be recognized when at least one processor determines a degree of certainty. As described elsewhere in this disclosure, the determination of the degree of certainty provides an indication of the confidence that the recognition hypothesis is correct. In other words, and for reference Figure 8, a degree of certainty provides an indication that the unknown individual 802 is the user 102. According to some disclosed embodiments, when the degree of certainty is not initially reached, at least one processor may analyze additional reflected signals to determine additional facial skin micromovements and reach the degree of certainty at least in part based on the analysis of the additional reflected signals. Figure 9 An example implementation of these embodiments is depicted (as described below).
[0262] Figure 9 A flowchart of an example process 900 for identifying an individual above a degree of certainty, performed by a processing device (e.g., processing device 400) of the speech detection system 100, is depicted. For illustrative purposes, certain components of the speech detection system 100 are referred to in the following description. However, it should be understood that other implementations are possible and other components may be used to implement the example process 900. It will also be readily understood that the example process 900 may be altered to modify the order of steps, delete steps, or further include additional steps.
[0263] When the processing device receives a reflection from the facial region, process 900 begins (block 902), then the processing device analyzes the reflection to determine a specific facial skin micromovement (block 904) and searches for a match between the determined specific facial skin micromovement and at least one reference facial skin micromovement (block 906). If no match is found (decision block 908), the processing device may initiate a second action (block 910), and the process continues by receiving additional reflected signals (block 912), analyzing them to determine additional facial skin micromovements, and searching for a match to identify the individual 802. If a match is found (decision block 908), the processing device may determine the degree of certainty of the match (block 914) and compare the determined degree of certainty with a threshold (decision block 916). If the degree of certainty is greater than the threshold, the processing device may initiate a first action (block 918), and the process continues by receiving additional reflected signals (block 912), analyzing (block 904), and searching (block 906). However, if the degree of certainty is less than the threshold, the processing device may initiate a second action (block 910).
[0264] According to some disclosed embodiments, at least one processor continuously compares new facial skin micromovements with multiple facial skin micromovements in a memory to determine an instantaneous level of certainty. As used herein, the phrase "continuous comparison" means continuously or periodically comparing new facial skin micromovements with multiple facial skin micromovements in a memory over a period of time (e.g., during a phone call). In this context, continuous comparison includes intervals between multiple comparisons, such as multiple times per second or multiple times per minute. The phrase "instantaneous level of certainty" refers to the confidence level in the identity of the individual associated with the new facial skin micromovements. For example, during a call with a banker, the system may periodically compare new facial skin micromovements to ensure that the same authorized individual remains online. According to some disclosed embodiments, when the instantaneous level of certainty is below a threshold, at least one processor is configured to initiate an associated action. The fact that the instantaneous level of certainty is below the threshold means there is a risk that someone other than the identified individual is the cause of the new facial skin micromovements. An associated action refers to an action associated with the fact that the instantaneous level of certainty is now below the threshold and may include a second action or stopping a first action. Specifically, in some embodiments, after initiating a first action, when the instantaneous level of certainty is below the threshold, at least one processor is configured to stop the first action. For example, the first action may be to authorize a transaction at a bank by talking to a banker over the phone and providing the banker with a continuous confirmation of the identity of the individual. However, once the instantaneous level of certainty drops below the threshold, which may indicate that someone other than the individual is talking to the banker, the transaction may be stopped. In some cases, the second action may include stopping the first action.
[0265] Reference Figure 9 , after initiating a first action at block 918, additional reflections are received and an analysis step (block 904) and a search step (block 906) are performed. If the determined instantaneous level of certainty associated with the additional reflections is below the threshold, the first action may be stopped by initiating a second action.
[0266] According to some disclosed embodiments, initiating a first action can be associated with an event, and at least one processor can continuously compare new facial skin micromovements during the event. In this context, the term "event" can refer to the occurrence of an action, activity, change in state, or any other type of detectable development or stimulus. The phrase "during the event" means any time from the time the event is detected until the end of the event. In one example, the event can be a purchase at a point of sale (POS), where the user wears a device to approve the transaction. In another example, the event can be associated with an online activity (e.g., a financial transaction, a betting session, an account access session, a gaming session, an exam, a lecture, or an educational session). In another example, the event can include maintaining a secure session with access to resources (e.g., files, folders, databases, computer programs, computer code, or computer settings).
[0267] Figure 10 A flowchart of an exemplary process 1000 for using facial skin micromovements to identify an individual in accordance with embodiments of the present disclosure is shown. In some disclosed embodiments, process 1000 can be executed by at least one processor (e.g., processing device 400 or processing device 460) to perform the operations or functions described herein. In some embodiments, some aspects of process 1000 can be implemented as software (e.g., program code or instructions) stored in a memory (e.g., memory device 402 or memory device 466) or a non-transitory computer-readable medium. In some embodiments, some aspects of process 1000 can be implemented as hardware (e.g., a dedicated circuit). In some embodiments, process 1000 can be implemented as a combination of software and hardware.
[0268] Reference Figure 10, process 1000 includes step 1002 of projecting light towards a facial region of an individual's head. For example, at least one processor may operate a wearable coherent light source (e.g., light source 410) to irradiate the facial region 108 (e.g., using multiple output beams 508). Process 1000 includes step 1004 of receiving coherent light reflections from the facial region and outputting an associated reflection signal. For example, at least one processor may operate at least one detector (e.g., at least one detector 412) to receive coherent light reflections (e.g., light reflection 300) from the facial region 108. Process 1000 includes step 1006 of analyzing the reflection signal to determine specific facial skin micromovements of the individual. For example, a light reflection processing module 706 and a silent reading decryption module 708 are used to determine specific facial skin micromovements. Process 1000 includes step 1008 of accessing a memory that associates multiple facial skin micromovements with the individual. Process 1000 includes step 1010 of searching for a match between the determined specific facial skin micromovement and at least one of the multiple facial skin micromovements in the memory. Process 1000 includes step 1012 of initiating an action based on a determination of whether a match is found. Specifically, if a match is recognized, a first action (e.g., first action 810A) is initiated, and if no match is recognized, a second action different from the first action (e.g., second action 810B) is initiated.
[0269] According to one embodiment, a speech detection system projects a pattern of light onto a user's facial skin (e.g., the cheek). Thereafter, the speech detection system may detect light reflections from various locations on the facial skin. Notably, the reflections associated with a specific region may be more relevant for extracting meaning (e.g., determining communication) than other regions. The specific region may be those regions that are located closer to specific facial muscles. Identifying the specific locations can be challenging because each user has unique facial features, and the position of the light source and / or detector relative to the user's face may change during each use and even during an ongoing operation. The following paragraphs describe systems, methods, and computer program products for identifying the locations of these specific regions, extracting meaning using light reflections from the specific regions, and ignoring light reflections from other regions to save processing resources.
[0270] Some disclosed embodiments include interpreting facial skin movement. The phrase "interpreting facial skin movement" refers to extracting meaning from detected skin movement, as described elsewhere in this disclosure. In one example, interpreting facial skin movement can include determining one or more spoken or mouthed words or determining an individual's facial expression (e.g., happy, sad, angry, afraid, surprised, disgusted, contemptuous, or other emotions) based on the facial skin movement. In another example, interpreting facial skin movement can include determining an individual's identity. These facial skin movements can be detectable, as described elsewhere in this disclosure.
[0271] Some disclosed embodiments include projecting light onto a plurality of facial region areas of an individual, where the plurality of areas includes at least a first area and a second area. The phrase "projecting" includes controlling a light source (e.g., a coherent light source) such that it emits light in a given direction (e.g., toward a portion of the face), as discussed elsewhere in this disclosure. The phrase "individual" includes a person using a speech detection system (or another individual onto whom the light source is projected), as described elsewhere in this disclosure. In the context of the face, the phrase "facial region area" or simply "area" includes a portion of an individual's face, as described elsewhere in this disclosure. For example, a facial region area can have a size of at least 1 cm 2 , at least 2 cm 2 , at least 4 cm 2 , at least 6 cm 2 , or at least 8 cm 2 . According to some disclosed embodiments, the projected light illuminates the plurality of facial region areas. For example, the plurality of areas includes 4, 8, 16, 32, or any other number of areas. In some cases, the projected light can include at least one light spot, as described elsewhere in this disclosure. At least one light spot can illuminate more than one facial region area, e.g., as Figure 3As shown, a single light spot 106 can illuminate different parts of the facial area 108. For example, the light spot 106 can include a first part 304A associated with a first facial muscle and a second part 304B associated with a second facial muscle. Alternatively, a single facial area partition can be illuminated by multiple light spots. Some of the multiple partitions can be spaced apart from each other, while other partitions among the multiple partitions can overlap each other. The term "spaced apart" can mean non-overlapping or separated by at least a certain distance. Thus, spaced-apart partitions can refer to two or more facial area partitions that do not overlap each other and even have a very small gap therebetween. For example, the statement that a first facial area partition is spaced apart from a second facial area partition can include a distance between the first area and the second area of at least 5 mm, at least 10 mm, at least 15 mm, or any other desired distance. In some embodiments, the distance can be less than 1 mm, or between 1 mm and 5 mm. In some cases, only a part of the facial area partition can be illuminated by the projected light. In other cases, all facial area partitions can be illuminated by the projected light. By way of example, Figure 11 and Figure 12 illustrates the use of multiple light spots to illuminate multiple facial area partitions of an individual. As shown, each of the facial area partitions 1100A and 1100B is shown being illuminated by more than one light spot.
[0272] Some disclosed embodiments include illuminating at least a part of a first partition and at least a part of a second partition with a common light spot. As used herein, the term "at least a part" and / or its grammatical equivalents can refer to any small part of the total amount. For example, "at least a part" can refer to at least about 1%, 5%, 10%, 20%, 40%, 65%, 90%, 95%, 99%, 99.9%, or 100% of the total amount, or any other small part. The term "common light spot" means that a single (common) light spot can cover some or all of the first partition and the second partition. The common light spot can illuminate at least a part of the first partition and the second partition. In one example, the common light spot can illuminate 30% of the first partition and 10% of the second partition. In another example, the common light spot can illuminate 100% of the first partition and 100% of the second partition. Controlling at least one coherent light source can include illuminating a continuous area on the face, the continuous area including the first partition and the second partition. As an example, as Figure 3 shown, a single light spot 106 can illuminate two or more facial area partitions (e.g., 304A and 304B).
[0273] Some disclosed embodiments include illuminating a first partition with a first set of light spots and illuminating a second partition with a second set of light spots different from the first set of light spots. The phrase "a set of light spots" refers to more than one light spot. The number of light spots in a set of light spots can range from 2 to 64 or more. For example, a set of light spots can include 4 light spots, 8 light spots, 16 light spots, 32 light spots, 64 light spots, or any number of light spots greater than two. As discussed elsewhere in this disclosure, there may be variations in the illumination characteristics between light spots or within a set of light spots. Illuminating a partition with a set of light spots can refer to illuminating some or all of a facial area with two or more light spots. In one example, a set of light spots can illuminate at least 15% of the partition, at least 40% of the partition, or at least 70% of the partition. The first partition can be irradiated by the first set of light spots, and the second partition can be irradiated by a second set of light spots different from the first set of light spots. As used herein, the phrase "different" means that the first set of light spots can be distinguished from the second set of light spots. For example, the first set of light spots can include at least one light spot not included in the second set of light spots. For example, Figure 11 and Figure 12 shows a first partition 1100A of the facial area illuminated by a first set of light spots 1108A and a second partition 1100B illuminated by a second set of light spots 1108B different from the first set of light spots.
[0274] Some disclosed embodiments include operating a coherent light source (as described elsewhere in this disclosure) located within a wearable housing (as described elsewhere in this disclosure) in a manner that enables illumination of multiple partitions of a facial area. As used herein, enabling illumination can refer to the process of controlling the light source to generate at least one light beam and directing at least one light beam towards multiple partitions of the facial area. For example, enabling illumination can also include utilizing a beam splitting element (as described elsewhere in this invention) configured to split an input light beam into a plurality of output light beams (as described elsewhere in this invention) extending over a portion of the face. In alternative embodiments, enabling illumination can include utilizing a plurality of light sources that generate corresponding sets of output light beams covering different corresponding sub-regions within a portion of the face. Figure 1 and FIG. 2 illustrates an example embodiment of a speech detection system (e.g., speech detection system 100) in which at least one partition of the facial area (e.g., facial area 108) is illuminated by a plurality of light spots (e.g., light spots 106). The plurality of light spots can be generated by an optical sensing unit 116 located within a wearable housing 110 and including at least one light source 410 and at least one light detector 412.
[0275] Some disclosed embodiments include operating a coherent light source (as described elsewhere in the present disclosure) positioned away from a wearable housing (as described elsewhere in the present disclosure) in a manner that can illuminate multiple facial region partitions (as described elsewhere in the present disclosure). The phrase "positioned away" indicates that two objects are separated from each other and have a physical distance between them such that they do not physically appear as a unified component. For example, the coherent light source can be part of a device other than the speech detection system and is located more than 1 cm away from the wearable housing of the speech detection system. As another example, the coherent light source can be located more than 3 cm away from the wearable housing of the speech detection system. It should be understood that the distances of 1 cm and 3 cm are exemplary and non-limiting, and other distances can be used. Figure 3 An exemplary embodiment of the speech detection system is shown, where multiple facial region zones (e.g., the first part 304A and the second part 304B of the facial region 108) are illuminated by a coherent light source (e.g., the non-wearable light source 302) positioned away from the wearable housing.
[0276] In some disclosed embodiments, the first partition is closer to at least one of the zygomaticus muscle or the risorius muscle than the second partition. The phrase "the first partition is closer to the muscle than the second partition" means that the distance from the first partition to the specific muscle is less than the distance from the second partition to the specific muscle. For example, the distance can be measured from the edge of the partition to the edge of the specific muscle, from the center of the partition to the center of the specific muscle, or any combination thereof. In this case, the center of the shape (i.e., the first partition, the second partition, or the specific muscle) can be the geometric center, which is the point corresponding to the average position of all points in the shape; the circumcenter, which is the center of the smallest circle that completely encloses the 2D shape; the incenter, which is the center of the inscribed circle that is tangent to all sides of the 2D shape; or any other reference point defined previously. As discussed, the first partition is closer to at least one of the zygomaticus muscle or the risorius muscle than the second partition. In other words, the disclosed embodiments capture two example use cases. The first example use case is that the first partition is closer to the zygomaticus muscle than the second partition. The second example use case is that the first partition is closer to the risorius muscle than the second partition. As an example, Figure 11 An implementation of the first and second example use cases is shown. Specifically, the first use case is shown with respect to individual 102A, and the second use case is shown with respect to individual 102B.
[0277] Figure 11Two example use cases for explaining facial skin movement are shown. In the two example use cases, multiple facial region partitions 1100 of an individual 102 can be illuminated by at least one light source (e.g., light source 410, not shown). The depicted multiple partitions include at least a first partition 1100A and a second partition 1100B. In a first example use case involving individual 102A, the first partition 1100A is closer to the zygomaticus muscle than the second partition 1100B, and in a second example use case involving individual 102B, the first partition 1100A is closer to the risorius muscle than the second partition 1100B.
[0278] Some disclosed embodiments include receiving reflections from multiple partitions. The term "receiving" can include obtaining, retrieving, acquiring, or otherwise gaining access to data or a signal. In some cases, receiving can include reading data from a memory and / or obtaining data from a computing device via a (e.g., wired and / or wireless) communication channel. In other cases, receiving can include detecting an electromagnetic wave (e.g., in the visible or invisible spectrum) and generating an output related to the measured properties of the electromagnetic wave. In a first embodiment, at least one processor can receive data indicating light reflected from multiple partitions from at least one detector. In a second embodiment, at least one detector can receive light rays reflected from multiple partitions. The term "reflection" refers to one or more light rays bouncing off a surface (e.g., an individual's face) or data obtained from one or more light rays bouncing off a surface. For example, reflection can include light detected by a light detector after the light has been deflected from an object. The light detected by the light detector can be generated by at least one coherent light source of the disclosed speech detection system and / or can be generated from a source other than the disclosed speech detection system. As an example, Figure 5A and Figure 5B the light detector 412 in is used to receive the reflection 300 of the light generated by the light source 410.
[0279] As an example, referring to Figure 11 the two use cases depicted in, the reflection image 1102A can represent the reflection received from the first partition 1100A, and the reflection image 1102B can represent the reflection received from the second partition 1100B. As shown, in the first example use case, the reflection image 1102A represents the reflection received from the partition closer to the zygomaticus muscle; and in the second example use case, the image 1102A represents the reflection received from the area closer to the risorius muscle.
[0280] Some disclosed embodiments include detecting a first facial skin movement corresponding to a reflection from a first partition and a second facial skin movement corresponding to a reflection from a second partition. In this context, the term "detecting" refers to the process of finding, identifying, or determining the presence of a light reflection (or a signal associated therewith). In one example, a change in the position of the facial skin can be detected. As discussed elsewhere in this disclosure, the detection process can involve using various techniques or arts to determine the presence of a pattern or event. In some cases, the process of detecting a facial skin movement can involve determining whether any movement has occurred and recording information representative of the detected movement. For example, at least one processor can detect a facial skin movement by applying light reflection analysis to the received reflection. In other cases, detecting a facial skin movement can include determining the time at which the facial skin movement occurs. In other cases, detecting a facial skin movement can include determining data representative of the facial skin movement (e.g., direction, speed, acceleration). The term "facial skin movement" broadly refers to any type of movement caused by the recruitment of underlying facial muscles. Facial skin movements include facial skin micromovements (as described elsewhere in this disclosure) and larger-scale skin movements that are generally visible and detectable to the naked eye without magnification (e.g., smiling, yawning, frowning). The term "facial skin movement corresponding to a reflection from a particular partition" means that the detected facial skin movement occurs in a particular partition of the face from which the reflection is received. For example, detecting a first facial skin movement corresponding to a reflection from a first partition means that the first facial skin movement can be detected by analyzing the reflection received from the first partition; and detecting a second facial skin movement corresponding to a reflection from a second partition means that the second facial skin movement can be detected by analyzing the reflection received from the second partition.
[0281] In some disclosed embodiments, detecting the first facial skin movement involves performing a first speckle analysis on the light reflected from the first partition, and detecting the second facial skin movement involves performing a second speckle analysis on the light reflected from the second partition. The term "performing" refers to the act of carrying out a task, activity, or function. The term "speckle analysis" can be understood as described elsewhere in this disclosure. According to this disclosure, performing speckle analysis can include detecting a speckle pattern or any other pattern in the signal received from the light reflected from the face partition. For example, performing speckle analysis can include identifying a secondary speckle pattern generated due to the reflection of coherent light from each partition. In other embodiments, in addition to or instead of speckle analysis, detecting a facial skin movement can involve performing pattern-based analysis or image-based analysis.
[0282] According to some disclosed embodiments, the first speckle analysis and the second speckle analysis occur simultaneously by at least one processor. The phrase "occur simultaneously" means that two or more events occur during an overlapping or coinciding time period, where one event starts and ends during the duration of another event, or where the later-occurring event starts before the completion of another event. In some cases, the two or more events can be speckle analysis (or any pattern-based analysis). To enable the first speckle analysis and the second speckle analysis to occur simultaneously, the at least one processor can include multiple processors or a multi-core processor that allows multiple speckle analyses to be executed simultaneously.
[0283] As an example, referring to Figure 11 the two use cases depicted in, the first facial skin movement 1104A can correspond to the reflection from the first partition 1100A, and the second facial skin movement 1104B can correspond to the reflection from the second partition 1100B. For example, in the first example use case, the first facial skin movement 1104A corresponds to the reflection received from a partition closer to the zygomaticus muscle; and in the second example use case, the second facial skin movement 1104B corresponds to the reflection received from a partition closer to the risorius muscle.
[0284] Some disclosed embodiments include determining that a reflection from a first partition closer to at least one of the zygomaticus major or the risorius muscles is a stronger communication indicator than a reflection from a second partition based on a difference between a first facial skin movement and a second facial skin movement. Determining means ascertaining. For example, based on the difference between the first and second facial skin movements, a processor can determine which is closer to the associated muscle. The difference between the first facial skin movement and the second facial skin movement can include any difference, variation, or dissimilarity between the first facial skin movement and the second facial skin movement. The difference between the first facial skin movement and the second facial skin movement can be determined using at least one of the following techniques: surface alignment, point-to-point comparison, surface registration, topological analysis, or any other technique for determining the difference between two data sets. For example, the difference between the first facial skin movement and the second facial skin movement can include differences in movement intensity, movement trajectory, movement speed, and / or various changes in facial skin topography. Based on the difference, at least one processor can determine that a reflection from the first partition is a stronger communication indicator than a reflection from the second partition. The term "communication" refers to the process of conveying information through various media such as spoken language, words, body language, gestures, or signals. For example, communication can include verbal cues (e.g., words, phrases, and language) and non-verbal cues (e.g., body language, facial expressions, gestures, and eye contact). The term "communication indicator" refers to a measure or symbol that reflects the information conveyed by an individual. For example, the statement that a reflection from the first partition is a stronger communication indicator than a reflection from the second partition means that it may be easier to determine what information an individual intends to convey and what the individual intends to communicate from the first facial skin movement than from the second facial skin movement. For example, a reflection from the first partition can be a stronger communication indicator than a reflection from the second partition because the facial skin micromovements determined from the reflection from the first partition can be associated with a higher speed, a higher displacement, or other higher parameters that indicate that the individual intends to convey information and / or the content of the information the individual intends to convey. According to the disclosed embodiments, in a first exemplary use case, when the first partition is closer to the zygomaticus major muscle, the first facial skin movement can reflect a movement with a speed of about 1 to 10 μm / ms, and the second facial skin movement can reflect a smaller movement (if any). In a second exemplary use case, when the first partition is closer to the risorius muscle, the first facial skin movement can reflect a movement of approximately 0.52 mm, and the second facial skin movement reflects a smaller movement (if any).
[0285] According to some disclosed embodiments, the difference between the first facial skin movement and the second facial skin movement includes a difference of less than 100 micrometers. The phrase "difference of less than 100 micrometers" means that the change between a first parameter representing the first facial skin movement and a second parameter representing the second facial skin movement is less than 100 micrometers. In one example, the first parameter may be the magnitude of a first displacement change vector associated with the first facial skin movement, and the second parameter may be the magnitude of a second displacement change vector associated with the second facial skin movement. The displacement change is a vector that quantifies the distance and direction change between two measurements of the facial skin. For example, the difference between the first facial skin movement and the second facial skin movement includes a difference of less than 50 micrometers, less than 10 micrometers, or less than 1 micrometer. In other embodiments, the difference between the first facial skin movement and the second facial skin movement includes a difference of less than 1 millimeter. Thus, the determination that the reflection from the first partition is a stronger AC indicator than the reflection from the second partition is based on a difference of less than 1 millimeter, less than 100 micrometers, less than 50 micrometers, less than 10 micrometers, or less than 1 micrometer.
[0286] Some disclosed embodiments include processing reflections from a first partition to determine communication based on determining that the reflections from the first partition are stronger indicators of communication. The term "processing" refers to an action of performing an operation or transformation on data or information to achieve a desired result. For example, processing can include systematically manipulating, analyzing, or altering an input to produce a meaningful output. The term "processing reflections" means extracting information from a signal representing the received reflections. For example, processing reflections can include actions such as: filtering, amplifying, modulating, and applying optical reflection analysis, as described elsewhere in this disclosure. Processing reflections from the first partition to determine communication based on the determination that the reflections from the first partition are stronger indicators of communication. The term "determining communication" means determining speech or facial expressions associated with non-verbal communication based on facial movements, as described elsewhere in this disclosure. According to this disclosure, reflections from the first partition can be processed to create an image of a speckle pattern. Even at a fast exposure time (such as 10 ms), the movement speed of the skin may be sufficient to cause the speckle pattern to change during each frame, causing bright pixels to blur and fade. The degree of speckle blur of a given speckle in a given frame (e.g., as manifested by the loss of contrast in the image) can indicate the instantaneous movement speed of the skin in a small partition of the cheek under the speckle. Processing reflections from the first partition can also include extracting quantitative image features from the image of the speckle pattern. A vector of these features extracted from consecutive image frames can be input into a neural network to determine communication. Details of the neural network architectures and training algorithms that can be used for this purpose are described elsewhere in this disclosure. Example features that can be extracted for the purpose of determining communication can include speckle contrast. Any suitable measure of contrast can be used for this purpose, e.g., the mean squared value of the luminance gradient taken over a partition of the speckle pattern. High contrast in the speckle pattern of a given speckle from the first partition can indicate that the corresponding location on the cheek is stationary, while reduced contrast can indicate movement. The contrast decreases as the movement speed increases. Such contrast features can generally be extracted from multiple speckles distributed over the first partition. Additionally or alternatively, other features can be extracted from the speckle image and input into the neural network. Examples of such features can include the total luminance of the speckle pattern and the orientation of the speckle pattern, e.g., as calculated by a Sobel filter. As an example, Figure 7 the silent reading decryption module 708 in Figure 7 can be used to process reflections from the first partition to determine communication.
[0287] According to some disclosed embodiments, the communication determined based on reflections from the first partition includes words expressed by an individual. "Determining words expressed by an individual" means understanding the words vocalized or silently read by an individual. The words can be determined by processing the signals generated by the reflections, as discussed elsewhere herein. As an example, Figure 11The word "hello" in represents a word expressed by individual 102A or individual 102B that can be determined based on the reflection from the first partition.
[0288] According to some disclosed embodiments, the communication determined based on the reflection from the first partition includes non-verbal cues of the individual. The phrase "non-verbal cue" refers to various forms of communication that occur without the use of spoken words. Some examples of non-verbal cues can include facial expressions, body language, gestures, eye contact, intonation, posture, and other subtle signals that convey meaning in interpersonal interactions. For example, non-verbal cues such as facial expressions can be used to convey basic emotions such as happiness, sadness, anger, fear, surprise, and disgust. As discussed elsewhere in this disclosure, at least one processor can determine non-verbal cues by analyzing reflection signals representing facial skin micromovements in the first facial partition. For example, Figure 11 the emoji in represents a non-verbal cue that can be determined based on the reflection from the first partition.
[0289] Some disclosed embodiments include determining that the reflection from the first partition is a stronger indicator of communication and ignoring the reflection from the second partition. In this context, the phrase "ignore the reflection" means that the processing action on the signal representing the reflection received from the second partition is less than the processing action on the signal representing the reflection received from the first partition. In one embodiment, the signal representing the reflection received from the second partition can be filtered, amplified, and analyzed to determine the second facial skin movement, but some quantitative features may not be extractable because communication may not be determinable from the signal representing the reflection received from the second partition. In another embodiment also involving "ignoring", during a first time frame, the reflections from both the first partition and the second partition can be processed to determine which partition is closer to the zygomaticus major muscle or the risorius muscle. Thereafter, during a subsequent second time frame, and when it is determined that the first partition is closer to the zygomaticus major muscle or the risorius muscle, the reflection from the second partition can be automatically discarded.
[0290] According to some disclosed embodiments, ignoring the reflection from the second partition includes omitting the use of the reflection from the second partition to determine communication. The phrase "omit the use" means not using the information associated with the reflection from the second partition when determining the meaning of communication.
[0291] As an example, refer to Figure 11The two use cases depicted can process the reflected image 1102A to determine the communication 1106 based on the facial skin movement 1104A associated with the zygomaticus or risorius muscles, and can ignore the reflected image 1102B, e.g., not use or omit the reflected image 1102B when determining the communication. As depicted, the determined communication can include at least one word 1106A (expressed silently or vocally by individual 102A or individual 102B) and / or at least one facial expression 1106B that serves as an example of a non-verbal cue.
[0292] Some disclosed embodiments include determining that a first partition is closer to subcutaneous tissue associated with cranial nerve V or cranial nerve VII than a second partition based on a difference between a first facial skin movement and a second facial skin movement. The phrase "subcutaneous tissue" refers to the layer of tissue that lies beneath the skin and above the muscles and bones. It consists of fat cells, connective tissue, blood vessels, nerves, and other structures. Cranial nerve V (also known as the trigeminal nerve) is the facial sensory nerve that controls the jaw muscles. Cranial nerve VII controls facial expressions and carries taste from the front of the tongue. Based on the difference between the first facial skin movement and the second facial skin movement (as described above), it can be determined that the first partition is closer to the subcutaneous tissue associated with cranial nerve V or cranial nerve VII than the second partition.
[0293] Some disclosed embodiments include operating a coherent light source in a manner that enables dual-mode illumination of multiple facial partitions. The phrase "coherent light source" can be understood as described elsewhere in this disclosure. Operating the coherent light source in this context refers to adjusting, supervising, indicating, permitting, and / or enabling the coherent light source to illuminate at least a portion of the face. For example, the coherent light source can be controlled to illuminate an area of the face in a specific illumination pattern when turned on in response to a trigger. Dual-mode illumination refers to the ability of the coherent light source to use at least two different illumination modes to illuminate an object. The phrase "illumination mode" refers to a specific configuration or setting of the coherent light source. Each of the two modes can be associated with different values of illumination parameters, such as light intensity, illumination pattern, pulse frequency, duty cycle, light flux. Figure 4 The light source 410 in is an example of a single-mode or multimode (e.g., dual-mode) light source.
[0294] In some disclosed embodiments, the first light intensity of the first illumination mode is different from the second light intensity of the second illumination mode. In some disclosed embodiments, the first illumination pattern of the first illumination mode is different from the second illumination pattern of the second illumination mode. Light intensity refers to the luminance level of illumination, and an illumination pattern refers to the arrangement, distribution, or sequence of coherent or incoherent light emitted from a light source or reflected from a surface. A light pattern can be created by a specific design, shape, or configuration of the light source to create a specific visual or non-visual effect on a portion of the face. Examples of illumination patterns can include a grid of light spots of the same size, a grid of light spots of various sizes, a single light spot, or any other pattern.
[0295] Some disclosed embodiments include analyzing the reflections associated with the first illumination mode to identify one or more light spots associated with a first partition, and analyzing the reflections associated with the second illumination mode to determine the AC. The phrase "identifying one or more light spots associated with a first partition" means determining which of the light spots projected by a coherent light source are located within the first partition. For example, identifying one or more light spots associated with a first partition can be implemented by comparing the light intensity at a specific location with the boundaries of the first partition, based on image analysis of an individual's face, or by any other processing method. In one example, the first illumination mode can include a first illumination pattern (e.g., 64 light spots), and the second illumination mode can include a second illumination pattern (e.g., 32 light spots). As an example, referring to Figure 11 the first exemplary use case depicted in, the first illumination mode can be used to identify eight light spots included within a first partition 1100A associated with the zygomaticus muscle. Thereafter, the second illumination mode (e.g., 4 light spots) can be used to illuminate the first partition 1100A in a manner that enables the AC to be determined from the received reflections.
[0296] According to some disclosed embodiments, the first partition is closer to the zygomaticus muscle than the second partition, and the plurality of partitions further includes a third partition that is closer to the risorius muscle than each of the first and second partitions. The phrases "plurality of partitions" and "closer to" can be understood as described elsewhere in this disclosure. As an example, referring to Figure 12, a plurality of facial sub - regions 1100 includes a first sub - region 1100A closer to the zygomaticus muscle than a second sub - region 1100B, and a third sub - region 1100C closer to the risorius muscle than each of the first sub - region 1100A and the second sub - region 1100B. In some disclosed embodiments, based on the determination that the individual 102C is engaged in silent speech, the processing device of the speech detection system may process the reflections from the first sub - region 1100A to determine communication and ignore the reflections from the second sub - region 1100B and the third sub - region 1100C. In other embodiments, based on the determination that the individual 102C is engaged in vocalized speech, the processing device of the speech detection system may process the reflections from the third sub - region 1100C to determine communication and ignore the reflections from the second sub - region 1100B and the first sub - region 1100A.
[0297] Some disclosed embodiments include analyzing the reflected light from the first sub - region when speech is generated in the presence of a perceivable vocalization (i.e., vocalized speech), and analyzing the reflected light from the third sub - region when speech is generated in the absence of a perceivable vocalization (i.e., silent speech). In other words, rather than monitoring the entire cheek and processing the reflections from multiple sub - regions, the speech detection system may process the reflections received from a subset of the cheek regions in these two sub - regions (e.g., only a few square millimeters or centimeters) to detect silent speech and vocalized speech. Additionally, when the multiple sub - regions are illuminated by multiple light sources (e.g., a laser diode array), only the light sources illuminating these two sub - regions may be actuated, thus reducing power consumption. If a relatively large movement of the speech detection system relative to the skin is detected, different groups of light sources may be actuated. In some disclosed embodiments, different processing modes may be applied to determine silent speech from vocalized speech. For example, during silent speech, the first sub - region closer to the zygomaticus muscle may exhibit a movement with a speed of about 1 to 10 μm / ms. Thus, the characteristics of the speckle image itself may change rapidly, and these characteristics may be analyzed to generate an output. However, during silent speech, the third sub - region closer to the risorius muscle may exhibit a movement of about 0.5 to 2 mm. Thus, due to the movement of the cheek, the position of the light spot on the cheek may shift laterally. In this case, the lateral movement of the light spot may indicate a change in the distance between the light spot and the speech detection system, and the speech detection system may thus be used as a kind of depth sensor. Two processing modes - speckle sensing and depth sensing - may be used separately to detect silent speech and vocalized speech respectively. Alternatively or additionally, these two processing modes may be used together to improve the accuracy and specificity of the measurement, for example, by applying the measurements of a given user's vocalized speech to learn the pattern of microscopic movements that will occur in the same user's silent speech.
[0298] Figure 13FIG. shows a flowchart of an exemplary process 1300 for identifying an individual using facial skin micro-movements according to an embodiment of the present disclosure. In some disclosed embodiments, process 1300 may be executed by at least one processor (e.g., processing device 400 or processing device 460) to perform the operations or functions described herein. In some disclosed embodiments, some aspects of process 1300 may be implemented as software (e.g., program code or instructions) stored in a memory (e.g., memory device 402 or memory device 466) or a non-transitory computer-readable medium. In some disclosed embodiments, some aspects of process 1300 may be implemented as hardware (e.g., a dedicated circuit). In some disclosed embodiments, process 1300 may be implemented as a combination of software and hardware.
[0299] Referring Figure 13 , process 1300 includes step 1302 of projecting light onto a plurality of facial regions of an individual. For example, at least one processor may operate a wearable coherent light source (e.g., light source 410) to illuminate at least a first region (e.g., first region 1100A) and a second region (e.g., second region 1100A). The first region may be closer to at least one of the zygomaticus muscle or the risorius muscle than the second region. Process 1300 includes step 1304 of receiving reflections from the plurality of regions. For example, at least one processor may operate at least one detector (e.g., at least one detector 412) to receive coherent light reflections (e.g., light reflection 300) from the plurality of regions 1100. Process 1300 includes step 1306 of detecting a first facial skin movement corresponding to the reflection from the first region and a second facial skin movement corresponding to the reflection from the second region. For example, at least one processor may use a light reflection processing module 706 to detect the first facial skin movement, and the second facial skin movement corresponds to the reflection from the second region. Process 1300 includes step 1308 of determining that the reflection from the first region is a stronger indicator of alternating current than the reflection from the second region. For example, the determination in step 1308 may be based on the difference between the first facial skin movement and the second facial skin movement. Process 1300 includes step 1310 of processing the reflection from the first region to determine alternating current and ignoring the reflection from the second region. For example, the determination in step 1310 may be based on the determination that the reflection from the first region is a stronger indicator of alternating current. At least one word 1106A and at least one facial expression 1106B are examples of the determined alternating current.
[0300] The embodiments for explaining facial skin movements discussed above may be implemented by, for example, software (e.g., as operations executed by code), as a method (e.g., Figure 13 the process 1300 shown in Figures 1 to 3implemented by a non - transitory computer - readable medium of the speech detection system 100) shown. When the implementation is realized as a system, the operations can be performed by at least one processor (e.g., Figure 4 the processing device 400 or the processing device 460) shown.
[0301] In some embodiments, an authentication or identity verification service provider uses biometrics (e.g., signals indicating the facial skin micromovements of an individual) for authentication purposes. For example, the authentication service provider can use the facial skin micromovements of an individual to verify the identity of the individual. The intensity and order of muscle activation (e.g., muscle fiber recruitment) on an individual's facial area vary among individuals. Muscle activation or recruitment is the process of activating motor neurons to produce various degrees of muscle contraction. An individual's skin micromovements can be affected by muscles, the structure of muscle fibers, the characteristics of the skin, subcutaneous characteristics (e.g., vascular structure, fat structure, hair structure, etc.), etc. The iris is an example of a visible muscle of an individual. The iris is the colored tissue in the front of the eye that contains the pupil in the center and helps control the size of the pupil to let more or less light into the eye. Although the iris of each individual is circular, the structure of each individual's iris can be unique and can be stable throughout an individual's life. The same is true for subcutaneous muscles and their activation. Facial skin micromovements can create a unique biometric signature of an individual, which can be used to identify the individual. For the sake of brevity, in the following discussion, facial skin micromovements can be abbreviated as facial micromovements. An institution that requires customer authentication (i.e., verification) can subscribe to the authentication service provided by the provider to authenticate an individual (e.g., a customer) before providing access to the services or facilities provided by the institution. Such institutions can include financial institutions (e.g., banks and brokerage services), subscription services (e.g., providing media content, research, or other information), online gaming sites, other online platforms, government agencies, and other organizations that require user authentication and verification, or any other entity or service that desires customer authentication. Authentication is the process of verifying or confirming an individual's identity.
[0302] Some disclosed embodiments include the authentication of an individual based on the facial micromovements of the individual. The verification can occur via a system, a computer - readable medium, or a method. The term "authentication" is the process of determining who an individual is. It can also refer to the process of confirming or denying whether an individual is the person the individual claims to be. For example, in some embodiments, the system of the present disclosure can determine who an individual is based on the facial micromovements of the individual. And in some embodiments, the system of the present disclosure can determine (e.g., confirm or deny) whether an individual is actually the person he / she claims to be based on the facial micromovements of the individual.
[0303] Figure 14is a schematic diagram of an exemplary embodiment that includes a system for providing authentication of an individual based on the individual's facial micro - movements. As Figure 14 (and Figures 1 to 4 ) shows, a detection system 100 associated with an individual 102 can use a communication network 126, for example, to directly or via a mobile communication device 120 detect signals indicative of (or representing) the individual's facial micro - movements and transmit them to a cloud server 122. In some embodiments, as described elsewhere in this disclosure, the server 122 can access a data structure 124 to determine, for example, the correlation between words and the individual's facial micro - movements. In some embodiments, the cloud server 122 can also be configured to authenticate the individual based on the received signals. In some embodiments, an authentication service provider (or identity verification service provider) can use a system such as server 122 for providing authentication of an individual based on the individual's facial micro - movements. In some embodiments, as Figure 14 shows, an institution 1400 and a speech detection system 100 associated with an individual 102 can communicate with each other and with the cloud server 122 using the communication network 126 to request and receive authentication of the individual.
[0304] Figure 15 、 Figure 16A and Figure 16B are simplified block diagrams showing different aspects of an exemplary system 1500 for providing authentication (or identity verification) based on an individual's facial skin micro - movements (or facial micro - movements). It should be noted that only the elements of the authentication system 1500 relevant to the following discussion are shown in these figures. Embodiments within the scope of this disclosure may include additional elements or fewer elements. As Figure 15 shows, the system 1500 includes a processor 1510 and a memory 1520. Although only one processor and one memory are shown in Figure 15 , in some embodiments, the processor 1510 can include more than one processor, and the memory 220 can include multiple devices. These multiple processors and memories can each have similar or different configurations and can be electrically connected or disconnected from each other. Although the memory 1520 is shown in Figure 15is shown as being separate from the processor 1510, but in some embodiments, the memory 1520 can be integrated with the processor 1510. In some embodiments, the memory 1520 can be located remote from the system 1500 and can be accessed by the system 1500. The memory 1520 can include any device for storing data and / or instructions, such as, for example, random access memory (RAM), read-only memory (ROM), hard disk, optical disk, magnetic media, flash memory, other permanent, fixed, or volatile memory. In some embodiments, the memory 1520 can be a non-transitory computer-readable storage medium storing instructions that, when executed by the processor 1510, cause the processor 1510 to perform authentication operations based on facial micro-movements. In some embodiments, some or all of the functions of the processor 1510 and the memory 1520 can be performed by remote processing devices and memory (e.g., the processing device 400 and the memory device 402 of the remote processing system 450, see Figure 4 ).
[0305] Some of the disclosed embodiments include receiving a reference signal in a trustworthy manner for verifying the correspondence between a particular individual and an account at an institution. The term "receiving" can include retrieving, obtaining, or otherwise gaining access to, for example, data. Receiving can include reading data from a memory and / or receiving data from a computing device via (e.g., wired and / or wireless) communication channels. At least one processor can receive data via synchronous and / or asynchronous communication protocols, such as by polling a memory buffer for data and / or by receiving data as an interrupt event. The term "signal" can refer to information encoded for transmission via a physical medium or wirelessly. Examples of signals can include signals in the electromagnetic radiation spectrum (e.g., AM or FM radio, Wi-Fi, Bluetooth, radar, visible light, lidar, IR, Zigbee, Z-wave, and / or GPS signals), sound or ultrasonic signals, electrical signals (e.g., voltage, current, or charge signals), electronic signals (e.g., as digital data), tactile signals (e.g., touch), and / or any other type of information encoded for transmission between two entities via a physical medium or wirelessly (e.g., via a communication network). In some embodiments, the signal can include or can represent "speckles", reflected image data, or light reflection analysis data (e.g., speckle analysis, pattern-based analysis, etc.) described elsewhere in this disclosure.
[0306] Receiving a signal in a "trustworthy" manner means receiving a reliable signal. For example, receiving a signal in a manner such that the authenticity and / or validity of the signal can be relied upon. In some embodiments, when receiving a signal in a trustworthy manner, there can be some degree of assurance that the signals are valid or as they are expected to be. In some embodiments, receiving a signal in a trustworthy manner can indicate that the signals are transmitted in a secure manner such that the signals may not be easily intercepted and / or decrypted by a third party. Generally, any known secure transmission method can be used to send and receive signals in a trustworthy manner. In some embodiments, receiving a signal in a trustworthy manner can refer to receiving an encrypted signal. Any currently known or later developed encryption technique (e.g., Wired Equivalent Privacy (WEP), Wi-Fi Protected Access (WPA), Wi-Fi Protected Access version 2 (WPA2), Wi-Fi Protected Access version 3 (WPA3), etc.) can be used to encrypt the signal. In some embodiments, the encrypted signal can include a key(s) that can be used to decrypt the encrypted signal by methods known in the art.
[0307] As used herein, the phrase "reference signal" means a signal that is used as a basis for determining something. For example, a reference signal can be a baseline signal for comparison purposes, e.g., to determine whether the characteristics of a signal have changed. In some embodiments, a reference signal can represent one or more attributes or characteristics of an individual. For example, a reference signal can represent one or more attributes / characteristics of an individual's facial micro-movements. In some embodiments, a reference signal can be (or can represent) a speckle pattern (e.g., Figure 6The reflected image 600) or another light reflection pattern output by the speech detection system 100 associated with the individual. In some embodiments, the reference signal may include or may represent one or more characteristics of the individual's facial micromovements. In some embodiments, the reference signal may be (or may include) a characteristic or feature extracted from the individual's light reflection pattern. In some embodiments, one or more algorithms may be used to extract these characteristics or features of the individual's facial micromovements embodied in the reference signal. These extracted features may include fiducial and / or non-fiducial features. Fiducial features may include measurable characteristics of the individual's facial micromovements (e.g., time or amplitude onset, peak (minimum or maximum), offset, interval, time difference between peaks, and other measurable characteristics). On the other hand, non-fiducial feature extraction may apply time and / or frequency analysis to obtain statistical features of the individual's facial micromovements. In some embodiments, the reference signal may represent multiple biometric signals of the individual (e.g., a combination of facial micromovements with one or more of the following: pulse, cardiac signal, ECG, temperature, pressure, or other biometric signals). It is also contemplated that, in some embodiments, the detected facial micromovement signal or the light reflection pattern itself output by the speech detection system 100 may be used as the reference signal for the individual.
[0308] A reference signal can be configured to enable verification of the correspondence between a specific individual and an account at an institution. The term "correspondence" refers to similarity, connection, equivalence, match, or link. For example, in some embodiments, the reference signal of a specific individual can be used to determine the equivalence, similarity, match, or link between the individual and an account at an institution (e.g., a customer's account). The institution can retain the biometric or other data of the customer in an associated manner, and this data or related data can be included in the reference signal. The term "institution" refers to any institution or organization without limitation. In some embodiments, the institution can be, for example, an organization that provides a certain type of service to a plurality of individuals who each have an account at the institution. In some embodiments, the institution can be a financial organization (e.g., a bank, a stock brokerage firm, a mutual fund, etc.) where multiple customers can have accounts (such as a cash account, a money market account, a stock account, an online account, a safe deposit box, etc.). In some embodiments, the institution can be a company associated with online activities (such as gaming activities, betting activities, exam / test providers, education / course providers, etc.), or a university or educational institution where multiple students have accounts (to access courses, bills, etc.). In some embodiments, the institution can be a healthcare provider (e.g., a hospital, a clinic, a testing laboratory, etc.) or an insurance provider (e.g., an insurance company) where multiple patients or customers have accounts, a company where multiple employees have accounts, etc. In other embodiments, the institution can be a government agency or body. The reference signal can be received from any source (e.g., an individual, an institution, etc.).
[0309] In some embodiments, the institution can engage a certification service provider and / or subscribe to a certification service to verify the identity of an individual (or customer) in connection with providing services to the individual (e.g., before allowing access to an account, etc.). The certification service provider can use a system such as Figure 15 , Figure 16A and Figure 16BThe system 1500) uses reference signals to verify the identity of an individual. In some embodiments, the system can access the reference signals of all customers of an institution (e.g., all account holders of a bank, all students enrolled in classes at a university, etc.). For example, in some embodiments, as shown in FIG. 16, the reference signals 1502 of all customers (e.g., account holders) of an institution 1400 (e.g., a bank) can be sent to the system 1500 (e.g., during enrollment). The system 1500 can securely store the correlations 1504 between the reference signals 1502 and the identities of different customers in a secure data structure (such as data structure 124) accessible to the system 1500. In some embodiments, the name and / or other identification information of the customer (account number or other information identifying the individual associated with the reference signal) can also be stored and associated with the reference data in the stored correlation 1504. As will be explained in more detail later, the system 1500 can use the stored reference signals and correlations to authenticate an individual. For example, as Figure 16B shown, when an individual participates in a transaction with the institution 1400 (e.g., attempts to access a customer's account), the institution 1400 can request 1506 the authentication service provider (or the system 1500) to authenticate the individual (e.g., verify the identity of the individual, confirm that the individual is the customer associated with the account, etc.). When an individual participates in a transaction, the system 1500 can receive the real-time facial micro-motion signal 1508 of the individual, and the system 1500 can compare 1512 the received real-time signal 1508 with the stored reference signal 1502 or correlation 1504 to determine whether the individual is a customer. For example, the system 1500 can compare the two signals to determine whether one or more characteristics of the received signal correspond to or sufficiently match the characteristics of the stored reference signal to determine whether the received signal is associated with the customer authorized to access the account.
[0310] According to some disclosed embodiments, a reference signal may be obtained based on reference facial micromovements detected using first coherent light reflected from a particular individual's face. The term "reference" in "reference facial micromovements" indicates that these facial micromovements are used to generate the reference signal. As explained elsewhere in this disclosure, "coherent light" includes light that is highly ordered and exhibits a high degree of spatial and temporal coherence. As detailed elsewhere in this disclosure, when coherent light irradiates an individual's facial skin, some of it is absorbed, some is transmitted, and some is reflected. The amount and type of the reflected light depend on the nature of the skin and the angle at which the light irradiates the skin. For example, coherent light irradiating a rough, undulating, or textured skin surface may be reflected or scattered in many different directions, resulting in a pattern of bright and dark regions called "speckle". In some embodiments, when coherent light is reflected from an individual's face, the optical reflection analysis performed on the reflected light may include speckle analysis or any pattern-based analysis to obtain information about the skin (e.g., facial skin micromovements) represented in the reflected signal. In some embodiments, the speckle pattern may occur as a result of the interference of coherent light waves that add together to give a composite wave with intensity variations. In some embodiments, the detected speckle pattern (or any other detected pattern) may be processed to generate reflected image data from which a reference signal may be generated.
[0311] As referenced elsewhere in this disclosure Figures 1 to 6 as explained, the speech detection system 100 associated with an individual may detect the individual's facial micromovements. For example, specifically referring to FIGS. 5 to Figure 7 , in some embodiments, the speech detection system 100 may analyze the reflection 300 of coherent light from the individual's facial region 108 to determine facial micromovements (e.g., the amount of skin movement, the direction of skin movement, the acceleration of skin movement, the speckle pattern, etc.) caused by the recruitment of muscle fibers 520 and output a signal representing the detected facial micromovements. In some embodiments, the determined facial skin micromovements may correspond to muscle activation.
[0312] According to some disclosed embodiments, a reference signal for authentication can correspond to muscle activation during the pronunciation of at least one word. The term "authentication" (and other forms of the term, such as authenticate) refers to determining the identity of an individual or determining whether the individual is actually the person the individual claims to be. In some embodiments, authentication is a security process that relies on unique characteristics of an individual to identify who they are or to verify that they are who they claim to be. For example, authentication can be a security measure that matches biometric characteristics of an individual, such as an individual who wishes to access a resource (e.g., a device, a system, a service). As used herein, the term "pronunciation" (or other forms, such as pronounce, pronunciation, etc.) refers to when an individual actually says (or vocalizes) at least one word (or syllable, etc.) or before the individual actually says the word (e.g., during silent speech or pre-articulation). As explained elsewhere in this disclosure, speech-related muscle activity occurs before vocalization (e.g., when there is no airflow from the lungs but facial muscles express the desired sound, when some air flows from the lungs but the word is expressed in a way that is not perceivable using an audio sensor, etc.). For example, reference Figure 15 , Figure 16A and Figure 16B , a reference signal 1502 that can be used to verify the correspondence between a particular individual and an account at an institution can correspond to a signal caused by muscle activation that occurs during or before the vocalization of at least one word (e.g., during silent speech). It should be noted that a real-time signal 1508 (described below) can also be generated in a similar manner.
[0313] Some disclosed embodiments include muscle activation associated with at least one specific muscle, the at least one specific muscle including the zygomaticus muscle, orbicularis oris muscle, risorius muscle, genioglossus muscle, or levator labii superioris alaeque nasi muscle. "Muscle activation" refers to the tension, force, and / or movement of a muscle. Such activation can occur when the brain recruits the muscle. In some embodiments, as explained elsewhere in the present disclosure, muscle activation or muscle recruitment is the process of activating motor neurons to produce muscle contraction. Also as explained elsewhere in the present disclosure, facial skin micromovements include various types of voluntary and involuntary movements (e.g., in the range of micrometers to millimeters and lasting from fractions of a second to several seconds) caused by muscle recruitment or muscle activation. Some muscles, such as the quadriceps (which is a powerful muscle group responsible for very quickly displaying force), have a high ratio of muscle fibers to motor neurons. Other muscles, such as the eye muscles, have a much lower ratio because they use more precise, delicate movements, resulting in small-scale skin deformations. As explained elsewhere in the present disclosure, the zygomaticus muscle, orbicularis oris muscle, risorius muscle, genioglossus muscle, and levator labii superioris alaeque nasi muscle can articulate specific points in the cheeks above the individual's oral cavity, the chin, the mid-jaw, the cheeks below the oral cavity, the high cheeks, and the posterior cheeks. In some embodiments, the reference signal for authentication can be based on facial micromovements (e.g., based on the reflection of coherent light) detected from an individual's face when the individual is engaged in normal activities (e.g., normal speech, silently reading something, etc.). In some embodiments, a reference signal can be generated based on facial skin micromovements when the individual utters or silently utters (pronounces, enunciates, vocalizes, etc.) a selected word, syllable, or phrase.
[0314] According to some disclosed embodiments, the authentication operation can further include presenting at least one word to be pronounced to a specific individual. As used herein, the term "present" generally means to make something known. For example, in some embodiments, a word can be presented to an individual by visually displaying the word to the individual, and the individual can attempt to pronounce the displayed word. In some embodiments, one or more words can also be presented to the individual audibly, and the individual can repeat or attempt to repeat the word, and a signal can be generated when the individual utters the presented word or before uttering the word. In some embodiments, one or more diagrams representing one or more words to be pronounced (e.g., dog, cat) can be presented to the individual.
[0315] For example, one or more words to be pronounced (words, sentences, etc.) can be presented to an individual, and a reference signal 1502 (and / or real-time signal 1508) can be generated based on the facial micromovements produced by the individual emitting the sound of one or more of the presented words or one or more syllables in the words. One or more words can be presented to the individual to pronounce in any manner and on any device. For example, refer to Figure 14, in some embodiments, the words used to generate the reference signal 1502 (and / or the real-time signal 1508) can be textually displayed on the display screen 1402 of the mobile communication device 120 to an individual, and when the user pronounces the displayed words, the reference signal 1502 (and / or the real-time signal 1508) can be generated. In some embodiments, at least one word can be presented to the user graphically. For example, an image (such as a picture, cartoon, etc.) representing a word (such as dog, cat, etc.) can be displayed to the individual, and when the individual pronounces the word represented by the image, the reference signal 1502 (and / or the real-time signal 1508) can be generated. Generally, any word (such as a random word) or multiple words to be pronounced can be presented to the individual.
[0316] According to some disclosed embodiments, presenting at least one word to be pronounced to a specific individual includes presenting at least one word textually. For example, presenting the word "dog" can be done by textually displaying the word "dog". In some embodiments, presenting the word "dog" can occur by graphically showing an image of a dog (a picture, cartoon, line drawing, or another similar graphical display). For example, one or more words (words, sentences, etc.) to be pronounced can be presented to the individual, and the reference signal 1502 (and / or the real-time signal 1508) can be generated based on the facial micro-movements generated by the individual pronouncing one or more of the presented words or one or more syllables of the words. One or more words to be pronounced can be presented to the individual in any manner and on any device. For example, in some embodiments, the words can be textually displayed on the display screen 1402 of the mobile communication device 120 to the individual, and when the user pronounces the displayed words, the reference signal 1502 (and / or the real-time signal 1508) can be generated. In some embodiments, at least one word can be presented to the user graphically. For example, an image (such as a picture, cartoon, etc.) representing a word (such as dog, cat, etc.) can be displayed to the individual, and when the individual pronounces the word represented by the image, the reference signal 1502 (and / or the real-time signal 1508) can be generated. Generally, any word (such as a random word) or multiple words to be pronounced can be presented to the individual.
[0317] According to some disclosed embodiments, presenting at least one word to be pronounced to a particular individual includes audibly presenting the at least one word. For example, one or more words can be presented to the individual by audibly emitting the words, for example, on a speaker. For example, referring to FIG. 16, when an individual establishes an account at an institution, one or more words to be pronounced can be presented to the individual, and a reference signal 1502 can be generated based on the resulting facial micro-movements. As another example, when conducting a transaction (e.g., setting up an account or attempting to access an account) with an institution 1400 using a mobile communication device 120, the words used to generate the reference signal 1502 (and / or the real-time signal 1508) can be audibly presented to the individual using the speaker of the device 120, the output unit 114 of the speech detection system 100, or another speaker. And when the user pronounces a word or one or more syllables of a word, the speech detection system 100 associated with the individual can generate the reference signal 1502 (and / or the real-time signal 1508) based on muscle activation.
[0318] It should be noted that although the mobile communication device 120 is described as being used to audibly, textually, and / or graphically display the words used to generate the reference signal 1502 and / or the real-time signal 1508 to the individual, this is merely exemplary. Generally, the words can be presented to the individual on any device. For example, in some embodiments, the words can be presented visually (e.g., textually, graphically, etc.) on the screen 1600 of any device accessible to the individual (see Figure 16B ) (e.g., the visual display of a smart phone, a tablet computer, a smart watch, a personal digital assistant, a desktop computer, a laptop computer, an Internet of Things (IoT) device, a dedicated terminal, a wearable communication device, VR / XR glasses, etc.). Similarly, the words can be presented to the individual audibly on any device (e.g., the speaker of any of the above devices, etc.). It is also contemplated that in some embodiments, instead of presenting the words used to generate the reference signal 1502 and / or the real-time signal 1508 to the user, a question or prompt for generating the words can be presented to the user (e.g., audibly, textually, graphically, etc.). For example, such as "what is your password?", "what is the city of your birth", etc., and the reference signal 1502 (and / or the real-time signal 1508) can be generated based on the response. In some embodiments, both the reference signal 1502 and the real-time signal 1508 can be generated by presenting the same word or syllable to be pronounced to the individual.
[0319] According to some disclosed embodiments, at least one presented word can be a password. Generally, a "password" can be any word or string. In some embodiments, a password can be a string, one or more words, or phrases that must be used to obtain permission to something. For example, when an individual sets up an account at an institution, the individual can be required to utter (e.g., vocalize or pre-vocalize) the sound of the password of the account, and a reference signal 1502 can be generated based on the resulting facial micromovements. As another example, in an embodiment where an individual attempts to access an account of a customer at a financial institution, the individual can be required to utter the sound of the password associated with the account, e.g., by presenting a query (e.g., "what is your password?"). And when the individual utters the sound of the password, a reference signal 1502 and / or a real-time signal 1508 can be generated based on the reflection of coherent light from the individual's face.
[0320] In some embodiments, the reference signal for authentication can correspond to muscle activation during the pronunciation of one or more syllables. For example, when an individual utters (vocalizes or pre-vocalizes) the sound of a syllable (e.g., a vowel or any other syllable), a reference signal can be generated. Although not required, in some embodiments, one or more syllables (e.g., vowels or any other characters) or one or more words containing syllables can be presented to the individual, and when the individual utters the one or more syllables, the system 1500 can generate a reference signal 1502 (and / or a real-time signal 1508) for authentication based on facial micromovements.
[0321] Some disclosed embodiments include storing a correlation between the identity of a particular individual and a reference signal reflecting facial micro-movements in a secure data structure. A "secure data structure" is a location where data or information can be stored securely without unauthorized access. Unauthorized access can include access by a member within an organization (e.g., an institution, an authentication service provider, etc.) who is not authorized to access the stored data or access by a member outside the organization. The data structure according to the present disclosure can include any collection of data values and the relationships between the data. The data can be stored linearly, horizontally, hierarchically, relationally, non-relationally, one-dimensionally, multi-dimensionally, operably, in an ordered manner, in an unordered manner, in an object-oriented manner, in a centralized manner, in a decentralized manner, in a distributed manner, in a customized manner, or in any manner that enables data access. As a non-limiting example, the data structure can include arrays, associative arrays, linked lists, binary trees, balanced trees, heaps, stacks, queues, sets, hash tables, records, tagged unions, ER models, and graphs. For example, the data structure can include an XML database, an RDBMS database, an SQL database, or a NoSQL alternative for data storage / search, such as MongoDB, Redis, Couchbase, Datastax Enterprise Graph, ElasticSearch, Splunk, Solr, Cassandra, Amazon DynamoDB, Scylla, HBase, and Neo4J. The data structure can be a component of the disclosed system or a remote computing component (e.g., a cloud-based data structure). The data in the data structure can be stored in contiguous or non-contiguous memory. Additionally, as used herein, a data structure does not require the information to be located in the same place. It can be distributed across multiple servers, e.g., it can be owned or operated by the same or different entities. Thus, the phrase "data structure" used herein in the singular includes plural data structures.
[0322] In some embodiments, the secure data structure can be a secure database. The stored information can be encrypted in the secure data structure. As explained elsewhere in this disclosure, the term "database" can be a collection of data that can be distributed or non-distributed. In some embodiments, the secure data structure can be a secure enclave (also referred to as a trusted execution environment). A secure enclave is a computing environment that provides isolation of code and data from the operating system by using hardware-based isolation or by isolating an entire virtual machine by placing a hypervisor within the trusted computing base (TCB). The trusted computing base (TCB) can be a computing system that provides a secure environment for operations. This includes its hardware, firmware, software, operating system, physical location, built-in security controls, and prescribed security and assurance processes. A hypervisor (also referred to as a virtual machine monitor or VMM) is software that creates and runs virtual machines (VMs). The hypervisor allows a host computer to support multiple guest VMs by virtually sharing its resources such as memory and processing. Even a user with physical or root access to the machine and the operating system may not be able to access the contents of the secure enclave or tamper with the execution of the code inside the enclave. By isolating application code and data and encrypting the memory, the secure enclave provides CPU hardware-level isolation and memory encryption on the server. The secure enclave is at the core of confidential computing. In some embodiments, a collection of security-related instruction code can be built into the processor to protect the stored data. The data in the secure enclave can be protected because the enclave is only decrypted on-the-fly within the processor and then only for the code and data running inside the enclave itself. With suitable software, the secure enclave can enable encryption of the stored data and provide full-stack security to the stored data. In some embodiments, the secure enclave support can be incorporated into one or more processors (such as processor 1510) of system 1500. In some embodiments, the secure data structure can include an encrypted key / value store. In some embodiments, the secure data structure can be on a dedicated chip, in a separate IC circuit, or on a portion of processor 1510. In some embodiments, the secure data structure can include remote authentication. For example, corresponding authentication keys can be stored locally on system 1500 and on a remote server, and access to the stored database can be provided based on a successful comparison of the two authentication keys.
[0323] According to some disclosed embodiments, the correlation between the identity of a particular individual and a reference signal (reflecting the facial micro-movements of that individual) can be stored in a secure data structure. "Correlation" refers to the relationship or connection between an individual's identity and the reference signal of that individual. For example, correlation is a measure expressing the degree of relatedness between the two. In some embodiments, a representation (or signature) of the received reference signal of an individual can be stored as the correlation. Although not required, in some embodiments, the stored signature can be a reduced-size version of the received reference signal. In some embodiments, an encrypted version of the signature can be stored in the secure data structure. In some embodiments, the "hash" of the received reference signal can be stored as the correlation. As those of ordinary skill in the art will recognize, a hash is a unique digital signature generated from an input signal (e.g., the received reference signal) using, for example, a commercial algorithm. The hash / encrypted signature of an individual can be stored as the correlation in, for example, a secure data structure to reduce the likelihood of unauthorized access to the data. In some embodiments, the correlation can be or include, for example, features or characteristics of the reference signal extracted using a feature extraction algorithm. In some embodiments, the correlation can include important information or landmarks in the reference signal (e.g., the position and orientation of peaks and / or valleys, the spatial and / or temporal gaps between peaks and / or valleys). In some embodiments, the encrypted reference signal itself can be stored as the correlation. Since the stored correlation is a representation of the facial micro-movements of an individual affected by the individual's personal traits (e.g., muscle fiber structure, blood vessel structure, tissue structure, etc.), the stored correlation can uniquely identify the individual corresponding to the reference signal. In some embodiments, the correlation can include the identity (e.g., name, account number, or other identification information) of the individual corresponding to or associated with the reference signal. In one exemplary embodiment, as Figure 15 shown, system 1500 stores the correlation 1504 of the reference signal 1502 of an individual in a secure data structure in the memory 1520. As Figure 16A and Figure 16B shown, in another exemplary embodiment, system 1500 stores the correlations 1504 of the reference signals 1502 of different individuals (e.g., Tom, Amy, Ron, etc.) in a secure data structure (e.g., data structure 124) in a remote database.
[0324] Some disclosed embodiments include, after storage, receiving, via an institution, a request to authenticate a particular individual. As previously mentioned, the term "authenticate" refers to determining the identity of an individual or determining whether the individual is actually the person (implicitly or explicitly) claimed to be. In some embodiments, authentication is a security process that relies on unique characteristics of an individual to identify who they are or to verify that they are who they claim to be. For example, authentication is a security measure that matches biometric characteristics of an individual, such as an individual who wishes to access a resource (e.g., a device, a system, a service). In some embodiments, access to a resource is granted only when the biometric characteristics of an individual match the biometric characteristics stored in a security data structure for that particular individual. Consistent with its common usage, the term "request" is a demand for something. In some embodiments, the request can be an electronic or digital signal. For example, in some embodiments, as Figure 15 , Figure 16A and Figure 16B shown, system 1500 can receive a request 1506 to authenticate an individual. In some embodiments, the request 1506 can originate from an institution (e.g., institution 1400) with which the individual is engaged in a transaction. In some embodiments, the individual can send the request 1506 to the institution (e.g., as part of a transaction), and the institution can forward the request to system 1500.
[0325] In some embodiments, institution 1400 can send the request 1506 to an authentication service provider to authenticate an individual when it receives (or in response to) a transaction request from the individual. Without limitation, a transaction can include any type of interaction between two parties (e.g., the individual and institution 1400). In some embodiments, a transaction between the individual and institution 1400 can include the individual requesting that the institution 1400 take some action (e.g., request information, request access to an account, request a funds transfer, etc.).
[0326] According to some disclosed embodiments, authentication is associated with financial transactions at an institution. As explained elsewhere in this disclosure, the term "transaction" refers to any type of interaction between two parties (e.g., an individual and an institution). For example, an individual may request access to a customer account in a financial institution (e.g., a bank, a stock brokerage, etc.), and in response to that request, before allowing the individual to access the account and conduct another transaction, the institution may request that an authentication service authenticate the individual (e.g., verify that the individual requesting access is the customer associated with the account). When an individual seeks to conduct any type of transaction, the institution may seek authentication. According to some embodiments, a financial transaction includes at least one of the following: transfer of funds, purchase of stocks, sale of stocks, access to financial data, or access to an account of a specific individual. For example, an individual may attempt to trade stocks from an account with a stock brokerage, transfer funds out of an account, or view a financial statement, and the brokerage may send a request to system 1500 to authenticate the individual.
[0327] Any type of institution may use the disclosed systems and authentication services. According to some embodiments, the institution is associated with online activities and, upon authentication, provides a particular individual with permission to perform the online activities. The term "online activities" may refer to any activity performed using the Inter...
Claims
1. A head-mounted system for identifying an individual using facial skin micromovements, the head-mounted system comprising: A wearable housing configured to be worn on the head of an individual; At least one coherent light source associated with the wearable housing and configured to project light towards a facial region of the head; At least one detector associated with the wearable housing and configured to receive coherent light reflections from the facial region and output an associated reflection signal; At least one processor configured to: Analyze the reflection signal to determine specific facial skin micromovements of the individual; Access a memory that associates multiple facial skin micromovements with the individual; Search the memory for a match between the determined specific facial skin micromovements and at least one of the multiple facial skin micromovements; If a match is identified, initiate a first action; and If no match is identified, initiate a second action different from the first action.
2. The head-mounted system according to claim 1, wherein, The first action implements at least one predetermined setting associated with the individual.
3. The head-mounted system according to claim 1, wherein, The first action unlocks a computing device, and the second action includes presenting a message indicating that the computing device remains locked.
4. The head-mounted system according to claim 1, wherein, The first action provides personal information, and the second action provides public information.
5. The head-mounted system according to claim 1, wherein, The first action authorizes a transaction, and the second action provides information indicating that the transaction is not authorized.
6. The head-mounted system according to claim 1, wherein, The first action allows access to an application, and the second action blocks access to the application.
7. The head-mounted system according to claim 1, wherein At least some of the specific facial skin micromovements in the facial region are micromovements less than 100 micrometers.
8. The head-mounted system according to claim 1, wherein, The specific facial skin micromovements correspond to pre-articulatory muscle recruitment.
9. The head-mounted system according to claim 1, wherein, The specific facial skin micromovements correspond to muscle recruitment during the pronunciation of at least one word.
10. The head-mounted system according to claim 9, wherein, The at least one word corresponds to a password.
11. The head-mounted system according to claim 1, wherein, The memory is configured to associate multiple facial skin movements with multiple individuals, and wherein the at least one processor is configured to distinguish the multiple individuals from each other based on reflection signals unique to each of the multiple individuals.
12. The head-mounted system according to claim 1, the head-mounted system further comprising an integrated audio output, and wherein, At least one of the first actions or at least one of the second actions includes outputting audio via the audio output.
13. The head-mounted system according to claim 1, wherein, The match is identified when the at least one processor determines a degree of certainty.
14. The head-mounted system according to claim 13, wherein, When the degree of certainty is not initially reached, the at least one processor is configured to analyze additional reflection signals to determine additional facial skin micromovements and reach the degree of certainty at least in part based on the analysis of the additional reflection signals.
15. The head-mounted system according to claim 13, wherein, The at least one processor is further configured to continuously compare new facial skin micromovements with the multiple facial skin micromovements in the memory to determine an instantaneous degree of certainty.
16. The head-mounted system according to claim 15, wherein, After initiating the first action, when the instantaneous degree of certainty is below a threshold, the at least one processor is configured to stop the first action.
17. The head-mounted system according to claim 15, wherein, When the instantaneous degree of certainty is below a threshold, the at least one processor is configured to initiate an associated action.
18. The head-mounted system according to claim 15, wherein, Initiating the first action is associated with an event, and the at least one processor is configured to continuously compare the new facial skin micromovements during the event.
19. A method for identifying an individual using facial skin micromovements, the method comprising: Operating a wearable coherent light source configured to project light towards a facial region of an individual's head; Operating at least one detector configured to receive coherent light reflections from the facial region and output an associated reflection signal; Analyzing the reflection signal to determine specific facial skin micromovements of the individual; Accessing a memory that associates multiple facial skin micromovements with the individual; Searching the memory for a match between the determined specific facial skin micromovements and at least one of the multiple facial skin micromovements; If a match is identified, initiating a first action; and If no match is identified, initiating a second action different from the first action.
20. A non-transitory computer-readable medium containing instructions that, when executed by at least one processor, cause the at least one processor to perform operations for identifying an individual using facial skin micromovements, the operations comprising: Operating a wearable coherent light source configured to project light towards a facial region of an individual's head; Operating at least one detector configured to receive coherent light reflections from the facial region and output an associated reflection signal; Analyzing the reflection signal to determine specific facial skin micromovements of the individual; Accessing a memory that associates multiple facial skin micromovements with the individual; Searching the memory for a match between the determined specific facial skin micromovements and at least one of the multiple facial skin micromovements; If a match is identified, initiating a first action; and If no match is identified, initiating a second action different from the first action.
21. A non-transitory computer-readable medium containing instructions that, when executed by at least one processor, cause the at least one processor to perform operations for interpreting facial skin movements, the operations comprising: Projecting light onto multiple facial regions of an individual, wherein the multiple regions include at least a first region and a second region, and the first region is closer to at least one of the zygomaticus major muscle or the risorius muscle than the second region; Receiving reflections from the multiple regions; Detecting a first facial skin movement corresponding to the reflection from the first region and a second facial skin movement corresponding to the reflection from the second region; Based on the difference between the first facial skin movement and the second facial skin movement, determining that the reflection from the first region, which is closer to at least one of the zygomaticus major muscle or the risorius muscle, is a stronger communication indicator than the reflection from the second region; Based on the determination that the reflection from the first region is a stronger communication indicator, processing the reflection from the first region to determine the communication and ignoring the reflection from the second region.
22. The non-transitory computer-readable medium according to claim 21, wherein, The first region and the second region are spaced apart.
23. The non-transitory computer-readable medium according to claim 21, wherein, The AC determined based on the reflection from the first region includes words expressed by the individual.
24. The non-transitory computer-readable medium according to claim 21, wherein, The communication determined based on the reflection from the first region includes non-verbal cues of the individual.
25. The non-transitory computer-readable medium according to claim 21, wherein, The operation further includes operating a coherent light source located within the wearable housing in a manner capable of illuminating the plurality of facial regions.
26. The non-transitory computer-readable medium according to claim 21, wherein The operation further includes operating a coherent light source positioned away from the wearable housing in a manner capable of illuminating the plurality of facial regions.
27. The non-transitory computer-readable medium according to claim 21, wherein, The operation further includes irradiating at least a portion of the first region and at least a portion of the second region with a common light spot.
28. The non-transitory computer-readable medium according to claim 21, wherein The operation further includes irradiating the first region with a first set of light spots and irradiating the second region with a second set of light spots different from the first set of light spots.
29. The non-transitory computer-readable medium according to claim 21, wherein, The operation further includes operating a coherent light source in a manner to achieve dual-mode illumination of the plurality of facial regions, analyzing the reflection associated with the first illumination mode to identify one or more light spots associated with the first region, and analyzing the reflection associated with the second illumination mode to determine the communication.
30. The non-transitory computer-readable medium according to claim 29, wherein, A first light intensity of the first illumination mode is different from a second light intensity of the second illumination mode.
31. The non-transitory computer-readable medium according to claim 29, wherein, A first illumination pattern of the first illumination mode is different from a second illumination pattern of the second illumination mode.
32. The non-transitory computer-readable medium according to claim 21, wherein, The operation further includes determining, based on a difference between the first facial skin movement and the second facial skin movement, that the first region is closer to subcutaneous tissue associated with cranial nerve V or cranial nerve VII than the second region.
33. The non-transitory computer-readable medium according to claim 21, wherein, The first region is closer to the zygomaticus muscle than the second region, and the plurality of light regions further includes a third region closer to the zygomaticus muscle than each of the first region and the second region.
34. The non-transitory computer-readable medium according to claim 33, wherein, The operation further includes analyzing the reflected light from the first region when speech is generated in the presence of a perceivable vocalization, and analyzing the reflected light from the third region when speech is generated in the absence of a perceivable vocalization.
35. The non-transitory computer-readable medium according to claim 21, wherein, The difference between the first facial skin movement and the second facial skin movement includes a difference of less than 100 microns, and the determination that the reflection from the first region is a stronger indicator of communication compared to the reflection from the second region is based on the difference of less than 100 microns.
36. The non-transitory computer-readable medium according to claim 21, wherein, Ignoring the reflection from the second region includes omitting the use of the reflection from the second region to determine the communication.
37. The non-transitory computer-readable medium according to claim 21, wherein, Detecting the first facial skin movement includes performing a first speckle analysis on the light reflected from the first region, and detecting the second facial skin movement includes performing a second speckle analysis on the light reflected from the second region.
38. The non-transitory computer-readable medium according to claim 37, wherein, The first speckle analysis and the second speckle analysis occur concurrently by the at least one processor.
39. A method of interpreting facial skin movement, the method comprising: Projecting light onto a plurality of facial regions of an individual, wherein the plurality of regions includes at least a first region and a second region, the first region being closer to at least one of the zygomaticus muscle or the risorius muscle than the second region; Receiving reflections from the plurality of regions; Detect a first facial skin movement corresponding to a reflection from the first region and a second facial skin movement corresponding to a reflection from the second region; Based on a difference between the first facial skin movement and the second facial skin movement, determine that a reflection from the first region, which is closer to at least one of the zygomaticus muscle or the risorius muscle, is a stronger communication indicator than a reflection from the second region; Based on the determination that the reflection from the first region is a stronger communication indicator, process the reflection from the first region to determine the communication and ignore the reflection from the second region.
40. A system for interpreting facial skin movements, the system comprising: At least one processor configured to: Project light onto a plurality of facial regions of an individual, wherein the plurality of regions includes at least a first region and a second region, and the first region is closer to at least one of the zygomaticus muscle or the risorius muscle than the second region; Receive reflections from the plurality of regions; Detect a first facial skin movement corresponding to a reflection from the first region and a second facial skin movement corresponding to a reflection from the second region; Based on a difference between the first facial skin movement and the second facial skin movement, determine that a reflection from the first region, which is closer to at least one of the zygomaticus muscle or the risorius muscle, is a stronger communication indicator than a reflection from the second region; Based on the determination that the reflection from the first region is a stronger communication indicator, process the reflection from the first region to determine the communication and ignore the reflection from the second region.
41. A non-transitory computer-readable medium comprising instructions that, when executed by at least one processor, cause the at least one processor to perform an authentication operation based on facial micromovements, the operation comprising: Receiving, in a trustworthy manner, a reference signal for verifying a correspondence between a particular individual and an account at an institution, the reference signal being based on reference facial micromovements detected using a first coherent light reflected from the face of the particular individual; Storing a correlation between the identity of the particular individual and the reference signal reflecting the facial micromovements in a secure data structure; After storing, receiving a request to authenticate the particular individual via the institution; Receiving a real-time signal indicative of a second coherent light reflection obtained from second facial micromovements of the particular individual; Comparing the real-time signal with the reference signal stored in the secure data structure to authenticate the particular individual; And Notifying the institution that the particular individual is authenticated upon authentication.
42. The non-transitory computer-readable medium according to claim 41, wherein, The authentication is associated with a financial transaction at the institution.
43. The non-transitory computer-readable medium according to claim 42, wherein, The financial transaction includes at least one of the following: transfer of funds, purchase of stocks, sale of stocks, access to financial data, or access to an account of the particular individual.
44. The non-transitory computer-readable medium according to claim 41, wherein, Receiving the real-time signal and comparing the real-time signal occur multiple times during the transaction, and wherein the operation further includes reporting a mismatch if a subsequent difference is detected after the notification.
45. The non-transitory computer-readable medium according to claim 44, wherein, The operation further includes determining a degree of certainty that an individual associated with the real-time signal is the specific individual.
46. The non-transitory computer-readable medium according to claim 45, wherein, When the degree of certainty is below a threshold, the operation further includes terminating the transaction.
47. The non-transitory computer-readable medium according to claim 45, wherein, The transaction is a financial transaction that includes providing access to the specific individual's account, and when the degree of certainty is below the threshold, the operation further includes preventing the individual associated with the real-time signal from accessing the specific individual's account.
48. The non-transitory computer-readable medium according to claim 41, wherein, The reference signal for authentication corresponds to muscle activation during the pronunciation of at least one word.
49. The non-transitory computer-readable medium according to claim 48, wherein, The muscle activation is associated with at least one specific muscle, and the at least one specific muscle includes: zygomaticus major, orbicularis oris, risorius, genioglossus, or levator labii superioris alaeque nasi.
50. The non-transitory computer-readable medium according to claim 48, wherein, The at least one word is a password.
51. The non-transitory computer-readable medium according to claim 48, wherein, The operation further includes presenting the at least one word to the specific individual for pronunciation.
52. The non-transitory computer-readable medium according to claim 51, wherein, Presenting the at least one word to the specific individual for pronunciation includes audibly presenting the at least one word.
53. The non-transitory computer-readable medium according to claim 51, wherein, Presenting the at least one word to the specific individual for pronunciation includes presenting the at least one word in text form.
54. The non-transitory computer-readable medium according to claim 41, wherein, The reference signal for authentication corresponds to muscle activation during the pronunciation of one or more syllables.
55. The non-transitory computer-readable medium according to claim 41, wherein, The institution is associated with an online activity, and at the time of authentication, provides the specific individual with permission to perform the online activity.
56. The non-transitory computer-readable medium according to claim 55, wherein, The online activity is at least one of the following: a financial transaction, a betting session, an account access session, a gaming session, an exam, a lecture, or an educational session.
57. The non-transitory computer-readable medium according to claim 41, wherein, The institution is associated with a resource, and at the time of authentication, provides the specific individual with access to the resource.
58. The non-transitory computer-readable medium according to claim 57, wherein, The resource is at least one of the following: a file, a folder, a data structure, a computer program, computer code, or computer settings.
59. A method for providing authentication based on facial micro-movements, the method comprising: Receiving, in a trustworthy manner, a reference signal for verifying the correspondence between a specific individual and an account at an institution, the reference signal being obtained based on reference facial micro-movements detected using a first coherent light reflected from the face of the specific individual; Storing, in a secure data structure, the correlation between the identity of the specific individual and the reference signal reflecting the facial micro-movements; After storing, receiving, via the institution, a request to authenticate the specific individual; Receiving a real-time signal indicating a second coherent light reflection obtained from second facial micro-movements of the specific individual; Comparing the real-time signal with the reference signal stored in the secure data structure to authenticate the specific individual; And At the time of authentication, notifying the institution that the specific individual is authenticated.
60. A system for providing authentication based on facial micro-movements, the system comprising: At least one processor configured to: Receiving, in a trustworthy manner, a reference signal for verifying the correspondence between a specific individual and an account at an institution, the reference signal being obtained based on reference facial micro-movements detected using a first coherent light reflected from the face of the specific individual; Store the correlation between the identity of the specific individual and the reference signal reflecting the facial micro-movement in a secure data structure; After storage, receive a request to authenticate the specific individual via the institution; Receive a real-time signal indicating a second coherent light reflection obtained from a second facial micro-movement of the specific individual; Compare the real-time signal with the reference signal stored in the secure data structure to authenticate the specific individual; And Upon authentication, notify the institution that the specific individual is authenticated.
61. A non-transitory computer-readable medium containing instructions that, when executed by at least one processor, cause the at least one processor to perform operations for continuous authentication based on facial skin micro-movements, the operations including: Receive a first signal during an ongoing electronic transaction, the first signal representing a coherent light reflection associated with a first facial skin micro-movement during a first time period; Use the first signal to determine the identity of a specific individual associated with the first facial skin micro-movement; Receive a second signal during the ongoing electronic transaction, the second signal representing a coherent light reflection associated with a second facial skin micro-movement, the second signal being received during a second time period after the first time period; Use the second signal to determine that the specific individual is also associated with the second facial skin micro-movement; Receive a third signal during the ongoing electronic transaction, the third signal representing a coherent light reflection associated with a third facial skin micro-movement, the third signal being received during a third time period after the second time period; Use the third signal to determine that the third facial skin micro-movement is not associated with the specific individual; and Based on the determination that the third facial skin micro-movement is not associated with the specific individual, initiate an action.
62. The non-transitory computer-readable medium according to claim 61, wherein, The ongoing electronic transaction is a phone call.
63. The non-transitory computer-readable medium according to claim 61, wherein, During the second time period, the operation further includes continuously outputting data confirming the association of the specific individual with the second facial skin micro-movement.
64. The non-transitory computer-readable medium according to claim 61, wherein, The action includes providing an indication that the third facial skin micro-movement was not caused by the specific individual.
65. The non-transitory computer-readable medium according to claim 61, wherein, The action includes performing a process for identifying another individual who caused the third facial skin micro-movement.
66. The non-transitory computer-readable medium according to claim 61, wherein, The first time period, the second time period, and the third time period are part of a single online activity associated with the ongoing electronic transaction.
67. The non-transitory computer-readable medium according to claim 66, wherein, The online activity is at least one of the following: a financial transaction, a betting session, an account access session, a gaming session, an exam, a lecture, or an educational session.
68. The non-transitory computer-readable medium according to claim 66, wherein, The online activity includes multiple sessions, and the operation further includes using the received signals associated with facial skin micro-movements to determine that the specific individual participates in each of the multiple sessions.
69. The non-transitory computer-readable medium according to claim 66, wherein, The action includes notifying an entity associated with the online activity that an individual other than the specific individual is now participating in the online activity.
70. The non-transitory computer-readable medium according to claim 76, wherein, The action includes blocking participation in the online activity until the identity of the specific individual is confirmed.
71. The non-transitory computer-readable medium according to claim 61, wherein, The first time period, the second time period, and the third time period are part of a secure session in which the resource can be accessed.
72. The non-transitory computer-readable medium according to claim 71, wherein, The resource is at least one of the following: a file, a folder, a database, a computer program, computer code, or computer settings.
73. The non-transitory computer-readable medium according to claim 71, wherein, The action includes notifying an entity associated with the resource that an individual other than the particular individual has obtained access to the resource.
74. The non-transitory computer-readable medium according to claim 71, wherein, The action includes terminating access to the resource.
75. The non-transitory computer-readable medium according to claim 61, wherein, The first time period, the second time period, and the third time period are part of a single communication session, and wherein the communication session is at least one of the following: a telephone call, a teleconference, a video conference, or a real-time virtual communication.
76. The non-transitory computer-readable medium according to claim 75, wherein, The action includes notifying an entity associated with the communication session that an individual other than the particular individual has joined the communication session.
77. The non-transitory computer-readable medium according to claim 61, wherein, Determining the identity of the particular individual includes accessing a memory that associates a plurality of reference facial skin micromovements with individuals, and determining a match between the first facial skin micromovement and at least one of the plurality of reference facial skin micromovements.
78. The non-transitory computer-readable medium according to claim 61, wherein, The operation further includes determining the first facial skin micromovement, the second facial skin micromovement, and the third facial skin micromovement by analyzing signals indicative of received coherent light reflections to identify temporal and intensity variations of the light spots.
79. A method for continuous authentication based on facial skin micromovements, the method comprising: Receiving a first signal during an ongoing electronic transaction, the first signal representing coherent light reflections associated with a first facial skin micromovement during a first time period; Using the first signal to determine the identity of a particular individual associated with the first facial skin micromovement; Receiving a second signal during the ongoing electronic transaction, the second signal representing coherent light reflections associated with a second facial skin micromovement, the second signal being received during a second time period after the first time period; Using the second signal to determine that the particular individual is also associated with the second facial skin micromovement; Receiving a third signal during the ongoing electronic transaction, the third signal representing coherent light reflections associated with a third facial skin micromovement, the third signal being received during a third time period after the second time period; Using the third signal to determine that the third facial skin micromovement is not associated with the particular individual; and Initiating an action based on the determination that the third facial skin micromovement is not associated with the particular individual.
80. A system for providing authentication based on facial micromovements, the system comprising: At least one processor configured to: Receive a first signal during an ongoing electronic transaction, the first signal representing coherent light reflections associated with a first facial skin micromovement during a first time period; Using the first signal to determine the identity of a particular individual associated with the first facial skin micromovement; Receive a second signal during the ongoing electronic transaction, the second signal representing coherent light reflections associated with a second facial skin micro - movement, the second signal being received during a second time period after the first time period; Use the second signal to determine that the specific individual is also associated with the second facial skin micro - movement; Receive a third signal during the ongoing electronic transaction, the third signal representing coherent light reflections associated with a third facial skin micro - movement, the third signal being received during a third time period after the second time period; Use the third signal to determine that the third facial skin micro - movement is not associated with the specific individual; and Initiate an action based on the determination that the third facial skin micro - movement is not associated with the specific individual.
81. A non - transitory computer - readable medium containing instructions that, when executed by at least one processor, cause the at least one processor to perform thresholding operations for interpreting facial skin micro - movements, the operations including: Detect the facial skin micro - movement in the absence of a perceivable vocalization associated with the facial micro - movement; Determine the intensity level of the facial skin micro - movement; Compare the determined intensity level with a threshold; When the intensity level is above the threshold, interpret the facial skin micro - movement; And When the intensity level drops below the threshold, ignore the facial skin micro - movement.
82. The non-transitory computer-readable medium according to claim 81, wherein, The operations further include implementing an adjustment of the threshold.
83. The non-transitory computer-readable medium according to claim 81, wherein, The threshold is changeable according to environmental conditions.
84. The non-transitory computer-readable medium according to claim 83, wherein, The environmental conditions include background noise level.
85. The non-transitory computer-readable medium according to claim 84, wherein, The operations further include receiving data indicating the background noise level and determining the value of the threshold based on the received data.
86. The non-transitory computer-readable medium according to claim 81, wherein, The threshold is changeable according to at least one physical activity participated in by the individual associated with the facial skin micro - movement.
87. The non-transitory computer-readable medium according to claim 86, wherein, The at least one physical activity includes walking, running, or breathing.
88. The non-transitory computer-readable medium according to claim 87, wherein, The operations further include receiving data indicating the at least one physical activity participated in by the individual and determining the value of the threshold based on the received data.
89. The non-transitory computer-readable medium according to claim 81, wherein, The threshold is customized for the user.
90. The non - transitory computer - readable medium according to claim 89, wherein the operations further include receiving a personalized threshold for the specific individual and storing the personalized threshold in a setting associated with the specific individual.
91. The non - transitory computer - readable medium according to claim 89, wherein the operations further include receiving a plurality of thresholds for a specific individual, each of the plurality of thresholds being associated with a different condition.
92. The non-transitory computer-readable medium according to claim 91, wherein, At least one of the different conditions includes the physical condition of the specific individual, the emotional condition of the specific individual, or the location of the specific individual.
93. The non-transitory computer-readable medium according to claim 92, wherein, The operations further include receiving data indicating the current condition of the specific individual and selecting one threshold from the plurality of thresholds based on the received data.
94. The non-transitory computer-readable medium according to claim 91, wherein, Interpretation of the facial skin micro - movement includes synthesizing speech associated with the facial skin micro - movement.
95. The non-transitory computer-readable medium according to claim 81, wherein, Interpretation of the facial skin micro - movement includes understanding and executing commands based on the facial skin micro - movement.
96. The non-transitory computer-readable medium according to claim 95, wherein, Executing the command includes generating a signal for triggering an action.
97. The non-transitory computer-readable medium according to claim 81, wherein, Determining the intensity level includes determining values associated with a series of micromovements over a period of time.
98. The non-transitory computer-readable medium according to claim 81, wherein, Facial micromovements having an intensity level that drops below the threshold can be interpreted but are still ignored.
99. A method for thresholded interpretation of facial skin micromovements, the method comprising: Detecting the facial micromovements in the absence of a perceivable vocalization associated with the facial micromovements; Determining the intensity level of the facial micromovements; Comparing the determined intensity level with a threshold; When the intensity level is above the threshold, interpreting the facial micromovements; And When the intensity level drops below the threshold, ignoring the facial micromovements.
100. A system for thresholded interpretation of facial skin micromovements, the system comprising: At least one processor configured to: Detect the facial micromovements in the absence of a perceivable vocalization associated with the facial micromovements; Determine the intensity level of the facial micromovements; Compare the determined intensity level with a threshold; When the intensity level is above the threshold, interpret the facial micromovements; And When the intensity level drops below the threshold, ignoring the facial micromovements.
101. A non-transitory computer-readable medium containing instructions that, when executed by at least one processor, cause the at least one processor to perform operations for establishing a silent conversation, the operations including: Establishing a wireless communication channel for enabling a silent conversation via a first wearable device and a second wearable device, wherein the first wearable device and the second wearable device each include a coherent light source and a light detector configured to detect facial skin micromovements from coherent light reflections; Detecting, by the first wearable device, first facial skin micromovements occurring in the absence of a perceivable vocalization; Transmitting, via the wireless communication channel, a first communication from the first wearable device to the second wearable device, wherein the first communication originates from the first facial skin micromovements and is transmitted through the second wearable device for presentation; Receiving, via the wireless communication channel, a second communication from the second wearable device, wherein the second communication originates from second facial skin micromovements detected by the second wearable device; and Presenting the second communication to a wearer of the first wearable device.
102. The non-transitory computer-readable medium according to claim 101, wherein, The first communication includes a signal reflecting the first facial skin micromovements.
103. The non-transitory computer-readable medium according to claim 101, wherein, The operations further include interpreting the first facial skin micromovements as words, and wherein the first communication includes transmission of the words.
104. The non-transitory computer-readable medium according to claim 101, wherein, Presenting the second communication to the wearer of the first wearable device includes synthesizing words derived from the second facial skin micromovements.
105. The non-transitory computer-readable medium according to claim 101, wherein, Presenting the second communication to the wearer of the first wearable device includes providing a text output reflecting words derived from the second facial skin micromovements.
106. The non-transitory computer-readable medium according to claim 101, wherein, Presenting the second communication to the wearer of the first wearable device includes providing a graphical output reflecting at least one facial expression derived from the second facial skin micromovements.
107. The non-transitory computer-readable medium according to claim 106, wherein, The graphical output includes at least one emoji.
108. The non-transitory computer-readable medium according to claim 101, wherein, The operation further includes determining that the second wearable device is located near the first wearable device.
109. The non-transitory computer-readable medium according to claim 108, wherein, The operation further includes automatically establishing the wireless communication channel between the first wearable device and the second wearable device.
110. The non-transitory computer-readable medium according to claim 108, wherein, The operation further includes presenting, via the first wearable device, a suggestion to establish a silent conversation with the second wearable device.
111. The non-transitory computer-readable medium according to claim 101, wherein, The operation further includes determining an intention of a wearer of the first wearable device to initiate a silent conversation with a wearer of the second wearable device, and automatically establishing the wireless communication channel between the first wearable device and the second wearable device.
112. The non-transitory computer-readable medium according to claim 111, wherein, The intention is determined from the first facial skin micro-movement.
113. The non-transitory computer-readable medium according to claim 101, wherein, The wireless communication channel is directly established between the first wearable device and the second wearable device.
114. The non-transitory computer-readable medium according to claim 101, wherein, The wireless communication channel is established from the first wearable device to the second wearable device via at least one intermediate communication device.
115. The non-transitory computer-readable medium according to claim 114, wherein, The at least one communication device includes at least one of the following: a first smart phone associated with a wearer of the first wearable device, a second smart phone associated with a wearer of the second wearable device, a router, or a server.
116. The non-transitory computer-readable medium according to claim 101, wherein, The first communication includes a signal reflecting a first word spoken in a first language, and the second communication includes a signal reflecting a second word spoken in a second language, and wherein presenting the second communication to the wearer of the first wearable device includes translating the second word into the first language.
117. The non-transitory computer-readable medium according to claim 101, wherein, The first communication includes details identifying the wearer of the first wearable device, and the second communication includes a signal identifying the wearer of the second wearable device.
118. The non-transitory computer-readable medium according to claim 101, wherein, The first communication includes a timestamp indicating when the first facial skin micro-movement is detected.
119. A method for establishing a silent conversation, the method comprising: Establishing a wireless communication channel for implementing a silent conversation via a first wearable device and a second wearable device, wherein both the first wearable device and the second wearable device include a coherent light source and a light detector, and the light detector is configured to detect facial skin micro-movements based on coherent light reflection; Detecting, by the first wearable device, a first facial skin micro-movement occurring without perceivable vocalization; Transmitting, via the wireless communication channel, a first communication from the first wearable device to the second wearable device, wherein the first communication originates from the first facial skin micro-movement and is transmitted for presentation to a wearer of the second wearable device; Receiving, via the wireless communication channel, a second communication from the second wearable device, wherein the second communication originates from a second facial skin micro-movement detected by the second wearable device; and Presenting the second communication to the wearer of the first wearable device.
120. A system for establishing a non-vocal conversation, the system comprising: At least one processor, the at least one processor being configured to: Establish a wireless communication channel for silent conversation via a first wearable device and a second wearable device, wherein both the first wearable device and the second wearable device include a coherent light source and a photodetector, and the photodetector is configured to detect facial skin micro-movements based on coherent light reflection; Detect, by the first wearable device, a first facial skin micro-movement that occurs without perceivable vocalization; Transmit a first communication from the first wearable device to the second wearable device via the wireless communication channel, wherein the first communication originates from the first facial skin micro-movement and is transmitted for presentation to the wearer of the second wearable device; Receive a second communication from the second wearable device via the wireless communication channel, wherein the second communication originates from a second facial skin micro-movement detected by the second wearable device; and Present the second communication to the wearer of the first wearable device.
121. A non-transitory computer-readable medium containing instructions that, when executed by at least one processor, cause the at least one processor to initiate a content interpretation operation before vocalizing the content of the interpretation, the operation including: Receiving a signal representing facial skin micro-movements; Determining, based on the signal, the at least one word to be spoken before vocalizing at least one word in the source language; Beginning an interpretation of the at least one word before vocalization of the at least one word; And Causing an interpretation of the at least one word to be presented when the at least one word is spoken.
122. The non-transitory computer-readable medium according to claim 121, wherein, The interpretation is to translate the at least one word from the source language into at least one target language different from the source language.
123. The non-transitory computer-readable medium according to claim 122, wherein, The interpretation of the at least one word includes transcribing the at least one word into text in the at least one target language.
124. The non-transitory computer-readable medium according to claim 122, wherein, The interpretation of the at least one word includes speech synthesis of the at least one word in the at least one target language.
125. The non-transitory computer-readable medium according to claim 122, wherein the operation further includes receiving a selection of the at least one target language.
126. The non-transitory computer-readable medium according to claim 125, wherein, The selection of the at least one target language includes a selection of multiple target languages, and wherein presenting an interpretation of the at least one word includes causing simultaneous presentation in the multiple languages.
127. The non-transitory computer-readable medium according to claim 121, wherein, The interpretation of the at least one word includes transcribing the at least one word into text in the source language.
128. The non-transitory computer-readable medium according to claim 127, wherein, Presenting an interpretation of the at least one word includes outputting a display of the transcribed text and a video of an individual associated with the facial skin micro-movements.
129. The non-transitory computer-readable medium according to claim 121, wherein, Receiving the signal occurs via at least one detector of coherent light reflection from the facial region of the person vocalizing the at least one word.
130. The non-transitory computer-readable medium according to claim 129, wherein, Presenting an interpretation of the at least one word occurs simultaneously with the vocalization of the at least one word by the person.
131. The non-transitory computer-readable medium according to claim 121, wherein, Presenting an interpretation of the at least one word includes outputting an audible presentation of the at least one word using a wearable speaker.
132. The non-transitory computer-readable medium according to claim 121, wherein, Presenting an interpretation of the at least one word includes transmitting an audio signal via a network.
133. The non-transitory computer-readable medium according to claim 121, wherein the operation further comprises: Determine at least one expected word to be spoken after the at least one word to be spoken, and start interpreting the at least one expected word before the utterance of the at least one word; and when the at least one word is spoken, present the interpretation of the at least one expected word after the presentation of the at least one word.
134. The non-transitory computer-readable medium according to claim 121, wherein, Presenting the interpretation of the at least one word includes transmitting a text translation of the at least one word over a network.
135. The non-transitory computer-readable medium according to claim 121, wherein, The operation further includes determining at least one non-verbal interjection based on the signal, and outputting a representation of the non-verbal interjection.
136. The non-transitory computer-readable medium according to claim 121, wherein, Determining at least one word based on the signal includes using speckle analysis to interpret the facial skin micro-movement.
137. The non-transitory computer-readable medium according to claim 121, wherein, The signal representing the facial skin micro-movement corresponds to muscle activation prior to the utterance of the at least one word.
138. The non-transitory computer-readable medium according to claim 137, wherein, The muscle activation is associated with at least one specific muscle, the at least one specific muscle including: zygomaticus major, orbicularis oris, risorius, genioglossus, or levator labii superioris alaeque nasi.
139. A method for initiating content interpretation prior to the utterance of content to be interpreted, the method comprising: Receiving a signal representing facial skin micro-movement; Determining the at least one word to be spoken based on the signal before the at least one word is spoken in the source language; Before the utterance of the at least one word, start interpreting the at least one word; And When the at least one word is spoken, present the interpretation of the at least one word.
140. A system for initiating content interpretation prior to the utterance of content to be interpreted, the system comprising: At least one processor configured to: Receive a signal representing facial skin micro-movement; Determine the at least one word to be spoken based on the signal before at least one word is spoken in the source language; Before the utterance of the at least one word, start interpreting the at least one word; And When the at least one word is spoken, present the interpretation of the at least one word.
141. A non-transitory computer-readable medium containing instructions that, when executed by at least one processor, cause the at least one processor to perform private voice assistant operations, the operations including: Receiving a signal indicating a specific facial skin micro-movement reflecting a private request for an assistant, wherein replying to the private request requires identifying a specific individual associated with the specific facial skin micro-movement; Accessing a data structure that maintains a correlation between the specific individual and a plurality of facial skin micro-movements associated with the specific individual; Searching the data structure for a match indicating a correlation between the stored identity of the specific individual and the specific facial skin micro-movement; In response to a determination that the match exists in the data structure, initiating a first action in response to the request, wherein the first action involves enabling access to unique information of the specific individual; and If the match is not identified in the data structure, initiate a second action different from the first action.
142. The non-transitory computer-readable medium according to claim 141, wherein, The second action includes providing non-private information.
143. The non-transitory computer-readable medium according to claim 141, wherein, The second action includes a notification of denying access to the unique information of the specific individual.
144. The non-transitory computer-readable medium according to claim 141, wherein, The second action includes blocking access to the unique information of the specific individual.
145. The non-transitory computer-readable medium according to claim 141, wherein, The second action includes attempting to authenticate the specific individual using additional data.
146. The non-transitory computer-readable medium according to claim 145, wherein, The additional data includes detected additional facial skin micromovements.
147. The non-transitory computer-readable medium according to claim 145, wherein, The additional data includes data other than facial skin micromovements.
148. The non-transitory computer-readable medium according to claim 141, wherein, When the match is not recognized, the operation further includes initiating an additional action for identifying an individual other than the specific individual.
149. The non-transitory computer-readable medium according to claim 148, wherein, In response to the identification of an individual other than the specific individual, the operation further includes initiating a third action in response to the request.
150. The non-transitory computer-readable medium according to claim 149, wherein, The third action involves enabling access to the unique information of the other individual.
151. The non-transitory computer-readable medium according to claim 141, wherein, The private request is for activating software code, the first action is activating the software code, and the second action is blocking the activation of the software code.
152. The non-transitory computer-readable medium according to claim 141, wherein, The private request is for confidential information, and the operation further includes determining that the specific individual has the authority to access the confidential information.
153. The non-transitory computer-readable medium according to claim 141, wherein, Receiving, accessing, and searching occur repeatedly during an ongoing session.
154. The non-transitory computer-readable medium according to claim 153, wherein, During a first time period of the ongoing session, the specific individual is identified and the first action is initiated, and wherein, during a second time period of the ongoing session, the specific individual is not identified, and any remaining first action is terminated to support the second action.
155. The non-transitory computer-readable medium according to claim 141, wherein, The operation further includes operating at least one coherent light source in a manner capable of illuminating a non-lip portion of the face of the individual making the private request, and wherein the signal is received via at least one detector of coherent light reflected from the non-lip portion of the face.
156. The non-transitory computer-readable medium according to claim 155, wherein, The at least one processor, the at least one coherent light source, and the at least one detector are integrated in a wearable housing configured to be supported by the ear of the individual.
157. The non-transitory computer-readable medium according to claim 155, wherein, The operation further includes analyzing the received signal to determine pre-articulatory muscle recruitment and determining the private request based on the determined pre-articulatory muscle recruitment.
158. The non-transitory computer-readable medium according to claim 155, wherein, The operation further includes determining the private request in the absence of a perceptible utterance of the private request.
159. A method of operating a private voice assistant, the method comprising: Receiving a signal indicative of specific facial skin micromovements reflecting a private request for the assistant, wherein replying to the private request requires identifying a specific individual associated with the specific facial skin micromovements; Accessing a data structure that maintains a correlation between the specific individual and a plurality of facial skin micromovements associated with the specific individual; Searching the data structure for a match indicating a correlation between the stored identity of the specific individual and the specific facial skin micromovements; In response to a determination that the match exists in the data structure, initiating a first action in response to the request, wherein the first action involves enabling access to the unique information of the specific individual; and If the match is not recognized in the data structure, initiating a second action different from the first action.
160. A system for operating a private voice assistant, the system comprising: At least one processor configured to: Receive a signal indicating a specific facial skin micro - movement reflecting a private request for the assistant, wherein responding to the private request requires identifying a specific individual associated with the specific facial skin micro - movement; Access a data structure that maintains a correlation between the specific individual and a plurality of facial skin micro - movements associated with the specific individual; Search the data structure for a match indicating a correlation between the stored identity of the specific individual and the specific facial skin micro - movement; In response to determining that the match exists in the data structure, initiate a first action in response to the request, wherein the first action involves enabling access to unique information of the specific individual; and If the match is not identified in the data structure, initiate a second action different from the first action.
161. A non - transitory computer - readable medium containing instructions that, when executed by at least one processor, cause the at least one processor to perform operations for determining a silent phoneme based on facial skin micro - movements, the operations comprising: Controlling at least one coherent light source in a manner that enables illumination of a first region of the face and a second region of the face; Performing a first pattern analysis on light reflected from the first region of the face to determine a first micro - movement of facial skin in the first region of the face; Performing a second pattern analysis on light reflected from the second region of the face to determine a second micro - movement of facial skin in the second region of the face; And Using the first micro - movement of facial skin in the first region of the face and the second micro - movement of facial skin in the second region of the face to determine at least one silent phoneme.
162. The non-transitory computer-readable medium according to claim 161, wherein, The execution of the second pattern analysis occurs after the execution of the first pattern analysis.
163. The non-transitory computer-readable medium according to claim 161, wherein, The execution of the second pattern analysis occurs simultaneously with the execution of the first pattern analysis.
164. The non-transitory computer-readable medium according to claim 161, wherein, The first region is spaced apart from the second region.
165. The non-transitory computer-readable medium according to claim 161, wherein, Determining the at least one silent phoneme includes determining a phoneme sequence, and wherein the operations further include extracting meaning from the phoneme sequence.
166. The non-transitory computer-readable medium according to claim 165, wherein, Each phoneme in the phoneme sequence is obtained from the first pattern analysis and the second pattern analysis.
167. The non-transitory computer-readable medium according to claim 165, wherein, The operations further include identifying at least one phoneme in the phoneme sequence as private and omitting the generation of an audio output reflecting the at least one private phoneme.
168. The non-transitory computer-readable medium according to claim 161, wherein, The operations further include determining both the first micro - movement and the second micro - movement during a common time period.
169. The non-transitory computer-readable medium according to claim 161, wherein, The operations further include receiving the first light reflection and the second light reflection via at least one detector, wherein the at least one detector and the at least one coherent light source are integrated in a wearable housing. The non-transitory computer-readable medium according to claim 161, wherein, Controlling the at least one coherent light source includes projecting different light patterns on the first region and the second region.
171. The non-transitory computer-readable medium according to claim 170, wherein, The different light patterns include a plurality of light spots such that a first region of the face is illuminated by at least a first light spot and a second region of the face is illuminated by at least a second light spot different from the first light spot.
172. The non-transitory computer-readable medium according to claim 161, wherein, Controlling the at least one coherent light source includes illuminating the first region and the second region with a common light spot.
173. The non-transitory computer-readable medium according to claim 161, wherein, The first micro - movement of the facial skin and the second micro - movement of the facial skin correspond to concurrent muscle recruitment, wherein a determined first micro - movement of the facial skin in the first region of the face corresponds to the recruitment of a first muscle selected from the following muscles: zygomaticus major, orbicularis oris, risorius or levator labii superioris alaeque nasi muscle, and a determined second micro - movement of the facial skin in the second region of the face corresponds to the recruitment of a second muscle different from the first muscle, the second muscle being selected from the following muscles: zygomaticus major, orbicularis oris, risorius or levator labii superioris alaeque nasi muscle.
174. The non-transitory computer-readable medium according to claim 161, wherein, The operation further includes accessing a default language of an individual associated with the facial skin micro - movement and extracting meaning from the at least one silent phoneme using the default language.
175. The non-transitory computer-readable medium according to claim 161, wherein, The operation further includes using synthetic speech to generate an audio output reflecting the at least one silent phoneme.
176. The non-transitory computer-readable medium according to claim 161, wherein, The at least one phoneme includes a sequence of phonemes, and wherein the operation further includes determining a prosody associated with the sequence of phonemes and extracting meaning based on the determined prosody.
177. The non-transitory computer-readable medium according to claim 161, wherein, The operation further includes determining an emotional state of an individual associated with the facial skin micro - movement and extracting meaning from the at least one silent phoneme and the determined emotional state. The non-transitory computer-readable medium according to claim 161, wherein, The operation further includes identifying at least one irrelevant phoneme as part of a filler and omitting the generation of an audio output reflecting the irrelevant phoneme.
179. A method for determining silent phonemes based on facial skin micro - movements, the method comprising: Controlling at least one coherent light source in a manner that enables illumination of a first region of a face and a second region of the face; Performing a first pattern analysis on light reflected from the first region of the face to determine a first micro - movement of the facial skin in the first region of the face; Performing a second pattern analysis on light reflected from the second region of the face to determine a second micro - movement of the facial skin in the second region of the face; And Using the first micro - movement of the facial skin in the first region of the face and the second micro - movement of the facial skin in the second region of the face to determine at least one silent phoneme.
180. A system for determining silent phonemes based on facial skin micro - movements, the system comprising: At least one processor, the at least one processor being configured to: Control at least one coherent light source in a manner that enables illumination of a first region of a face and a second region of the face; Perform a first pattern analysis on light reflected from the first region of the face to determine a first micro - movement of the facial skin in the first region of the face; Perform a second pattern analysis on light reflected from the second region of the face to determine a second micro - movement of the facial skin in the second region of the face; And Determine at least one silent phoneme using the first micromovement of the facial skin in the first region of the face and the second micromovement of the facial skin in the second region of the face.
181. A non-transitory computer-readable medium comprising instructions that, when executed by at least one processor, cause the at least one processor to perform operations for generating a synthetic representation of a facial expression, the operations including: Controlling at least one coherent light source in a manner capable of illuminating a portion of the face; Receiving an output signal from a light detector, where the output signal corresponds to the reflection of coherent light from the portion of the face; Applying speckle analysis to the output signal to determine facial skin micromovements based on speckle analysis; Using the determined facial skin micromovements based on speckle analysis to identify at least one word that is pre-articulated or articulated during a period of time; Using the determined facial skin micromovements based on speckle analysis to identify at least one change in the facial expression during the period of time; and During the period of time, outputting data for causing a virtual representation of the face to mimic the at least one change in the facial expression in combination with an audio rendering of the at least one word.
182. The non-transitory computer-readable medium according to claim 181, wherein, Controlling the at least one coherent light source in a manner that enables illumination of the portion of the face includes projecting a light pattern onto the portion of the face.
183. The non-transitory computer-readable medium according to claim 182, wherein, The light pattern includes a plurality of light spots.
184. The non-transitory computer-readable medium according to claim 182, wherein, The portion of the face includes cheek skin.
185. The non-transitory computer-readable medium according to claim 182, wherein, The portion of the face does not include the lips.
186. The non-transitory computer-readable medium according to claim 181, wherein, The output signal from the light detector is emitted from a wearable device.
187. The non-transitory computer-readable medium according to claim 181, wherein, The output signal from the light detector is emitted from a non-wearable device.
188. The non-transitory computer-readable medium according to claim 181, wherein, The determined facial skin micromovements based on speckle analysis are associated with the recruitment of at least one of the following muscles: zygomaticus major, orbicularis oris, genioglossus, risorius, or levator labii superioris alaeque nasi.
189. The non-transitory computer-readable medium according to claim 181, wherein, The at least one change in the facial expression during the period of time includes speech-related facial expressions and non-speech-related facial expressions. The non-transitory computer-readable medium according to claim 189, wherein, The virtual representation of the face is associated with an avatar of the individual from whom the output signal is derived, and where mimicking the at least one change in the facial expression includes causing a visual change in the avatar that reflects at least one of the speech-related facial expression and the non-speech-related facial expression. The non-transitory computer-readable medium according to claim 190, wherein, The visual change in the avatar involves changing the color of at least a portion of the avatar. The non-transitory computer-readable medium according to claim 181, wherein, The audio rendering of the at least one word is based on a recording of an individual. The non-transitory computer-readable medium according to claim 181, wherein, The audio rendering of the at least one word is based on synthetic speech.
194. The non-transitory computer-readable medium according to claim 193, wherein, The synthetic speech corresponds to the speech of the individual from whom the output signal is derived. The non-transitory computer-readable medium according to claim 193, wherein, The synthetic speech corresponds to a template speech selected by the individual from whom the output signal is derived. The non-transitory computer-readable medium according to claim 181, wherein, The operations further include determining, at least in part based on the facial skin micromovements, the emotional state of the individual from whom the output signal is derived, and enhancing the virtual representation of the face to reflect the determined emotional state. The non-transitory computer-readable medium according to claim 181, wherein, The operations further include receiving a selection of a desired emotional state, and enhancing the virtual representation of the face to reflect the selected emotional state. The non-transitory computer-readable medium according to claim 181, wherein, The operation further includes identifying an undesired facial expression, and wherein the output data for causing the virtual representation omits data for causing the undesired facial expression.
199. A method of generating a synthetic representation of a facial expression, the method comprising: Controlling at least one coherent light source in a manner capable of illuminating a portion of a face; Receiving an output signal from a light detector, wherein the output signal corresponds to the reflection of coherent light from the portion of the face; Applying speckle analysis to the output signal to determine facial skin micromovements based on the speckle analysis; Using the determined facial skin micromovements based on the speckle analysis to identify at least one word pre-spoken or spoken during a period of time; Using the determined facial skin micromovements based on the speckle analysis to identify at least one change in facial expression during the period of time; and During the period of time, outputting data for causing a virtual representation of the face to mimic the at least one change in facial expression in combination with an audio rendering of the at least one word.
200. A system for generating a synthetic representation of a facial expression, the system comprising: At least one processor configured to: Control at least one coherent light source in a manner enabling illumination of a portion of a face; Receive an output signal from a light detector, wherein the output signal corresponds to the reflection of coherent light from the portion of the face; Apply speckle analysis to the output signal to determine facial skin micromovements based on the speckle analysis; Use the determined facial skin micromovements based on the speckle analysis to identify at least one word pre-spoken or spoken during a period of time; Use the determined facial skin micromovements based on the speckle analysis to identify at least one change in facial expression during the period of time; and During the period of time, outputting data for causing a virtual representation of the face to mimic the at least one change in facial expression in combination with an audio rendering of the at least one word.
201. A non-transitory computer-readable medium containing instructions that, when executed by at least one processor, cause the at least one processor to perform operations for attention-related interactions based on facial skin micromovements, the operations including: Determining facial skin micromovements of an individual based on the reflection of coherent light from a facial region of the individual; Using the facial skin micromovements to determine a specific level of focus of the individual; Receiving data associated with an expected interaction with the individual; Accessing a data structure associating information reflecting alternative levels of focus with different presentation modalities; Based on the specific level of focus and the relevance information, determining a specific presentation modality for the expected interaction; and And Associating the specific presentation modality with the expected interaction for subsequent interaction with the individual. The non-transitory computer-readable medium according to claim 201, wherein The operation further includes generating an output reflecting the expected interaction according to the determined specific presentation modality. The non-transitory computer-readable medium according to claim 201, wherein, The operation further includes: operating at least one coherent light source in a manner capable of illuminating a non-lip portion of the individual's face, and receiving a signal indicative of the coherent light reflection from the non-lip portion of the face.
204. The non-transitory computer-readable medium according to claim 203, wherein, The operation further includes performing speckle analysis on the coherent light reflection from the non-lip portion of the face to determine facial skin micromovements. The non-transitory computer-readable medium according to claim 201, wherein, The specific level of attentiveness is a category of attentiveness. The non-transitory computer-readable medium according to claim 201, wherein, The specific level of attentiveness includes the magnitude of attentiveness. The non-transitory computer-readable medium according to claim 201, wherein, The specific level of attentiveness reflects the degree to which the individual is engaged in an activity including at least one of conversation, thinking, or resting. The non-transitory computer-readable medium according to claim 207, wherein, The operation further includes determining the degree to which the individual is engaged in the activity based on facial skin micromovements corresponding to the recruitment of at least one muscle in a muscle group, the at least one muscle including: the zygomaticus major, orbicularis oris, risorius, or levator labii superioris alaeque nasi muscle. The non-transitory computer-readable medium according to claim 201, wherein, The received data associated with the expected interaction includes an incoming call, and wherein the associated different presentation manners include notifying the individual of the incoming call and routing the incoming call to voicemail. The non-transitory computer-readable medium according to claim 201, wherein, The received data associated with the expected interaction includes an incoming text message, and wherein the associated different presentation manners include presenting the text message to the individual in real time and delaying the presentation of the text message to a later time.
211. The non-transitory computer-readable medium according to claim 201, wherein, Determining the specific presentation manner for the expected interaction includes determining how to notify the individual of the expected interaction.
212. The non-transitory computer-readable medium according to claim 211, wherein, Determining how to notify the individual of the expected interaction is at least partially based on the identities of a plurality of electronic devices currently used by the individual.
213. The non-transitory computer-readable medium according to claim 201, wherein, The received data associated with the expected interaction indicates the importance level of the expected interaction, and wherein the specific presentation manner is determined at least partially based on the importance level.
214. The non-transitory computer-readable medium according to claim 201, wherein, The received data associated with the expected interaction indicates the urgency level of the expected interaction, and wherein the specific presentation manner is determined at least partially based on the specific urgency level.
215. The non-transitory computer-readable medium according to claim 201, wherein, The specific presentation manner includes delaying the presentation of content until a detected period of low attentiveness, and wherein the operation further includes detecting low attentiveness at a subsequent time and presenting the content at the subsequent time.
216. The non-transitory computer-readable medium according to claim 201, wherein, The operation further includes using the facial skin micromovements to determine that the individual is engaged in a conversation with another individual, determining whether the expected interaction is related to the conversation, and wherein the specific presentation manner is determined at least partially based on the relevance of the expected interaction to the conversation.
217. The non-transitory computer-readable medium according to claim 216, wherein, The operation further includes using the facial skin micromovements to determine the topic of the conversation, and wherein determining that the expected interaction is related to the conversation is based on the received data associated with the expected interaction and the topic of the conversation.
218. The non-transitory computer-readable medium according to claim 216, wherein, When it is determined that the expected interaction is related to the conversation, a first presentation manner is used for the expected interaction, and when it is determined that the expected interaction is not related to the conversation, a second presentation manner is used for the expected interaction.
219. A method for attention-related interaction based on facial skin micromovements, the method comprising: Determine the facial skin micromovements of the individual based on the reflection of coherent light from the facial region of the individual; Use the facial skin micromovements to determine the specific attentiveness of the individual; Receive data associated with an expected interaction with the individual; Access a data structure that associates information reflecting alternative attentiveness levels with different presentation modalities; Based on the specific attentiveness and the relevant information, determine a specific presentation modality for the expected interaction; And Associate the specific presentation modality with the expected interaction for subsequent interaction with the individual.
220. A system for attention-related interaction based on facial skin micromovements, the system comprising: At least one processor configured to: Determine the facial skin micromovements of the individual based on the reflection of coherent light from the facial region of the individual; Use the facial skin micromovements to determine the specific attentiveness of the individual; Receive data associated with an expected interaction with the individual; Access a data structure that associates information reflecting alternative attentiveness levels with different presentation modalities; Based on the specific attentiveness and the relevant information, determine a specific presentation modality for the expected interaction; And Associate the specific presentation modality with the expected interaction for subsequent interaction with the individual.
221. A non-transitory computer-readable medium containing instructions that, when executed by at least one processor, cause the at least one processor to perform a speech synthesis operation based on detected facial skin micromovements, the operation comprising: Based on the reflection of light from the facial region of a first individual, determine the specific facial skin micromovements of the first individual while conversing with a second individual; Access a data structure that associates facial micromovements with words; Perform a lookup in the data structure for a specific word associated with the specific facial skin micromovements; Obtain an input associated with the preferred speech consumption characteristics of the second individual; Adopt the preferred speech consumption characteristics; and Use the adopted preferred speech consumption characteristics to synthesize an audible output of the specific word.
222. The non-transitory computer-readable medium according to claim 221, the operation further comprising presenting a user interface for changing the preferred speech consumption characteristics to at least one of the first individual and the second individual.
223. The non-transitory computer-readable medium according to claim 221, wherein, Obtaining an input associated with the preferred speech consumption characteristics of the second individual includes receiving the input from the first individual.
224. The non-transitory computer-readable medium according to claim 221, wherein, Obtaining an input associated with the preferred speech consumption characteristics of the second individual includes receiving the input from the second individual. The non-transitory computer-readable medium according to claim 221, wherein Obtaining an input associated with the preferred speech consumption characteristics of the second individual includes retrieving information about the second individual.
226. The non-transitory computer-readable medium according to claim 225, wherein, Obtaining an input associated with the preferred speech consumption characteristics of the second individual includes determining the information based on image data captured by an image sensor worn by the first individual. The non-transitory computer-readable medium according to claim 221, wherein, The input associated with the preferred speech consumption characteristics of the second individual indicates the age of the second individual. The non-transitory computer-readable medium according to claim 221, wherein, The input associated with the preferred speech consumption characteristics of the second individual indicates environmental conditions associated with the second individual.
229. The non-transitory computer-readable medium according to claim 221, wherein, An input associated with the preferred speech consumption characteristic of the second individual indicates a hearing impairment of the second individual. The non-transitory computer-readable medium according to claim 221, wherein, The second individual is one of a plurality of individuals, and wherein the operation further includes obtaining additional input from the plurality of individuals and classifying the plurality of individuals based on the additional input.
231. The non-transitory computer-readable medium according to claim 221, wherein, Employing the preferred speech consumption characteristic includes preset speech synthesis control for expected facial micromovements.
232. The non-transitory computer-readable medium according to claim 221, wherein, The input associated with the preferred speech consumption characteristic includes a preferred speech rate, and wherein the synthesized audible output of the specific word occurs at the preferred speech rate.
233. The non-transitory computer-readable medium according to claim 221, wherein, The input associated with the preferred speech consumption characteristic includes a speech volume, and wherein the synthesized audible output of the specific word occurs at the preferred speech volume.
234. The non-transitory computer-readable medium according to claim 221, wherein, The input associated with the preferred speech consumption characteristic includes a target language of speech other than language associated with the specific facial skin micromovement, and wherein the synthesized audible output of the specific word occurs in the target language of the speech. The non-transitory computer-readable medium according to claim 221, wherein, The input associated with the preferred speech consumption characteristic includes a preferred voice, and wherein the synthesized audible output of the specific word occurs in the preferred voice.
236. The non-transitory computer-readable medium according to claim 235, wherein, The preferred voice is at least one of a celebrity voice, an accented voice, or a gender-based voice. The non-transitory computer-readable medium according to claim 221, wherein The operation further includes presenting a first synthesized version of the expected speech based on the facial micromovement, and presenting a second synthesized version of the speech based on the facial micromovement in combination with the preferred speech consumption characteristic.
238. The non-transitory computer-readable medium according to claim 237, wherein, Presenting the first synthesized version and the second synthesized version to the first individual occurs sequentially.
239. A method for performing speech synthesis according to detected facial micromovements, the method comprising: Determining specific facial skin micromovements of a first individual conversing with a second individual based on light reflection from a facial region of the first individual; Accessing a data structure associating facial micromovements with words; Performing a lookup of a specific word associated with the specific facial skin micromovement in the data structure; Obtaining an input associated with the preferred speech consumption characteristic of the second individual; Employing the preferred speech consumption characteristic; and Using the employed preferred speech consumption characteristic to synthesize an audible output of the specific word.
240. A system for performing speech synthesis according to detected facial micromovements, the system comprising: At least one processor configured to: Determine specific facial skin micromovements of a first individual conversing with a second individual based on light reflection from a facial region of the first individual; Access a data structure associating facial micromovements with words; Perform a lookup of a specific word associated with the specific facial skin micromovement in the data structure; Obtain an input associated with the preferred speech consumption characteristic of the second individual; Employ the preferred speech consumption characteristic; and Use the employed preferred speech consumption characteristic to synthesize an audible output of the specific word.
241. A non - transitory computer - readable medium containing instructions that, when executed by at least one processor, cause the at least one processor to perform operations for pre - vocal personal presentation, the operations including: Receiving a reflected signal corresponding to light reflected from an individual's facial region; Using the received reflected signal to determine the individual's specific facial skin micro - movements in the absence of a perceivable vocalization associated with a specific facial skin micro - movement; Accessing a data structure that associates facial skin micro - movements with words; Performing a lookup of a specific non - vocal word associated with the specific facial skin micro - movement in the data structure; And Causing the specific non - vocal word to be audibly presented to the individual before the individual vocalizes the specific word.
242. The non-transitory computer-readable medium according to claim 241, wherein, The operations further include recording data associated with the specific non - vocal word for future use.
243. The non-transitory computer-readable medium according to claim 242, wherein, The data includes at least one of the audible presentation of the specific non - vocal word or the text presentation of the specific non - vocal word.
244. The non-transitory computer-readable medium according to claim 241, wherein, The light reflected from the individual's facial region includes coherent light reflection. The non-transitory computer-readable medium according to claim 243, wherein, The operations further include adding punctuation to the text presentation.
246. The non-transitory computer-readable medium according to claim 241, wherein, The operations further include adjusting the speed of the audible presentation of the specific non - vocal word based on input from the individual.
247. The non-transitory computer-readable medium according to claim 241, wherein, The operations further include adjusting the volume of the audible presentation of the specific non - vocal word based on input from the individual.
248. The non-transitory computer-readable medium according to claim 241, wherein, Causing the audible presentation includes outputting an audio signal to a personal hearing device configured to be worn by the individual.
249. The non-transitory computer-readable medium according to claim 248, wherein, The operations further include operating at least one coherent light source in a manner that can illuminate the individual's facial region, wherein the at least one coherent light source is integrated with the personal hearing device. The non-transitory computer-readable medium according to claim 241, wherein The audible presentation of the specific non - vocal word is a synthesis of a selected voice.
251. The non-transitory computer-readable medium according to claim 250, wherein, The selected voice is a synthesis of the individual's voice.
252. The non-transitory computer-readable medium according to claim 250, wherein, The selected voice is a synthesis of the voice of another individual other than the individual associated with the facial skin micro - movement. The non-transitory computer-readable medium according to claim 241, wherein, The specific non - vocal word corresponds to a vocalizable word in a first language, and the audible presentation includes a synthesis of a vocalizable word in a second language different from the first language.
254. The non-transitory computer-readable medium according to claim 253, wherein, The operations further include associating the specific facial skin micro - movement with multiple vocalizable words in the second language and selecting the most appropriate vocalizable word from the multiple vocalizable words, wherein the audible presentation includes the most appropriate vocalizable word in the second language. The non-transitory computer-readable medium according to claim 241, wherein, The operations further include determining that the intensity of a part of the specific facial skin micro - movement is below a threshold and providing associated feedback to the individual. The non-transitory computer-readable medium according to claim 241, wherein, The audible presentation of the specific non - vocal word is provided to the individual at least 20 milliseconds before the individual vocalizes the specific word.
257. The non-transitory computer-readable medium according to claim 241, wherein, The operations further include stopping the audible presentation of the specific non - vocal word in response to a detected trigger.
258. The non-transitory computer-readable medium according to claim 257, wherein The operations further include detecting the trigger based on the determined facial skin micro - movement of the individual.
259. A method for personal pre - vocal presentation, the method including: Receive a reflected signal corresponding to light reflected from an individual's facial region; In the absence of a perceivable vocalization associated with a particular facial skin micromovement, use the received reflected signal to determine the particular facial skin micromovement of the individual; Access a data structure that associates facial skin micromovements with words; Perform a lookup of a particular non-vocal word associated with the particular facial skin micromovement in the data structure; And Cause the particular non-vocal word to be audibly presented to the individual before the individual vocalizes the particular word.
260. A system for pre-vocal personal presentation, the system comprising: At least one processor configured to: At least one processor configured to: Receive a reflected signal corresponding to light reflected from an individual's facial region; In the absence of a perceivable vocalization associated with a particular facial skin micromovement, use the received reflected signal to determine the particular facial skin micromovement of the individual; Access a data structure that associates facial skin micromovements with words; Perform a lookup of a particular non-vocal word associated with the particular facial skin micromovement in the data structure; And Cause an audible presentation of the particular non-vocal word to the individual before the individual vocalizes the particular word.
261. A non-transitory computer-readable medium containing instructions that, when executed by at least one processor, cause the at least one processor to perform operations for determining facial skin micromovements, the operations including: Control at least one coherent light source for projecting a plurality of light spots on an individual's facial region, wherein the plurality of light spots includes at least a first light spot and a second light spot spaced apart from the first light spot; Analyze the reflected light from the first light spot to determine a change in the reflection of the first light spot; Analyze the reflected light from the second light spot to determine a change in the reflection of the second light spot; Based on the determined changes in the first light spot reflection and the second light spot reflection, determine the facial skin micromovement; Interpret the facial skin micromovement obtained from analyzing the first light spot reflection and analyzing the second light spot reflection; and Generate an output of the interpretation.
262. The non-transitory computer-readable medium according to claim 261, wherein, The plurality of light spots further includes a third light spot and a fourth light spot, wherein each of the third light spot and the fourth light spot is spaced apart from each other and from the first light spot and the second light spot.
263. The non-transitory computer-readable medium according to claim 262, wherein, Based on the determined changes in the first light spot reflection and the second light spot reflection and the changes in the third light spot reflection and the fourth light spot reflection, determine the facial skin micromovement.
264. The non-transitory computer-readable medium according to claim 261, wherein, The plurality of light spots includes at least 16 spaced-apart light spots. The non-transitory computer-readable medium according to claim 261, wherein, The plurality of light spots are projected on a non-lip region of the individual.
266. The non-transitory computer-readable medium according to claim 261, wherein, The change in the first light spot reflection and the change in the second light spot reflection correspond to concurrent muscle recruitment.
267. The non-transitory computer-readable medium according to claim 266, wherein, Both the first light spot reflection and the second light spot reflection correspond to the recruitment of a single muscle selected from the following muscles: zygomaticus major, orbicularis oris, genioglossus, risorius, or levator labii superioris alaeque nasi.
268. The non-transitory computer-readable medium according to claim 266, wherein, The first spot reflection corresponds to the recruitment of a muscle selected from the following muscles: zygomaticus major, orbicularis oris, risorius, genioglossus, or levator labii superioris alaeque nasi muscle; and the second spot reflection corresponds to the recruitment of another muscle selected from the following muscles: zygomaticus major, orbicularis oris, risorius, genioglossus, or levator labii superioris alaeque nasi muscle.
269. The non-transitory computer-readable medium according to claim 261, wherein, The at least one coherent light source is associated with a detector, and wherein the at least one coherent light source and the detector are integrated within a wearable housing. The non-transitory computer-readable medium according to claim 261, wherein, Determining the facial skin micromovement includes analyzing the change in the first spot reflection relative to the change in the second spot reflection. The non-transitory computer-readable medium according to claim 261, wherein, The determined facial skin micromovement in the facial region includes micromovements less than 100 micrometers. The non-transitory computer-readable medium according to claim 261, wherein, The interpretation includes the emotional state of the individual. The non-transitory computer-readable medium according to claim 261, wherein, The interpretation includes at least one of the individual's heart rate or respiratory rate.
274. The non-transitory computer-readable medium according to claim 261, wherein, The interpretation includes the identity of the individual. The non-transitory computer-readable medium according to claim 261, wherein, The interpretation includes words.
276. The non-transitory computer-readable medium according to claim 275, wherein, The output includes a textual presentation of the words. The non-transitory computer-readable medium according to claim 275, wherein, The output includes an audible presentation of the words. The non-transitory computer-readable medium according to claim 275, wherein, The output includes metadata indicating a facial expression or prosody associated with the words.
279. A method for determining facial skin micromovement, the method comprising: Controlling at least one coherent light source for projecting a plurality of spots onto a facial region of an individual, wherein the plurality of spots includes at least a first spot and a second spot spaced apart from the first spot; Analyzing the reflected light from the first spot to determine a change in the first spot reflection; Analyzing the reflected light from the second spot to determine a change in the second spot reflection; Determining the facial skin micromovement based on the determined changes in the first spot reflection and the second spot reflection; Interpreting the facial skin micromovement resulting from analyzing the first spot reflection and analyzing the second spot reflection; and Generating an output of the interpretation.
280. A system for determining facial skin micromovement, the system comprising: At least one processor configured to: Control at least one coherent light source for projecting a plurality of spots onto a facial region of an individual, wherein the plurality of spots includes at least a first spot and a second spot spaced apart from the first spot; Analyze the reflected light from the first spot to determine a change in the first spot reflection; Analyze the reflected light from the second spot to determine a change in the second spot reflection; Determine facial skin micromovement based on the determined changes in the first spot reflection and the second spot reflection; Interpret the facial skin micromovement resulting from analyzing the first spot reflection and analyzing the second spot reflection; and Generate an output of the interpretation.
281. A non-transitory computer-readable medium containing instructions that, when executed by at least one processor, cause the at least one processor to perform operations for interpreting speech disorders based on facial movements, the operations including: Receiving a signal associated with a specific facial skin movement of an individual with a speech disorder that affects the way the individual pronounces a plurality of words; Access a data structure that includes a correlation between the plurality of words and a plurality of facial skin movements, the plurality of facial skin movements corresponding to the way the individual pronounces the plurality of words; Based on the received signal and the correlation, identify a specific word associated with the specific facial skin movement; And Generate an output of the specific word for presentation, wherein the output is different from how the individual pronounces the specific word.
282. The non-transitory computer-readable medium according to claim 281, wherein, The facial skin movement is a facial skin micromovement.
283. The non-transitory computer-readable medium according to claim 282, wherein, The signal is received from a sensor that detects light reflection from a non-lip portion of the individual's face.
284. The non-transitory computer-readable medium according to claim 283, wherein, The facial skin micromovement corresponds to the recruitment of at least one muscle in a muscle group that includes: the zygomaticus major, the genioglossus, the orbicularis oris, the risorius, or the levator labii superioris alaeque nasi muscle. The non-transitory computer-readable medium according to claim 281, wherein, The signal is received from an image sensor configured to measure incoherent light reflection.
286. The non-transitory computer-readable medium according to claim 281, wherein, The data structure is personalized for the individual's unique facial skin movements. The non-transitory computer-readable medium according to claim 281, wherein, The operation further includes using a training model to populate the data structure.
288. The non-transitory computer-readable medium according to claim 281, wherein, The specific facial skin movement is associated with the enunciation of the specific word, and wherein the enunciation of the specific word is in a non-canonical manner. The non-transitory computer-readable medium according to claim 281, wherein, The output of the specific word is audible and is used to correct the speech disorder of the individual. The non-transitory computer-readable medium according to claim 289, wherein, The speech disorder is stuttering, and wherein the correction includes outputting the specific word spoken in a non-stuttering form.
291. The non-transitory computer-readable medium according to claim 289, wherein, The speech disorder is hoarseness, and wherein the correction includes outputting the specific word in a non-hoarse form. The non-transitory computer-readable medium according to claim 289, wherein, The speech disorder is low volume, and wherein the correction includes outputting the specific word at a volume higher than the volume at which the specific word was spoken. The non - transitory computer - readable medium according to claim 281, wherein, The output of the specific word is textual. The non-transitory computer-readable medium according to claim 293, wherein The operation further includes adding punctuation to the textual output of the specific word. The non-transitory computer-readable medium according to claim 281, wherein The data structure includes data associated with at least one recording of the individual previously pronouncing the specific word.
296. The non-transitory computer-readable medium according to claim 281, wherein, The identified specific word associated with the specific facial skin movement is unvoiced.
297. The non-transitory computer-readable medium according to claim 281, wherein, The specific facial skin movement is associated with the silent reading of the specific word, and wherein the generated output includes a private audible presentation of the silently read word to the individual.
298. The non-transitory computer-readable medium according to claim 281, wherein, The specific facial skin movement is associated with the silent reading of the specific word, and wherein the generated output includes a non-private audible presentation of the silently read word.
299. A method for interpreting a speech disorder based on facial movement, the method comprising: Receiving a signal associated with a specific facial skin movement of an individual having a speech disorder that affects the way the individual pronounces a plurality of words; Accessing a data structure that includes a correlation between the plurality of words and a plurality of facial skin movements, the plurality of facial skin movements corresponding to the way the individual pronounces the plurality of words; Based on the received signal and the correlation, identifying a specific word associated with the specific facial skin movement; And Generating an output of the specific word for presentation, wherein the output is different from how the individual pronounces the specific word.
300. A system for interpreting speech disorders based on facial movements, the system comprising: at least one processor configured to: receive a signal associated with specific facial skin movements of an individual with a speech disorder, the speech disorder affecting the way the individual pronounces a plurality of words; access a data structure containing correlations between the plurality of words and a plurality of facial skin movements, the plurality of facial skin movements corresponding to the way the individual pronounces the plurality of words; identify, based on the received signal and the correlations, a plurality of specific words associated with the plurality of specific facial skin movements; and generate an output of the specific words for presentation, wherein the output is different from how the individual pronounces the specific words.
301. A non-transitory computer-readable medium containing instructions that, when executed by at least one processor, cause the at least one processor to perform operations for continuous verification of communication authenticity based on light reflections from facial skin, the operations comprising: generating a first data stream representing a communication of an object, the communication having a duration; generating, based on facial skin light reflections captured during the duration of the communication, a second data stream for authenticating the identity of the object; transmitting the first data stream to a destination; transmitting the second data stream to the destination; and wherein the second data stream is related to the first data stream in such a way that, once the second data stream is received at the destination, the second data stream can be used to repeatedly check that the communication originated from the object during the duration of the communication.
302. The non-transitory computer-readable medium according to claim 301, wherein, Checking that the communication originated from the object includes verifying that all words in the communication originated from the object. The non-transitory computer-readable medium according to claim 301, wherein, Checking that the communication originated from the object includes verifying, at regular time intervals during the duration of the conversation, that the speech captured at the regular time intervals originated from the object.
304. The non-transitory computer-readable medium according to claim 301, wherein, The first data stream and the second data stream are mixed in a common integrated data stream. The non-transitory computer-readable medium according to claim 301, wherein, The destination is a social networking service, and the second data stream enables the social networking service to post the communication with an authenticity indicator. The non-transitory computer-readable medium according to claim 301, wherein, The destination is an entity conducting a real-time transaction with the object, and the second data stream enables the entity to verify the identity of the object in real time during the duration of the communication.
307. The non-transitory computer-readable medium according to claim 306, wherein, Verifying the identity includes verifying the name of the object. The non-transitory computer-readable medium according to claim 306, wherein Verifying the identity includes verifying, at least at periodic intervals throughout the communication, that the object uttered the words presented in the communication. The non-transitory computer-readable medium according to claim 301, wherein The operations further include determining a biometric signature of the object based on light reflections associated with facial skin captured prior to the communication, and wherein the identity of the object is determined using the authenticated facial skin light reflections and the biometric signature. The non-transitory computer-readable medium according to claim 309, wherein, The biometric signature is determined based on the microvein pattern in the facial skin.
311. The non-transitory computer-readable medium according to claim 309, wherein, The biometric signature is determined based on a sequence of facial skin micromovements associated with the phonemes uttered by the object. The non-transitory computer-readable medium according to claim 301, wherein, The second data stream indicates a vivid state of the object, and transmitting the second data stream enables verification of the communication authenticity based on the vivid state of the object. The non-transitory computer-readable medium according to claim 301, wherein, The first data stream indicates an expression of the object, and the second data stream enables confirmation of the expression. The non-transitory computer-readable medium according to claim 301, wherein, The operation further includes storing, in a data structure, facial skin micromovements of the object that identify a spoken or pre-spoken passphrase, and identifying the object based on the spoken or pre-spoken passphrase. The non-transitory computer-readable medium according to claim 301, wherein, The operation further includes storing, in a data structure, a profile of the object based on a pattern of facial skin micromovements, and identifying the object based on the pattern. The non-transitory computer-readable medium according to claim 301, wherein, The first data stream is based on a signal associated with sound captured by a microphone during a duration of the communication. The non-transitory computer-readable medium according to claim 301, wherein, The first data stream and the second data stream are determined based on signals from the same light detector.
318. The non-transitory computer-readable medium according to claim 317, wherein, Generating the first data stream representing the communication of the object includes: reproducing speech based on the confirmed facial skin light reflection.
319. A method for continuous verification of communication authenticity based on light reflection from facial skin, the method comprising: generating a first data stream representing a communication of an object, the communication having a duration; generating, based on facial skin light reflection captured during the duration of the communication, a second data stream for authenticating the identity of the object; transmitting the first data stream to a destination; transmitting the second data stream to the destination; and wherein the second data stream is related to the first data stream in such a way that once the second data stream is received at the destination, the second data stream can be used to repeatedly check during the duration of the communication that the communication originated from the object.
320. A system for determining facial skin micromovements, the system comprising: at least one processor configured to: generate a first data stream representing a communication of an object, the communication having a duration; generate, based on facial skin light reflection captured during the duration of the communication, a second data stream for authenticating the identity of the object; transmit the first data stream to a destination; transmit the second data stream to the destination; and wherein the second data stream is related to the first data stream in such a way that when the second data stream is received at the destination, the second data stream can be used to repeatedly check during the communication that the communication originated from the object.
321. A head-mounted system for noise suppression, the head-mounted system comprising: a wearable housing configured to be worn on a head of a wearer; at least one coherent light source associated with the wearable housing and configured to project light toward a facial region of the head; at least one detector associated with the wearable housing and configured to receive coherent light reflections from the facial region associated with facial skin micromovements and output an associated reflected signal; at least one processor configured to: Analyze the reflected signal to determine speech timing based on the facial skin micromovements in the facial region; Receive an audio signal from at least one microphone, the audio signal comprising ambient sounds and sounds of words spoken by the wearer; Based on the speech timing, correlate the reflected signal with the received audio signal to determine the portion of the audio signal associated with the words spoken by the wearer; And Output the determined portion of the audio signal associated with the words spoken by the wearer while omitting the output of other portions of the audio signal that do not contain words spoken by the wearer.
322. The head - mountable system according to claim 321, wherein, The at least one processor is further configured to record the determined portion of the audio signal. The head-wearable system according to claim 321, wherein, The at least one processor is further configured to determine the other portions of the audio signal that are not associated with the words spoken by the wearer.
324. The head - mountable system according to claim 321, wherein, The other portions of the audio signal include ambient noise.
325. The head - mountable system according to claim 321, wherein, The at least one processor is further configured to determine that the other portions of the audio signal include speech of at least one person other than the wearer.
326. The head - mountable system according to claim 325, wherein, The at least one processor is further configured to record the speech of the at least one person. The head - mountable system according to claim 325, wherein, The at least one processor is further configured to receive an input indicating the wearer's desire to output the speech of the at least one person, and output the portion of the audio signal associated with the speech of the at least one person.
328. The head-wearable system according to claim 325, wherein, The at least one processor is further configured to identify the at least one person, determine the relationship of the at least one person to the wearer, and automatically output the portion of the audio signal associated with the speech of the at least one person based on the determined relationship. The head-wearable system according to claim 321, wherein, The at least one processor is further configured to analyze the audio signal and the reflected signal to identify non-verbal interjections of the wearer, and omit the non-verbal interjections from the output.
330. The head - mountable system according to claim 321, wherein, Outputting the determined portion of the audio signal includes synthesizing the vocalizations of the words spoken by the wearer.
331. The head - mountable system according to claim 330, wherein, The synthesized vocalizations mimic the wearer's voice.
332. The head - mountable system according to claim 330, wherein, The synthesized vocalizations mimic the voice of a specific individual other than the wearer. The head - mountable system according to claim 330, wherein, The synthesized vocalizations include a translated version of the words spoken by the wearer.
334. The head-wearable system according to claim 321, wherein, The at least one processor is further configured to analyze the reflected signal to identify the intention to speak, and activate at least one microphone in response to the identified intention.
335. The head - mountable system according to claim 321, wherein, The at least one processor is further configured to analyze the reflected signal to identify pauses in the words spoken by the wearer, and deactivate at least one microphone during the identified pauses.
336. The head-mounted system according to claim 321, wherein, At least one microphone is part of a communication device configured to be wirelessly paired with the head-mounted system.
337. The head-wearable system according to claim 321, wherein, At least one microphone is integrated with the wearable housing, and the wearable housing is configured such that when worn, the at least one coherent light source presents an aiming direction for illuminating at least a portion of the wearer's cheek.
338. The head - mountable system according to claim 337, wherein, A first portion of the wearable housing is configured to be placed in the wearer's ear canal, and a second portion is configured to be placed outside the ear canal, and at least one microphone is included in the second portion.
339. A method for noise suppression using facial skin micro - movements, the method comprising: Operating a wearable coherent light source configured to project light towards a facial area of a wearer's head; Operating at least one detector configured to receive coherent light reflections of the facial area associated with facial skin micro - movements and output associated reflection signals; Analyzing the reflection signals to determine speech timing based on the facial skin micro - movements in the facial area; Receiving an audio signal from at least one microphone, the audio signal comprising ambient sounds and sounds of words spoken by the wearer; Based on the speech timing, correlating the reflection signals with the received audio signal to determine the portion of the audio signal associated with the words spoken by the wearer; And Outputting the determined portion of the audio signal associated with the words spoken by the wearer while omitting the output of other portions of the audio signal that do not contain words spoken by the wearer.
340. A non - transitory computer - readable medium comprising instructions that, when executed by at least one processor, cause the at least one processor to perform operations for noise suppression using facial skin micro - movements, the operations comprising: Operating a wearable coherent light source configured to project light towards a facial area of a wearer's head; Operating at least one detector configured to receive coherent light reflections of the facial area associated with facial skin micro - movements and output associated reflection signals; Analyzing the reflection signals to determine speech timing based on the facial skin micro - movements in the facial area; Receiving an audio signal from at least one microphone, the audio signal comprising ambient sounds and sounds of words spoken by the wearer; Based on the speech timing, correlating the reflection signals with the received audio signal to determine the portion of the audio signal associated with the words spoken by the wearer; And Outputting the determined portion of the audio signal associated with the words spoken by the wearer while omitting the output of other portions of the audio signal that do not contain words spoken by the wearer.
341. A non - transitory computer - readable medium comprising instructions that, when executed by at least one processor, cause the at least one processor to perform operations for providing a private response to a silent question, the operations comprising: Receiving a signal indicating a specific facial micro - movement in the absence of a perceivable vocalization; Accessing a data structure that associates facial micro - movements with words; In the data structure, using the received signal to perform a lookup of a specific word associated with the specific facial micro - movement; Determining a query based on the specific word; Accessing at least one data structure to perform a lookup of a response to the query; And Generating a covert output comprising the response to the query.
342. The non-transitory computer-readable medium according to claim 341, wherein, The received signal is obtained via a head - mounted light detector and derived from skin micro - movements of facial portions outside the mouth. The non-transitory computer-readable medium according to claim 342, wherein, The head-mounted light detector is configured to detect incoherent light reflections from the facial portion.
344. The non-transitory computer-readable medium according to claim 342, wherein, The operation further includes controlling at least one coherent light source in a manner that enables illumination of the facial portion, and wherein the head-mounted light detector is configured to detect coherent light reflections from the facial portion. The non-transitory computer-readable medium according to claim 342, wherein, The covert output includes an audible output to the wearer of the head-mounted light detector via at least one earpiece. The non-transitory computer-readable medium according to claim 342, wherein, The covert output includes a text output to the wearer of the head-mounted light detector. The non-transitory computer-readable medium according to claim 342, wherein, The covert output includes a tactile output to the wearer of the head-mounted light detector.
348. The non-transitory computer-readable medium according to claim 341, wherein, The facial micromovements correspond to muscle activations of at least one of the following muscles: zygomaticus major, orbicularis oris, risorius, genioglossus, or levator labii superioris alaeque nasi.
349. The non-transitory computer-readable medium according to claim 341, wherein, The operation further includes receiving image data, and wherein the query is determined based on the inaudible expression of the specific word and the image data. The non-transitory computer-readable medium according to claim 349, wherein, The image data is obtained from a wearable image sensor. The non-transitory computer-readable medium according to claim 349, wherein, The image data reflects the identity of a person, the query is directed to the name of the person, and the covert output includes the name of the person.
352. The non-transitory computer-readable medium according to claim 349, wherein, The image data reflects the characteristics of an edible product, the query is directed to a list of allergens contained in the edible product, and the covert output includes the list of allergens.
353. The non-transitory computer-readable medium according to claim 349, wherein, The image data reflects the characteristics of an inanimate object, the query is directed to details about the inanimate object, and the covert output includes the requested details about the inanimate object.
354. The non-transitory computer-readable medium according to claim 341, wherein, The operation further includes using the specific facial micromovement to attempt to authenticate an individual associated with the specific facial micromovement.
355. The non-transitory computer-readable medium according to claim 354, wherein: When the individual is authenticated, the operation further includes providing a first response to the query, the first response including private information; And When the individual is not authenticated, the operation further includes providing a second response to the query, the second response omitting the private information.
356. The non-transitory computer-readable medium according to claim 354, wherein, The operation further includes accessing personal data associated with the individual and using the personal data to generate the covert output including the response to the query.
357. The non-transitory computer-readable medium according to claim 356, wherein, The personal data includes at least one of the following: the age of the individual, the gender of the individual, the current location of the individual, the occupation of the individual, the home address of the individual, the educational level of the individual, or the health status of the individual. The non-transitory computer-readable medium according to claim 341, wherein, The operation further includes using the facial micromovement to determine the emotional state of the individual associated with the facial micromovement, and wherein the response to the query is determined at least in part based on the determined emotional state.
359. A method for providing a private response to a silent question, the method comprising: Receiving, in the absence of a perceivable vocalization, a signal indicating a specific facial micromovement; Accessing a data structure associating facial micromovements with words; In the data structure, performing a lookup of a specific word associated with the specific facial micromovement using the received signal; Determining a query based on the specific word; Access at least one data structure to perform a lookup for a response to the query; And Generate a covert output including the response to the query.
360. A system for providing a private response to a silent question, the system comprising: At least one processor configured to: Receive a signal indicating a specific facial micro - movement without perceivable vocalization; Access a data structure associating facial micro - movements with words; In the data structure, perform a lookup for a specific word associated with the specific facial micro - movement using the received signal; Determine a query based on the specific word; Access at least one data structure to perform a lookup for a response to the query; And Generate a covert output including the response to the query.
361. A non - transitory computer - readable medium containing instructions that, when executed by at least one processor, cause the at least one processor to perform control commands based on facial skin micro - movements, the operations including: Operate at least one coherent light source in a manner capable of illuminating a non - lip portion of the face; Receive a specific signal representing coherent light reflection associated with a specific non - lip facial skin micro - movement; Access a data structure associating multiple non - lip facial skin micro - movements with control commands; In the data structure, identify a specific control command associated with the specific signal, the specific signal being associated with the specific non - lip facial skin micro - movement; And Execute the specific control command.
362. The non-transitory computer-readable medium according to claim 361, wherein, The facial skin micro - movement corresponds to a non - vocal expression of at least one word associated with the specific control command.
363. The non-transitory computer-readable medium according to claim 361, wherein, The facial skin micro - movement corresponds to the recruitment of at least one specific muscle.
364. The non-transitory computer-readable medium according to claim 366, wherein, The at least one specific muscle includes: the zygomaticus major muscle, the orbicularis oris muscle, the risorius muscle, or the levator labii superioris alaeque nasi muscle. The non-transitory computer-readable medium according to claim 361, wherein, The facial skin micro - movement includes a sequence of facial skin micro - movements from which the specific control command is derived.
366. The non-transitory computer-readable medium according to claim 361, wherein, The facial skin micro - movement includes involuntary micro - movements.
367. The non-transitory computer-readable medium according to claim 366, wherein, The involuntary micro - movements are triggered by an individual thinking of uttering the specific control command.
368. The non-transitory computer-readable medium according to claim 366, wherein, The involuntary micro - movements are not obvious to the human eye. The non-transitory computer-readable medium according to claim 361, wherein Operating the at least one coherent light source includes determining an intensity or a light pattern for illuminating the non - lip portion of the face. The non-transitory computer-readable medium according to claim 361, wherein, The specific signal is received at a rate between 50 Hz and 200 Hz.
371. The non-transitory computer-readable medium according to claim 361, wherein, The operation further includes analyzing the specific signal to identify temporal and intensity changes of speckles generated by light reflection from the non - lip portion of the face.
372. The non-transitory computer-readable medium according to claim 361, wherein, The operation further includes processing data from at least one sensor to determine the context of the specific non - lip facial skin micro - movement, and determining an action to initiate based on the specific control command and the determined context. The non-transitory computer-readable medium according to claim 361, wherein, The specific control command is configured to cause an audible translation of a word from a source language to at least one target language different from the source language.
374. The non-transitory computer-readable medium according to claim 361, wherein, The specific control command is configured to cause an action in a media player application. The non-transitory computer-readable medium according to claim 361, wherein, The specific control command is configured to cause an action associated with an incoming call. The non-transitory computer-readable medium according to claim 361, wherein, The specific control command is configured to cause an action associated with an ongoing call.
377. The non-transitory computer-readable medium according to claim 361, wherein, The specific control command is configured to cause an action associated with the text message.
378. The non-transitory computer-readable medium according to claim 361, wherein, The specific control command is configured to cause the activation of a virtual personal assistant.
379. A method for performing a control command based on facial skin micromovements, the method comprising: Operating at least one coherent light source in a manner capable of illuminating a non-lip portion of the face; Receiving a specific signal representing coherent light reflection associated with a specific non-lip facial skin micromovement; Accessing a data structure that associates multiple non-lip facial skin micromovements with control commands; Identifying, in the data structure, a specific control command associated with the specific signal, the specific signal being associated with the specific non-lip facial skin micromovement; And Performing the specific control command.
380. A head-mounted system for performing a control command based on facial skin micromovements, the head-mounted system comprising: A wearable housing configured to be worn on an individual's head; At least one coherent light source associated with the wearable housing and configured to illuminate a non-lip portion of the individual's face; At least one detector associated with the wearable housing and configured to receive a specific signal representing coherent light reflection associated with a specific non-lip facial skin micromovement; At least one processor configured to: Access a data structure that associates multiple non-lip facial skin micromovements with control commands; Identifying, in the data structure, a specific control command associated with the specific signal, the specific signal being associated with the specific non-lip facial skin micromovement; And Performing the specific control command.
381. A non-transitory computer-readable medium containing instructions that, when executed by at least one processor, cause the at least one processor to initiate an operation of detecting changes in neuromuscular activity over time, the operation comprising: Establishing a baseline of neuromuscular activity based on coherent light reflection associated with historical skin micromovements; Receiving a current signal representing coherent light reflection associated with an individual's current skin micromovement; Identifying a deviation of the current skin micromovement from the baseline of the neuromuscular activity; And Outputting an indicator of the deviation.
382. The non-transitory computer-readable medium according to claim 381, wherein, The operation further comprises establishing the baseline based on a historical signal representing previous coherent light reflection associated with a person other than the individual. The non-transitory computer-readable medium according to claim 381, wherein, The operation further comprises establishing the baseline based on a historical signal representing previous coherent light reflection associated with the individual. The non-transitory computer-readable medium according to claim 383, wherein, The historical signal is based on skin micromovements that occurred over a period of more than one day. The non-transitory computer-readable medium according to claim 383, wherein, The historical signal is based on skin micromovements that occurred at least one year before receiving the current signal.
386. The non-transitory computer-readable medium according to claim 381, wherein, The operation further comprises receiving the current signal from the wearable light detector when the individual is wearing the wearable light detector.
387. The non-transitory computer-readable medium according to claim 386, wherein, The operation further includes controlling at least one wearable coherent light source in a manner capable of illuminating a portion of the individual's face, and wherein the current signal is associated with coherent light reflection from the portion of the face illuminated by the at least one wearable coherent light source.
388. The non-transitory computer-readable medium according to claim 381, wherein, The current skin micromovement corresponds to the recruitment of at least one of the zygomaticus major, orbicularis oris, genioglossus, risorius, or levator labii superioris alaeque nasi muscles. The non-transitory computer-readable medium according to claim 381, wherein, The operation further includes receiving the current signal from a non-wearable light detector. The non-transitory computer-readable medium according to claim 389, wherein The coherent light reflection associated with the current skin micromovement is received from skin other than facial skin.
391. The non-transitory computer-readable medium according to claim 390, wherein, The skin other than facial skin is from the individual's neck, wrist, or chest.
392. The non-transitory computer-readable medium according to claim 381, wherein, The operation further includes receiving an additional signal associated with the individual's skin micromovement during a period of time prior to the current skin micromovement, determining a trend of change in the neuromuscular activity of the individual based on the current signal and the additional signal, and wherein the indicator indicates the trend of change.
393. The non-transitory computer-readable medium according to claim 381, wherein, The operation further includes determining a possible cause of the deviation of the current skin micromovement from a baseline of the neuromuscular activity, and wherein the indicator indicates the possible cause.
394. The non-transitory computer-readable medium according to claim 393, wherein, The operation further includes outputting an additional indicator of the possible cause of the deviation. The non-transitory computer-readable medium according to claim 393, wherein, The operation further includes receiving data indicative of at least one environmental condition, and wherein determining the possible cause of the deviation is based on the at least one environmental condition and the identified deviation.
396. The non-transitory computer-readable medium according to claim 393, wherein, The operation further includes receiving data indicative of at least one physical condition of the individual, and wherein determining the possible cause of the deviation is based on the at least one physical condition and the identified deviation.
397. The non-transitory computer-readable medium according to claim 393, wherein, The possible cause corresponds to at least one physical condition, the at least one physical condition including: affected, fatigued, or under stress.
398. The non-transitory computer-readable medium according to claim 393, wherein, The possible cause corresponds to at least one health condition including heart attack, multiple sclerosis (MS), Parkinson's disease, epilepsy, or stroke.
399. A method for detecting changes in neuromuscular activity over time, the method comprising: establishing a baseline of neuromuscular activity based on coherent light reflection associated with historical skin micromovements of an individual; receiving a signal representative of coherent light reflection associated with a current skin micromovement of the individual; identifying a deviation of the current skin micromovement from the baseline of the neuromuscular activity; and outputting an indicator of the deviation.
400. A system for detecting changes in neuromuscular activity over time, the system comprising: at least one processor configured to: establish a baseline of neuromuscular activity based on coherent light reflection associated with historical skin micromovements of an individual; receive a signal representative of coherent light reflection associated with a current skin micromovement of the individual; identifying a deviation of the current skin micromovement from the baseline of the neuromuscular activity; and outputting an indicator of the deviation.
401. A dual-purpose head-mounted system for projecting graphic content and for interpreting non-verbal speech, the head-mounted system comprising: a wearable housing configured to be worn on an individual's head; At least one light source associated with the wearable housing and configured to project light in a graphic pattern onto a facial region of the individual, wherein the graphic pattern is configured to visually convey information; A sensor for detecting a portion of the light reflected from the facial region; At least one processor configured to: Receive an output signal from the sensor; Determine facial skin micromovements associated with non-verbal expressions based on the output signal; and Process the output signal to interpret the facial skin micromovements. The head-mounted system according to claim 401, wherein The at least one processor is further configured to receive a selection of the graphic pattern and control the at least one light source to project the selected graphic pattern. The head-mounted system according to claim 401, wherein, The graphic pattern is composed of a plurality of light spots for determining the facial skin micromovements through speckle analysis.
404. The head-mounted system according to claim 401, wherein, The projected light is configured to be visible to an individual other than the individual via the human eye.
405. The head-mounted system according to claim 401, wherein, The projected light is visible via an infrared sensor. The head-mounted system according to claim 401, wherein, The projected light source includes a laser. The head-mounted system according to claim 401, wherein, The at least one processor is not configured to change the graphic pattern over time. The head-mounted system according to claim 401, wherein, The at least one processor is configured to receive position information and change the graphic pattern based on the received position information. The head-mounted system according to claim 401, wherein, The graphic pattern includes a scrolling message, and the at least one processor is configured to scroll the message. The head-mounted system according to claim 401, wherein, The at least one processor is further configured to detect a trigger and display the graphic pattern in response to the trigger. The head-mounted system according to claim 401, wherein Processing the output signal to interpret the facial skin micromovements includes determining the speech of the non-verbalized expression based on the facial skin micromovements. The head-mounted system according to claim 411, wherein The at least one processor is configured to determine the graphic pattern based on the speech of the non-verbalized expression. The head-mounted system according to claim 401, wherein, Processing the output signal to interpret the facial skin micromovements includes determining an emotional state based on the facial skin micromovements. The head-mounted system according to claim 413, wherein, The at least one processor is configured to determine the graphic pattern based on the determined emotional state. The head-mounted system according to claim 401, the head-mounted system further comprising an integrated audio output, and wherein, The at least one processor is configured to initiate an action that involves outputting audio via the audio output. The head-mounted system according to claim 401, wherein, The at least one processor is configured to identify a trigger and modify the pattern based on the trigger. The head-mounted system according to claim 416, wherein, The at least one processor is configured to analyze the facial skin micromovements to identify the trigger. The head-mounted system according to claim 416, wherein, Modifying the pattern includes stopping the projection of the graphic pattern.
419. A non-transitory computer-readable medium containing instructions that, when executed by at least one processor, cause the at least one processor to perform operations for projecting graphic content and for interpreting non-verbal speech, the operations including: Operating a wearable light source configured to project light in a graphic pattern onto a facial region of an individual, wherein the graphic pattern is configured to visually convey information; Receiving an output signal from a sensor corresponding to a portion of the light reflected from the facial region; Determining facial skin micromovements associated with non-verbal expressions based on the output signal; and Processing the output signal to interpret the facial skin micromovements.
420. A method for projecting graphic content and for interpreting non-verbal speech, the method including: Project light in a graphic pattern onto a facial region of an individual, wherein the graphic pattern is configured to visually convey information; Receive light reflected from the facial region; Determine skin micromovements associated with non-verbal expressions based on the reflected light; and Process the output signal to interpret the facial skin micromovements.
421. A head-mounted system for interpreting facial skin micromovements, the head-mounted system comprising: A housing configured to be worn on the head of a wearer; At least one detector integrated with the housing and configured to receive light reflections from a facial region of the head and output an associated reflected signal; At least one microphone associated with the housing and configured to capture sound generated by the wearer and output an associated audio signal; And At least one processor in the housing configured to use both the reflected signal and the audio signal to generate an output corresponding to a word expressed by the wearer.
422. The head-mounted system according to claim 421, the head-mounted system further comprising at least one light source integrated with the housing and configured to project coherent light toward the facial region of the head. The head-mounted system according to claim 421, wherein, The at least one processor is configured to receive the spoken form of a word and determine at least one of the words before the at least one word is spoken. The head-mounted system according to claim 421, wherein, The word expressed by the wearer includes at least one word expressed in a non-spoken manner, and the at least one processor is configured to determine the at least one word without using the audio signal. The head-mounted system according to claim 421, wherein The at least one processor is configured to use the reflected signal to identify one or more words expressed without a perceivable utterance. The head-mounted system according to claim 425, wherein, The at least one processor is configured to use the reflected signal to determine a specific facial skin micromovement and associate the specific facial skin micromovement with a reference skin micromovement corresponding to the word. The head-mounted system according to claim 426, wherein, The at least one processor is configured to use the audio signal to determine the reference skin micromovement.
428. The head-mounted system according to claim 421, the head-mounted system further comprising a speaker integrated with the housing and configured to generate an audio output. The head-mounted system according to claim 421, wherein, The output includes an audible presentation of the word expressed by the wearer. The head-mounted system according to claim 429, wherein The audible presentation includes the synthesis of speech of an individual other than the wearer. The head-mounted system according to claim 429, wherein, The audible presentation includes the synthesis of speech of the wearer. The head-mounted system according to claim 431, wherein, The word expressed by the wearer is in a first language, and the generated output includes the word spoken in a second language. The head-mounted system according to claim 431, wherein, The at least one processor is configured to use the audio signal to determine the speech of the individual for synthesizing the spoken word without a perceivable utterance. The head-mounted system according to claim 421, wherein, The output includes a text presentation of the word expressed by the wearer. The head-mounted system according to claim 434, wherein, The at least one processor is configured to transmit the text presentation of the word over a wireless communication channel to a remote computing device. The head-mounted system according to claim 421, wherein, The at least one processor is configured to cause the generated output to be transmitted to a remote computing device to execute a control command corresponding to a word expressed by the wearer. The head-mounted system according to claim 421, wherein, The at least one processor is further configured to analyze the reflected signal to determine facial skin micromovements corresponding to the recruitment of at least one specific muscle. The head-mounted system according to claim 437, wherein, The at least one specific muscle includes the zygomaticus major, orbicularis oris, risorius, or levator labii superioris alaeque nasi muscle.
439. A method for interpreting facial skin micromovements, the method comprising: Receiving coherent light reflection from a facial region associated with an individual's facial skin micromovements; Outputting a reflected signal associated with the light reflection; Capturing sound produced by the individual; Outputting an audio signal associated with the captured sound; And Using both the reflected signal and the audio signal to generate an output corresponding to a word expressed by the individual.
440. A non-transitory computer-readable medium containing instructions that, when executed by at least one processor, cause the at least one processor to perform operations for interpreting facial skin micromovements, the operations comprising: Receiving coherent light reflection from a facial region associated with an individual's facial skin micromovements; Outputting a reflected signal associated with the light reflection; Capturing sound produced by the individual; Outputting an audio signal associated with the captured sound; And Using both the reflected signal and the audio signal to generate an output corresponding to a word expressed by the individual.
441. A non-transitory computer-readable medium containing instructions that, when executed by at least one processor, cause the at least one processor to initiate a training operation to interpret facial skin micromovements, the operations comprising: Receiving a first signal representing pre-articulatory facial skin micromovements during a first time period; Receiving a second signal representing sound during a second time period after the first time period; Analyzing the sound to identify a word spoken during the second time period; Associating the word spoken during the second time period with the pre-articulatory facial skin micromovements received during the first time period; Storing the association; Receiving a third signal representing facial skin micromovements received without vocalization during a third time period; Using the stored association to identify a language associated with the third signal; and Outputting the language. The non-transitory computer-readable medium according to claim 441, wherein The operations further include identifying additional associations of additional words spoken during an additional extended time period with additional pre-articulatory facial skin micromovements detected during the additional extended time period, and using the additional associations to train a neural network. The non-transitory computer-readable medium according to claim 441, wherein The output language includes an indication of the word spoken during the second time period. The non-transitory computer-readable medium according to claim 441, wherein, The output language includes an indication of at least one word different from the word spoken during the second time period. The non-transitory computer-readable medium according to claim 444, wherein, The at least one word includes a phoneme sequence similar to the at least one word spoken during the second time period. The non-transitory computer-readable medium according to claim 441, wherein, The first signal is associated with a first individual, and the third signal is associated with a second individual. The non-transitory computer-readable medium according to claim 441, wherein, The first signal and the third signal are associated with the same individual. The non-transitory computer-readable medium according to claim 447, wherein The operation further includes continuously updating a user profile associated with the individual using the correlation. The non-transitory computer-readable medium according to claim 441, wherein The correlation is stored in a cloud-based data structure. The non-transitory computer-readable medium according to claim 441, wherein, The operation further includes accessing a voice signature of an individual associated with the facial skin micromovement, and wherein analyzing the voice to identify words spoken during the second time period is based on the voice signature.
451. The non-transitory computer-readable medium according to claim 441, wherein, The second time period begins less than 350 milliseconds after the first time period.
452. The non-transitory computer-readable medium according to claim 451, wherein, The third time period begins at least one day after the second time period.
453. The non-transitory computer-readable medium according to claim 441, wherein, The first signal is based on coherent light reflection, and wherein the operation further includes controlling at least one coherent light source to project coherent light onto a facial region of an individual and receiving the light reflection from the facial region of the individual.
454. The non-transitory computer-readable medium according to claim 453, wherein, The first signal is received from a light detector, and wherein the light detector and the coherent light source are part of a wearable component.
455. The non-transitory computer-readable medium according to claim 454, wherein, The second signal representing the voice is received from a microphone that is part of the wearable component.
456. The non-transitory computer-readable medium according to claim 441, wherein, Outputting the language includes presenting textually words associated with the third signal. The non-transitory computer-readable medium according to claim 441, wherein, The operation further includes: when a degree of certainty for identifying the language associated with the third signal is lower than a threshold, processing additional signals captured during a fourth time period after the third time period to increase the degree of certainty. The non-transitory computer-readable medium according to claim 441, wherein, The operation further includes receiving a fourth signal representing additional pre-articulatory facial skin micromovement during a fourth time period, receiving a fifth signal representing a voice during a fifth time period after the fourth time period, and using the fourth signal to identify words spoken during the fifth time period.
459. A method for interpreting facial skin micromovement, the method comprising: Receiving a first signal representing pre-articulatory facial skin micromovement during a first time period; Receiving a second signal representing a voice during a second time period after the first time period; Analyzing the voice to identify words spoken during the second time period; Associating words spoken during the second time period with the pre-articulatory facial skin micromovement received during the first time period; Storing the correlation; Receiving a third signal representing facial skin micromovement received without vocalization during a third time period; Using the stored correlation to identify the language associated with the third signal; and Outputting the language.
460. A system for interpreting facial skin micromovement, the system comprising: At least one processor configured to: Receive a first signal representing pre-articulatory facial skin micromovement during a first time period; Receive a second signal representing a voice during a second time period after the first time period; Analyze the voice to identify words spoken during the second time period; Associate words spoken during the second time period with the pre-articulatory facial skin micromovement received during the first time period; Store the correlation; Receive a third signal representing facial skin micromovement received without vocalization during a third time period; Use the stored correlations to identify the language associated with the third signal; and Output the language.
461. A multifunctional earphone, the multifunctional earphone comprising: A housing capable of being mounted on the ear; A speaker, the speaker integrated with the housing capable of being mounted on the ear for presenting sound; A light source, the light source integrated with the housing capable of being mounted on the ear to project light towards the skin of the wearer's face; A light detector, the light detector integrated with the housing capable of being mounted on the ear and configured to receive a reflection from the skin, the reflection corresponding to a facial skin micromovement indicating a pre-articulated word of the wearer; And Wherein the multifunctional earphone is configured to simultaneously present the sound through the speaker, project the light towards the skin, and detect the received reflection indicating the pre-articulated word.
462. The multifunctional earphone according to claim 461, wherein, At least a portion of the housing capable of being mounted on the ear is configured to be placed in the ear canal.
463. The multi-functional earphone according to claim 461, wherein, At least a portion of the housing capable of being mounted on the ear is configured to be placed above or behind the ear.
464. The multifunctional earphone according to claim 461, the multifunctional earphone further comprising at least one processor configured to output an audible simulation of the pre-articulated word obtained from the reflection via the speaker. The multifunctional earphone according to claim 464, wherein The audible simulation of the pre-articulated word includes the synthesis of speech of an individual other than the wearer. The multifunctional earphone according to claim 464, wherein, The audible simulation of the pre-articulated word includes synthesizing the pre-articulated word in a first language other than a second language different from the pre-articulated word.
467. The multifunctional earphone according to claim 461, the multifunctional earphone further comprising a microphone integrated with the housing capable of being mounted on the ear to receive audio indicating the speech of the wearer.
468. The multi-functional earphone according to claim 461, wherein, The light source is configured to project a pattern of coherent light towards the skin of the wearer's face, the pattern including a plurality of light spots. The multifunctional earphone according to claim 461, wherein, The light detector is configured to output an associated reflection signal indicating muscle fiber recruitment. The multifunctional earphone according to claim 469, wherein, The recruited muscle fibers include at least one of zygomatic muscle fibers, orbicularis oris muscle fibers, risorius muscle fibers, or levator labii alaeque nasi muscle fibers.
471. The multifunctional earphone according to claim 461, the multifunctional earphone further comprising at least one processor configured to analyze the light reflection to determine the facial skin micromovement.
472. The multi-functional earphone according to claim 471, wherein, The analysis includes speckle analysis.
473. The multifunctional earphone according to claim 471, further comprising a microphone integrated with the ear mountable housing to receive audio indicative of the speech of the wearer, and wherein, The at least one processor is configured to use the audio received via the microphone and the reflection received via the light detector to associate the facial skin micromovement with the spoken word, and train a neural network to determine subsequent pre-articulated words based on subsequent facial skin micromovements. The multifunctional earphone according to claim 471, wherein The at least one processor is configured to identify a trigger for activating the microphone in the determined facial skin micromovement.
475. The multi-functional earphone according to claim 471, wherein the multi-functional earphone further comprises a pairing interface for pairing with a communication device, and wherein, The at least one processor is configured to transmit the audible simulation of the pre-articulated word to the communication device.
476. The multifunctional earphone according to claim 471, further comprising a pairing interface for pairing with a communication device, and wherein, The at least one processor is configured to transmit the text presentation of the pre-articulated word to the communication device.
477. The multifunctional earphone according to claim 461, wherein, The light source is configured to project coherent light toward the skin of the wearer's face. The multifunctional earphone according to claim 461, wherein, The light source is configured to project incoherent light toward the skin of the wearer's face.
479. A method of operating a multifunctional earphone, the method comprising: operating a speaker integrated with a housing mountable on an ear to present sound, the housing mountable on the ear being associated with the multifunctional earphone; operating a light source integrated with the housing mountable on the ear to project light toward the skin of the wearer's face; and operating a light detector integrated with the housing mountable on the ear, the light detector being configured to receive a reflection from the skin, the reflection corresponding to a facial skin micromovement indicative of a pre-articulated word of the wearer; and simultaneously presenting the sound via the speaker, projecting the light toward the skin, and detecting the received reflection indicative of the pre-articulated word.
480. A non-transitory computer-readable medium comprising instructions that, when executed by at least one processor, cause the at least one processor to perform operations for operating a multifunctional earphone, the operations comprising: operating a speaker integrated with a housing mountable on an ear to present sound, the housing mountable on the ear being associated with the multifunctional earphone; operating a light source integrated with the housing mountable on the ear to project light toward the skin of the wearer's face; operating a light detector integrated with the housing mountable on the ear, the light detector being configured to receive a reflection from the skin, the reflection corresponding to a facial skin micromovement indicative of a pre-articulated word of the wearer; and simultaneously presenting the sound via the speaker, projecting the light toward the skin, and detecting the received reflection indicative of the pre-articulated word.
481. A driver for integration with a software program and for enabling a neuromuscular detection device to interface with the software program, the driver comprising: an input processing program for receiving an inaudible muscle activation signal from the neuromuscular detection device; a lookup component for mapping a specific inaudible activation signal in the inaudible activation signals to a corresponding command in the software program; a signal processing module for receiving the inaudible muscle activation signal from the input processing program, providing the specific signal in the inaudible muscle activation signal to the lookup component, and receiving an output as the corresponding command; and a communication module for passing the corresponding command to the software program, thereby enabling control within the software program based on inaudible muscle activity detected by the neuromuscular detection device.
482. The driver according to claim 481, wherein The input processing program, the lookup component, the signal processing module, and control code are embedded in the software program. The driver according to claim 481, wherein The input processing program, the lookup component, the signal processing module, and control code are embedded in the neuromuscular detection device. The driver according to claim 481, wherein The input processing program, the lookup component, the signal processing module, and the control code are embedded in an application programming interface (API). The driver according to claim 483, wherein, The neuromuscular detection device includes a light source configured to project light onto the skin, a light detector configured to sense the reflection of light from the skin, and at least one processor configured to generate the inaudible muscle activation signal based on the sensed light reflection. The driver according to claim 485, wherein, The sensed reflection of light from the skin corresponds to the micromovements of the skin. The driver according to claim 481, wherein, The lookup component is pre-populated based on training data associating the inaudible muscle activation signal with the corresponding command.
488. The driver according to claim 481, the driver including a training module for determining the correlation between the inaudible muscle activation signal and the corresponding command and for populating the lookup component. The driver according to claim 481, wherein, The lookup component includes a lookup table. The driver according to claim 481, wherein, The lookup component includes an artificial intelligence data structure. The driver according to claim 481, wherein The neuromuscular detection device includes a light source for projecting light onto the skin, a light detector configured to sense the reflection of light from the skin, and at least one processor configured to generate the inaudible muscle activation signal based on the sensed light reflection. The driver according to claim 491, wherein, The light source is configured to output coherent light.
493. The driver according to claim 492, wherein, The at least one processor is configured to generate the inaudible muscle activation signal based on speckle analysis of the reflection of the received coherent light. The driver according to claim 481, wherein, The lookup component is further configured to map some of the specific inaudible activation signals in the inaudible activation signal to text. The driver according to claim 494, wherein, The text corresponds to silent reading performance in the inaudible muscle activation signal. The driver according to claim 494, wherein, The lookup component is further configured to map some of the specific inaudible muscle activation signals in the inaudible muscle activation signal to a command for causing at least one of a visual output of the text or an audible synthesis of the text.
497. The driver according to claim 481, the driver further including a return path output for transmitting data to the neuromuscular detection device. The driver according to claim 497, wherein The data is configured to cause at least one of audio, tactile, or text output via the neuromuscular detection device.
499. The driver according to claim 481, the driver further including a detection and correction routine for detecting and correcting errors occurring during data transmission.
500. The driver according to claim 481, the driver including a configuration management routine for allowing the driver to be configured into an application outside the software program.
501. A non-transitory computer-readable medium containing instructions that, when executed by at least one processor, cause the at least one processor to perform context-driven facial micromovement operations, the operations including: Receiving, during a first time period, a first signal representing a first coherent light reflection associated with a first facial skin micromovement; Analyzing the first coherent light reflection to determine a first plurality of words associated with the first facial skin micromovement; Receive first information indicating a first context condition in which the first facial skin micro - movement occurs; Receive a second signal representing a second coherent light reflection associated with a second facial skin micro - movement during a second time period; Analyze the second coherent light reflection to determine a second plurality of words associated with the second facial skin micro - movement; Receive second information indicating a second context condition in which the second facial skin micro - movement occurs; Access a plurality of control rules associating a plurality of actions with a plurality of context conditions, wherein a first control rule prescribes a form of private presentation based on the first context condition, and a second control rule prescribes a form of non - private presentation based on the second context condition; When the first information is received, implement the first control rule to privately output the first plurality of words; and When the second information is received, implement the second control rule to non - privately output the second plurality of words.
502. The non-transitory computer-readable medium according to claim 501, wherein, The first information indicating the first context condition includes an indication that the first facial skin micro - movement is associated with a private thought.
503. The non-transitory computer-readable medium according to claim 501, wherein, The first information indicating the first context condition includes an indication that the first facial skin micro - movement is made in a private situation.
504. The non-transitory computer-readable medium according to claim 501, wherein, The first information indicating the first context condition includes an indication that the individual generating the facial micro - movement is looking down.
505. The non-transitory computer-readable medium according to claim 501, wherein, The second information indicating the second context condition includes an indication that the second facial skin micro - movement is made during a phone call. The non-transitory computer-readable medium according to claim 501, wherein, The second information indicating the second context condition includes an indication that the second facial skin micro - movement is made during a video conference. The non-transitory computer-readable medium according to claim 501, wherein, The second information indicating the second context condition includes an indication that the second facial skin micro - movement is made during a social interaction. The non-transitory computer-readable medium according to claim 501, wherein, At least one of the first information and the second information indicates the activity of the individual generating the facial micro - movement, and the operation further includes implementing the first control rule or the second control rule based on the activity. The non-transitory computer-readable medium according to claim 501, wherein, At least one of the first information and the second information indicates the location of the individual generating the facial micro - movement, and the operation further includes implementing the first control rule or the second control rule based on the location. The non-transitory computer-readable medium according to claim 501, wherein, At least one of the first information and the second information indicates the type of focus of the individual generating the facial micro - movement on the computing device, and the operation further includes implementing the first control rule or the second control rule based on the type of focus.
511. The non-transitory computer-readable medium according to claim 501, wherein, Privately outputting the first plurality of words includes generating an audio output for a personal voice - generating device. The non-transitory computer-readable medium according to claim 501, wherein, Privately outputting the first plurality of words includes generating a text output for a personal text - generating device. The non-transitory computer-readable medium according to claim 501, wherein, Non - privately outputting the second plurality of words includes transmitting an audio output to a mobile communication device.
514. The non-transitory computer-readable medium according to claim 501, wherein, Non - privately outputting the second plurality of words includes presenting a text output on a shared display.
515. The non-transitory computer-readable medium according to claim 501, wherein, The operation further includes determining a trigger for switching between a private output mode and a non - private output mode.
516. The non-transitory computer-readable medium according to claim 515, wherein, The operation further includes: receiving third information indicating a change in the context condition, and wherein the trigger is determined according to the third information.
517. The non-transitory computer-readable medium according to claim 515, wherein, The operation further includes determining the trigger based on the first plurality of words or the second plurality of words. The non-transitory computer-readable medium according to claim 515, wherein, The operation further includes receiving an output mode selection from an associated user interface and determining the trigger based on the output mode selection.
519. A method for generating context-driven facial micro-motion output, the method comprising: Receiving, during a first time period, a first signal representing a first coherent light reflection associated with a first facial skin micro-motion; Analyzing the first coherent light reflection to determine a first plurality of words associated with the first facial skin micro-motion; Receiving first information indicating a first context condition in which the first facial skin micro-motion occurs; Receiving, during a second time period, a second signal representing a second coherent light reflection associated with a second facial skin micro-motion; Analyzing the second coherent light reflection to determine a second plurality of words associated with the second facial skin micro-motion; Receiving second information indicating a second context condition in which the second facial skin micro-motion occurs; Accessing a plurality of control rules associating a plurality of actions with a plurality of context conditions, wherein a first control rule specifies a form of private presentation based on the first context condition, and a second control rule specifies a form of non-private presentation based on the second context condition; When the first information is received, implementing the first control rule to privately output the first plurality of words; and When the second information is received, implementing the second control rule to non-privately output the second plurality of words.
520. A system for generating context-driven facial micro-motion output, the system comprising: At least one processor configured to: Receiving, during a first time period, a first signal representing a first coherent light reflection associated with a first facial skin micro-motion; Analyzing the first coherent light reflection to determine a first plurality of words associated with the first facial skin micro-motion; Receiving first information indicating a first context condition in which the first facial skin micro-motion occurs; Receiving, during a second time period, a second signal representing a second coherent light reflection associated with a second facial skin micro-motion; Analyzing the second coherent light reflection to determine a second plurality of words associated with the second facial skin micro-motion; Receiving second information indicating a second context condition in which the second facial skin micro-motion occurs; Accessing a plurality of control rules associating a plurality of actions with a plurality of context conditions, wherein a first control rule specifies a form of private presentation based on the first context condition, and a second control rule specifies a form of non-private presentation based on the second context condition; When the first information is received, implementing the first control rule to privately output the first plurality of words; and When the second information is received, implementing the second control rule to non-privately output the second plurality of words.
521. A non - transitory computer - readable medium containing instructions that, when executed by at least one processor, cause the at least one processor to perform operations for extracting a reaction to content based on facial skin micromovements, the operations including: During a period when an individual consumes content, determining the facial skin micromovements of the individual based on the reflection of coherent light from the facial region of the individual; Determining at least one specific micro - expression based on the facial skin micromovements; Accessing at least one data structure containing correlations between multiple micro - expressions and multiple non - verbalized perceptions; Based on the at least one specific micro - expression and the correlations in the data structure, determining a specific non - verbalized perception of the content consumed by the individual; And Initiating an action associated with the specific non - verbalized perception.
522. The non-transitory computer-readable medium according to claim 521, wherein, The at least one specific micro - expression is not perceptible to the human eye. The non-transitory computer-readable medium according to claim 521, wherein, The facial skin micromovements for determining the at least one specific micro - expression correspond to the recruitment of at least one muscle from a muscle group including the following muscles: zygomaticus major, genioglossus, orbicularis oris, risorius, or levator labii superioris alaeque nasi muscle. The non-transitory computer-readable medium according to claim 521, wherein, The at least one specific micro - expression includes a sequence of micro - expressions associated with the specific non - verbalized perception. The non-transitory computer-readable medium according to claim 524, wherein, The operations further include determining the degree of the specific non - verbalized perception based on the sequence of micro - expressions, and determining the action to be initiated based on the degree of the specific non - verbalized perception. The non-transitory computer-readable medium according to claim 521, wherein, The at least one data structure includes past non - verbalized perceptions of previously consumed content, and wherein the operations further include determining the degree of the specific non - verbalized perception relative to the past non - verbalized perceptions, and determining the action to be initiated based on the degree of the specific non - verbalized perception.
527. The non-transitory computer-readable medium according to claim 521, wherein, The non - verbalized perception includes the emotional state of the individual. The non-transitory computer-readable medium according to claim 521, wherein, The operations further include determining the action to be initiated based on the consumed content and the specific non - verbalized perception.
529. The non-transitory computer-readable medium according to claim 521, wherein, The initiated action includes transmitting a message reflecting the correlation between the specific non - verbalized perception and the consumed content. The non-transitory computer-readable medium according to claim 521, wherein The initiated action includes storing the correlation between the specific non - verbalized perception and the consumed content in a memory. The non-transitory computer-readable medium according to claim 521, wherein, The action includes determining additional content to be presented to the individual based on the specific non - verbalized perception and the consumed content.
532. The non-transitory computer-readable medium according to claim 531, wherein, The consumed content has a first type, and the additional content has a second type different from the first type. The non-transitory computer-readable medium according to claim 521, wherein, The consumed content is part of a chat with at least one other individual, and the action includes generating a visual representation of the specific non - verbalized perception in the chat.
534. The non-transitory computer-readable medium according to claim 521, wherein, The action includes selecting an alternative way to present the consumed content. The non-transitory computer-readable medium according to claim 521, wherein, The action changes based on the type of the consumed content. The non-transitory computer-readable medium according to claim 521, wherein, The operations further include operating at least one wearable coherent light source in a non - lip - illuminating manner that enables illumination of the individual's face, and receiving a signal indicating the reflection of coherent light from the non - lip portion of the face. The non-transitory computer-readable medium according to claim 536, wherein, The facial skin micromovements are determined based on speckle analysis of the coherent light reflection. The non-transitory computer-readable medium according to claim 521, wherein, The reflection of the coherent light is received by a wearable light detector.
539. A method for extracting a reaction to content based on facial skin micromovements, the method comprising: During a period in which an individual consumes content, determining the facial skin micromovements of the individual based on the reflection of coherent light from a facial region of the individual; Determining at least one specific microexpression based on the facial skin micromovements; Accessing a data structure that includes correlations between multiple microexpressions and multiple non-verbal perceptions; Determining a specific non-verbal perception of the content consumed by the individual based on the at least one specific microexpression and the correlations in the data structure; And Initiating an action associated with the specific non-verbal perception.
540. A system for extracting a reaction to content based on facial skin micromovements, the system comprising: At least one processor configured to: During a period in which an individual consumes content, determine the facial skin micromovements of the individual based on the reflection of coherent light from a facial region of the individual; Determine at least one specific microexpression based on the facial skin micromovements; Access a data structure that includes correlations between multiple microexpressions and multiple non-verbal perceptions; Determine a specific non-verbal perception of the content consumed by the individual based on the at least one specific microexpression and the correlations in the data structure; And Initiate an action associated with the specific non-verbal perception.
541. A non-transitory computer-readable medium containing instructions that, when executed by at least one processor, cause the at least one processor to perform operations for removing noise from facial skin micromovement signals, the operations comprising: During a period in which an individual participates in at least one non-verbal related physical activity, operating a light source in a manner that enables illumination of the facial skin region of the individual; Receiving a signal representing light reflection from the facial skin region; Analyzing the received signal to identify a first reflection component indicative of pre-articulatory facial skin micromovements and a second reflection component associated with the at least one non-verbal related physical activity; And Filtering out the second reflection component to enable interpretation of words indicative of the pre-articulatory facial skin micromovements from the first reflection component. The non-transitory computer-readable medium according to claim 541, wherein, The light source is a coherent light source. The non-transitory computer-readable medium according to claim 541, wherein, The second reflection component is the result of walking. The non-transitory computer-readable medium according to claim 541, wherein, The second reflection component is the result of running. The non-transitory computer-readable medium according to claim 541, wherein, The second reflection component is the result of breathing. The non-transitory computer-readable medium according to claim 541, wherein, The second reflection component is the result of blinking and is based on neural activation of at least one orbicularis oculi muscle. The non-transitory computer-readable medium according to claim 541, wherein, When the individual simultaneously participates in a first physical activity and a second physical activity, the operations further include identifying a first portion of the second reflection component associated with the first physical activity and a second portion of the second reflection component associated with the second physical activity, and filtering out the first portion of the second component and the second portion of the second component from the first component to enable interpretation of words indicative of the pre-articulatory facial skin micromovements associated with the first component. The non-transitory computer-readable medium according to claim 541, wherein, The operations further include receiving data from a mobile communication device that indicates the at least one non-verbal related physical activity.
549. The non-transitory computer-readable medium according to claim 548, wherein, The mobile communication device does not have a light sensor for detecting the light reflection. The non-transitory computer-readable medium according to claim 548, wherein, The data received from the mobile communication device includes at least one of the following: data indicating the heart rate of the individual, data indicating the blood pressure of the individual, or data indicating the movement of the individual.
551. The non-transitory computer-readable medium according to claim 541, wherein, The operation further includes presenting the word in a synthesized voice. The non-transitory computer-readable medium according to claim 541, wherein, The signal is received from a sensor associated with the wearable housing, and wherein the instructions further include analyzing the signal to determine the at least one non-verbal related body activity.
553. The non-transitory computer-readable medium according to claim 552, wherein, The sensor is an image sensor configured to capture an image of at least one event in the environment of the individual, and wherein the at least one processor is configured to determine that the event is associated with the at least one non-verbal related body activity. The non-transitory computer-readable medium according to claim 541, wherein, The operation further includes using a neural network to identify the second reflection component associated with the at least one non-verbal related body activity.
555. The non-transitory computer-readable medium according to claim 541, wherein, The pre-articulatory facial skin micromovement corresponds to one or more involuntary muscle fiber recruitments.
556. The non-transitory computer-readable medium according to claim 555, wherein, The involuntary muscle fiber recruitment is the result of the individual thinking of saying the word.
557. The non-transitory computer-readable medium according to claim 555, wherein, The one or more muscle fiber recruitments include the recruitment of at least one of zygomaticus fibers, orbicularis oris fibers, genioglossus fibers, risorius fibers, or levator labii superioris alaeque nasi fibers. The non-transitory computer-readable medium according to claim 541, wherein, The signal is received at a rate between 50 Hz and 200 Hz.
559. A method for removing noise from a facial skin micromovement signal, the method comprising: During a period when an individual is engaged in at least one non-verbal related body activity, operating a light source in a manner that enables illumination of a facial skin area of the individual; Receiving a signal representing light reflection from the facial skin area; Analyzing the received signal to identify a first reflection component indicative of pre-articulatory facial skin micromovement and a second reflection component associated with the at least one non-verbal related body activity; And Filtering out the second reflection component so as to enable interpretation of a word from the first reflection component indicative of the pre-articulatory facial skin micromovement.
560. A system for determining facial skin micromovement, the system comprising: At least one processor, the at least one processor being configured to: During a period when an individual is engaged in at least one non-verbal related body activity, operate a light source in a manner that enables illumination of a facial skin area of the individual; Receive a signal representing light reflection from the facial skin area; Analyze the received signal to identify a first reflection component indicative of pre-articulatory facial skin micromovement and a second reflection component associated with the at least one non-verbal related body activity; And Filter out the second reflection component so as to enable interpretation of a word from the first reflection component indicative of the pre-articulatory facial skin micromovement.
Citation Information
Cited By
Noninvasive blood sugar concentration detection method based on near infrared spectrum detection
CN120983033A
Non-invasive blood glucose concentration detection method based on near-infrared spectrum detection
CN120983033B
Intelligent teaching evaluation and diagnosis system and method based on multi-mode audio and video analysis
CN120997010A
Wide-spectrum snapshot type hyperspectral fusion imaging method and system and medium
CN122415334A
A wide-band snapshot hyperspectral fusion imaging method, system and medium
CN122415334B