Systems and methods for speech detection

The system detects facial micromovements to interpret silent speech, addressing the challenge of real-time communication and translation by synthesizing audible output from subvocalized words, enhancing interaction with virtual assistants.

WO2026053177A1PCT designated stage Publication Date: 2026-03-12Q (CUE) LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-08
Publication Date
2026-03-12

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively detect and interpret silent speech, which occurs when air flow from the lungs is absent but facial muscles articulate sounds, limiting applications in real-time communication and translation.

Method used

A system and method for detecting facial skin micromovements using sensors and image processing algorithms to identify subvocalized words, synthesizing audible output, and resolving ambiguities by combining audio and non-audio signals.

Benefits of technology

Enables real-time detection and interpretation of silent speech, facilitating applications such as live subtitle generation and real-time translation without the need for vocalization, enhancing communication and interaction with virtual assistants.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000038_0001
    Figure IMGF000038_0001
  • Figure IMGF000039_0001
    Figure IMGF000039_0001
  • Figure 00000127_0000
    Figure 00000127_0000
Patent Text Reader

Abstract

Systems, methods, and computer readable medium are disclosed for resolving detected speech ambiguities. Resolving the detected speech ambiguities includes receiving audio signals representing a plurality of words vocalized by an individual; determining an ambiguity in the audio signals; during receiving of the audio signals, operating at least one sensor directed towards a non-lip region of a head of the individual; receiving, from the at least one sensor, non-audio signals indicative of neuromuscular activity associated with the non-lip region; analyzing the non-audio signals to resolve the ambiguity through an identification of at least one phoneme corresponding to the ambiguity; and generating a hybrid output of the plurality of vocalized words, wherein the hybrid output includes a first portion derived from the audio signals and a second portion derived from the non-audio signals, the second portion including a representation of the at least one phoneme.
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Docket No. 16198.0057-00304SYSTEMS AND METHODS FOR SPEECH DETECTIONCROSS REFERENCES TO RELATED APPLICATIONS

[0001] This application claims the benefit of priority of U.S. Provisional Patent Application No. 63 / 843,590, filed on July 14, 2025, and U.S. Provisional Patent Application No. 63 / 692,344, filed on September 9, 2024, all of which are incorporated herein by reference in their entirety.TECHNICAL FIELD

[0002] The present disclosure generally relates to the field of speech detection, and more particularly to ways for processing information from facial neuromuscular activity that occurs during speech to generate output.BACKGROUND

[0003] The human brain and neural activity are complex and involve many subsystems. One of those subsystems is the facial region used by humans for communication with others. From birth, humans are trained to activate craniofacial muscles to articulate sounds. Even before full language ability evolves, babies use facial expressions, including micro -expressions, to convey deeper information about themselves. After language abilities are learned, however, speech is the main technique that humans use to communicate.

[0004] The normal process of vocalized speech uses multiple groups of muscles and nerves, from the chest and abdomen, through the throat, and up through the mouth and face. To utter a given phoneme, motor neurons activate muscle groups in the face, larynx, and mouth in preparation for propulsion of air flow out of the lungs, and these muscles continue moving during speech to create words and sentences. Without this air flow, no sounds are emitted from the mouth. Silent speech occurs when the air flow from the lungs is absent, while the muscles in the face, larynx, and mouth articulate the desired sounds or move in a manner enabling interpretation.

[0005] Some of the disclosed embodiments are directed to providing a new approach for extracting meaning from neuromuscular activity, one that detects facial skin micromovements that occur during subvocalization, such as silent speech.SUMMARY

[0006] Embodiments consistent with the present disclosure provide systems, methods, and devices for detection and usage of facial movements.

[0007] Some disclosed embodiments may include systems, methods, and non -transitory computer readable media for synthesizing voice from subvocalized speech. These embodiments may involve operating at least one sensor for detecting neuromuscular activity associated with a non -lip region of aAttorney Docket No. 16198.0057-00304 head of an individual; receiving, from the at least one sensor, subvocalization signals indicative of the neuromuscular activity associated with the non-lip region, wherein the subvocalization signals are captured in an absence of perceptible vocalization; processing the received subvocalization signals to determine a plurality of subvocalized words including a first subvocalized word and a second subvocalized word; determining a first voice characteristic from a first set of the subvocalization signals associated with the first subvocalized word; determining a second voice characteristic from a second set of subvocalization signals associated with the second subvocalized word, the second voice characteristic differing from the first voice characteristic; and synthesizing audible output of the plurality of subvocalized words, wherein the first subvocalized word is audibly outputted in a manner reflecting the first voice characteristic and the second subvocalized word is audibly outputted in a manner reflecting the second voice characteristic.

[0008] Some disclosed embodiments may include systems, methods, and non -transitory computer readable media for speech detection. These embodiments may involve operating at least one sensor to detect neuromuscular activity, wherein the at least one sensor is integrated with a head -mountable housing and configured for direction towards a wearer’s head; operating a speaker for outputting audio; and limiting the speaker from generating an audio output in at least one frequency range that causes vibratory interference with operation of the at least one sensor.

[0009] Some disclosed embodiments may include systems, methods, and non -transitory computer readable media for resolving detected speech ambiguities. These embodiments may involve, receiving audio signals representing a plurality of words vocalized by an individual; determining an ambiguity in the audio signals; during receiving of the audio signals, operating at least one sensor directed towards a non-lip region of a head of the individual; receiving, from the at least one sensor, non -audio signals indicative of neuromuscular activity associated with the non-lip region; analyzing the nonaudio signals to resolve the ambiguity through an identification of at least one phoneme corresponding to the ambiguity; and generating a hybrid output of the plurality of vocalized words, wherein the hybrid output includes a first portion derived from the audio signals and a second portion derived from the non-audio signals, the second portion including a representation of the at least one phoneme.

[0010] The foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the claims.BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The accompanying drawings, which are incorporated in and constitute a part of this disclosure, illustrate various disclosed embodiments. In the drawings:

[0012] Fig. 1 is a schematic illustration of a user using a first example speech detection system, consistent with some embodiments of the present disclosure.Attorney Docket No. 16198.0057-00304

[0013] Fig. 2A is a schematic illustration of a user using a second example speech detection system, consistent with some embodiments of the present disclosure.

[0014] Fig. 2B is a perspective view of a user using a third example speech detection system, consistent with some embodiments of the present disclosure.

[0015] Fig. 3 is a schematic illustration of a user using a fourth example speech detection system, consistent with some embodiments of the present disclosure.

[0016] Fig. 4 is a block diagram illustrating some of the components of a speech detection system and a remote processing system, consistent with some embodiments of the present disclosure .

[0017] Fig. 5A and 5B are schematic illustrations of part of the speech detection system as it detects facial skin micromovements, consistent with some embodiments of the present disclosure.

[0018] Fig. 6 is a schematic illustration of a reflection image associated with light reflections received from an area of facial region associated with a single spot, consistent with some embodiments of the present disclosure.

[0019] Fig. 7 is a block diagram of a memory consistent with the disclosed embodiments.

[0020] Fig. 8 is an illustration of two example use cases for interpreting facial skin movements from light reflections, consistent with some embodiments of the present disclosure.

[0021] Fig. 9 is an illustration of another example use case for interpreting facial skin movements from light reflections, consistent with some embodiments of the present disclosure.

[0022] Figs. 10A-10C is an exemplary schematic diagram of various systems configured to detect subvocalized speech and associated voice characteristics, consistent with some disclosed embodiments.

[0023] Fig. 11 is an exemplary flowchart illustrating a method for detecting subvocalized speech and associated voice characteristics, consistent with some disclosed embodiments.

[0024] Fig. 12A is an exemplary schematic diagram of a system for detecting subvocalized speech and associated voice characteristics, consistent with some disclosed embodiments.

[0025] Fig. 12B is an exemplary schematic diagram of a speech detection system, consistent with some disclosed embodiments.

[0026] Fig. 13 is a schematic illustration of a user operating a wearable system, consistent with some embodiments of the present disclosure.

[0027] Fig. 14 is a set of two graphical representations of an audio output, consistent with some disclosed embodiments.

[0028] Fig. 15A is a graph of the effect of a filter on the audio output of Fig. 14, consistent with some disclosed embodiments.

[0029] Fig. 15B is a graph of a modified version of audio output originating from the audio output of Fig. 14, consistent with some disclosed embodiments.Attorney Docket No. 16198.0057-00304

[0030] Fig. 16 is a flowchart of a method for speech detection, consistent with some disclosed embodiments.

[0031] Fig. 17 is a flow chart of a process for resolving detected speech ambiguities, consistent with some disclosed embodiments.

[0032] Fig. 18 is word chart example of a hybrid output generated after resolving detected speech ambiguities, consistent with some disclosed embodiments.

[0033] Fig. 19 is a graphical representation of examples of possible causes of detected speech ambiguities, consistent with some disclosed embodiments.

[0034] Fig. 20 is a flowchart of a method for resolving detected speech ambiguities, consistent with some disclosed embodiments.DETAILED DESCRIPTION

[0035] The following detailed description includes references to the accompanying drawings. Wherever possible, the same reference numbers are used in the drawings and the description to refer to the same or similar parts. While several illustrative embodiments are described herein, modifications, adaptations, and other implementations are possible. For example, substitutions, additions, or modifications may be made to the components illustrated in the drawings, and the illustrative methods described herein may be modified by substituting, reordering, removing, or adding steps to the disclosed methods. Accordingly, the following detailed description is not limited to the disclosed embodiments and examples. Instead, the proper scope is defined by the appended claims.

[0036] Various terms used in the specification and claims may be defined or summarized differently when discussed in connection with differing disclosed embodiments. It is to be understood that the definitions, summaries and explanations of terminology in each instance apply to all instances, even when not repeated, unless the transitive definition, explanation, or summary would result in inoperability of an embodiment. It is also to be understood that once a term is defined herein, in the absence of an inherent inconsistency, that definition applies to all other uses of the term herein. Moreover, the exemplary embodiments of the figures and their description are not to be considered definitions of claim terms, but rather are non -limiting examples used to illustrate specific embodiments.

[0037] Throughout, this disclosure mentions “embodiments” and “disclosed embodiments,” which refer to examples of inventive ideas, concepts, and / or manifestations described herein. Many related and unrelated embodiments are described throughout this disclosure. The fact that some “disclosed embodiments” are described as exhibiting a feature or characteristic does not mean that other disclosed embodiments necessarily share that feature or characteristic.Attorney Docket No. 16198.0057-00304

[0038] This disclosure employs open-ended permissive language, indicating for example, that some embodiments “may” employ, involve, or include specific features. The use of the term “may,” and other open-ended terminology, is intended to indicate that although not every embodiment may employ the specific disclosed feature, at least one embodiment employs the specific disclosed feature.

[0039] Differing embodiments of this disclosure may involve systems, methods, and / or computer readable media containing instructions. A system refers to at least two interconnected or interrelated components or parts that work together to achieve a common objective, function, or subfimction. A method refers to at least two steps, actions, or techniques to be followed to complete a task or a subtask, to reach an objective, or to arrive at a next step. Computer-readable media containing instructions refers to any storage mechanism that contains program code instmctions, for example to be executed by a computer processor. Examples of computer-readable media are further described elsewhere in this disclosure. Instmctions may be written in any type of computer programming language, such as an interpretive language (e.g., scripting languages such as HTML and JavaScript), a procedural or functional language (e.g., C or Pascal that may be compiled for converting to executable code), an object-oriented programming language (e.g., Java or Python), a logical programming language (e.g., Prolog or Answer Set Programming), and / or any other programming language. Instmctions executed by at least one processor may include implementing one or more program code instmctions in hardware, in software (including in one or more signal processing and / or application specific integrated circuits), in firmware, or in any combination thereof, as described earlier. Causing a processor to perform operations may involve causing the processor to calculate, execute, or otherwise implement one or more arithmetic, mathematic, logic, reasoning, or inference steps.

[0040] Some disclosed embodiments may involve detecting facial skin micromovements. The term “facial skin micromovements” broadly refers to skin motions on the face that may be detectable using a sensor, but which might not be readily detectable to the naked eye. The facial skin micromovements include various types of movements, including involuntary movements caused by muscle recruitments and other types of small-scale skin deformations that fall within the range of micrometers to millimeters and fractions of a second to several seconds in duration. In some cases, the facial skin micromovements are part of a larger-scale skin movement visible to the naked eye (e.g., a smile may involve many facial skin micromovements). In other cases, the facial skin micromovements are not part of any laiger-scale skin movement visible to the naked eye. While such micromovements may occur over a multi-square millimeter facial area, they may occur in a surface area of the facial skin of less than one square centimeter, less than one square millimeter, less than 0.1 square millimeter, less than 0.01 square millimeter, or an even smaller area. In some embodiments, the facial skin micromovements correspond to one or more muscle recruitments in a facial region of a head of an individual. The facial region may include specific anatomical areas, for example: a part of the cheek above the mouth, a part of the cheek below the mouth, a part of the mid -jaw, a part of the cheek belowAttorney Docket No. 16198.0057-00304 the eye, a neck, a chin, and other areas associated with specific muscle recruitments that may cause facial skin micromovements. In some embodiments, the specific muscles may be connected to skin tissue and not to any bone. In particular, the specific muscles may be in a subcutaneous tissue associated with cranial nerve V or cranial nerve VII. As is discussed herein in greater detail, first facial skin micromovement 522A and second facial skin micromovement 522B in Fig. 5A and are non -limiting examples of facial skin micromovements, consistent with the present disclosure.

[0041] When specific muscles contract, the muscles pull on the facial skin and cause movements of the facial skin. Some of the movements that occur when the specific muscles contract may be micromovements. By way of example, the specific muscles that may cause facial skin micromovements in the context of the present disclosure may broadly be split into four groups: orbital, nasal, oral, and tongue. The orbital group of facial muscles contains two muscles associated with the eye socket. These muscles control the movements of the eyelids to protect the cornea from damage. They are both innervated by cranial nerve VII. The nasal group of facial muscles is associated with movements of the nose and the skin around it. There are three muscles in this group, and they are also all innervated by cranial nerve VII. The oral group is the most important group of the facial expressors: responsible for movements of the mouth and lips. Such movements are required in singing and whistling and add emphasis to vocal communication. The oral group of muscles consists of the orbicularis oris, buccinator, and various smaller muscles. In a specific embodiment, a disclosed system may monitor facial skin micromovements that correspond to recruitment of the buccinator muscle. The buccinator muscle is located between the mandible and maxilla relatively deep compared to other muscles of the face. The tongue group of muscles consists of four intrinsic muscles (e.g., the superior longitudinal muscle, the inferior longitudinal muscle, the vertical muscle, and the transverse muscle) used to change the shape of the tongue; and four extrinsic muscles (e.g., the genioglossus, the hyoglossus, the styloglossus, and the palatoglossus) used to change the position of the tongue. Any of the tongue muscles listed above may cause movements of the tongue that may be detected by analyzing detected facial skin micromovements. As is discussed herein in greater detail, muscle fiber 520 in Figs. 5A and 5B is a non -limiting example of a facial muscle that causes micromovements of the facial skin, consistent with the present disclosure.

[0042] Consistent with the present disclosure, facial skin micromovements may be detected during subvocalization. The term “during subvocalization” refers to any speech -related activity that takes place without utterance, before utterance, or preceding an imperceptible utterance. In one embodiment, the speech-related activity may include silent speech (i.e., when air flow from the lungs is absent but the facial muscles articulate the desired sounds). In another embodiment, the speech- related activity may include speaking soundlessly (i.e., when some air flows from the lungs, but words are articulated in a manner that is not perceptible using an audio sensor). In yet another embodiment, the speech -related activity may include prevocalization muscle recruitments (i.e., subvocalization thatAttorney Docket No. 16198.0057-00304 occurs prior to an onset of vocalization is sometimes referred to herein as prevocalization). In some cases, the prevocalization facial skin micromovements may be triggered by voluntary muscle recruitments that occur when certain craniofacial muscles start to vocalize words. In other cases, the prevocalization facial skin micromovements may be triggered by involuntary facial muscle recruitments that the individual makes when certain craniofacial muscles prepare to vocalize words. By way of example, the involuntary facial muscle recruitments may occur between 0.1 seconds to 0.5 seconds before the actual vocalization. In some cases, a suggested system may use the detected facial skin micromovement that occurs during subvocalization to identify words that are about to be vocalized. Determining words that the user intends to say before they are actually vocalized may have many benefits because the system does not have to wait for the user to vocally articulate the words to start processing the words. In one example, a disclosed system may generate subtitles for live broadcasts without delays. In another example, a disclosed system may translate what the user is saying in real-time to a different language. Additionally, because the disclosed system can detect words before they are vocalized, the actual vocalization of these words is not a requirement. Thus, facial skin micromovements that occur during subvocalization may be detected in an absence of perceptible vocalization. Movement of facial skin or muscles in an absence of vocalization but which nevertheless conveys speech-related information is referred to herein as silent speech. Detecting silent speech may have various usages, including but not limited to enabling silent communicating with other users, initiating a command, or enabling interaction with a virtual personal assistance. As is discussed herein in greater detail, subvocalization deciphering module 708 in Fig. 7 is a non -limiting example of a software module used for deciphering some subvocalization facial skin micromovements .

[0043] In some embodiments, the detection of the facial skin micromovements occurs using a speech detection system. While the shorthand “speech detection system” is employed, it is to be understood that the system may alternatively or additionally be configured to detect non -speech commands, expressions, or emotions. The system may also be used for user authentication. The speech detection system may include any device of a group of devices operatively coupled together. As used herein, the term “system” includes any device or a group of devices operatively connected together and configured to perform a function. In some embodiments, the system may include a computer (e.g., a desktop computer, a laptop computer, a server, a smart phone, a portable digital assistant (PDA), or a similar device) or plurality of computers or servers operatively connected together (e.g., using wires or wirelessly) to share information and / or data. The computer(s) may include special purpose computers (e.g., hardwired and coded to perform desired functions) or may include general purpose computers (e.g., using software to perform any desired function). In some embodiments, the system may include a cloud server. As described elsewhere in this disclosure, a cloud server may be a computer platform that provides services via a network, such as the Internet. InAttorney Docket No. 16198.0057-00304 one embodiment, the speech detection system may include a wearable housing, a coherent light source or a non -coherent light source, a light detector, and a processor. However, the specific list of components mentioned above is not intended to limit systems covered by the present disclosure. As will be appreciated by a person skilled in the art having the benefit of this disclosure, numerous variations and / or modifications may be made to the example speech detection system. For example, not all components may be essential for the detection of facial skin micromovements in all cases. Moreover, the components may be rearranged into a variety of configurations while providing the functionality of various disclosed embodiments. In some cases, a speech detection system according to some embodiments of the disclosure does not have to be wearable, but could be aimed at a skin from a location not connected to a human body. A wearable or a non -wearable system may project coherent light towards a facial region of a user, analyze reflected light, and determine facial skin micromovements. Alternatively, in other cases, a speech detection system according to some embodiments of the disclosure does not have to include a coherent light source. Specifically, the light detector may be an ultra-high resolution image sensor (e.g., more than 120 megapixel) or any other sensor capable of facial micromovement detection, and the detection of the facial skin micromovements may be accomplished using one or more image processing algorithms. As is discussed herein in greater detail, speech detection systems 100 in Figs. 1-3 are non-limiting examples of a speech detection system, consistent with the present disclosure. As illustrated in these examples, the system includes a wearable housing 110, a light source 410, a light detector 412, and a processing device 400.

[0044] Some disclosed embodiments involve a wearable housing configured to be worn on a head of an individual. The term “wearable housing” broadly includes any structure or enclosure designed for connection to a human head, such as in a manner configured to be worn by a user. Such a wearable housing may be configured to contain or support one or more electronic components or sensors. In one example, the wearable housing is configured for association with a pair of glasses. In another example, the wearable housing is associated with an earbud. The wearable housing may have a cross-section that is button-shaped, P-shaped, square, rectangular, rounded rectangular, or any other regular or irregular shape capable of being worn by a user. Such a structure may permit the wearable housing to be worn on, in, or around a body part associated with a head of the user (e.g., on the ear, in the ear, around the neck). The wearable housing may be made of plastic, metal, composite, a combination of two or more of plastic, metal and composite, or other suitable material. Consistent with disclosure embodiments, the housing may be worn on an ear. There are several ways in which the housing can be attached to the ear: In-the-ear (ITE): the housing may be inserted directly into the ear canal and held in place by the shape of the ear. Examples include earbuds and earplugs. In some cases, the housing may be custom-made to fit the specific shape of an individual's ear and seated in the ear bowl. Behind -the -ear (BTE): the housing may be seated behind the ear and with a small tubeAttorney Docket No. 16198.0057-00304 that runs to the ear canal. Examples include hearing aids and Bluetooth headsets. Over-the-ear (OTE): the housing may be seated on top of the ear and held in place by a headband or other support. Examples include structures like headphones and earmuffs. Over-the-head (OTH): the housing may be held in place by a headband that goes over the top of the head. In other embodiments, the wearable housing may be attached to a secondary device such as a glasses (sun or corrective vision glasses), a hat, a helmet, a visor, or any other type of head wearable devices. In some cases, the wearable housing may be attached to a secondary device using at least one adaptor. Specifically, the at least one adaptor may be configured to enable the individual to wear the speech detection system in two or more different ways. For example, a single adapter may enable the wearable housing to be attached to glasses and to an earbud. As is discussed herein in greater detail, wearable housings 110 in Fig. 1 and Fig. 2A are non-limiting examples of a wearable housing, consistent with the present disclosure.

[0045] Some embodiments involve a coherent light source configured to project light towards a facial region of the user. Other embodiments involve a non -coherent light source configured to project light towards a facial region of the user. As used herein, the term “light source” broadly refers to any device configured to emit light. The term “coherent light” includes light that is highly ordered and exhibits a high degree of spatial and temporal coherence. This may occur, for example, when the light waves are in phase with each other and have a uniform frequency and wavelength, resulting in a beam of light that is highly directional and has restricted outward spread out as it travels. Alternatively, coherent light may include a scenario when light waves have constant phase difference. In some examples, coherent light may be produced by a coherent light source, such as lasers and other types of light sources that have a narrow spectral range and a high degree of monochromaticity (i.e., the light consists of a single wavelength). In contrast, incoherent light may be produced by a non -coherent light source such as incandescent bulbs and natural sunlight, which have a broad spectral range and a low degree of monochromaticity.

[0046] By way of example, coherent light may include many waves of the same frequency, having different phases and amplitudes, not necessarily in the same time and locations. To control the interference, light phase information may be required to be recognized in advance. In one embodiment, the coherent light source may be a laser such as a solid-state laser, laser diode, a high- power laser, Quantum -Cascade Laser (QCLs), or an alternative light source such as a light emitting diode (LED)-based light source. In addition, the coherent light source may emit light in differing formats, such as light pulses, continuous wave (CW), quasi -CW, and so on. For example, one type of light source that may be used is a vertical -cavity surface -emitting laser (VCSEL). Another type of light source that may be used is an external cavity diode laser (ECDL). In some examples, the light source may include a laser diode configured to emit light at a wavelength between about 650 nm and 1150 nm. Alternatively, the coherent light source may include a laser diode configured to emit light at a wavelength between about 800 nm and about 1020 nm, between about 850 nm and about 950 nm, orAttorney Docket No. 16198.0057-00304 between about 1300 nm and about 1700 nm. Unless indicated otherwise, the terms “about” and “substantially the same,” with regard to a numeric value, may include a variance of up to 5% with respect to the stated value. As is discussed herein in greater detail, light source 410 in Fig. 4 and in Figs. 5 A and 5B are non 4 uniting examples of a light source, consistent with the present disclosure. In the context of this disclosure, it should be recognized that the use of a coherent light source is intended as a non-limiting example implementation in the context of speech detection systems, methods, and computer readable media. Many of the embodiments described herein may be practiced with coherent light or non -coherent light, and the reference to either by way of example, is not intended to be limiting. For example, even when not explicitly stated, the described and claimed speech detection systems, methods, and computer program products may be configured to measure non -coherent light reflections for detecting facial skin micromovements.

[0047] Some embodiments involve at least one detector configured to receive light reflections from a facial region of the user. The term “light detector,” or simply “detector,” broadly refers to any device, element, or system capable of measuring one or more properties (e.g., power, frequency, phase, pulse timing, pulse duration, or other characteristics) of electromagnetic waves and to generate an output relating to the measured property or properties. Examples of detectors consistent with this disclosure may include: a light sensitive sensor, an imaging sensor, a phase detector, a MEMS sensor, a wavemeter, a spectrometer, a spectrophotometer, a homodyne detector, or a heterodyne detector. In some embodiments, the at least one detector may be configured to detect coherent light reflections. Additionally or alternatively, the at least one detector may be configured to detect non-coherent light reflections. The at least one detector may include a plurality of detectors constructed from a plurality of detecting elements. The at least one detector may include a light detector of different types. The at least one detector may include multiple detectors of the same type which may differ in other characteristics (e.g., sensitivity, size). Combinations of several types of detectors may be used for different reasons. Consistent with some embodiments, the at least one detector may measure any form of reflection and of scattering of light, including secondary speckle patterns, different types of specular reflections, diffuse reflections, speckle interferometry, and any other form of light scattering. In some embodiments, the at least one detector is configured to output associated reflection signals from the detected coherent light reflections. In the context of this disclosure, the term “reflection signals” broadly refers to any form of data retrieved from the at least one light detector in response to the light reflections from the facial region. The reflection signals may be any electronic representation of a property determined from the light reflections, or raw measurement signals detected by the at least one light detector. As is discussed herein in greater detail, light detector 412 in Fig. 4 and in Figs. 5 A and 5B are non-limiting examples of a light detector, consistent with the present disclosure.

[0048] Some embodiments involve at least one processor configured to use the reflection signals from the detector and determine the facial skin micromovements. The term “at least one processor”Attorney Docket No. 16198.0057-00304 may involve any physical device or group of devices having electric circuitry that performs a logic operation on an input or inputs. For example, the at least one processor may include one or more integrated circuits (IC), including an application -specific integrated circuit (ASIC), microchips, microcontrollers, microprocessors, all or part of a central processing unit (CPU), graphics processing unit (GPU), digital signal processor (DSP), field -programmable gate array (FPGA), server, virtual server, or other circuits suitable for executing instructions or performing logic operations. The instructions executed by at least one processor may, for example, be pre-loaded into a memory integrated with or embedded into the controller or may be stored in a separate memory. The memory may include a Random Access Memory (RAM), a Read-Only Memory (ROM), a hard disk, an optical disk, a magnetic medium, a flash memory, other permanent, fixed, or volatile memory, or any other mechanism capable of storing instructions. In some embodiments, the at least one processor may include more than one processor. Each processor may have a similar construction, or the processors may be of differing constructions that are electrically connected or disconnected from each other. For example, the processors may be separate circuits or integrated in a single circuit. When more than one processor is used, the processors may be configured to operate independently or collaboratively and may be co-located or located remotely from each other. The processors may be coupled electrically, magnetically, optically, acoustically, mechanically, or by other means that permit them to interact. As is discussed herein in greater detail, processing unit 112 in Fig. 1 and processing device 400 in Fig. 4 are non-limiting examples of at least one processor, consistent with the present disclosure

[0049] In some embodiments, the at least one processor may determine the facial skin micromovements by applying a light reflection analysis. The term “light reflection analysis” involves the evaluation of properties of a surface by analyzing patterns of light scattered off the surface. When light strikes a surface (e.g., the facial skin), some of it is absorbed, some are transmitted, and some are reflected. The amount and type of light that is reflected depends on the properties of the surface and the angle at which the light strikes it. In one example, when a non -coherent light source is used, the light reflection analysis may include scattering analysis which involves measuring the scattering of light from the surface (e.g., the facial skin). In another example, when a coherent light source is used, the light reflection analysis may include a speckle analysis or any pattern -based analysis. By way of example, coherent light shining onto a rough, contoured, or textured surface may be reflected or scattered in many different directions, resulting in a pattern of bright and dark areas called “speckles.” Such analysis may be performed using a computer (e.g., including a processor) to identify a speckle pattern and derive information about a surface (e.g., facial skin) represented in reflection signals received from at least one light detector. A speckle pattern may occur as the result of the interference of coherent light waves added together to give a resultant wave whose intensity varies. The detected speckle pattern or any other detected pattern may then be processed to generate reflection image data. As is discussed herein in greater detail, light reflections processing module 706 depicted in Fig. 7 is aAttorney Docket No. 16198.0057-00304 non-limiting example of a software module used for determining facial skin micromovements by applying a light reflection analysis.

[0050] Consistent with the present disclosure, the reflection image data may be processed by any image processing algorithms, including classic and / or artificial neural network (ANN) based algorithms such as Convolutional Neural Network (CNN), Recurrent Neural Networks (RNN). In some examples, the reflection image data may be preprocessed by transforming the image data using a transformation function to obtain a transformed speckle image. For example, the transformed reflection image data may include one or more convolutions of the speckle image. The transformation function may include one or more image filters, such as low-pass filters, high-pass filters, band-pass filters, all-pass filters, and so forth. In some examples, the transformation function may comprise a nonlinear function. In some examples, the reflection image data may be preprocessed by smoothing at least parts of the reflection image data, for example using Gaussian convolution, using a median filter, and so forth. In some examples, the reflection image data may be preprocessed to obtain a different representation of the reflection image data. For example, reflection image data may comprise: a representation of at least part of the reflection image data in a frequency domain; a Discrete Fourier Transform of at least part of the reflection image data;a Discrete Wavelet Transform of at least part of the reflection image data;atime / frequency representation of at least part of the reflection image data;a representation of at least part of the reflection image data in a lower dimension; a lossy representation of at least part of the reflection image data;a lossless representation of at least part of the reflection image data;a time-ordered series of any of the above; any combination of the above. In some examples, the reflection image data may be preprocessed to extract edges, and the preprocessed reflection image data may comprise information based on and / or related to the extracted edges. In some examples, the reflection image data may be preprocessed to extract features from the reflection image data. Some examples of such features may comprise information related to edges, comers, blobs, ridges, Scale Invariant Feature Transform (SIFT) features, temporal features, and more.

[0051] In some embodiments, performing light reflection analysis may include evaluating the reflection image data and / or the preprocessed reflection image data using one or more mles, functions, procedures, artificial neural networks, object detection algorithms, visual event detection algorithms, action detection algorithms, motion detection algorithms, background subtraction algorithms, inference models, and so forth. Some non-limiting examples of such inference models may include: an inference model preprogrammed manually; a classification model; a regression model; a result of training algorithms, such as machine learning algorithms and / or deep learning algorithms, on training examples, where the training examples may include examples of data instances, and in some cases, a data instance may be labeled with a corresponding desired label and / or result; and so forth. In some embodiments, performing speckle analysis may comprise analyzing pixels, voxels, point cloud, range data, etc. included in the reflection image data.Attorney Docket No. 16198.0057-00304

[0052] Some embodiments may involve analyzing the reflection image data to decipher speech.The process of deciphering the speech from the reflection image data may involve identifying patterns or recognizing signatures in the reflection image data. For example, know data, patterns, or signatures may be associated with certain phenomes, combinations of phonemes, words, combinations of words, or any other speech-related component. By recognizing such information in the reflection image data, speech may be deciphered. Such recognition and / or deciphering may be aided by machine learning. For example, machine learning models or algorithms may be employed to recognize and / or understand speech or commands. Some non -limiting examples of machine learning algorithms that may be used include classification algorithms, data regressions algorithms, image segmentation algorithms, visual detection algorithms (such as object detectors, motion detectors, edge detectors, etc.), visual recognition algorithms (such as object recognition, etc.), speech recognition algorithms, mathematical embedding algorithms, natural language processing algorithms, support vector machines, random forests, nearest neighbors algorithms, deep learning algorithms, artificial neural network algorithms, convolutional neural network algorithms, recursive neural network algorithms, linear machine learning models, non-linear machine learning models, ensemble algorithms, and so forth. For example, a trained machine learning algorithm may include an inference model, such as a predictive model, a classification model, a regression model, a clustering model, a segmentation model, an artificial neural network (such as a deep neural network, a convolutional neural network, a recursive neural network, etc.), a random forest, a support vector machine, and so forth. In some examples, the training examples may include example inputs together with the desired outputs corresponding to the example inputs. Further, in some examples, training machine learning algorithms using the training examples may generate a trained machine learning algorithm, and the trained machine learning algorithm may be used to estimate outputs for inputs not included in the training examples. In some examples, engineers, scientists, processes, and machines that train machine learning algorithms may further use validation examples and / or test examples. For example, validation examples and / or test examples may include example inputs together with the desired outputs corresponding to the example inputs, a trained machine learning algorithm and / or an intermediately trained machine learning algorithm may be used to estimate outputs for the example inputs of the validation examples and / or test examples, the estimated outputs may be compared to the corresponding desired outputs, and the trained machine learning algorithm and / or the intermediately trained machine learning algorithm may be evaluated based on a result of the comparison. In some examples, a machine learning algorithm may have parameters and hyper parameters, where the hyper parameters are set manually by a person or automatically by a process external to the machine learning algorithm (such as a hyper parameter search algorithm), and the parameters of the machine learning algorithm are set by the machine learning algorithm according to the training examples. In some implementations, the hyper-parameters are set according to the training examples and theAttorney Docket No. 16198.0057-00304 validation examples, and the parameters are set according to the training examples and the selected hyper-parameters.

[0053] In some examples, deciphering the speech from the reflection image data may involve a trained machine learning algorithm that is used as an inference model that when provided with an input generates an inferred output. For example, a trained machine learning algorithm may include a classification algorithm, the input may include a sample, and the inferred output may include a classification of the sample. In another example, a trained machine learning algorithm may include a regression model, the input may include a sample, and the inferred output may include an inferred value for the sample. In yet another example, a trained machine learning algorithm may include a clustering model, the input may include a sample, and the inferred output may include an assignment of the sample to at least one cluster. In an additional example, a trained machine learning algorithm may include a classification algorithm, the input may include an image, and the inferred output may include a classification of an item depicted in the image. In yet another example, a trained machine learning algorithm may include a regression model, the input may include an image, and the inferred output may include an inferred value for an item depicted in the image (such as an estimated facial skin motion, and so forth). In an additional example, a trained machine learning algorithm may include an image segmentation model, the input may include an image, and the inferred output may include a segmentation of the image. In yet another example, a trained machine learning algorithm may include an object detector, the input may include an image, and the inferred output may include one or more detected objects in the image and / or one or more locations of objects within the image. In some examples, the trained machine learning algorithm may include one or more formulas and / or one or more functions and / or one or more rules and / or one or more procedures, the input may be used as input to the formulas and / or functions and / or mles and / or procedures, and the inferred output may be based on the outputs of the formulas and / or functions and / or mles and / or procedures (for example, selecting one of the outputs of the formulas and / or functions and / or mles and / or procedures, using a statistical measure of the outputs of the formulas and / or functions and / or rules and / or procedures, and so forth). As is discussed herein in greater detail, reflection image 600 in Fig. 6 is a non -limiting example of a visualization of reflection image data, consistent with the present disclosure.

[0054] In some embodiments, artificial neural networks may be configured to analyze inputs and generate corresponding outputs. Some non-limiting examples of such artificial neural networks may include shallow artificial neural networks, deep artificial neural networks, feedback artificial neural networks, feed-forward artificial neural networks, autoencoder artificial neural networks, probabilistic artificial neural networks, time -delay artificial neural networks, convolutional artificial neural networks, recurrent artificial neural networks, long / short term memory artificial neural networks, and so forth. In some examples, an artificial neural network may be configured manually. For example, a structure of the artificial neural network may be selected manually, a type of an artificial neuron of theAttorney Docket No. 16198.0057-00304 artificial neural network may be selected manually, a parameter of the artificial neural network (such as a parameter of an artificial neuron of the artificial neural network) may be selected manually, and so forth. In some examples, an artificial neural network may be configured using a machine learning algorithm. For example, a user may select hyper-parameters for the artificial neural network and / or the machine learning algorithm, and the machine learning algorithm may use the hyper-parameters and training examples to determine the parameters of the artificial neural network, for example using back propagation, using gradient descent, using stochastic gradient descent, using mini -batch gradient descent, and so forth. In some examples, an artificial neural network may be created from two or more other artificial neural networks by combining the two or more other artificial neural networks into a single artificial neural network.

[0055] Disclosed embodiments may include and / or access a data structure or data. A data structure consistent with the present disclosure may include any collection of data values and relationships among them. By way of example, a data structure may contain correlations of facial micromovements with words or phonemes, and the at least one processor may perform a lookup in the data structure of particular words or phenomes associated with detected facial skin micromovements. The data may be stored linearly, horizontally, hierarchically, relationally, non-relationally, uni-dimensionally, multidimensionally, operationally, in an ordered manner, in an unordered manner, in an object- oriented manner, in a centralized manner, in a decentralized manner, in a distributed manner, in a custom manner, or in any manner enabling data access. By way of non -limiting examples, data structures may include an array, an associative array, a linked list, a binary tree, a balanced tree, a heap, a stack, a queue, a set, a hash table, a record, a tagged union, ER model, and a graph. For example, a data structure may include an XML database, an RDBMS database, an SQL database, or NoSQL alternatives for data storage / search such as, for example, MongoDB, Redis, Couchbase, Datastax Enterprise Graph, Elastic Search, Splunk, Solr, Cassandra, Amazon DynamoDB, Scylla, HBase, and Neo4J. A data structure may be a component of the disclosed system or a remote computing component (e.g., a cloud-based data structure). Data in the data structure may be stored in contiguous or non-contiguous memory. Moreover, a data structure, as used herein, does not require information to be co-located. It may be distributed across multiple servers, for example, servers that may be owned or operated by the same or different entities. Thus, the term “data structure” as used herein in the singular is inclusive of plural data structures. As is discussed herein in greater detail, data structure 124 in Fig. 1 and data structures 422 and 464 in Fig. 4 are non -limiting examples of a data structure, consistent with the present disclosure.

[0056] Consistent with the present disclosure, at least one processor may generate output associated with the determined facial skin micromovements. The term “generating an output” broadly refers to emitting a command, emitting data, and / or causing any type of electronic device to initiate an action. In some embodiments, the output may be sound (e.g., delivered via a speaker configured to fitAttorney Docket No. 16198.0057-00304 in the ear of the user), and the sound may be an audible presentation of words associated with silent or prevocalized speech. In one example, the audible presentation of words may include an answer to a question that the user silently asked a virtual personal assistance. In another example, the audible presentation of words may include synthesized speech (e.g., artificial production of human speech). According to other disclosed embodiments, the output may be directed to a display (e.g., a visual display such as a computer monitor, television, mobile communications device, VR or XR glasses, or any other device that enables visual perception) and the generated output may include graphics, images, or textual presentations of words associated with prevocalized or vocalized speech (e.g., subtitles). The textual presentation of the words may be presented at the same time words are vocalized. In other embodiments, the output may be directed to a communications device associated with the user and the generated output may be any data exchanged with the communications device. The term “communications device “ is intended to include all possible types of devices capable of exchanging data using a network configured to convey data. In some examples, the communications device may include a smartphone, a tablet, a smartwatch, a personal digital assistant, a desktop computer, a laptop computer, an Internet of Things (loT) device, a dedicated terminal, a wearable communications device, and any other device that enables data communications. As is discussed herein in greater detail, output determination module 712 in Fig. 7 is a non -limiting example of a software module used for generating output associated with the determined facial skin micromovements .

[0057] Disclosed embodiments may involve exchanging data (e.g., textual data) using a network. The term “communications network,” or simply “network,” may include any type of physical or wireless computer networking arrangement used to exchange data. For example, a network may be the Internet, a private data network, a virtual private network using a public network, a Wi-Fi network, a LAN or WAN network, a combination of one or more of the foregoing, and / or other suitable connections that may enable information exchange among various components of the system. In some embodiments, a network may include one or more physical links used to exchange data, such as Ethernet, coaxial cables, twisted pair cables, fiber optics, or any other suitable physical medium for exchanging data. A network may also include a public switched telephone network (“PSTN”) and / or a wireless cellular network. A network may be a secured network or an unsecured network. In other embodiments, one or more components of the system may communicate directly through a dedicated communication network. Direct communications may use any suitable technologies, including, for example, BLUETOOTH™, BLUETOOTH LE™ (BLE), Wi-Fi, near-field communications (NFC), or other suitable communication methods that provide a medium for exchanging data and / or information between separate entities. As is discussed herein in greater detail, communications network 126 shown in Fig. 1, is a non -limiting example of a communications network, consistent with the present disclosure.Attorney Docket No. 16198.0057-00304

[0058] As used herein, a non -transitory computer-readable storage medium (or similar constructs such as a non-transitory computer-readable media) refers to any type of physical memory on which information or data readable by at least one processor can be stored. Examples include Random Access Memory (RAM), Read-Only Memory (ROM), volatile memory, nonvolatile memory, hard drives, CD ROMs, DVDs, flash drives, disks, any other optical data storage medium, any physical medium with patterns of holes, markers, or other readable elements, a PROM, an EPROM, a FLASH- EPROM or any other flash memory, NVRAM, a cache, a register, any other memory chip or cartridge, and networked versions of the same. The terms “memory” and “computer-readable storage medium” may refer to multiple structures, such as a plurality of memories or computer-readable storage mediums located within a wearable device or at a remote location. Additionally, one or more computer-readable storage mediums can be utilized in implementing a computer-implemented method. Accordingly, the term computer-readable storage medium should be understood to include tangible items and exclude carrier waves and transient signals.

[0059] Reference is now made to Fig. 1, which illustrates a user 102 using a speech detection system consistent with some embodiments of the present disclosure. Fig. 1 is a single exemplary representation, and it is to be understood that some illustrated elements might be omitted, and others may be added within the scope of this disclosure. In the illustrated example implementation, a speech detection system 100 may be mountable on a head of user 102. Specifically, speech detection system 100 (also referred to herein simply as “the system”) may have the form and appearance of an over- the-ear clip-on headset. Alternatively, the system may be head -mountable in one of many other ways within the scope of this disclosure, including an in-ear bud, integration into or connectable to a temple of glasses, a head band, or any other mechanism capable of securing the system or a portion thereof to a human head. Speech detection system 100 may be configured to direct projected light 104 (e.g., coherent light) toward respective locations on the face of user 102, thus creating an array of light spots 106 extending over a facial region 108 of the face. Facial region 108 may have an area of at least 1 cm2, at least 2 cm2, at least 4 cm2, at least 6 cm2, or at least 8 cm2. In some embodiments, the size of facial region 108 may be determined to enable sensing the motion of different parts of the facial muscles. In the depicted example, only one beam of projected light 104 is illustrated, however, it is contemplated that every spot projected towards facial region 108 may be associated with a corresponding light beam or with one or more light beams. In other embodiments, the light source may project light in a manner other than an array of spots. For example, a region of the face may be uniformly or non-uniformly illuminated.

[0060] For embodiments that are head -worn, speech detection system 100 may include a wearable housing 110 configured to be worn on a head of user 102. Wearable housing 110 may include or be associated with a processing unit 112 configured to interpret facial skin micromovements; an output unit 114 configured to fit into the user’s ear and to present audible and / or vibrational output; andAttorney Docket No. 16198.0057-00304 optical sensing unit 116 configured to project light toward a non -lip part of the face of user 102 and to detect reflections of the projected light. In the illustrated example, optical sensing unit 116 may be connected to output unit 114 by an arm 118 and thus may be held in a location in proximity to and / or facing the user’s face. According to some disclosed embodiments, optical sensing unit 116 does not contact the user’s skin at facial region 108, but rather optical sensing unit 116 may be held at a certain distance from the skin surface of facial region 108. The distance of optical sensing unit 116 from the skin surface may be at least 5 mm, at least 7.5 mm, at least 10 mm, at least 15 mm, or at least 20 mm.

[0061] Optical sensing unit 116 may be configured to receive reflections of light 104 from facial region 108 and to output associated reflection signals. Specifically, the reflection signals may be indicative of light patterns (e.g., secondary speckle patterns) that may arise due to reflection ofthe coherent light from each of spots 106 within a field of view of speech detection system 100. To cover a sufficiently laige facial region 108, the detector of speech detection system 100 may have a wide field of view, for example, the field of view may have an angular width of at least 60°, at least 70°, or at least 90°. Within this field of view, speech detection system 100 may sense and process the signals reflective of light patterns in all of spots 106 or only a certain subset of spots 106. For example, processing unit 112 may select a subset of spots 106 determined to give the largest amount of useful and reliable information with respect to the relevant movements of the skin surface of user 102 and may avoid processing data from other spots 106. Additional details ofthe structure and operation of optical sensing unit 116 are described below with reference to Fig. 5.

[0062] Consistent with the present disclosure, speech detection system 100 may be capable of detecting facial skin micromovements of user 102 and extract meaning from the detected movements, even without vocalization of speech or utterance of any other sounds by user 102. The extracted meaning may be an identification of user 102 wearing speech detection system 100, an identification of a subvocalization by a user, such as a word silently spoken by user 102, an identification of a word vocally spoken by user 102, an identification of a phoneme silently spoken by user 102, or an identification of a phoneme vocally spoken by user 102. Similarly, the extract meaning may include an identification of a heart rate of user 102, an identification of a breathing rate of user 102, and / or other characteristics associated with verbal or non-verbal communication by user 102. In one example, speech detection system 100 may generate output signals that include data associated with an identification information, a UI command, synthesized audio signal, a textual transcription, or any combination thereof. In one example, the synthesized audio signal may be played back to user 102 via a speaker in output unit 114. This playback may be useful in giving user 102 feedback with respect to the speech output.

[0063] Consistent with the present disclosure, speech detection system 100 may exchange data (e.g., output signals) with a variety of communications devices associated with users, for example, a mobile communications device 120 or a server 122. The term “communications device” is intended toAttorney Docket No. 16198.0057-00304 include all possible types of devices capable of exchanging data using a digital communications network, an analog communication network, or any other communications network configured to convey data. In some examples, the communications device may include a wearable communications device, such as a smartphone, a tablet, a smartwatch, a personal digital assistant, a laptop computer, an loT device, a dedicated terminal, industrial machinery, a vehicle, a smart house, an appliance, or any other electronic device capable of exchanging information or data with another electronic device. In other examples, the communications device may include a non -wearable communications device, such as a desktop computer, a smart home hub, a router, a server, or any other network -connected equipment. In some cases, a processing device of mobile communications device 120 or server 122 may supplement or replace some functions of processing unit 112 of speech detection system 100. In some embodiments, the output signals generated by speech detection system 100 may be transmitted via a communication link to mobile communications device 120 or to a cloud server. The term “cloud server” refers to a computer platform that provides services via a network, such as the Internet. In the example embodiment illustrated in Fig. 1, a server 122 may use one or more virtual machines that may not correspond to individual pieces of hardware. For example, computational and / or storage capabilities may be implemented by allocating appropriate portions of desirable computation / storage power from a scalable repository, such as a data center or a distributed computing environment. In one example configuration, server 122 may be a cloud server that determines neural activity of user 102 based on facial skin micromovements. In one example, server 122 may implement the methods described herein using customized hard-wired logic, one or more Application Specific Integrated Circuits (ASICs), Field Programmable Gate Arrays (FPGAs), firmware, and / or program logic which, in combination with the computer system, cause server 122 to be a special -purpose machine.

[0064] In some embodiments, server 122 may access data structure 124 to determine, for example, correlations between words and a plurality of facial movements. Data structure 124 may utilize a volatile or non-volatile, magnetic, semiconductor, tape, optical, removable, non -removable, other type of storage device or tangible or non -transitory computer-readable medium, or any medium or mechanism for storing information. Data structure 124 may be part of server 122 or separate from server 122, as shown. When data structure 124 is not part of server 122, server 122 may exchange data with data structure 124 via a communication link. Data structure 124 may include one or more memory devices that store data and instructions used to perform one or more features of the disclosed methods. In one embodiment, data structure 124 may include any of a plurality of suitable data structures, ranging from small data structures hosted on a workstation to large data structures distributed among data centers. Data structure 124 may also include any combination of one or more data structures controlled by memory controller devices (e.g., servers) or software. Consistent with the present disclosure, speech detection system 100 may communicate with mobile communications device 120 or server 122 using a communications network 126 as defined above.Attorney Docket No. 16198.0057-00304

[0065] Reference is now made to Fig. 2A, which illustrates another example implementation of speech detection system 100, in accordance with the present disclosure. In this example, wearable housing 110 may be integrated with or otherwise attached to a pair of glasses 200 having a frame 202. In this example implementation, glasses 200 may include nasal electrodes 204 and temporal electrodes 206 attached to frame 202 and contacting the user’s skin surface. Electrodes 204 and 206 may receive body surface electromyogram (sEMG) signals, which provide additional information regarding the activation of the user’s facial muscles. Speech detection system 100 may use the electrical activity sensed by electrodes 204 and 206 together with the output of optical sensing unit 116 in generating, for example, the synthesized audio signals. Additionally or alternatively, speech detection system 100 may include one or more additional optical sensing units 208, similar to optical sensing unit 116, for sensing skin movements in other areas of the user’s face, such as eye movement. These additional optical sensing units may be used together with or instead of optical sensing unit 116. In the illustrated example, optical sensing unit 116 may illuminate a first facial region 108A and optical sensing unit 208 may illuminate a second facial region 108B. First facial region 108A and second facial region 108B may be nonoverlapping.

[0066] In some disclosed embodiments, the speech detection system may be incorporated with, integrated with, or otherwise attached to an extended reality appliance. As used herein, the term “extended reality appliance” may include any type of device or system that enables a user to perceive and / or interact with an extended reality environment. The term “extended reality environment,” refers to all types of real-and-virtual combined environments and human -machine interactions at least partially generated by computer technology. One non -limiting example of an extended reality environment may be a Virtual Reality (VR) environment. A virtual reality environment may be an immersive simulated non-physical environment which provides the user with the perception of being present in the virtual environment. Another non -limiting example of an extended reality environment may be an Augmented Reality (AR) environment. An augmented reality environment may involve live direct or indirect views of a physical real -world environment enhanced with virtual computergenerated perceptual information, such as virtual objects with which the user may interact. Another non-limiting example of an extended reality environment is a Mixed Reality (MR) environment. A mixed reality environment may be a hybrid of physical real-world and virtual environments, in which physical and virtual objects may coexist and interact in real time. Examples of the extended reality appliance may include VR headsets, AR headsets, MR headsets, smart glasses, and wearable projection devices.

[0067] Reference is now made to Fig. 2B, illustrating another example implementation of speech detection system 100, in accordance with some embodiments of the present disclosure. In the depicted example, speech detection system 100 may be part of an extended reality appliance 250. Extended reality appliance 250 may include all the sensors discussed above with reference to glasses 200 andAttorney Docket No. 16198.0057-00304 more. For example, extended reality appliance 250 may include one or more of a gyroscope, an accelerometer, a magnetometer, an image sensor, a depth sensors, an infrared sensors, a proximity sensor, and / or any other sensor configured to measure one or more properties associated with the individual wearing extended reality appliance 250 and to generate an output relating to the measured property or properties. In some cases, speech detection system 100 may use the input from any one of the sensors of extended reality appliance 250 to determine the vocalized or subvocalized words that user 102 articulated. The term “determining” may refer to establishing or arriving at a conclusive outcome as a result of a reasoned, learned, calculated or logical process. For example, speech detection system 100 may use input from an image sensor of extended reality appliance 250 together with data from optical sensing unit 116 (See Fig. 1) to extract meaning of facial movements. In other cases, extended reality appliance 250 may generate output that includes a visual and / or audible presentation associated with the words detected by the speech detection system 100. For example, user 102 may interact with extended reality appliance 250 using silent commands.

[0068] Reference is now made to Fig. 3, which illustrates another example implementation of speech detection system 100, in accordance with the present disclosure. In the implementation illustrated in Fig. 3, speech detection system 100 may be integrated with mobile communications device 120. Specifically, mobile communications device 120 may include a light detector configured to detect reflections 300 of light from facial region 108. In this example, the light projected to facial region 108 originates from a non-wearable light source 302 that may be a coherent light source or non -coherent light source. In some configurations, non-wearable light source 302 may be included in mobile communications device 120. Alternatively, non-wearable light source 302 may be separated from mobile communications device 120.

[0069] Consistent with the present disclosure, and as depicted in Fig. 3, the pattern of the light projected to facial region 108 may be a single spot 106 large enough to illuminate different portions of facial region 108. For example, spot 106 may include a first portion 304A associated with a first facial muscle and a second portion 304B associated with a second facial muscle. Thereafter, a processing device of mobile communications device 120 may apply a light reflection analysis on received reflections 300 to determine facial skin micromovements. In particular, the processing device of mobile communications device 120 may determine first facial skin micromovements of first portion 304A and second facial skin micromovements of second portion 304B. The processing device may use both the first facial skin micromovements and the second facial skin micromovements to extract meaning (e.g., determine speech or a command, or to authenticate user 102) and to generate output. The example implementation of speech detection system 100 illustrated in Fig. 3 may be used when the extracted meaning includes a continuous authentication of user 102. Specifically, speech detection system 100 may provide an authentication service that uses biometrics of facial micromovements for continuous authentication during usage of mobile communications device 120.Attorney Docket No. 16198.0057-00304

[0070] Fig. 4 is a block diagram of an exemplary configuration of speech detection system 100 and an exemplary configuration of remote processing system 450. It is to be noted that Fig. 4 is a representation of just one embodiment, and it is to be understood that some illustrated elements might be omitted and others added within the scope of this disclosure. In the depicted embodiment, speech detection system 100 comprises processing unit 112 that includes a processing device 400 and a memory device 402; output unit 114 that includes a speaker 404, a light indicator 406, and a haptic feedback device 408; optical sensing unit 116 that includes at least one light source 410 and at least one light detector 412; an audio sensor 414, a power source 416, one or more additional sensors 418, network interface 420, and data structure 422. Speech detection system 100 may directly or indirectly access a bus 424 (or any other communication mechanism) that interconnects the above-mentioned subsystems and components for transferring information and commands within speech detection system 100. Some of the subsystems and components listed above are referred to herein in the singular but in alternative configurations may be plural. For example, in some configurations speech detection system 100 may include multiple light sources 410 or multiple light detectors 412.

[0071] Processing device 400, shown in Fig. 4, may constitute any physical device or group of devices having electric circuitry that performs a logic operation on an input or inputs. The instmctions executed by at least one processor may, for example, be pre-loaded into a memory integrated with or embedded into processing device 400, or may be stored in a separate memory (e.g., memory device 402 or data structure 422). As described above, the processing device may include more than one processor. Each processor may have a similar construction, or the processors may be of differing constructions that are electrically connected or disconnected from each other. For example, the processors may be separate circuits or integrated in a single circuit. When more than one processor is used, the processors may be configured to operate independently or collaboratively and may be co- located or located remotely from each other. The processors may be coupled electrically, magnetically, optically, acoustically, mechanically, or by other means that permit them to interact. Consistent with the present disclosure, at least some of the functionalities described below regarding processing device 400 may be executed by a processing device of remote processing system 450.

[0072] Memory device 402, shown in Fig. 4, may include high-speed random-access memory and / or non-volatile memory, such as one or more magnetic disk storage devices, one or more optical storage devices, and / or flash memory (e.g., NAND, NOR). Consistent with the present disclosure, the components of memory device 402 may be distributed in more than one unit of speech detection system 100 and / or in more than one memory device. In particular, memory device 402 may be used to store a software product and / or data stored on a non-transitory computer-readable medium. As described above, the terms “memory” and “computer-readable storage medium” may refer to multiple structures, such as a plurality of memories or computer-readable storage mediums located within speech detection system 100 or at a remote location (e.g., at remote processing system 450).Attorney Docket No. 16198.0057-00304Additionally, one or more computer-readable storage mediums can be utilized in implementing a computer-implemented method. Examples of software modules stored in memory device 402 are described below with reference to Fig. 7.

[0073] Output unit 114, shown in Fig. 4, may cause output from a variety of output devices, such as speaker 404, light indicator 406, and a haptic feedback device 408. Examples of speaker 404 may include or may be incorporated with a loudspeaker, earbuds, audio headphones, a hearing aid type device, a bone conduction headphone, and any other device capable of converting an electrical audio signal into a corresponding sound. In some embodiments, speaker 404 may be configured to let only user 102 to listen to the generated audio signals. Alternatively, speaker 404 may be configured to emit sound into the open air for anyone nearby to hear. Light indicator 406 may include one or more light sources, for example, a LED array associated with different colors. Light indicator 406 may be used to indicate the battery status of speech detection system 100 orto indicate its operational mode. Haptic feedback device 408 may include a vibrating motor, linear actuator, vibrational transducer, or any other force feedback device that provide tactile or haptic cues or can convert an electrical signal into corresponding vibrations or force applications.

[0074] Optical sensing unit 116, shown in Fig. 4, may include light source 410 and light detector 412. Light source 410 may project coherent light or non -coherent light to facial region 108. As discussed above, light source 410 may be a laser such as a solid-state laser, laser diode, a high -power laser, or an alternative light source such as a light emitting diode (LED) -based light source. In addition, the light source 410, may emit light in differing formats, such as light pulses, continuous wave (CW), quasi-CW, and so on. In one embodiment, light source 410 may be an infrared laser diode configured to emit an input beam of coherent radiation. Light source 410 may be associated with a beam-splitting element, such as a Dammann grating or another suitable type of diffractive optical element (DOE), for splitting an input beam into multiple output beams, which form respective spots 106 at a matrix of locations extending over facial region 108. In another embodiment (not shown in the figures) light source 410 may include multiple laser diodes or other emitters, which generate respective groups of the output beams, covering different respective sub-areas within facial region 108. In one embodiment, processing unit 112 may select and actuate only a subset of the emitters, without actuating all the emitters. For example, to reduce the power consumption of speech detection system 100, processing unit 112 may actuate only one emitter or a subset consisting of two or more emitters that illuminates a specific area on the user’s face that has been found to give the most useful information for generating the desired speech output.

[0075] Light detector 412, shown in Fig. 4, may be used to detect reflections from facial region 108 indicative of facial skin movements. As discussed above, a light detector may be capable of measuring properties of coherent or non -coherent light, such as power, frequency, phase, pulse timing, pulse duration, and other properties. In some embodiments, light detector 412 may include anAttorney Docket No. 16198.0057-00304 array of detecting elements, for example, a set of a charge -coupled device (CCD) sensors and / or a set of complementary metal -oxide semiconductor (CMOS) sensors, with objective optics for imaging facial region 108 onto the array. Due to the small dimensions of optical sensing unit 116 and its proximity to the skin surface, light detector 412 may have a sufficiently wide field of view to detect many of spots 106 at a high angle of at least 60°, at least 70°, or at least 90°. Light detector 412 may be configured to generate an output relating to the measured properties of the detected light. Consistent with the present disclosure, the output of light detector 412 may include any form of data determined in response to the received light reflections from facial region 108. In some embodiments, the output may include reflection signals that include electronic representation of one or more properties determined from the coherent or non-coherent light reflections. In other embodiments, the output may include raw measurements detected by at least one light detector 412.

[0076] In some embodiments, light detector 412 may measure one of more optical attributes associated with skin changes. The term “skin changes” refers to any detectable movements, alterations, or modifications that occurred to the skin. Such skin changes may include changes in the epidermis (i.e., the outermost layer of the skin), changes in the dermis (i.e., the middle layer of the skin), changes in the hypodermis (i.e., the deepest layer of the skin), and changes in deeper muscle tissues. The optical attributes may be measured without contacting the skin of user 102 . Examples of one of more optical attributes of the reflected light that may be measured by light detector 412 may include intensity, frequency, reflection, angle, sharpness, bidirectional reflectance distribution function, color, brightness, glossiness, transparency, opacity, surface texture, surface relief, surface movement, and other optical attributes derivable from analysis of light reflections. The output of light detector 412 may be used to determine information associated with skin changes. In some embodiments, the information associated with those skin changes may be derived from changes in a distance from the skin to the detector as the skin moves, and in other embodiments the changes may not be derived from variations in the distance of the skin from light detector 412. For example, the determined speed or angular speed of the changes of the facial skin may be determined by detecting the changes of non-distance measurements (e.g., image sharpness) overtime. Thus, in one nonlimiting example, optical attributes may be detected from random intensity variations observed when coherent light interacts with a rough or scattering surface, such as human skin. In another non -limiting example, optical attributes may be detected based on the interference of light waves, such as when interference patterns are used to measure the phase difference or amplitude changes between two or more optical paths.

[0077] In some embodiments, optical sensing unit 116 may not require reference to parameters of the light source, such as the light source’s wavelength, intensity, or coherence, and may not require a reference beam (typically used with a beam-splitter) to measure the one or more optical attributes of the reflected light. For example, optical sensing unit 116 may use a single beam to illuminate the skinAttorney Docket No. 16198.0057-00304 and then process the light reflections returned to light detector 412. While some speech detection systems may include a single pixel sensor (e.g., a photo diode), in other embodiments, light detector 412 may include one or more multi -pixel sensors (e.g., each pixel sensor includes more than 4 megapixels, more than 10 megapixels, or more than 10 megapixels) that enables producing an image providing spatial information beyond a single point. For example, a reflection image depicted in Fig. 6 may be produced from the output of light detector 412. As described throughout the disclosure, output of light detector 412 may be analyzed using image processing methods to determine patterns of light scattered off a surface. For example, features of secondary speckles may be determined.

[0078] In some non-limiting examples, optical sensing unit 116 may use a diffractive element to split the outbound beam to multiple beams and may not rely on superposition of coherent light waves to cause interference. In some non-limiting examples, optical sensing unit 116 may be arranged such that light detector 412 may be positioned along a different optical axis from light source 410. In other non-limiting examples, aligning the light source and the sensor along the same optical axis is needed may be used for maintaining coherence, achieving path length matching, ensuring spatial overlap, and preserving the sensitivity and accuracy of the interference patterns. However, since some implementations of light detector 412 detect a reflection image and not a distance to a point, optical sensing unit 116 may include a first optical axis for outbound light and a second optical axis, not aligned with the first optical axis, for inbound light. In some embodiments, light detector 412 is configured to measure both sub-microbic speed and depth changes in the ranges of 5-500 microns. In alternative embodiments, light detector 412 is configured to measure changes that are less than a micron. All the examples provided in this paragraph are alternatives and may be implemented in the many alternative embodiments provided herein, depending on the specifics of implementation.

[0079] Audio sensor 414, shown in Fig. 4, may include one or more audio sensors configured to capture audio by converting sounds to digital information. Some examples of audio sensors may include microphones, unidirectional microphones, bidirectional microphones, cardioid microphones, omnidirectional microphones, onboard microphones, wired microphones, wireless microphones, or any combination of the above. Audio sensor 414 may be configured to capture sounds uttered by user 102, thereby enabling user 102 to use speech detection system 100 as a conventional headphone when desired. Additionally or alternatively, audio sensor 414 may be used in conjunction with the silent speech sensing capabilities of speech detection system 100. In one embodiment, the audio signals output by audio sensor 414 can be used in changing the operational state of speech detection system 100. For example, processing unit 112 may generate the speech output only when audio sensor 414 does not detect vocalization of words by user 102. In another embodiment, audio sensor 414 may be used in a calibration procedure, in which optical sensing unit 116 detects micromovements of the skin while user 102 utters certain phonemes or words. Processing unit 112 may compare the reflection signals output by light detector 412 to the sounds sensed by audio sensor 414 to calibrate opticalAttorney Docket No. 16198.0057-00304 sensing unit 116. This calibration may include prompting user 102 to shift the position of optical sensing unit 116 to align the optical components in the desired position relative to facial region 108. In yet another embodiment, audio sensor 414 enables on-the-fly training of a neural network of speech detection system 100. For example, speech detection system 100 may be configured to correlate facial skin micromovements with words using audio signals concurrently captured with the micromovements. After recognizing recorded words, speech detection system 100 can perform a look -back to identify facial micromovement that preceded articulation of those words, thereby training speech detection system 100. In a similar way, speech detection systems can be used to train on expressions, commands, user recognition, and emotions.

[0080] Power source 416, shown in Fig. 4, may provide electrical energy to power speech detection system 100. A power source may include any device or system that can store, dispense, or convey electric power, including, but not limited to, one or more batteries (e.g., a lead-acid battery, a lithium-ion battery, a nickel-metal hydride battery, a nickel-cadmium battery), one or more capacitors, one or more connections to external power sources, one or more power convertors, or any combination of the foregoing. With reference to the example illustrated in Fig. 4, power source 416 may be mobile, which means that speech detection system 100 can be wearable. The mobility of the power source enables user 102 to use speech detection system 100 in a variety of situations. In other embodiments, power source 416 may be associated with a connection to an external power source (such as an electrical power grid) that may be used to charge power source 416.

[0081] Additional sensors 418, shown in Fig. 4, may include a variety of sensors, for example, image sensors, motion sensors, environmental sensors, Electromyography (EMG) sensors, resistive sensors, ultrasonic sensors, proximity sensors, biometric sensors, or other sensing devices configured to facilitate related functionalities. For example, speech detection system 100 may include one or more image sensors configured to capture visual information from the environment of user 102 by converting light (not emitted from light source 410) to image data. Consistent with the present disclosure, an image sensor may be included in any device or system capable of detecting and converting optical signals in the near-infrared, infrared, visible, and / or ultraviolet spectmms into electrical signals. Examples of image sensors may include digital cameras, semiconductor charge - coupled devices (CCDs), active pixel sensors in complementary metal -oxide semiconductor (CMOS), or N-type metal -oxide-semiconductor (NMOS, Live MOS). The electrical signals may be used to generate image data. Consistent with the present disclosure, the image data may include pixel data streams, digital images, digital video streams, data derived from captured images, and data that may be used to construct one or more 3D images, a sequence of 3D images, 3D videos, or a virtual 3D representation. The image data acquired by the one or more image sensors may be transmitted by wired or wireless transmission to processing unit 112 or to remote processing system 450 .Attorney Docket No. 16198.0057-00304

[0082] Speech detection system 100 may also include one or more motion sensors configured to measure motion of user 102. Specifically, a motion sensor may perform at least one of the following: detect motion of user 102, measure the velocity of user 102, measure the acceleration of user 102, or measure any other action that involves movement. In some embodiments, the motion sensor may include one or more accelerometers configured to detect changes in acceleration (e.g., proper acceleration) and / or to measure acceleration of speech detection system 100. In some embodiments, the motion sensor may include one or more gyroscopes configured to detect changes in the orientation of speech detection system 100 and / or to measure information related to the orientation of spe ech detection system 100. In some embodiments, the motion sensors may include one or more using image sensors, LIDAR sensors, radar sensors, or proximity sensors. For example, by analyzing captured images, processing device 400 may determine the motion of speech detection system 100, for example, using ego-motion algorithms. In addition, the processing device may determine the motion of objects in the environment of speech detection system 100, for example, through object tracking.

[0083] Speech detection system 100 may also include one or more environmental sensors of different types configured to capture data reflective of the environment of user 102. In some embodiments, the environmental sensor may include one or more chemical sensors configured to perform at least one of the following: measure chemical properties in the environment of user 102, measure changes in the chemical properties in the environment of user 102, detect the present of chemicals in the environment of user 102, and / or measure the concentration of chemicals in the environment of user 102. Examples of measurable chemical properties include pH level, toxicity, and temperature. Examples of chemicals or phenomena that may be measured include electrolytes, particular enzymes, particular hormones, particular proteins, smoke, carbon dioxide, carbon monoxide, oxygen, ozone, hydrogen, and hydrogen sulfide. In other embodiments, the environmental sensor may include one or more temperature sensors configured to detect changes in the temperature of the environment of user 102 and / or to measure the temperature of the environment of user 102. In other embodiments, the environmental sensor may include one or more barometers configured to detect changes in the atmospheric pressure in the environment of user 102 and / or to measure the atmospheric pressure in the environment of user 102. In other embodiments, the environmental sensor may include one or more light sensors configured to detect changes in the ambient light in the environment of user 102.

[0084] Network interface 420, shown in Fig. 4, may provide two-way data communications to a network, such as communications network 126. In one embodiment, network interface 420 may include an Integrated Services Digital Network (ISDN) card, cellular modem, satellite modem, or a modem to provide a data communication connection over the Internet. As another example, network interface 420 may include a Wireless Local Area Network (WLAN) card. In another embodiment,Attorney Docket No. 16198.0057-00304 network interface 420 may include an Ethernet port connected to radio frequency receivers and transmitters and / or optical (e.g., infrared) receivers and transmitters. The specific design and implementation of network interface 420 may depend on the communications network or networks over which speech detection system 100 is intended to operate. For example, in some embodiments, speech detection system 100 may include network interface 420 designed to operate over a GSM network, a GPRS network, an EDGE network, a Wi-Fi or WiMax network, and a Bluetooth network. In any such implementation, network interface 420 may be configured to send and receive electrical, electromagnetic, or optical signals that carry digital data streams or digital signals representing various types of information.

[0085] Data structure 422, shown in Fig. 4, may include any hardware, software, firmware, or combination thereof for storing and facilitating the retrieval of information from a database . The term “database” may be understood to include a collection of data that may be distributed or non- distributed. A database may include a database management system that controls the organization, storage and retrieval of data contained within the database. As described above, the data included in the database may be stored linearly, horizontally, hierarchically, relationally, non-relationally, uni- dimensionally, multidimensionally, operationally, in an ordered manner, in an unordered manner, in an object-oriented manner, in a centralized manner, in a decentralized manner, in a distributed manner, in a custom manner, or in any manner enabling data access. In disclosed embodiments, data structure 422 may include correlations of facial micromovements with words, commands, emotions, expressions, and / or biological conditions. The at least one processor may perform a lookup in the data structure to thereby interpret the detected facial skin micromovements. In accordance with one embodiment, at least some of the data stored in data structure 422 may alternatively or additionally be stored in remote processing system 450.

[0086] Consistent with the present disclosure, speech detection system 100 may be configured to communicate with a remote processing system 450 (e.g., mobile communications device 120 or server 122). Remote processing system 450 may directly or indirectly accesses a bus 452 (or other communication mechanism) interconnecting subsystems and components for transferring information within remote processing system 450. For example, bus 452 may interconnect a memory interface 454, a network interface 456, a power source 458, a processing device 460, one or more additional sensors 462, a data structure 464, and memory device 466.

[0087] Memory interface 454, shown in Fig. 4, may be used to access a software product and / or data stored on a non -transitory computer-readable medium or on other memory devices, such as memory devices 402, 466, data structure 422, or data structure 464. Memory device 466 may contain software modules to execute processes consistent with the present disclosure. In some embodiments, memory device 466 may include a shared memory module 472, a node registration module 473, a load balancing module 474, one or more computational nodes 475, an internal communication moduleAttorney Docket No. 16198.0057-00304476, an external communication module 477, and a database access module (not shown). Modules 472-477 may contain software instructions for execution by at least one processor (e.g., processing device 460) associated with remote processing system 450. Shared memory module 472, node registration module 473, load balancing module 474, computational node 475, and external communication module 477 may cooperate to perform various operations.

[0088] Shared memory module 472 may allow information sharing between remote processing system 450 and other devices related to one or more speech detection systems 100. In some embodiments, shared memory module 472 may be configured to enable processing device 460 to access, retrieve, and store data. For example, using shared memory module 472, processing device 460 may perform at least one of: executing software programs stored on memory devices 402, 466, data structure 422, or data structure 464; storing information in memory devices 402, 466, Data structure 422, or data structure 464; or retrieving information from memory devices 402, 466, data structure 422, or data structure 464.

[0089] Node registration module 473 may be configured to track the availability of one or more computational nodes 475. In some examples, node registration module 473 may be implemented as: a software program, such as a software program executed by one or more computational nodes 475, a hardware solution, or a combined software and hardware solution. In some implementations, node registration module 473 may communicate with one or more computational nodes 475, for example, using internal communication module 476. In some examples, one or more computational nodes 475 may notify node registration module 473 of their status, for example, by sending messages: at startup, at shutdown, at constant intervals, at selected times, in response to queries received from node registration module 473, or at any other determined times. In some examples, node registration module 473 may query about the status of one or more computational nodes 475, for example, by sending messages: at startup, at constant intervals, at selected times, or at any other determined times.

[0090] Load balancing module 474 may be configured to divide the workload among one or more computational nodes 475. In some examples, load balancing module 474 may be implemented as a software program, such as a software program executed by one or more of the computational nodes 475, a hardware solution, or a combined software and hardware solution. In some implementations, load balancing module 474 may interact with node registration module 473 to obtain information regarding the availability of one or more computational nodes 475. In some implementations, load balancing module 474 may communicate with one or more computational nodes 475, for example, using internal communication module 476. In some examples, one or more computational nodes 475 may notify load balancing module 474 of their status, for example, by sending messages: at startup, at shutdown, at constant intervals, at selected times, in response to queries received from load balancing module 474, or at any other determined times. In some examples, load balancing module 474 mayAttorney Docket No. 16198.0057-00304 query about the status of one or more computational nodes 475, for example, by sending messages: at startup, at constant intervals, at pre-selected times, or at any other determined times.

[0091] Internal communication module 476 may be configured to receive and / or to transmit information from one or more components of remote processing system 450. For example, control signals and / or synchronization signals may be sent and / or received through internal communication module 476. In one embodiment, input information for computer programs, output information of computer programs, and / or intermediate information of computer programs may be sent and / or received through internal communication module 476. In another embodiment, information received though internal communication module 476 may be stored in memory device 466 or in data structure 464. For example, information retrieved from data structure 464 may be transmitted using internal communication module 476. In another example, reference signals reflecting facial micromovements of user 102 may be stored in data structure 464 and accessed using internal communication module 476.

[0092] External communication module 477 may be configured to receive and / or to transmit information from one or more speech detection systems 100. For example, control signals may be sent and / or received through external communication module 477. In one embodiment, information received though external communication module 477 may be stored in memory device 466, in data structure 464, and / or any memory device in the one or more speech detection systems 100. In another embodiment, information retrieved from data structure 464 may be transmitted using external communication module 477 to speech detection system 100 or to any entity with whom user 102 communicates. For example, when user 102 communicate with a financial institution (e.g., a bank) information retrieved from data structure 464 may be transmitted to enable authentication of user 102. In another embodiment, sensor data may be transmitted and / or received using external communication module 477. Examples of such input data may include data received from speech detection system 100, information captured from the environment of user 102 using one or more sensors such as additional sensors 418 and additional sensors 462.

[0093] In some embodiments, aspects of modules 472-477 may be implemented in hardware, in software (including in one or more signal processing and / or application specific integrated circuits), in firmware, or in any combination thereof, executable by one or more processors, alone, or in various combinations with each other. Specifically, modules 472-477 may be configured to interact with each other and / or other modules of speech detection system 100 to perform functions consistent with disclosed embodiments. Memory device 466 may include additional modules and instructions or fewer modules and instructions.

[0094] Network interface 456, power source 458, processing device 460, additional sensors 462, and data structure 464, shown in Fig. 4, may share similar functionality with the functionality of corresponding elements in speech detection system 100, as described above. The specific design andAttorney Docket No. 16198.0057-00304 implementation of the above-mentioned components may vary based on the implementation of remote processing system 450. In addition, remote processing system 450 may include more or fewer components. For example, when remote processing system 450 is a mobile communications device associated with user 102 (e.g., mobile communications device 120) it may include a speaker, a microphone, and additional sensors.

[0095] The components and arrangements of speech detection system 100 and remote processing system 450 as illustrated in Fig. 4 are not intended to limit the disclosed embodiments. As will be appreciated by a person skilled in the art having the benefit of this disclosure, numerous variations and / or modifications may be made to the depicted configuration of speech detection system 100 and remote processing system 450. For example, not all components may be essential for the operation of an input unit in all cases. Any component may be located in any appropriate part of speech detection system 100 or remote processing system 450. Moreover, the components may be rearranged into a variety of configurations while providing the functionality of the disclosed embodiments. For example, some speech detection systems may not include all of the elements as shown in speech detection system 100 and in remote processing system 450. Other speech detection systems may include additional components and still fall within the scope of this disclosure.

[0096] Figs. 5 A and 5B include two schematic illustrations of optical sensing unit 116 as it detects facial skin micromovements in accordance with some embodiments of the present disclosure. The two schematic illustrations show a simplified scenario before muscle recruitment and after muscle recruitment. As depicted, optical sensing unit 116 may include an illumination module 500, a detection module 502, and, optionally, audio sensor 414. As discussed above and illustrated in Fig. 5, optical sensing unit 116 may be configured not to contact the user’s skin at facial region 108, but rather may be held at a distance D from the skin surface of facial region 108. The distance D of optical sensing unit 116 from the skin surface may be at least 5 mm, at least 7.5 mm, at least 10 mm, at least 15 mm, or at least 20 mm.

[0097] In the depicted embodiment, illumination module 500 includes light source 410 (e.g., an infrared laser diode) configured to generate an input light beam 504. Illumination module 500 further includes a beam-splitting element 506, such as a Dammann grating or another suitable type of diffractive optical element (DOE), configured to split input light beam 504 into multiple output beams 508, which form respective spots 106A-106E at a pattern (e.g., a matrix of locations) extending over facial region 108. In an alternative embodiment (not shown in the figure), illumination module 500 may include multiple light sources 410, which generate respective groups of output beams 508, covering different respective sub-areas within facial region 108. In this alternative embodiment, processing unit 112 may select and actuate only a subset of the multiple light sources, without actuating all of them. For example, to reduce the power consumption of speech detection system 100,Attorney Docket No. 16198.0057-00304 processing unit 112 may actuate only one light source or a group of two or more light sources that illuminate a part of facial region 108.

[0098] Detection module 502 may include light detector 412, which may include an array 510 of optical sensors (e.g., an array of CMOS image sensors) with objective optics 512 for obtaining reflections 300 of coherent light from facial region 108. Because of the small dimensions of optical sensing unit 116 and its proximity to the skin surface, detection module 502 may be configured to have a wide field of view to acquire reflections from many spots 106 at a high angle. As mentioned above, the field of view of light detector 412 may have an angular width of at least 60°, at least 70°, or at least 90°. Due to the roughness of the skin surface, the light patterns at spots 106 can be detected at these high angles, as well.

[0099] Speech detection system 100 may analyze light reflections 300 to determine facial skin micromovements resulting from recmitment of muscle fiber 520. Determining the facial skin micromovements may include determining an amount of the skin movement, determining a direction of the skin movement, and / or determining an acceleration of the skin movement. The determined facial skin micromovements may include voluntary and / or involuntary recmitment of muscle fiber 520. Muscle fiber 520 may be part of: a zygomaticus muscle, an orbicularis oris muscle, a risorius muscle, genioglossus muscle, or a levator labii superioris alaeque nasi muscle. Processing device 400 may be configured to perform a first speckle analysis on light reflected from a first region of face in proximity to spot 106A to determine that the first region moved by a distance dl, i.e., first facial skin micromovement 522A; and perform a second speckle analysis on light reflected from a second region of face in proximity to spot 106E to determine that the second region moved by a distance d2, i.e., second facial skin micromovement 522B. Thereafter, processing device 400 may use the determined movements of the first region and the second region to ascertain at least one spoken word. Consistent with disclosed embodiments, distances dl and d2 may be less than 1000 micrometers, less than 100 micrometers, less than 10 micrometers, or less.

[0100] Fig. 6 is a schematic illustration of a reflection image 600 associated with light reflections 300 received from an area of facial region 108 associated with a single spot 106 (e.g., spot 106A depicted in Fig. 5). In disclosed embodiments, processing device 400 may receive reflection signals indicative of coherent light reflections from facial region 108. The reflection signals may be represented by reflection image 600. Thereafter, processing device 400 may determine the facial skin micromovements by applying a light reflection analysis. When light source 410 is a coherent light source, the light reflection analysis may include a speckle analysis or any pattern-based analysis. Such analysis may be performed by processing device 400 or processing device 460 to identify a speckle pattern and derive thereof movement of a corresponding area of facial region 108.

[0101] In the depicted example, a speckle 602 appears in reflection image 600 after recruitment of muscle fiber 520. The detected speckle or any other detected pattern may then be processed toAttorney Docket No. 16198.0057-00304 generate reflection image data. With reference to the example discussed above, assuming reflection image 600 reflects spot 106A, the reflection image data may include data indicating that the first region moved by a distance dl . In some cases, the reflection image data may be processed by any image processing algorithms (e.g., CNN and RNN) to determine skin movements of at least two areas within facial region 108. Thereafter, processing device 400 may use one or more machine learning (ML) algorithms and artificial intelligence (Al) algorithms to decipher the reflection image data and to extract meaning from the facial skin micromovement.

[0102] As shown in Fig. 7, memory device 700 may contain software modules to execute processes consistent with the present disclosure. In particular, memory device 700 may include an illumination control module 702, a sensors communication module 704, a light reflections processing module 706, an artificial neural network (ANN) training module 710, a subvocalization deciphering module 708, an output determination module 712, and a database structure access module 714. The disclosed embodiments are not limited to any particular configuration of memory device 700. Further, processing device 400 and / or processing device 460 may execute the instructions stored in any of modules 702-714 included in memory device 700. It is to be understood that references in the following discussions to a processing device may refer to processing device 400 of speech detection system 100 and processing device 460 of remote processing system 450 individually or collectively. Accordingly, steps of any of the following processes associated with modules 702-714 may be performed by one or more processors associated with speech detection system 100.

[0103] Consistent with disclosed embodiments, illumination control module 702, sensors communication module 704, light reflections processing module 706, subvocalization deciphering module 708, ANN training module 710, output determination module 712, and database access module 714 may cooperate to perform various operations. For example, illumination control module 702 may determine light characteristics for illuminating facial region 108. Sensors communication module 704 may receive coherent light reflections from facial region 108 and output associated reflection signals. Light reflections processing module 706 may process the reflection signals to determine facial skin micromovements. Subvocalization deciphering module 708 and database access module 714 may cooperate to extract meaning (e.g., determine silently spoken words) from the facial skin micromovements. In some cases, ANN training module 710 may use the determined silently spoken words and the determined facial skin micromovements to train an artificial network. Output determination module 712 may generate a presentation of the determined words.

[0104] Illumination control module 702 may regulate the operation of light source 410 to illuminate facial region 108. In some embodiments, illumination control module 702 may determine values for characteristics of projected light 104 such as light intensity, pulse frequency, duty cycle, illumination pattern, light flux, or any other optical characteristic. In a specific embodiment, as long as user 102 is not speaking, speech detection system 100 may operate in a first illumination modeAttorney Docket No. 16198.0057-00304(e.g., low frame rate) to conserve power of its battery. While speech detection system 100 operates at this first illumination mode, it may process the images to detect at least one trigger in the reflection signals (e.g., a movement of the face) indicative of speech. When such trigger is detected, illumination control module 702 may cause the coherent light source to operate in a second illumination mode (e.g., high frame rate) to enable detection of changes in the coherent light patterns (e.g., speckle) that occur due to silent speech. Illumination control module 702 may also be configured to change one or more characteristics of projected light 104 based on various types of triggers. The various types of triggers may be detected by analysis of data from sensors communication module 704.

[0105] Sensors communication module 704 may regulate the operation of light detector 412, audio sensor 414, and additional sensors 418 to receive captured measurements from one or more sensors, integrated with, or connected to, speech detection system 100. In one embodiment, sensors communication module 704 may use the signals received from one or more sensors to generate sensor data associated with user 102. In one example, sensors communication module 704 may receive reflection signals from light detector 412 and may generate a first data stream of reflections images from which the facial skin micromovements in the facial region may be determined. In another example, sensors communication module 704 may receive audio signals from audio sensor 414 and may generate a second data stream from which the words vocally spoken by user 102 may be determined. In another example, sensors communication module 704 may receive motion signals from a motion sensor included in additional sensors 418 and generate a third data stream from which an activity that user 102 is engaged with may be determined. Sensors communication module 704 may convey the sensor data to other software modules for processing.

[0106] Light reflections processing module 706 may process the sensor data received from sensors communication module 704 in preparation for speech deciphering. In one embodiment, light reflections processing module 706 may receive from sensors communication module 704 reflection signals indicative of coherent light reflections from facial region 108 that originates from light detector 412. The reflection signals may be represented by a reflection image (e.g., reflection image 600) that can be processed by at least one image processing algorithm to extracts the skin motion at a set of pre-selected locations on the face of user 102. The number of locations to inspect may be an input to the image processing algorithm. In some cases, the locations on the skin that are extracted for coherent light processing may be taken from a list of points of interest. The list of points of interest specifies anatomical locations that correspond with the zygomaticus muscle, the orbicularis oris muscle, the risorius muscle, genioglossus muscle, or the levator labii superioris alaeque nasi muscle. In plain language, the list of points of interest may include specific points in the cheek above mouth, in the chin, in mid-jaw, in the cheek below mouth, in the high cheek, and in the back of the cheek. Consistent with the present disclosure, the list of points of interest may be dynamically updated withAttorney Docket No. 16198.0057-00304 more points on the face that are extracted during a training phase. The entire set of locations may be ordered in descending order such that any subset of the list (in order) minimizes the word error rate (WER) with respect to the chosen number of locations that are inspected. In another embodiment, light reflections processing module 706 may crop each of the coherent light spots that were extracted from the raw image frames around the coherent light spots, and the algorithm process only the cropped images. Typically, the process of coherent light spot processing involves reducing by two the order of magnitude of a size of full frame image pixels (of ~1.5MP) that are received from sensors communication module 704, with a very short exposure. Exposure may be dynamically set and adapted to be able to capture only coherent light reflections and not skin segments. The cropped images of the coherent light spots may depict coherent light patterns. In other embodiments, light reflections processing module 706 may apply an image processing algorithm on the reflection image. For example, light reflections processing module 706 may improve the images’ contrast, by removing noise using a threshold to determine black pixels and computing a characteristic metric of the coherent light, such as scalar speckle energy measure, e.g., an average intensity. In addition, light reflections processing module 706 may analyze changes in time in the reflections pattern (e.g., in average speckle intensity). Alternatively, other metrics may be used such as the detection of specific coherent light patterns. Thereafter, light reflections processing module 706 may assign a sequence of values of the characteristic metric of the coherent light, which may be calculated frame-by-frame and aggregated to generate reflection image data indicative of facial skin micromovements. Light reflections processing module 706 may convey the reflection image data indicative of facial skin micromovements to other software modules for processing.

[0107] Subvocalization deciphering module 708 may use machine learning (ML) algorithms and artificial intelligence (Al) algorithms to decipher the reflection image data indicative of facial skin micromovements received from light reflections processing module 706 . Consistent with the present disclosure, deciphering the reflection image data may include extracting meaning from the detected facial skin micromovements. In one embodiment, subvocalization deciphering module 708 may use a trained ANN to correlate words with the facial skin micromovements. Different types of ANNs may be used, such as a classification NN that eventually outputs words, and a sequence -to-sequence NN which outputs a sentence (word sequence). In some embodiments, during normal speech of the user, speech detection system 100 may simultaneously sample the voice of user 102 and the facial movements. Automatic speech recognition (ASR) and Natural Language Processing (NLP) algorithms may be applied by subvocalization deciphering module 708 on the actual voice, and the outcome of these algorithms may be used for optimizing the parameters of the algorithms used by subvocalization deciphering module 708. These parameters may include the weights of the various neural networks, as well as the spatial distribution of laser beams for optimal performance. In addition, subvocalization deciphering module 708 may limit the output of the algorithms to a pre -defined word set mayAttorney Docket No. 16198.0057-00304 significantly increase the accuracy of word detection in cases of ambiguity, i.e., when two different words result in similar micromovements on the facial skin. The used word set can be personalized overtime, adjusting the dictionary to the actual words used by the specific user, with their respective frequency and context. In addition, subvocalization deciphering module 708 may use the context of a conversation between user 102 and a callee. The context may be determined from the input of the words and sentences extraction algorithms to increase the accuracy by eliminating out-of-context options. The context of the conversation may be understood by applying Automatic speech recognition (ASR) and Natural Language Processing (NLP) algorithms on the side of user 102 and on the side of the callee.

[0108] ANN training module 710 may be used to train an ANN to perform silent speech deciphering, in accordance with embodiments of the disclosure. To train an ANN such as the one that may be used by subvocalization deciphering module 708 may require several thousands of examples. To achieve this, ANN training module 710 may rely on a large group of people (e.g., a group of reference human subjects). In one example, subvocalization deciphering module 708 may perform fine adjustments to the ANN such that it is customized to user 102. In this manner, within minutes or less of wearing speech detection system 100, subvocalization deciphering module 708 may be ready for deciphering the facial skin micromovements. ANN training module 710 can be used to train two different ANN types: a classification neural network that eventually outputs words, and a sequence - to-sequence neural network which outputs a sentence (word sequence). To do so, ANN training module 710 may upload from a memory training data, such as silent speech data received from light reflections processing module 706 that was gathered from multiple reference human subjects. The silent speech data may be collected from a wide variety of people (people of varying ages, genders, ethnicities, physical disabilities, etc.). It is to be noted that the number of examples required for learning and generalization may be task -dependent. For word / utterance prediction (within a closed group) at least several thousands of examples may be gathered. Thereafter, ANN training module 710 may augment the image processed training data to get more artificial data for the training process. In particular, the augmented data may include image processed coherent light patterns, with some of the image processing steps described herein. The data augmentation process may include the steps of (i) time dropout, where amplitudes at random time points are replaced by zeros;(ii) frequency dropout, where the signal is transformed into the frequency domain, and random frequency chunks are filtered out;(iii) clipping, where the maximum amplitude of the signal at random time points is clamped. This clipping may add a saturation effect to the data;(iv) noise addition, where Gaussian noise is added to the signal, and speed change, where the signal is resampled to achieve a slightly lower or slightly faster signal.

[0109] The augmented dataset may go through a feature extraction process. In this process, ANN training module 710 may compute time domain silent speech features. For this purpose, for example,Attorney Docket No. 16198.0057-00304 each signal may be split into low and high frequency components, x low and x high, and windowed to create time frames, for example, using a frame length of 27ms and shift of 10 ms. For each of the frame five time-domain features and the nine frequency domain features, a total of 14 features per signal may be computed. Specifically, the time-domain features may be represented as follows:where ZCR is the zero-crossing rate. In addition, in this example, the magnitude values used are from a 16-point short Fourier transform, i.e., frequency domain features and all features are normalized to zero mean unit variance.

[0110] Thereafter, ANN training module 710 may split the data into training, validation, and test sets. The training set may be the data used to train the model. Hyperparameter tuning may be done using the validation set, and final evaluation may be done using the test set. The model architecture may be task dependent. Two different examples describe training two networks for two conceptually different tasks. A first task may include signal transcription, i.e., translating silent speech to text by generating a word, a phoneme, or a letter. This first task may be addressed by using a sequence -to- sequence model. A second task may include predicting a word or an utterance, i.e., categorizing utterances uttered by users into a single category within a closed group. This second task may be addressed by using a classification model. The disclosed sequence -to -sequence model may be composed of an encoder, which may transform the input signal into high level representations (embeddings), and a decoder, which produces linguistic outputs (i.e., characters or words) from the encoded representations. The input entering the encoder may be a sequence of feature vectors. In one example, the input may enter the first layer of the encoder, a temporal convolution layer, which may down-sample the data to achieve a good performance. The model may use an order of a hundred of such convolution layers.

[0111] In some embodiments, the outputs from the temporal convolution layer at each time step may be passed to three layers of bidirectional recurrent neural networks (RNN). ANN training module 710 may employ long short-term memory (LTSM) as units in each RNN layer. Each RNN state may be a concatenation of the state of the forward RNN with the state of the backward RNN. The decoder RNN may be initialized with the final state of the encoder RNN (concatenation of the final state of the forward encoder RNN with the first state of the backward encoder RNN). At each time step, the decoder RNN may receive as input the preceding word, encoded one-hot and embedded in a 150- dimensional space with a fully connected layer. The decoder RNN output may be projected through a matrix into the space of words or phonemes (depending on the training data). The sequence -to- sequence model may condition the next step prediction on the previous prediction. During learning, a log probability may be maximized:Attorney Docket No. 16198.0057-00304where y<i is the ground truth of the previous prediction. The classification neural network may be composed of the encoder as in the sequence -to -sequence network and an additional fully connected classification layer on top of the encoder output. The output may be projected into the space of closed words and the scores may be translated into probabilities for each word in the dictionary. The results of the above entire procedure may include two types of trained ANNs, expressed in computed coefficients. The coefficients may be stored in a data structure associated with speech detection system 100 (e.g., data structure 422 and data structure 464). In day-to-day use, ANN training module 710 may receive up to date coefficients for the trained ANN. The first ANN task may be the signal transcription, i.e., translating silent speech to text by wo rd / phoneme / letter generation. The second ANN task may be word / utterance prediction, i.e., categorizing utterances uttered by users into a single category within a closed group.

[0112] Output determination module 712 may regulate the operation of output unit 114 and the operation of network interface 420 to generate output using speaker 404, light indicator 406, haptic feedback device 408, and / or to send data to a remote computing device. In some embodiments, the output generated by output determination module 712 may include various types of output associated with silent speech determined from detected facial skin micromovements. Specifically, output determination module 712 may synthesize vocalization of words determined from the facial skin movements by subvocalization deciphering module 708. The synthesis may emulate a voice of user 102 or emulate a voice of someone other than user 102 (e.g., a voice of a celebrity or preselected template voice). The vocalization of the words may be presented via speaker 404 or transmitted to the remote computing device via network interface 420. Alternatively, output determination module 712 may generate a textual output from the facial skin movements by subvocalization deciphering module 708. The textual output may be transmitted to the remote computing device via network interface 420. According to another embodiment, the output generated by output determination module 712 may relate to the operation of speech detection system 100. In some cases, light indicator 406 may include a light indicator that shows the battery status of speech detection system 100. For example, the light indicator may start to blink when speech detection system 100 has a low battery. Additional examples of the types of output that may be generated by output determination module 712 are described throughout the present disclosure.

[0113] Database access module 714 may cooperate with data structures 422 and 464 to retrieve stored data. The retrieved data may include, for example, correlations between a plurality of words and a plurality of facial skin movements, correlations between a specific individual and a plurality of facial skin micromovements associated with the specific individual, and more. As described above, subvocalization deciphering module 708 may use a trained ANN to perform silent speechAttorney Docket No. 16198.0057-00304 deciphering. The trained ANN may use data stored in data structures 422 and 464 to extract meaning from detected facial skin micromovements. Data structures 422 and 464 may include separate databases, including, for example, a vector database, raster database, tile database, viewport database, and / or a user input database. The data stored in data structures 422 and 464 may be received from modules 702-712 or other components of speech detection system 100. Moreover, the data stored in data structures 422 and 464 may be provided as input using data entry, data transfer, or data uploading.

[0114] Modules 702-714 may be implemented in software, hardware, firmware, a mix of any of those, or the like. Processing devices of speech detection system 100 and remote processing system 450 may be configured to execute the instmctions of modules 702-714. In some embodiments, aspects of modules 702-714 may be implemented in hardware, in software (including in one or more signal processing and / or application specific integrated circuits), in firmware, or in any combination thereof, executable by one or more processors, alone, or in various combinations with each other. Specifically, modules 702-714 may be configured to interact with each other and / or other modules associated with speech detection system 100 to perform functions consistent with disclosed embodiments.

[0115] In accordance with one implementation, a speech detection system projects a pattern of light on facial skin (e.g., a cheek) of a user. Thereafter, the speech detection system may detect light reflections from various locations of the facial skin. Notably, reflections associated with specific areas may be more relevant for extracting meaning (e.g., determining communication) than other areas. The specific areas may be those that are located closer to particular facial muscles. Identifying the specific locations may pose challenges because each user has unique facial features, and the position of the light source and / or detector relative to the user’s face may change during every usage and even during ongoing operations. The following paragraphs describes systems, methods, and computer program products for identifying the locations of those specific areas, using the light reflections from the specific areas to extract meaning, and ignoring light reflections from other areas to conserve processing resources.

[0116] Some disclosed embodiments involve interpreting facial skin movements. The term “interpreting facial skin movements” refers to extracting meaning from detected skin movements, as described elsewhere in this disclosure. In one example, interpreting facial skin movements may include determining one or more vocalized or subvocalized words from the facial skin movements or determining a facial expression (e.g., happy, sad, anger, fear, surprise, disgust, contempt, or other emotion) of the individual. In another example, interpreting facial skin movements may include determining an identity of the individual. These facial skin movements may be detectable as described elsewhere in this disclosure.

[0117] Some disclosed embodiments involve projecting light on a plurality of facial region areas of an individual, wherein the plurality of areas includes at least a first area and a second area. The termAttorney Docket No. 16198.0057-00304“projecting” includes controlling a light source (e.g., a coherent light source) such that it emits light in a given direction (e.g., toward a portion of the face), as discussed elsewhere in this disclosure. The term “individual” includes a person who uses a speech detection system (or another person to whom the light source is projected), as described elsewhere in this disclosure. The term “facial region area” or simply “area” in the context of the face includes a portion of the face of the individual, as described elsewhere in this disclosure. For example, a facial region area may have a size of at least 1 cm2, at least 2 cm2, at least 4 cm2, at least 6 cm2, or at least 8 cm2. Consistent with some disclosed embodiments, the projected light illuminates a plurality of facial region areas. For example, the plurality of areas includes 4, 8, 16, 32, or any othernumbers of areas. In some cases, the projected light may include at least one spot, as described elsewhere in this disclosure. The at least one spot may illuminate more than one facial region area, for example, as illustrated in Fig. 3, a single spot 106 may illuminate different portions of facial region 108. For example, spot 106 may include a first portion 304A associated with a first facial muscle and a second portion 304B associated with a second facial muscle. Alternatively, a single facial region area may be illuminated by multiple light spots. Some of the plurality of areas may be spaced apart from each other while others of the plurality of areas may be overlapping with each other. The term “spaced apart” may refer to being non- overlapping or separated by at least some distance. Thus, spaced apart areas may refer to two or more facial region areas that do not overlap with each other and have even a very small gap in between. For example, stating that a first facial region area is spaced apart from a second facial region area may include distances between the first and second region of at least 5 mm, at least 10 mm, at least 15 mm, or any other desired distance. In some embodiments the distance may be less than 1 mm, or between 1mm and 5mm. In some cases, only a portion of a facial region area may be illuminated by the projected light. In other cases, all of the facial region areas may be illuminated by the projected light. By way of example, Figs. 8 and 12 illustrate illuminating plurality of facial region areas of an individual using a plurality of spots. As illustrated, each of areas 800A and 800B are illustrated by more than one light spot.

[0118] Some disclosed embodiments involve illuminating at least a portion of the first area and at least a portion of the second area with a common light spot. As used herein, the term “at least a portion” and / or grammatical equivalents thereof can refer to any fraction of a whole amount. For example, “at least a portion” can refer to at least about 1%, 5%, 10%, 20%, 40%, 65%, 90%, 95%, 99%, 99.9%, or 100% of a whole amount, or any other fraction. The term “common light spot” means that a single (common) light spot may cover some or all of the first area and the second area. The common light spot may illuminate at least a portion of the first area and the second area. In one example, the common light spot may illuminate 30% of the first area and 10% of the second area. In another example, the common light spot may illuminate 100% of the first area and 100% of the second area. Controlling the at least one coherent light source may include illuminating a continuousAttorney Docket No. 16198.0057-00304 area on the face that includes the first area and the second area. By way of one example, as illustrated in Fig. 3 single light spot 106 may illuminate two or more facial areas (e.g., 304A and 304B).

[0119] Some disclosed embodiments involve illuminating the first area with a first group of spots and illuminating the second area with a second group of sports distinct from the first group of spots . The term “group of spots” refers to more than one light spot. The number of spots in the group of spots may range from two to 64 or more. For example, the group of spots may include 4 spots, 8 spots, 16 spots, 32 spots, 64 spots, or any number of spots greater than two. There may be variations in illumination characteristics between spots or within the group of spots, as discussed elsewhere in this disclosure. Illuminating an area with a group of spots may refer to illuminating some or all of a facial area region by two or more spots. In one example, the group of spots may illuminate at least 15% of the area, at least 40% of the area, or at least 70% of the area. A first area may be illuminated by a first group of spots and a second area may be illuminated by a second group of spots distinct from the first group of spots. In this context, the term “distinct” means that the first group of spots is distinguishable from the second group of spots. For example, the first group of spots may include at least one spot not included in the second group of spots. By way of example, Figs. 8 and 12 illustrate a first area facial area 800A illuminated by a first group of spots 808A and a second area 800B illuminated by a second group of sports 808B distinct from the first group of spots.

[0120] Some disclosed embodiments involve operating a coherent light source (as described elsewhere in this disclosure) located within a wearable housing (as described elsewhere in this disclosure) in a manner enabling illumination of the plurality of facial region areas. Enabling illumination, as used herein, may refer to a process of controlling a light source to generate at least one light beam and directing the at least one light beam toward the plurality of facial region areas. For example, enabling illumination may also include utilizing a beam-splitting element (as described elsewhere in this disclosure) configured to split an input beam into multiple output beams (as described elsewhere in this disclosure) extending over a portion of a face. In an alternative embodiment, enabling illumination may include utilizing multiple light sources which generate respective groups of output beams, covering different respective sub-areas within a portion of a face. Figs. 1 and 2 illustrate an example implementation of speech detection system (e.g., speech detection system 100) in which at least one facial region area (e.g., facial region 108) is illuminated by a plurality of light spots (e.g., light spots 106). In some embodiments, the plurality of light spots may be generated by optical sensing unit 116 that includes at least one light source 410 and at least one light detector 412 and located in a wearable housing 110.

[0121] Some disclosed embodiments involve operating a coherent light source (as described elsewhere in this disclosure) located remote from a wearable housing (as described elsewhere in this disclosure) in a manner enabling illumination of the plurality of facial region areas (as described elsewhere in this disclosure). The term “located remote” indicates that two objects are separated fromAttorney Docket No. 16198.0057-00304 each other and with a physical distance between them such that they do not appear physically as a unified component. For example, the coherent light source may be part of device other than the speech detection system and located more than 1 cm from a wearable housing of the speech detection system. As another example, the coherent light source may be located more than 3 cm from a wearable housing of the speech detection system. It should be understood that the distances 1 cm and 3 cm are exemplary and nonlimiting and other distances may be used. Fig. 3 illustrate an example implementation of speech detection system in which a plurality of facial region areas (e.g., first portion 304A of facial region 108 and second portion 304B of facial region 108 ) are illuminated by a coherent light source located remote from the wearable housing (e.g., a non -wearable light source 302).

[0122] In some disclosed embodiments, the first area is closer to at least one of a zygomaticus muscle or a risorius muscle than the second area. The phrase “a first area is closer to a muscle than a second area” means that a distance of the first area to a specific muscle is less than a distance of the second area to a specific muscle. For example, the distances may be measured from an edge of an area to an edge of specific muscle, from a center of an area to a center of a specific muscle, or any combination thereof. In this context, the center of a shape (i.e., the first area, the second area, or a specific muscle) may be a geometric center, which is the point which corresponds to the mean position of all the points in shape; a circumscribed center, which is the center of the smallest circle that completely encloses the 2D shape; an incenter, which is the center of the inscribed circle that is tangent to all sides of the 2D shape, or any other reference point previously defined. As discussed, the first area is closer to at least one of a zygomaticus muscle or a risorius muscle than a second area. In other words, the disclosed embodiments capture two example use cases, the first example use case is that the first area is closer to the zygomaticus muscle than the second area. The second example use case is that the first area is closer to the risorius muscle than the second area. By way of example, Fig. 8 illustrates one implementation of the first and second example use cases. Specifically, the first use case is illustrated with regards to user 102 A and the second use case is illustrated with regards to user 102 B.

[0123] Fig. 8 illustrates two example use cases for interpreting facial skin movements. In both example use cases, a plurality of facial areas 800 of user 102 may be illuminated by at least one light source (e.g., light source 410, not shown). The depicted plurality of areas includes at least a first area 800 A and a second area 800B. In the first example use case involving user 102 A, first area 800 A is closer to the zygomaticus muscle than second area 800B, and in the second example use case involving user 102 B, first area 800A is closer to the risorius muscle than second area 800B.

[0124] Some disclosed embodiments involve receiving reflections from the plurality of areas. The term “receiving” may include obtaining, retrieving, acquiring, or otherwise gaining access to data or signals. In some cases, receiving may include reading data from memory and / or obtaining data from aAttorney Docket No. 16198.0057-00304 computing device via a (e.g., wired and / or wireless) communications channel. In other cases, receiving may include detecting electromagnetic waves (e.g., in the visible or invisible spectrum) and generating an output relating to measured properties of the electromagnetic waves. In a first embodiment, at least one processor may receive data indicative of light reflected from the plurality of areas from at least one detector. In a second embodiment, at least one detector may receive light rays reflected from the plurality of areas. The term “reflections” refers to one or more light rays bouncing off a surface (e.g., the individual’s face) or data derived from the one or more light rays bouncing off the surface. For example, the reflections may include light detected by a light detector after it was deflected from an object. The light detected by the light detector may be generated by at least one coherent light source of the disclosed speech detection system and / or may be generated from sources other than the disclosed speech detection system. By way of one example, light detector 412 in Figs. 5A and 5B is employed to receive reflections 300 that originated from light generated by light source 410.

[0125] By way of example with reference to the two uses cases depicted in Fig. 8, a reflection image 802A may represent the reflections received from the first area 800 A, and reflection image 802B may represent the reflections received from the second area 800B. As illustrated, in the first example use case, reflection image 802A represents the reflections received from an area closer to the zygomaticus muscle; and in the second example use case, reflection image 802A represents the reflections received from an area closer to the risorius muscle.

[0126] Some disclosed embodiments involve detecting first facial skin movements corresponding to reflections from the first area and second facial skin movements corresponding to reflections from the second area. The term “detecting” in this context refers to the process of discovering, identifying, or determining the existence of light reflections (or signals associated therewith). In one example, a change in the position of facial skin may be detected. As discussed elsewhere in this disclosure, the detection process may involve using various techniques or technologies to determine the existence of the pattern or the event. In some cases, the process of detecting facial skin movement may involve determining if there is any movement that occurred and recording information representing the detected movement. For example, at least one processor may detect facial skin movements by applying a light reflection analysis on received reflections. In other cases, detecting facial skin movements may include determining times in which facial skin movements occurred. In other cases, detecting facial skin movements may include determining data representing the facial skin movements (e.g., direction, velocity, acceleration). The term “facial skin movements” broadly refers to any type of movements prompted by recruitment of underlying facial muscles. The facial skin movements include facial skin micromovements — as described elsewhere in this disclosure — and larger-scale skin movements generally visible and detectable to the naked eye without the need for magnification (e.g., a smile, a yawn, a frown). The term “the facial skin movements corresponding to reflections from aAttorney Docket No. 16198.0057-00304 specific area” means that the detected facial skin movements took place in a specific area of the face from which reflections were received. For example, detecting first facial skin movements corresponding to reflections from the first area means that the first facial skin movements may be detected by analyzing reflections received from the first area; and detecting second facial skin movements corresponding to reflections from the second area means that the second facial skin movements may be detected by analyzing reflections received from the second area.

[0127] In some disclosed embodiments, detecting the first facial skin movements involves performing a first speckle analysis on light reflected from the first area, and detecting the second facial skin movements involves performing a second speckle analysis on light reflected from the second area. The term “performing” refers to the act of carrying out a task, activity, or function. The term “speckle analysis” may be understood as described elsewhere in this disclosure. Consistent with the present disclosure, performing a speckle analysis may include detecting a speckle pattern, or any other patterns in signals received from a light reflected from a facial region area. For example, performing a speckle analysis may include identifying secondary speckle patterns that arise due to reflection of the coherent light from each area. In other embodiments, detecting facial skin movements may involve performing a pattern-based analysis or an image-based analysis additionally or alternatively from performing a speckle analysis.

[0128] Consistent with some disclosed embodiments, the first speckle analysis and the second speckle analysis occur concurrently by the at least one processor, the term “occur concurrently” means that two or more events occur during coincident or overlapping time periods, either where one begins and ends during the duration of the other, or where a later one starts before the completion of the other. In some cases the two or more events may be speckle analyses (or any pattern -based analysis). In order for the first speckle analysis and the second speckle analysis to occur concurrently, the at least one processor may include a plurality of processors or a multi -core processor that allows multiple speckle analyses to be executed simultaneously.

[0129] By way of example with reference to the two uses cases depicted in Fig. 8, first facial skin movements 804A may correspond to reflections from the first area 800A and second facial skin movements 804B may correspond to reflections from the second area 800B . For example, in the first example use case, first facial skin movements 804A correspond to reflections received from an area closer to the zygomaticus muscle; and in the second example use case, second facial skin movements 804B correspond to reflections received from an area closer to the risorius muscle.

[0130] Some disclosed embodiments involve determining, based on differences between the first facial skin movements and the second facial skin movements, that the reflections from the first area closer to the at least one of a zygomaticus muscle or a risorius muscle are a stronger indicator of communication than the reflections from the second area. Determining refers to ascertaining. For example, from the differences between the first and second facial skin movements, the processor mayAttorney Docket No. 16198.0057-00304 determine which is closer to the associated muscle. The differences between the first facial skin movements and the second facial skin movements may include any distinctions, variations, or dissimilarities between the first facial skin movements and the second facial skin movements. The differences between the first facial skin movements and the second facial skin movements may be determined using at least one of the following techniques: surface alignment, point-to-point comparison, surface registration, topological analysis, or any other technique for determining differences between two data sets. For example, the differences between the first facial skin movements and the second facial skin movements may include differences in the movement intensity, movement trajectory, the movement speed, and / or various changes in topography the facial skin. Based on the differences, the at least one processor may determine that reflections from a first area are a stronger indicator of communication than the reflections from a second area. The term “communication” refers to the process of conveying information through various mediums, such as spoken language, words, body language, gestures, or signals. For example, the communication may include verbal cues (e.g., words, phrases, and language) and non-verbal cues (e.g., body language, facial expressions, gestures, and eye contact). The term “indicator of communication” refers to a measure or sign reflective of an information conveyed by the individual. For example, the statement that reflections from the first area are a stronger indicator of communication than the reflections from a second area means that it may be easier to determine that the individual intends to convey information and what communication the individual intends to convey from the first facial skin movements than from the second facial skin movements. For example, the reflections from the first area may be a stronger indicator of communication than the reflections from a second area because the facial skin micromovements determined from the reflections from the first area may be associated with a higher velocity, a higher displacement, or a higher other parameter indicating that the individual intents to convey information and / or the content of the information that the individual intends to convey. Consistent with disclosed embodiments, in the first example use case, when the first area is closer to the zygomaticus muscle, the first facial skin movements may reflect movements with a velocity on the order of one to ten pm / ms, and the second facial skin movements may reflect smaller movements, if any. In the second example use case, when the first area is closer to the risorius muscle, the first facial skin movements may reflect movements on the order of 0.5 -2 mm, and the second facial skin movements reflect smaller movements, if any.

[0131] Consistent with some disclosed embodiments, the differences between the first facial skin movements and the second facial skin movements include differences of less than 100 microns. The term “differences of less than 100 microns” means that the changes between a first parameter that represents the first facial skin movements and a second parameter that represents second facial skin movements is less than 100 microns. In one example, the first parameter may be a magnitude of a first displacement change vector associated with the first facial skin movements and a second parameterAttorney Docket No. 16198.0057-00304 may be a magnitude of a second displacement change vector associated with the second facial skin movements. A displacement change is a vector that quantifies the distance and direction changes between two measurements of the facial skin. For example, the differences between the first facial skin movements and the second facial skin movements include differences of less than 50 microns, less than 10 microns, or less than 1 micron. In other embodiments, the differences between the first facial skin movements and the second facial skin movements include differences of less than 1 millimeter. Accordingly, the determination that the reflections from the first area are a stronger indicator of communication than the reflections from the second area is based on the differences of less than 1 millimeter, less than 100 microns, less than 50 microns, less than 10 microns, or less than 1 micron.

[0132] Some disclosed embodiments involve, based on the determination that the reflections from the first area are a stronger indicator of communication, processing the reflections from the first area to ascertain the communication. The term “processing” refers to the act of performing operations or transformations on data or information to achieve a desired outcome. For example, processing may include manipulating, analyzing, or altering inputs in a systematic way to produce meaningful outputs. The term “processing reflections” means extracting information from signals representing the received reflections. For example, processing reflections may include actions, such as filtering, amplifying, modulating, and applying light reflection analysis as described elsewhere in this disclosure. Based on the determination that the reflections from the first area are a stronger indicator of communication, the reflections from the first area are processed to ascertain the communication. The term “ascertain the communication” means determining speech or facial expressions associated with non-verbal communication from facial movements, as described elsewhere in this disclosure. Consistent with the present disclosure, the reflections from the first area may be processed to create images of speckle patterns. Even at fast exposure times, such as 10 ms, the velocity of motion of the skin may be sufficient to make the speckle pattern change during each frame so that the bright pixels are blurred and washed out. The degree of speckle blur of a given spot in a given frame, as manifested by the loss of contrast in the image, for example, may be indicative of the instantaneous velocity of motion of the skin in the small area of the cheek under the spot. Processing the reflections from the first area may also include extracting quantitative image features from the images of speckle patterns. Vectors of these features, extracted from successive image frames, may be input to a neural network in order to ascertain the communication. Details of neural network architectures and training algorithms that may be used for this purpose are described elsewhere in this disclosure. An example feature that may be extracted for the purpose of ascertaining the communication may include speckle contrast. Any suitable measure of contrast may be used for this purpose, for example, the mean square value of the luminance gradient taking over the area of the speckle pattern. High contrast in the speckle pattern of a given spot from the first area may be indicative that the corresponding location ofAttorney Docket No. 16198.0057-00304 the cheek is stationary, while reduced contrast may be indicative of motion. The contrast decreases with increasing velocity of motion. Contrast features of this sort may be typically extracted from multiple spots distributed over the first area. Additionally, or alternatively, other features may be extracted from the speckle images and input to the neural network. Examples of such features may include total brightness of the speckle pattern and orientation of the speckle pattern, for instance, as computed by a Sobel filter. By way of one example, subvocalization deciphering module 708 in Fig. 7 may be used for processing the reflections from the first area to ascertain the communication.

[0133] Consistent with some disclosed embodiments, the communication ascertained from the reflections from the first area includes words articulated by the individual. “Ascertaining words articulated by the individual” refers to understanding words that are either vocalized or sub vocalized by the individual. By processing the signals resulting from reflections, words can be ascertained as discussed elsewhere herein. By way of example, the word “Hello” in Fig. 8 represents the words articulated by user 102 A or user 102 B that may be ascertained from the reflections from the first area.

[0134] Consistent with some disclosed embodiments, the communication ascertained from the reflections from the first area includes non-verbal cues of the individual. The term “non-verbal cues” refers to the various forms of communication that occur without the use of spoken words. Some examples of non-verbal cues may include facial expressions, body language, gestures, eye contact, tone of voice, postures, and other subtle signals that convey meaning in interpersonal interactions. For example, non-verbal cues, such as facial expressions, may be used to communicate basic emotions like happiness, sadness, anger, fear, surprise, and disgust. As discussed elsewhere in this disclosure, the at least one processor may determine a non-verbal cue by analyzing reflection signals representing facial skin micromovements in the first facial area. By way of example, the emoji in Fig. 8 represents the non-verbal cues that may be ascertained from the reflections from the first area.

[0135] Some disclosed embodiments involve, based on the determination that the reflections from the first area are a stronger indicator of communication, ignoring the reflections from the second area. In this context, the term “ignoring the reflections” means that the processing actions on the signals representing the received reflections from the second area are less than the processing actions on the signals representing the received reflections from the first area. In one embodiment, signals representing the received reflections from the second area may be filtered, amplified, and analyzed to determine the second facial skin movements, but some quantitative features may not be extracted because the communication may not be ascertained from signals representing the received reflections from the second area. In another embodiment which also involves “ignoring,” during a first time frame, reflections from both the first area and the second area may be processed to determine which area is closer to the zygomaticus muscle or the risorius muscle. Thereafter, during a subsequentAttorney Docket No. 16198.0057-00304 second time frame, and upon determining that the first area is closer to the zygomaticus muscle or the risorius muscle, reflections from the second area may be automatically discarded.

[0136] According to some disclosed embodiments, ignoring the reflections from the second area includes omitting use of the reflections from the second area to ascertain the communication. The term “omitting use” refers to not using information associated with reflections from the second area when determining the meaning of the communication.

[0137] By way of example with reference to the two uses cases depicted in Fig. 8, reflection image802A may be processed to ascertain communication 806 from first facial skin movements 804A associated with the zygomaticus muscle or the risorius muscle, and reflection image 802B may ignored, e.g., not used or omitted in ascertaining the communication. As depicted, the ascertained communication may include at least one word 806A (articulated silently or vocally by user 102 A or user 102 B) and / or at least one facial expression 806B that serves as an example of a non-verbal cue.

[0138] Some disclosed embodiments involve determining, based on differences between the first facial skin movements and the second facial skin movements, that the first area is closer than the second area to the subcutaneous tissue associated with cranial nerve V or with cranial nerve VII. The term “subcutaneous tissue” refers to the layer of tissue located beneath the skin and above the underlying muscles and bones. It is composed of fat cells, connective tissue, blood vessels, nerves, and other structures. Cranial nerve V, also known as the trigeminal nerve, is a sensory nerve for the face that control of jaw muscles. Cranial nerve VII controls facial expressions and carries taste sensation from the front of the tongue. Based on differences between the first facial skin movements and the second facial skin movements (as described above), a determination may be made that the first area is closer than the second area to the subcutaneous tissue associated with cranial nerve V or with cranial nerve VII.

[0139] Some disclosed embodiments involve operating a coherent light source in a manner enabling bi-mode illumination of the plurality of facial region areas. The term “coherent light source” may be understood as described elsewhere in this disclosure. Operating a coherent light source in this context refers to regulating, supervising, instmcting, allowing, and / or enabling the coherent light source to illuminate at least part of a face. For example, the coherent light source may be controlled to illuminate a region of a face in a specific mode of illumination when turned on in response to a trigger. Bi-mode illumination refers to a capability of the coherent light source to illuminate an object using at least two different modes of illumination. The term “mode of illumination” refers to a specific configuration or settings of the coherent light source. Each of the two modes may be associated with different values of illumination parameters, such as light intensity, illumination pattern, pulse frequency, duty cycle, light flux. Light source 410 in Fig. 4 is one example of either a single mode or multi-mode (e.g., bi-mode) light source.Attorney Docket No. 16198.0057-00304

[0140] In some disclosed embodiments, a first light intensity of the first mode of illumination differs from a second light intensity of the second mode of illumination. In some disclosed embodiments, a first illumination pattern of the first mode of illumination differs from a second illumination pattern of the second mode of illumination. Light intensity refers to a brightness level of an illumination and an illumination pattern refers to an arrangement, distribution, or sequence of coherent or non-coherent light emitted from a source or reflected off a surface. The light pattern may be created by a specific design, shape, or configuration of light sources to create a particular visual or non -visual effect on the portion of the face. Examples of illumination patterns may include a grid of light spots having the same size, a grid of light spots having the various sizes, a single light spot, or any other pattern.

[0141] Some disclosed embodiments involve analyzing reflections associated with a first mode of illumination to identify one or more light spots associated with the first area, and analyzing reflections associated with a second mode of illumination to ascertain the communication. The term “identifying one or more light spots associated with the first area” means determining which of the light spots projected by the coherent light source are located in the first area. For example, identifying the one or more light spots associated with the first area may be implemented by comparing light intensity at a particular location with boundaries of the first area, based on image analysis of the face of the individual, or by any other processing method. In one example, the first mode of illumination may include a first illumination pattern (e.g., 64 light spots) and the second mode of illumination may include a second illumination pattern (e.g., 32 light spots). By way of example, with reference to the first example use case depicted in Fig. 8, the first mode of illumination may be used to identify eight light spots included within first area 800 A associated with the zygomaticus muscle. Thereafter, the second mode of illumination (e.g., 4 light spots) may be used to illuminate first area 800A in a manner that enables ascertaining the communication from received reflections.

[0142] Consistent with some disclosed embodiments, the first area is closer than the second area to the zygomaticus muscle, and the plurality of areas further include a third area closer to the risorius muscle than each of the first area and second area. The terms “plurality of areas” and “closer to” may be understood as described elsewhere in this disclosure. By way of example with reference to Fig. 9, the plurality of facial areas 800 includes the first area 800A closer to the zygomaticus muscle than second area 800B, and a third area 800C closer to the risorius muscle than each of the first area 800A and second area 800B. In some disclosed embodiments, based on a determination that user 102C is engaged in silent speech, a processing device of the speech detection system may process the reflections from the first area 800A to ascertain the communication, and ignore the reflections from the second area 800B and the third area 800C. In other embodiments, based on a determination that user 102 C is engaged in voiced speech, a processing device of the speech detection system mayAttorney Docket No. 16198.0057-00304 process the reflections from third area 800C to ascertain the communication, and ignore the reflections from the second area 800B and the first area 800A.

[0143] Some disclosed embodiments involve analyzing reflected light from the first area when speech is generated with perceptible vocalization (i.e., voiced speech) and analyzing reflected light from the third area when speech is generated in an absence of perceptible vocalization (i.e., silent speech). In other words, rather than monitoring the entire cheek and processing reflections from a plurality of areas, the speech detection system may process reflections received from a subset of the cheek area (e.g., only a few square millimeters or centimeters) in these two areas to detect both silent and voiced speech.

[0144] Furthermore, when the plurality of areas are illuminated by multiple light sources (e.g., an array of laser diodes) only the light sources that illuminate these two areas may be actuated, thus reducing power consumption. If a laige movement of the speech detection system relative to the skin is detected, a different set of light sources may be actuated.

[0145] In some disclosed embodiments, different modes of processing may be applied to ascertain silent speech from voiced speech. For example, during silent speech, the first area being closer to the zygomaticus muscle may exhibit movements with a velocity on the order of one to ten pm / ms. Therefore, features of the images of the speckles themselves may change rapidly, and these features may be analyzed to generate an output. But during voiced speech, the third area being closer to the risorius muscle may exhibit movements on the order of 0.5-2 mm. Thus, the locations of the spots on the cheek may shift laterally due to the movement of the cheek. In this case, the lateral movements of the spots may be indicative of changes in the distance of the spots from the speech detection system, which may thus function as a sort of depth sensor. The two processing modes — speckle sensing and depth sensing — may be used individually in detecting silent and voiced speech, respectively. Alternatively, or additionally, these two processing modes may be used together to improve the precision and specificity of measurement, for example, by applying measurements of voiced speech by a given user to learn the patterns of microscopic movement that will occur in silent speech by the same user.

[0146] Disclosed speech detection systems may enable synthesizing not only words, but also vocal characteristics such as intonation, pitch, pace, intensity, prosody, articulation, and volume. For example, in some implementations, a system may enable such synthesizing. Some disclosed speech detection systems can determine the vocal characteristics from facial skin micromovements, similarly to ways described above for determining words from facial skin micromovements.

[0147] In one aspect of this disclosure, methods, systems, and software are provided for performing voice synthesizing operations from detected facial skin micromovements. The phrase “performing voice synthesizing operations from detected facial skin micromovements” refers to creating artificial speech from the movements detected on the skin of the face. This may include, forAttorney Docket No. 16198.0057-00304 example, the conversion of facial skin micromovements into audible speech, the generation of synthetic speech based on the detected movements, or the creation of real-time subtitles for silent communication. It might also enable gesture -based command input, emotion recognition from facial dynamics, or silent language translation using micromovement data.

[0148] Some disclosed embodiments involve a non -transitory computer readable medium containing instructions that when executed by at least one processor cause the at least one processor to perform operations for synthesizing voice from subvocalized speech. The term “voice” refers audible characteristics of spoken language. Voice may include, for example, a combination of auditory qualities, including but not limited to pitch, tone, timbre, rhythm, and articulation, that characterize spoken language and that may convey both linguistic content and expressive nuance. The term “subvocalized speech” or “subvocalized words” refers to silent articulation of words, phrases, or sentences that involves internal speech or invisible, inaudible movement of the speech -related muscles and nerves. The words of subvocalized speech may be completely inaudible or while nominally audible, may not be loud enough to be understood by a person nearby. While there is no audible output, disclosed systems, methods, and media can detect and process the underlying physiological signals to interpret and synthesize expressive, naturalistic speech from silent cues. This offers new possibilities for accessibility, privacy, and human -computer interaction.

[0149] Some disclosed embodiments involve operating at least one sensor for detecting neuromuscular activity associated with a non -lip region of a head of an individual. The term “operating” refers to the execution or functioning, in the current context, of a sensor. The sensor may be part of a speech detection system, and operating may include capturing information with the sensor, for use in analyzing, and / or synthesizing signals derived from facial micromovements. For example, operating may involve initiating the sensors, and collecting physiological data. Ultimately that data might be interpreted to produce synthesized speech or outputs for user interaction. The term “sensor” refers to any device or component capable of detecting physical, optical, electrical, or mechanical phenomena. For example, a sensor might detect changes associated with neuromuscular or speech-related activity in non-lip regions of the head. By way of other examples, a sensor may involve a light-based detectorthat measures reflected light from the cheek, a mechanical deformation sensor embedded in a wearable patch, or an electrode configured to sense electrical signals from cranial nerves. The term “neuromuscular activity” refers to activation of muscles and nerves. Such activation may facilitate movement or expression in various regions of the head and face. For example, neuromuscular activity may involve subtle muscle contractions, nerve impulses, and tissue responses that occur during speech articulation, facial gestures, or silent internal communication. These physiological processes produce micromovements and electrical signals that can be detected and analyzed to interpret both spoken and subvocalized language.Attorney Docket No. 16198.0057-00304

[0150] The term “individual” refers to any person. Such a person may be a user who interacts with or wears the speech detection system and whose neuromuscular activity is being measured and processed. For example, an individual may involve a person silently articulating words for communication in a noise -sensitive environment, or someone using the system for accessibility or privacy -focused applications. The phrase “operating at least one sensor” refers to an action of controlling or managing one or more sensors. This may involve, for example, turning on a sensor, adjusting the intensity of the sensor, receiving one or more signals from the sensor, and / or directing the sensor towards a specific area. The phrase “non -lip region of a head” refers to an area of the head that is not the lips. For example, the non -lip region of a head of an individual may include any anatomical area of the head that excludes the lips. This may encompass, for example, cheeks, chin, jawline, temples, forehead, ear canal, neck, and other craniofacial regions. These regions may be associated with muscle groups or nerve pathways that contribute to speech articulation or expression, either alongside or in the absence of vocalization. The non -lip region may include subcutaneous tissue, dermal layers, or bone-adjacent areas, and may be involved in both voluntary and involuntary neuromuscular activity. In particular, the non-lip region may be associated with a subcutaneous tissue associated with cranial nerve V or cranial nerve VII.

[0151] Some disclosed embodiments involve a sensorthat includes a light source and a light detector configured to detect reflections of light projected from the light source, wherein the reflections are indicative of non-lip facial skin micromovements associated with the neuromuscular activity. The term “light source” refers to any device or component that emits light. Examples of light sources include a laser diode, LED, or other optical emitter, used to project light. The taigeted light, in this context, is directed to a non-lip region of the face for the purpose of detecting micromovements. The term “light detector” may refer to a sensor or photodetector capable of receiving and measuring reflected light. The light may be reflected from the skin onto a component such as a photodiode, or phototransistor, or any other component capable of detecting light.. Such a component may be used to sense changes in the reflected light caused by subtle facial micromovements during subvocalized speech. The light source may emit infrared, visible, or near-infrared light toward a region of the face excluding the lips, such as, for example, the cheek, temple, jawline, or forehead. The light detector may be positioned to receive reflected light from the skin surface and measure changes in intensity, angle, or pattern of the reflected light. These changes may correspond to subtle micromovements of the skin caused by muscle contractions or nerve activity during sub vocalization. The term “reflected light” refers to light rays that bounce off a surface rather than being absorbed or transmitted through it. The reflected light may take the form of any optical signal. For example, light that returns from the skin surface after being emitted by the light source may include specular reflections, diffuse reflections, scattered light, or any combination thereof. The reflections may be analyzed to determine displacement, vibration, or deformation of the skin surface, which may be indicative ofAttorney Docket No. 16198.0057-00304 neuromuscular activity associated with silent speech. The phrase “non -lip facial skin micromovements” may refer to any minute movement or deformation of the skin in facial regions excluding the lips. These micromovements may result from muscle engagement, nerve activation, or other physiological processes involved in subvocalization. The micromovements may be imperceptible to the naked eye but detectable through optical sensing techniques. As used herein, nonlip facial skin micromovements may include any measurable change in skin position, tension, or contour in non-lip regions such as, for example, the cheeks, temples, jawline, or forehead.

[0152] Some disclosed embodiments include a wearable housing that incorporates the light source and detector. A wearable housing refers to a structure or enclosure designed to be worn on the body. For example, a wearable housing may include the frame of eyeglasses, smart glasses, or an appendage to either; , a head-mounted device; or a cheek -mounted patch. The system may use time-of-flight measurements, structured light, or photoplethysmography to detect micromovements and convert them into subvocalization signals. These signals may then be processed to determine subvocalized words and associated voice characteristics, enabling expressive speech synthesizing from silent input.

[0153] Some disclosed embodiments involve a sensor that includes a sensor configured to detect mechanical deformation caused by muscle engagement, wherein the detected deformation is indicative of facial muscle activity in the non-lip region associated with the neuromuscular activity. The term “mechanical deformation” refers to any physical change in tissue structure or position resulting from activation of underlying muscles. This may include compression, expansion, torsion, bending, or vibration of skin, cartilage, or subcutaneous layers. The mechanical deformation may occur in facial regions excluding the lips, such as, for example, the cheeks, temples, jawline, forehead, or ear canal, and may be associated with silent articulation of speech. The deformation may be detected, for example, using strain gauges, piezoelectric sensors, capacitive sensors, resistive sensors, accelerometers, or any other modality capable of measuring mechanical change. The sensor may be embedded in a wearable device, such as a patch, band, headset, or earbud, and may be positioned to monitor specific regions of the head where silent speech -related muscle activity occurs. The term “facial muscle activity” refers to any engagement of muscles located in the face. These muscles may contribute to speech articulation or expression and may include the buccinator, masseter, temporalis, orbicularis oculi, frontalis, or other craniofacial muscles. The facial muscle activity may be voluntary or involuntary, and may occur during subvocalization, internal speech, or imagined articulation. Facial muscle activity may include any detectable movement, tension, or deformation of tissue that reflects neuromuscular engagement in the context of silent or sub vocalized speech.

[0154] Some disclosed embodiments include a mechanical deformation sensor integrated into a wearable housing. For example, the sensor may be embedded in a cheek -mounted patch or temple- worn band that detects subtle shifts in skin tension during subvocalized speech. The sensor may transmit deformation data to a processing unit, which analyzes the signals to identify subvocalizedAttorney Docket No. 16198.0057-00304 words and associated voice characteristics. The system may operate in real-time and may be calibrated to distinguish between speech-related deformations and unrelated facial movements.

[0155] Some disclosed embodiments involve a sensor that includes an electrode configured to detect electrical signals transmitted through cranial nerves, wherein the detected electrical signals are indicative of facial nerve activity in the non-lip region associated with the neuromuscular activity. The term electrode refers to a conductor or conductive interface through which electricity enters or leaves a medium. An electrode may serve as the interface between electrical circuits and the materials involved in electrochemical or electronic processes. The phrase “electrode configured to detect electrical signals transmitted through cranial nerves” refers to any conductive element or interface designed to sense neural electrical activity associated with nerve pathways within or near the head or. The electrode may be configured to measure bioelectrical signals, such as action potentials or field potentials, generated by neural activity that propagate through cranial nerves during subvocalization. These signals may be captured from the surface of the skin, from within the ear canal, or from other regions of the head using surface electrodes, intradermal electrodes, or in-ear electrodes. The electrode may be made of metal, conductive polymer, or other biocompatible material, and may be integrated into a wearable device such as, e.g., a headset, earbud, glasses frame, or adhesive patch. The electrode may be configured to detect signals from one or more cranial nerves, including but not limited to the facial nerve (cranial nerve VII), trigeminal nerve (cranial nerve V), vestibulocochlear nerve (cranial nerve VIII), or glossopharyngeal nerve (cranial nerve IX). The electrical signals may be processed to extract features such as amplitude, frequency, timing, or waveform shape, which may be indicative of neuromuscular activity related to silent speech. The term “facial nerve activity” refers to any neural signaling that activates muscles or tissues in facial regions. This may include nerve impulses that control the cheeks, jaw, temples, forehead, or ear canal. The activity may be voluntary or involuntary and may occur during internal speech, imagined articulation, or silent rehearsal of words. Facial nerve activity may include any detectable electrical signal that reflects the engagement of cranial nerves in the context of subvocalized speech.

[0156] Some disclosed embodiments include a wearable electrode positioned to detect cranial nerve signals from the temple, jawline, or ear canal. Any electrode configured to be worn is a wearable electrode. For example, an in-ear electrode may detect electrical activity from the facial or vestibulocochlear nerve during subvocalization. The signals may be transmitted to a processing unit that analyzes the data to identify subvocalized words and determine associated voice characteristics. The system may operate in real-time and may be calibrated to distinguish speech -related neural activity from unrelated physiological signals.

[0157] Some disclosed embodiments involve receiving, from the at least one sensor, subvocalization signals indicative of the neuromuscular activity associated with the non-lip region, wherein the subvocalization signals are captured in an absence of perceptible vocalization.Attorney Docket No. 16198.0057-00304The term “receiving” refers to obtaining, gathering, capturing, acquiring, and / or detecting. For example, signals can be received from a sensor that monitors physiological activity. Receiving may involve gathering data outputs from light detectors, mechanical sensors, or electrodes as they measure micromovements, deformations, or electrical signals in the non -lip regions of the head. Receiving may include continuously or incrementally monitoring these signals, recording their variations, and transmitting the collected information to one or more processors for further analysis. The term “subvocalization signals” refers to any measurable signal associated with silent or imperceptable speech, regardless of the anatomical origin of the signal within the head or body. For example, subvocalization signals may be electrical, mechanical, optical, acoustic, thermal, or any other form of detectable output generated by muscle movement, nerve transmission, tissue deformation, or other physiological activity. The signals may originate from facial muscles, cranial nerves, ear canal structures, or other head -based anatomical features involved in speech articulation or expression. Receiving subvocalization signals may refer to capturing and collecting the physiological signal s produced when an individual silently articulates words or phrases — that is, when they internally "speak" without generating any audible sound. These signals may originate from neuromuscular activity in various non-lip regions of the head, including areas such as the cheeks, jawline, temples, forehead, ear canal, and other craniofacial sites. Subvocalization signals may encompass micromovements of the skin or underlying tissue caused by muscle engagement, mechanical deformations detected in the facial or neck area, and electrical signals generated by nerve activity during silent articulation. The signals may be received from off-the-shelf or specialized sensors that detect and record the minute, sometimes imperceptible, physical or electrical changes associated with silent or internal speech. The signals may be generated, for example, by the individual's body as the individual engages in subvocalization — internally articulating words or phrases without producing any audible voice. These signals may arise from neuromuscular activity such as muscle contractions in the face, neck, or head (excluding the lips), nerve signals transmitted through cranial nerves like the facial or trigeminal nerves, and / or subtle skin movements or mechanical deformations in non-lip regions. Even in the absence of sound, signals may occur, for example, during mental rehearsal of words, silent reading, or imagined speech. Detection and collection of these signals may depend on sensors designed to monitor such subtle physiological changes. Light-based sensors may employ a light source and light detector to measure reflections from the skin, capturing micromovements as the individual subvocalizes. Mechanical deformation sensors may be capable of sensing changes in skin or tissue shape, tension, or displacement resulting from muscle activity in non-lip facial regions. Electrodes may be used to pick up electrical signals from cranial nerves through electrical contact with the skin, ear canal, or other parts of the head. These sensors may be integrated into wearable devices — such as smart glasses, ear buds, patches placed on the cheek or temple, or headsets — and the collected data may then be transmitted to one or more processors which may analyze the signals to,Attorney Docket No. 16198.0057-00304 for example, decode the words and expressive characteristics of the intended, though unspoken, speech.

[0158] The term “captured” refers to recording or collecting physiological signals, such as micromovements, electrical impulses, or tissue deformations, generated during silent or subvocalized speech. For example, capturing may involve using specialized sensors — such as optical detectors, mechanical deformation sensors, or electrodes — to gather real-time data from non-lip regions of the face or head. The phrase “in the absence of perceptible vocalization” refers to situations in which an individual does not produce any audible speech or sound through the act of speaking, or situations where the sound is so soft that it is imperceptible to those nearby. This condition may occur during silent articulation, internal speech, or imagined verbal expression, where words, phrases, or sentences are formed in the mind or with barely noticeable movements, but no sound is emitted that can be heard by others. The absence of perceptible vocalization may be intentional, such as when communicating silently, reading to oneself, or rehearsing language mentally, and simply means that there is no outward, detectable vocal output. This may include whispering, mouthing words silently, internal speech (i.e., thinking words), or any other form of speech -related activity that does not result in sound waves perceptible to human hearing or conventional microphones. The absence of perceptible vocalization may be determined, e.g., by the lack of acoustic energy, lack of airflow through the vocal tract, or other indicators of silence.

[0159] At least one sensor may monitor neuromuscular activity (e.g., muscle movement, skin movement, and / or electrical signals) from various non-lip regions of the head, enabling robust identification and analysis of subvocalized words or phrases. In some embodiments, the at least one sensor may be calibrated to function in environments where audible speech is not desirable or possible, such as noise -sensitive settings, covert communications, or accessibility applications. In some embodiments, at least one sensor may be controlled to detect only signals generated without perceptible vocalization, ensuring, for example, that only silent or internal speech produces actionable data. In other embodiments, all detected signals may be received and / or processed.

[0160] Some embodiments of the disclosure involve receiving reflection signals associated with light projected on the non-lip region, wherein the reflection signals are indicative of non-lip facial skin micromovements. The term “receiving” may be understood as described elsewhere in this disclosure. The term “reflection signals” refers to the result of light being reflected off the non-lip region of the face. As previously discussed, such signals may be the result of skin micromovements. The reflection signal may, for example, be the result of detecting changes in the intensity or angle of the reflected light that correspond to small movements of the skin. The term “indicative of’ refers to an ability to convey information about something. For example, light reflections from a non-lip portion of the face may convey information about associated skin micromovements. The information may be a property or signal that reflects the presence or intensity of a specific physiological activity.Attorney Docket No. 16198.0057-00304For example, reflection signals indicative of non -lip facial skin micromovements may involve subtle changes in light detected from regions such as the cheek or temple, which correspond to underlying muscle contractions associated with silent articulation.

[0161] Some disclosed embodiments involve processing the received subvocalization signals to determine a plurality of subvocalized words including a first subvocalized word and a second subvocalized word. The term “processing” refers to performing an action on signals. In this example, processing may be used to identify one or more words that have been subvocalized, or spoken without producing audible sound. The processing may include analyzing the received subvocalization signals (e.g., reflection signals). Analysis may be based on historical data where prior signals were correlated to spoken words, phenomes, or other portions of speech. The prior analysis may have been accomplished using artificial intelligence.

[0162] Additionally or alternatively, Al may be employed in real time to decipher or analyze the reflection signals to determine one or more words subvocalized or spoken without producing words perceptible to a bystander. The subvocalized word may be formed through neuromuscular activity associated with silent speech, internal speech, imagined speech, or non -audible articulation. The articulation may involve movement or activation of muscles, nerves, or tissues in the head or neck region, including but not limited to the cheeks, jaw, tongue, throat, temples, or ear canal. A subvocalized word may include a complete word (e.g., “hello”), a partial word or syllable (e.g., “hel”), a phoneme (e.g., / h / ), a linguistic constmct such as a command or keyword, or a word formed in response to a prompt or internal intention. The generation of a subvocalized word may occur during silent reading, mental rehearsal, covert communication, or any other context in which the individual intends to produce speech without audible output. The word may be detected through sensors that measure electrical signals, mechanical deformations, optical reflections, or other physiological indicators of speech -related neuromuscular activity. A subvocalized word may include any internally generated linguistic unit that is intended to be communicated or processed, regardless of whether it is audibly expressed. The terms first and second in the context of words is intended to indicate that there are two separate words as opposed to a single word.

[0163] Some disclosed embodiments involve analyzing subvocalization signals received from sensors positioned near, on, or attached to the head. As previously described, these signals may be light reflections or electo -physiological signals associated with micromovements of the cheek or temple. The signals may include features such as amplitude, frequency, timing, or spatial distribution, and these features may be used to identify subvocalized words. In some embodiments, an artificial neural network (ANN) training module may be used to learn associations between signal patterns and specific words. The system may be trained on a dataset of known subvocalized speech to improve accuracy and generalization. The output determination module may then classify incoming signals as corresponding to a first subvocalized word and a second subvocalized word, enabling the system toAttorney Docket No. 16198.0057-00304 reconstruct the intended speech content. The processing may also involve temporal sequencing, where the system identifies the order in which subvocalized words are produced, allowing for sentence -level reconstruction. In other embodiments, the system may incorporate contextual data (e.g., emotional or environmental cues) to refine word identification and improve semantic accuracy.

[0164] Some disclosed embodiments involve determining a first voice characteristic from a first set of the subvocalization signals associated with the first subvocalized word. The term “determining” may be understood as described elsewhere in this disclosure. The term “voice characteristic” refers to a feature or attribute of a voice, such as its pitch, volume, tone, or timbre. A voice characteristic may include any attribute that affects the auditory quality of speech. Voice characteristics may define how a spoken word is perceived in terms of its expressive, acoustic, and prosodic features. These characteristics may include, but are not limited to, intonation, which refers to the rise and fall of pitch across a spoken phrase and may convey emotion or emphasis; pitch, which is the perceived frequency of the voice and may vary based on speaker identity or emotional state; pace, which refers to the speed at which words are spoken and may reflect urgency, hesitation, or fluency; intensity, which relates to the loudness or energy of speech and may indicate emphasis or emotional arousal; prosody, which encompasses the rhythm, stress, and intonation patterns that contribute to the natural flow and expressiveness of speech; articulation, which refers to the clarity and precision of speech sounds and may be influenced by physical or cognitive conditions; and volume, which is the amplitude of the spoken output and may vary based on context, intent, or environmental factors. A voice characteristic may include any measurable or inferable feature that contributes to the expressive or perceptual quality of synthesized or spoken speech. For example, a voice characteristic may be the unique sound of a person's voice that allows it to be recognized. The determination of a voice characteristic may be based on signal amplitude, frequency, timing, waveform shape, or other measurable parameters. It may also involve machine learning models trained to infer expressive features from physiological data. The system may determine the voice characteristic in real-time or post-processing, and may use it to synthesize speech that reflects the speaker’s intended expression, emotion, or emphasis. The phrase “a set of the subvocalization signals associated with a subvocalized word” refers to a group of subvocalization signals that are related to a specific subvocalized word. This may include, for example, the signals that are received when a person subvocalizes the word “hello.” The first set of the subvocalization signals may include any portion of the captured data that is temporally or spatially associated with the first subvocalized word, and may be derived from micromovements, muscle contractions, nerve signals, or other neuromuscular activity.

[0165] The system may analyze micromovement amplitude, timing, or frequency to infer pitch or intensity. In other embodiments, the system may detect emotional cues embedded in the neuromuscular activity and use those cues to determine prosody or articulation. The voiceAttorney Docket No. 16198.0057-00304 characteristic may be stored, transmitted, or used immediately to synthesize speech output that mimics the user’s intended vocal expression.

[0166] Some embodiments of the disclosure involve determining a second voice characteristic from a second set of subvocalization signals associated with the second subvocalized word, the second voice characteristic differing from the first voice characteristic. The term “differing from” refers to the action of being different or distinct in some way. This may involve, for example, the second voice characteristic having a different pitch, volume, or tone than the first voice characteristic, or any other variation in expressive or acoustic features between two subvocalized words. The difference may be in one or more dimensions, including but not limited to: a change in pitch (e.g., low to high), a change in pace (e.g., fast to slow), a change in intensity (e.g., soft to forceful), a change in prosody (e.g., monotone to expressive), a change in articulation (e.g., slurred to crisp), or a change in volume (e.g., quiet to loud). The difference may be intentional, contextual, emotional, or physiological, and may be inferred from signal parameters such as amplitude, frequency, timing, waveform shape, or neural activation patterns. A difference in voice characteristics may include any measurable or inferable variation in how two subvocalized words would be audibly expressed.

[0167] For example, some disclosed embodiments include a processing unit configured to analyze a second set of subvocalization signals corresponding to a second subvocalized word. For example, the system may use modules such as an ANN training module, speech deciphering module, and output determination module to extract expressive features from the second signal set. These features may be used to determine a second voice characteristic, such as a lower pitch or faster pace, that differs from the first voice characteristic previously determined. The system may compare the two sets of signals to identify differences in amplitude, timing, frequency, or other parameters. These differences may be used to synthesize speech output that reflects expressive variation between the first and second subvocalized words. In some embodiments, the system may use contextual data, emotional cues, or user preferences to guide the determination of voice characteristics and ensure that the synthesized speech reflects the speaker’s intended expression.

[0168] Some disclosed embodiments involve synthesizing audible output of the plurality of subvocalized words, wherein the first subvocalized word is audibly outputted in a manner reflecting the first voice characteristic and the second subvocalized word is audibly outputted in a manner reflecting the second voice characteristic. The term “synthesizing” refers to the process of generating or producing a coherent voice output by combining or integrating different sound elements or speech units, or any process by which a system generates sound signals that represent the internally articulated words detected from subvocalization. The synthesizing may occur in real-time or near- real-time, and may be based on stored voice models, dynamic signal processing, or machine learning algorithms trained to convert physiological signals into speech. For example, synthesizing may include combining phonemes, pitch, tone, and intonation to create a natural -sounding voice, orAttorney Docket No. 16198.0057-00304 merging prerecorded audio segments to produce continuous speech. This process may involve manipulating various audio parameters to achieve a desired vocal quality or to replicate a specific voice pattern and thereby cause a generation of audible speech from the sub vocalized words. The term “audible output” refers to any sound that can be heard or any expressive rendering of each word in accordance with its associated voice features. The audible output may be produced using a speaker, headphone, bone conduction device, or any other sound -emitting component. It will be understood, however, that the output may also be non-audible (e.g., text output may be provided in addition or instead of the audible output). As an example, the first word may be synthesized with a higher pitch and slower pace, while the second word may be synthesized with a lower pitch and faster pace. The system may use voice synthesizing engines capable of modulating intonation, pitch, pace, intensity, prosody, articulation, and volume to reflect the speaker’s intended expression. The voice characteristics may be applied individually to each word or across a sequence of words to preserve natural speech flow and emotional nuance. This may involve, for example, the sound of a voice speaking words. The phrase “a manner reflecting a voice characteristic” refers to a way of producing sound that mirrors or represents a particular feature or quality of a voice. This may involve, for example, speaking in a high pitch or a low volume. For instance, a sequence of neuromuscular actions exhibiting gradually rising amplitude and increased tempo might reflect an expressive crescendo, signaling excitement or anticipation in the intended speech. Alternatively, consistently low -frequency and low-amplitude micromovements can be indicative of a subdued, monotone delivery, as might occur when conveying calmness, seriousness, or fatigue.

[0169] In some disclosed embodiments, each of the first voice characteristic and the second voice characteristic may include at least one of intonation, pitch, pace, intensity, prosody, articulation, or volume, and wherein the first voice characteristic and the second voice characteristic differ from each other by at least one of the intonation, pitch, pace, intensity, prosody, articulation, or volume . The term “intonation” refers to the pattern of pitch changes in spoken language, which conveys emphasis, questions, emotions, or intent. The term “pitch” refers to the perceived highness or lowness of a voice, determined by the frequency of vocal fold vibrations. For example, pitch may involve raising one’s voice at the end of a question or lowering it to signal seriousness. The term “pace” refers to the speed at which words are spoken, affecting how quickly or slowly speech is delivered. For example, pace may involve speaking rapidly in excitement or more slowly for emphasis or clarity. The term “intensity” refers to the strength or energy of a voice, contributing to how forceful or soft speech sounds. For example, intensity may involve increasing vocal power to express urgency or lowering it for subtlety. The term “prosody” refers to the rhythm, tress, and melody of speech, shaping its expressiveness and natural flow. For example, prosody may involve varying stress on syllables to highlight meaning or using a sing-song pattern when reading aloud. The term “articulation” refers to the clarity and precision with which speech sounds are formed and pronounced. For example,Attorney Docket No. 16198.0057-00304 articulation may involve crisply pronouncing each consonant for clear communication or blending sounds for a more relaxed tone. The term “volume” refers to the loudness or softness of spoken words, which may be determined by the amplitude of sound waves. For example, volume may involve speaking softly in a quiet setting or projecting one’s voice to address a crowd.

[0170] As an example, some disclosed embodiments involve an individual, equipped with a device configured to detect neuromuscular activity in non-lip facial regions, silently articulating the word “hello” without producing audible sound. The system may identify “hello” as the subvocalized word through the capture of moderate amplitude micromovements in the cheek and temple areas. The frequency of these micromovements may conform to typical speech tempo, with a brief temporal separation observed between syllables. The system may further determine the voice characteristic associated with this articulation, assigning a slightly elevated pitch indicative of a friendly or inquisitive effect, a normal pace, and a moderate intensity. The system may subsequently synthesize an audible output of “hello,” wherein the generated speech exhibits a bright and welcoming tone consistent with the underlying neuromuscular features.

[0171] As another example, some disclosed embodiments involve a user intending to issue the command “Help!” utilizing subvocalization to express urgency. The system or device may detect this word by registering high amplitude micromovements at the jawline, a heightened frequency of muscle contractions, and a rapid, forceful timing sequence. These parameters may result in the determination of voice characteristics, specifically a high intensity and pitch, and a swift pace. The audible output generated by the system may involve a sharply pronounced and urgent “Help!” reflecting the user’s intended emotional emphasis through signal -derived prosodic features.

[0172] As yet another example, some disclosed embodiments involve a user subvocalizing the filler “um” while mentally rehearsing speech. The system may identify “um” as the relevant linguistic unit, with low amplitude and slow micromovements detected. The frequency and timing parameters may suggest a drawn-out pace and low intensity, and a downward inflection of pitch typical of hesitation. The synthesized speech output may be a soft and extended “um,” accurately reflecting the user’s contemplative or uncertain state as inferred from physiological signal data.

[0173] As another example, some disclosed embodiments involve an individual subvocalizing the phrase “Yes, absolutely!” with the intention of conveying firmness in the word “Yes” and enthusiasm in the word “absolutely,” and the device or system may process each word independently. For “Yes,” moderate amplitude, steady timing, and a lower pitch may be detected, resulting in the assignment of a firm and deliberate voice characteristic. For “absolutely,” the captured features may include higher amplitude, increased frequency, and a rising pitch contour, which the system may interpret as expressive and enthusiastic. The synthesized output may be a phrase rendered audibly with “Yes” spoken in a steady, firm tone, followed by an emphatic and eneigetic “absolutely! ,” thereby preserving the intended expressive transition between the subvocalized words.Attorney Docket No. 16198.0057-00304

[0174] As another example, some disclosed embodiments involve a user silently articulating the phrase “oh no” upon reading a message, intending to convey worry. The system may detect “oh no” as the subvocalized output and may register low amplitude for “oh,” increasing for “no.” Pitch analysis may further reveal a downward trend for “oh” and a sharp rise for “no,” with timing parameters indicating a slow initial pace and a rapid progression for the second word. The system may determine that “oh” should be synthesized as subdued, while “no” is rendered with an alarmed quality, producing an audible “oh no” that transitions from a worried to a heightened emotional state, reflecting the underlying neuromuscular signals.

[0175] For example, some disclosed embodiments include a speech synthesizing module configured to generate audible output based on the detected subvocalized words and their associated voice characteristics. For example, the system may use a combination of a speech deciphering module, output determination module, and illumination control module to convert processed signals into expressive speech. The synthesized output may be personalized to the user’s voice profile, emotional state, or contextual intent, and may be used for communication, accessibility, or interaction with digital systems. In some embodiments, the system may store the synthesized output or transmit it to external devices or networks for further use.

[0176] As an example, some disclosed embodiments may include a wearable sensor system configured to detect neuromuscular activity in non-lip regions of the head. For example, the system may include a light source and detector aimed at the cheek or temple to measure micromovements caused by muscle contractions. These micromovements may be indicative of silent speech or subvocalization. In other embodiments, sensors may be positioned near the ear canal to detect deformations associated with cranial nerve activity. The sensor may be optical, mechanical, electrical, or acoustic, and may be embedded in a wearable housing such as an earbud, glasses frame, head- mounted device, or other wearable embedded device.

[0177] By way of non-limiting example, reference is made to Figure 10A, which is an exemplary schematic diagram of a system for detecting and processing subvocalized speech and voice characteristics, consistent with some disclosed embodiments. The system includes a head of an individual 1002a with one or more non-lip regions 1008a. A wearable device 1006a is positioned on the head ofthe individual 1002a and includes sensor(s) 1010a and electrode(s) 1012a configured to detect neuromuscular activity associated with various non-lip regions 1008a. Sensor(s) 1010a and electrode(s) 1012a are connected to a processor 1022a, which is designed to receive and process subvocalization signals and voice characteristics indicative of neuromuscular activity data 1004a associated with neuromuscular activity even in the absence of perceptible vocalization. Processor 1022a is further connected to an output 1024a, which is configured to synthesize output based on the subvocalized words and voice characteristics determined by processor 1022a. As discussed elsewhere herein, the synthesized output may reflect different voice characteristics associated with differentAttorney Docket No. 16198.0057-00304 subvocalized words, as determined by processor 1022a from the signals received from sensor(s) 1010a and electrode(s) 1012a. The system may be used in various applications, such as silent communication in noise-sensitive environments, assistive technology for individuals with speech impairments, or covert communication in security settings. For example, a user may subvocalize commands to control a smart home system without disturbing others, or a person with vocal cord paralysis may use the system to communicate with synthesized speech that reflects their intended tone and emotion.

[0178] By way of non-limiting example, reference is made to Figure 10B, which is an exemplary schematic diagram of a system for detecting and processing subvocalized speech and voice characteristics, consistent with some disclosed embodiments. The system includes a non-lip region 1008b of an individual's head, from which neuromuscular activity data 1004b is obtained. A sensor 1016 of a non-wearable device 1026 (e.g., a mobile device) is positioned to detect reflections 1014 from the non-lip region 1008b, which are indicative of neuromuscular activity associated with subvocalized speech or voice characteristics. Sensor 1016 is connected to a processor 1022b, which is configured to receive and process the neuromuscular activity data 1004b to determine subvocalized words and their associated voice characteristics. An output 1024b is connected to the processor 1022b and is configured to synthesize audible speech (or other speech) based on the determined subvocalized words and voice characteristics. As discussed elsewhere herein, the synthesized output may reflect different voice characteristics for different subvocalized words. The system may be integrated with various non-wearable devices, such as, e.g., smartphones, speakers, displays, tablets, laptops, computers, televisions, monitors, cameras, whiteboards, printers, and other similar devices which are not worn directly by a user. For instance, a user may silently dictate messages on their smartphone in public spaces, or a person may use the system to enhance their smart speaker by providing visual transcriptions of their own subvocalized speech.

[0179] By way of non-limiting example, reference is made to Figure 10C, which is an exemplary schematic diagram of a system for detecting and processing subvocalized speech and voice characteristics, consistent with some disclosed embodiments. The system includes a head of an individual 1002c having one or more non-lip regions 1008c. A wearable device 1006c (e.g., smart glasses) is positioned on the head of the individual 1002c and includes sensors 1010c configured to detect neuromuscular activity associated with the non-lip region 1008c. Sensors 1010c are connected to a processor 1022c, which is configured to receive and process neuromuscular activity data 1004c captured by sensors 1010c. Processor 1022c is further connected to an output 1024c, which is configured to synthesize output based on the subvocalized words and other characteristics determined by processor 1022c. As discussed elsewhere herein, the synthesized output may reflect different voice characteristics associated with different subvocalized words. The system may be used in various scenarios, such as language learning applications where users can practice pronunciation silently, or inAttorney Docket No. 16198.0057-00304 virtual reality environments where users can communicate with avatars using subvocalized speech that reflects their intended emotional state.

[0180] Some disclosed embodiments involve analyzing a first set of subvocalization signals to determine a sequence of non -lip neuromuscular actions indicative of the first subvocalized word, and determining the first voice characteristic based on a value of at least one parameter associated with the sequence of non-lip neuromuscular activities. The term “analyzing” refers to considering, parsing, or understanding. Subvocalization signals can be analyzed to identify a temporal or spatial pattern of muscle engagement, nerve activation, or tissue deformation that corresponds to the silent articulation of a specific word. The temporal or spatial partem may include multiple discrete neuromuscular events, such as contractions, relaxations, or signal spikes, that occur in a defined order and reflect the internal generation of speech. For example, analyzing a first set of subvocalization signals may involve detecting the sequence and timing of muscle contractions in the cheek and temple regions, measuring their amplitude and frequency, and mapping these features onto recognized speech patterns to accurately infer the intended word or phrase.

[0181] The phrase “sequence of non-lip neuromuscular actions” refers to any ordered set of physiological events, changes, or movements occurring in the skin, muscles or nerves. These events may be obtained from regions of the head excluding the lips, and may include micromovements of the cheek, temple, jawline, or ear canal, electrical signals transmitted through cranial nerves, or mechanical deformations caused by muscle engagement. The sequence may, for example, collectively represent the articulation of a word, a phrase, a syllable, an emotional exclamation, a hesitation marker such as “uh” or “um,” or even a paralinguistic cue like a sigh, gasp, or laugh. The sequence may be identified using time-series analysis, pattern recognition, or machine learning models trained to associate signal patterns with specific linguistic units. The phrase “value of at least one parameter” may refer to any measurable feature extracted from the signal sequence. The value can be used to infer expressive qualities of speech. Parameters may include amplitude, frequency, timing, duration, rate of change, or spatial distribution of the neuromuscular activity. For example, a high amplitude and rapid sequence may indicate a forceful or urgent tone, while a low amplitude and slow sequence may suggest a calm or hesitant expression. The system may use one or more of these parameters to determine the first voice characteristic, such as pitch, pace, or intensity, and apply it to the synthesized output of the first subvocalized word.

[0182] In some disclosed embodiments, the parameters may include at least one of an amplitude associated with facial micromovements, a frequency of facial micromovements, timing of facial micromovements, an amplitude of muscle contraction, a frequency of muscle contraction, timing of muscle contraction, an amplitude of electrical signals transmitted through facial nerves, a frequency of electrical signals transmitted through facial nerves, or a timing of electrical signals transmitted through cranial nerves. The term “amplitude” refers to the magnitude or strength of a signal, such asAttorney Docket No. 16198.0057-00304 the size of a muscle movement or the power of an electrical impulse; for example, greater amplitude might indicate a more forceful contraction or a louder vocal expression. The term “frequency” refers to how often a particular event or signal occurs within a specific period of time, such as the number of micromovement cycles per second; for example, a higher frequency may correspond to a faster speech tempo or more rapid muscle activation. The term “timing” refers to the moment or duration at which an event or signal occurs in relation to other events, such as when a muscle activates during the articulation of a word; for example, timing may be used to coordinate the sequence of neuromuscular activities necessary for producing expressive speech.

[0183] For example, the system may analyze the amplitude associated with facial micromovements, which may indicate the strength or intensity of muscle engagement during subvocalization. The frequency of facial micromovements may reflect the rate or rhythm of articulation, while the timing of facial micromovements may provide insight into pacing or prosody. In addition to micromovement data, the system may evaluate amplitude, frequency, and timing of muscle contractions, which may be detected using, e.g., mechanical deformation sensors or electromyographic techniques. These parameters may help distinguish between different expressive tones, such as uigency, hesitation, or emphasis. For example, a rapid sequence of high -amplitude contractions may suggest a forceful or excited vocal delivery.

[0184] The system may also analyze electrical signals transmitted through facial nerves, including parameters such as an amplitude of electrical signals transmitted through facial nerves, a frequency of electrical signals transmitted through facial nerves, or a timing of electrical signals transmitted through cranial nerves, to infer voice characteristics. These signals may be captured using electrodes positioned near the temple,jawline, or ear canal, and may reflect neural activation patterns associated with speech articulation. The timing of electrical signals transmitted through cranial nerves may be particularly useful for identifying prosodic features such as stress patterns or intonation contours.

[0185] Some disclosed embodiments may involve extracting and analyzing parameters in realtime. The term “real-time” refers to the immediate or near-instantaneous processing and analysis of data as it is acquired, allowing for response and adaption without perceptible delay. For example, realtime may involve capturing neuromuscular signals from wearable sensors, analyzing the data streams on-the-fly, and generating expressive speech output or transcriptions within milliseconds of the user’s silent articulation.

[0186] Some disclosed embodiments involve determining that a first voice characteristic has a first manifestation when the value of at least one parameter associated with a sequence of non -lip neuromuscular actions is greater than a threshold, and a second manifestation when the value of that parameter is less than the threshold. In some embodiments, this threshold -based modulation allows the system to dynamically adjust the expressive quality of synthesized speech based on the intensity, timing, or frequency of the detected physiological signals.Attorney Docket No. 16198.0057-00304

[0187] The term “manifestation” refers to the observable expression or output of a detected characteristic, such as how a particular voice quality is rendered in synthesized speech or text. For example, manifestation may involve a pitch, pace, or intensity of the spoken output, or the visual presentation of transcribed words — such as emphasizing text with bold, italics, or different font sizes — to reflect nuances like urgency, emotion, or emphasis present in the user’s subvocalized input. The term “first manifestation” refers to one manifestation and a second manifestation refers to another manifestation, different in at least one of quality or time from the first manifestation. A manifestation may be a specific expressive rendering of a voice characteristic, such as a high pitch, fast pace, or strong intensity. In one example, a “second manifestation” may represent a different rendering, such as a low pitch, slow pace, or soft intensity. Predefined or adaptive thresholds may be used to determine which manifestation to apply. For example, if the amplitude of facial micromovements exceeds a certain value, the system may synthesize the word with a louder volume or a more emphatic tone — e.g., a particular manifestation of the voice characteristic. If the amplitude is below the threshold, the manifestation might be a softer volume or a more neutral tone in the output. Different manifestations may involve adjusting loudness, emphasis, or emotion in the synthesized speech based on a measured parameter.

[0188] The term “value” refers to any measurable or quantifiable feature, attribute, metric, or characteristic such as those derived from detected neuromuscular activity or related signals. For example, a value may involve the amplitude of a muscle contraction, the frequency of micromovements in the cheek or temple, or the precise timing of neuromuscular events during subvocalized articulation. The term “threshold” refers to a predefined value or criterion that serves as a dividing line for interpreting signal parameters; when the value of a measured parameter exceeds or falls below this threshold, it may trigger a processor to cause a selection or modification of a corresponding voice characteristic in the synthesized speech output. The threshold may, for example, be static, user-defined, context-sensitive, or learned overtime through machine learning. Thresholds may vary based on user profiles, emotional states of a user, environmental conditions, or application - specific requirements.

[0189] Some alternative embodiments involve the generation of a textual output that mirrors first voice characteristic associated with the first word and second voice characteristic associated with the second word. For instance, this textual output may be a transcription of the verbal expressions articulated by the individual. The presentation of the word, such as font size, bold or non -bold formatting, underlined or non -underlined, italicized or non -italicized, may be determined to add emphasis, contingent on the identified voice characteristics. Furthermore, the textual output may incorporate punctuation marks, such as a period, an exclamation point, or a question mark. The inclusion and type of these punctuation marks may be determined based on the identified voice characteristics. For example, an exclamation point may be used to indicate a raised voice orAttorney Docket No. 16198.0057-00304 heightened emotion, while a question mark may be used when the vocal inflection rises at the end of a sentence, indicating a question.

[0190] Some disclosed embodiments involve determining an emotional condition based on the neuromuscular activity, and wherein determining the first voice characteristic is further based on the determined emotional condition. The term “emotional condition” refers to any psychological or affective state that can be inferred from physiological signals associated with silent or subvocalized speech, such as changes in muscle activity, micromovement frequency, timing patterns, or other subvocalization signals. The emotional condition can be, for example, excitement, fear, sadness, emphasis, intensity, curiosity, confusion, skepticism, uigency, enthusiasm, surprise, skepticism, hesitation, suspense, trailing thoughts, neutrality, finality, calmness, anger, disbelief, dramatic pause, interruption, side thoughts, sarcasm, clarification, irony, shouting, happiness, fear, excitement, embarrassment, disgust, or contempt. . The emotional condition may be a physiological or affective state such as an overall physical or emotional condition of an individual at a given moment in time. For example, a physiological or affective state may involve experiences such as stress, fatigue, excitement, happiness, or anxiety, each of which may influence the manner in which neuromuscular signals are produced during subvocalization.

[0191] An emotional condition may be inferred from signal patterns such as increased muscle tension, rapid micromovement frequency, elevated electrical activity, or changes in timing and rhythm. For example, a high-frequency, high-amplitude signal pattern may suggest excitement or urgency, while a slow, low-amplitude pattern may indicate calmness or fatigue. Some disclosed embodiments may analyze subvocalization signals — such as micromovements, muscle contractions, or cranial nerve activity — to infer the emotional state of the individual. This emotional condition may then be used to modulate expressive features of the synthesized speech, ensuring that the output reflects not only the linguistic content but also the speaker’s effective intent.

[0192] For example, some disclosed embodiments may utilize emotional context to guide the selection or modulation of expressive speech features. For instance, if it is determined that the user is expressing anger, the first subvocalized word may be synthesized with a sharper articulation and higher intensity. If the user is expressing sadness, a slower pace and lower pitch may be used instead. The emotional condition may serve as an input to a voice synthesizing engine, neural network, or rule-based algorithm that adjusts voice characteristics accordingly.

[0193] Some disclosed embodiments involve determining a change in the emotional condition and modifying the second voice characteristic based on the change in the emotional condition. For example, physiological signals such as facial micromovements, muscle contractions, or cranial nerve activity can be monitored overtime and transitions between emotional states can be detected. These transitions may be subtle or pronounced and may occur between any of the emotional conditionsAttorney Docket No. 16198.0057-00304 described in claim 10, including happiness, sadness, anger, fear, excitement, embarrassment, surprise, disgust, or contempt.

[0194] The phrase “determining a change in the emotional condition” refers to identifying a shift in the user’s affective state or physiological state. This may involve comparing current signal patterns to previously recorded baselines, detecting deviations in amplitude, frequency, or timing, or using machine learning models trained to recognize emotional transitions. For example, a shift from calm to excited may be indicated by an increase in signal intensity and frequency, while a shift from anger to sadness may involve a decrease in amplitude and a change in rhythm.

[0195] The phrase “modifying the second voice characteristic based on the determined change in the emotional condition” refers to dynamically adjusting the expressive features of the second subvocalized word to reflect the updated emotional context. For instance, if a processor detects a transition from neutral to excited, it may cause an increase in the pitch and pace of the second word. If the emotional state shifts from happy to disappointed, the system may lower the pitch and reduce the intensity. The modification may be applied in real-time and may affect one or more voice characteristics, including intonation, pitch, pace, intensity, prosody, articulation, or volume.

[0196] Some disclosed embodiments involve determining a physical condition based on the neuromuscular activity and wherein determining the first voice characteristic is further based on the determined physical condition. As discussed previously in connection with an emotional condition, a physical condition may be ascertained in a similar manner based on muscle contractions, micromovements, or cranial nerve activity. This physical condition may then be used to modulate expressive features of the synthesized speech, ensuring that the output reflects not only the linguistic content and emotional tone, but also the speaker’s physiological state.

[0197] The term “physical condition” refers to any bodily state or physiological status. Such conditions may be inferred from signal patterns associated with silent speech. Physical conditions may include pain, fatigue, anxiety, stress, respiratory abnormality, intoxication, pain, breathlessness, drug usage, age, gender, or any other health -related, condition-related, or performance-related states. These conditions may be inferred from changes in signal amplitude, timing irregularities, muscle tension levels, or patterns of cranial nerve activation. For example, elevated muscle tension and irregular timing may suggest stress or anxiety, while reduced signal strength and slower micromovement frequency may indicate fatigue.

[0198] As an example, a system may utilize physiological context to guide the selection or modulation of expressive speech features. For instance, if the system determines that the user is experiencing pain, it may synthesize the first subvocalized word with a strained or subdued tone. If the user is fatigued, the system may apply a slower pace and lower intensity. The physical condition may serve as an input to a voice synthesizing engine, neural network, or rule -based algorithm that adjusts voice characteristics accordingly.Attorney Docket No. 16198.0057-00304

[0199] In some disclosed embodiments the determined physical condition includes conditions such as pain, fatigue, anxiety, stress, or respiratory abnormality. Physical conditions may be inferred from signal patterns captured during subvocalization, such as changes in muscle tension, micromovement frequency, cranial nerve activity, or timing irregularities. Each physical condition may be associated with distinct physiological signatures that the system can detect and classify using signal processing techniques, biometric profiling, or machine learning models.

[0200] For example, pain may be indicated by elevated muscle tension, abrupt signal spikes, or irregular timing in neuromuscular activity. Fatigue may manifest as reduced signal amplitude, slower micromovement frequency, or delayed response timing. Anxiety may be reflected in rapid, high- frequency micromovements and increased variability in cranial nerve signals. Stress may produce sustained muscle contractions and elevated electrical activity, while respiratory abnormality may be inferred from timing disruptions or irregular patterns in signals associated with speech -related breathing coordination.

[0201] Physical conditions may be used to modulate voice characteristics such as pitch, pace, intensity, and articulation. For instance, if the detected physical condition is fatigue, the speech may be synthesized with a slower pace and softer tone. If the condition is stress, the output may be rendered with sharper articulation and increased intensity. The physical condition may be determined in real-time and used to personalize the synthesized speech output, enhancing its realism and responsiveness to the user’s physiological state.

[0202] Some disclosed embodiments include at least one processor configured to classify physical states based on sensor data analyzed to decipher speech and one or more of the conditions discussed herein. For example, neuromuscular signals may be mapped to physical condition labels, which are then used to guide the selection of voice characteristics for each subvocalized word. The mapping may occur through Al or based on historical data.

[0203] Some disclosed embodiments involve determining a change in the physical condition and modifying the second voice characteristic based on the change in the physical condition. The physiological signals may be monitored over time — such as facial micromovements, muscle contractions, or cranial nerve activity — and transitions between physical states can thereby be detected. These transitions may occur between any of the conditions described herein, including pain, fatigue, anxiety, stress, or respiratory abnormality, and may reflect shifts in the user’s physiological status during subvocalization.

[0204] The phrase “determining a change in the physical condition” refers to identifying a variation in the user’s bodily state. This may involve comparing current signal patterns to previously recorded baselines, detecting deviations in amplitude, timing, or frequency, or using machine learning models trained to recognize physiological transitions. For example, a shift from low to high muscleAttorney Docket No. 16198.0057-00304 tension may indicate increasing stress, while a drop in signal intensity and frequency may suggest the onset of fatigue.

[0205] The phrase “modifying the second voice characteristic based on the change in the physical condition” refers to dynamically adjusting the expressive features of the second subvocalized word to reflect the updated physiological context. For instance, if a transition from a neutral state to pain is detected, the second word may be synthesized with a strained or subdued tone. If the condition shifts from fatigue to alertness, the pace and intensity of the output may be increased. The modification may affect one or more voice characteristics, including pitch, pace, intensity, prosody, articulation, or volume, and may be applied in real-time or near-real-time.

[0206] Some disclosed embodiments involve receiving auxiliary data, using the auxiliary data to determine a context for the plurality of subvocalized words, and determining the first voice characteristic is further based on the determined context. The term “auxiliary data” refers to any supplemental information that is not part of the neuromuscular signal stream but is relevant to interpreting or expressing subvocalized speech. For example external or environmental information may be sensed or determined alongside neuromuscular signals to enhance the expressiveness and relevance of synthesized speech. This auxiliary data may provide situational awareness, user-specific preferences, or interaction cues that influence how subvocalized words are rendered audibly. Auxiliary data may include interaction data (e.g., user-device interactions, gestures, gaze tracking), location data (e.g., GPS coordinates, indoor positioning), environmental data (e.g., ambient noise levels, lighting conditions, temperature), biometric data (e.g., heart rate, skin conductivity), or contextual metadata (e.g., time of day, calendar events, social setting). This data may be collected from sensors, user input, networked systems, or external databases.

[0207] The phrase “context for the plurality of subvocalized words” refers to any inferred situational framework that helps interpret the meaning, tone, or delivery of silent speech. Context may include the physical environment, social setting, emotional state, task being performed, or identity of the person being addressed. For example, if the auxiliary data indicates that the user is in a quiet library, the system may synthesize speech with a softer volume and slower pace. If the context suggests a high-stress emergency, the system may apply a louder, faster, and more urgent tone.

[0208] “Determining the first voice characteristic is further based on the determined context” refers to the use of contextual information to guide the selection or modulation of expressive speech features. Rule-based logic, adaptive models, or machine learning algorithms may be employed to map contextual cues to voice characteristics such as pitch, pace, intensity, prosody, articulation, or volume. This enables speech to be produced and output in a manner that is not only linguistically accurate but also situationally appropriate.

[0209] Some disclosed embodiments include a processing unit configured to receive and interpret auxiliary data such as from situational sensors or data in the cloud. Contextual inputs provided via theAttorney Docket No. 16198.0057-00304 auxiliary data, combined with neuromuscular signals can be used to determine expressive parameters for each subvocalized word or phrase, or expressed thought. This approach supports context -aware voice synthesizing that can adapt to the user’s environment, activity, and intent.

[0210] In some embodiments, the auxiliary data may include at least one of interaction data, location data, or environmental data, and the determined context may include at least one of a location, an environmental condition, or an identity of a person with whom the individual is interacting. This context may then be used to inform the determination of voice characteristics for synthesized output of subvocalized speech.

[0211] The term “interaction data” refers to any information related to the user’s engagement with devices, systems, or other individuals. This may include touch inputs, gesture recognition, eye tracking, device usage patterns, conversational history, or biometric feedback. Interaction data may be collected from wearable sensors, mobile devices, cameras, microphones, or other input systems. The term “location data” refers to any information that identifies the physical or virtual position of the user. This may include GPS coordinates, indoor positioning signals, proximity to known landmarks or devices, or virtual environment identifiers in augmented or virtual reality settings. Location data may be used to infer situational context, such as whether the user is in a public space, private setting, or professional environment. The term “environmental data” refers to any information about the surrounding conditions in which the user is operating. This may include ambient noise levels, lighting conditions, temperature, humidity, air quality, or presence of other individuals. Environmental data may be collected using onboard sensors, external monitoring systems, or networked devices.

[0212] The term “determined context” refers to any inferred situational framework derived from auxiliary data. This may include a specific location (e.g., office, vehicle, outdoors), an environmental condition (e.g., noisy, dimly lit, crowded), or the identity of a person with whom the individual is interacting (e.g., colleague, family member, virtual assistant). The system may use this context to tailor voice characteristics such as volume, pace, pitch, or articulation to suit the situation. For example, if the user is interacting with a known contact in a quiet room, the system may synthesize speech with a softer tone and slower pace. If the user is outdoors in a noisy environment, the system may increase volume and articulation clarity.

[0213] Some disclosed embodiments involve receiving input from an individual and personalizing a manner in which the first subvocalized word is audibly outputted based on the received input. Such received input may include preferences, adjustments, or expressive styles that influence how an individual’s silent speech is rendered into audible output. This personalization may occur during initial setup, through real-time interaction, or via adaptive learning based on historical usage.

[0214] The phrase “input from the individual” refers to any form of user-provided data, command, or feedback that guides the behavior of the system. This input may include explicit settings (e.g., preferred pitch range, speaking rate, emotional tone), real-time corrections (e.g., adjusting volume orAttorney Docket No. 16198.0057-00304 articulation), biometric feedback (e.g., facial expressions, gestures), or contextual cues (e.g., task type, audience). Input may be provided through touch interfaces, voice commands, gesture recognition, eye tracking, or other interaction modalities. The term “personalizing” refers to modifying one or more voice characteristics — such as pitch, pace, intensity, prosody, articulation, or volume — based on the received input. For example, a user may configure the system to always speak with a calm tone, or to emphasize certain words based on context. The personalization may also include stylistic preferences such as accent, gendered voice, emotional expressiveness, or speech clarity. The system may store these preferences and apply them consistently across synthesized outputs, or may adapt them dynamically based on ongoing input. For example, a user with an accent may provide input guiding the system to soften or remove the accent. Or a user may instmct the system to translate words spoken in mixed language into a selected language.

[0215] Some disclosed embodiments involve storing the first voice characteristic associated with the first subvocalized word and storing the second voice characteristic associated with the second subvocalized word as reference data for use in future synthesization. The system may capture and retain expressive speech parameters linked to specific subvocalized words, enabling consistent, personalized, or context-aware voice synthesizing in subsequent interactions.

[0216] The term “storing” refers to the process of saving expressive attributes — such as pitch, pace, intensity, prosody, articulation, or volume — alongside their corresponding silent speech inputs. These associations may be stored in a local memory, cloud -based repository, or distributed database, and may be indexed by word identity, user profile, timestamp, or contextual metadata. Some disclosed embodiments involve accessing reference data and using this data to personalize the manner in which the first subvocalized word is audibly outputted. The term “reference data” refers to a curated set of stored voice characteristics — including parameters such as pitch, pace, intensity, prosody, articulation, or volume — that are systematically paired with their respective subvocalized words. For example, if a user frequently subvocalizes the word “hello” with a cheerful tone and moderate speed, the system can save these expressive parameters as reference data, allowing future synthesized outputs of “hello” to automatically reflect the same vocal qualities for a natural and consistent user experience. The term “future synthesization” refers to the use of previously stored voice characteristics to inform or replicate the expressive rendering of subvocalized words in later sessions. For example, if a user consistently subvocalizes the word “yes” with a rising intonation and moderate pace, the system may retrieve and apply those characteristics automatically in future instances. This reference data may also support adaptive learning, allowing the system to refine its synthesizing models based on historical usage patterns, emotional trends, or user feedback.

[0217] Some disclosed embodiments include a memory device or data structure configured to store and retrieve voice characteristic mappings. For example, the system may use a memory device, data structure, and output determination to associate subvocalized words with expressive parameters andAttorney Docket No. 16198.0057-00304 apply them during speech synthesizing. This enables continuity, personalization, and efficiency in voice generation, particularly in applications involving frequent or repetitive silent speech inputs.

[0218] By way of non-limiting example, reference is made to Figure 11, which is an exemplary flowchart of a method 1100 for synthesizing voice from subvocalized speech, consistent with some disclosed embodiments. It will be appreciated that the steps of the disclosed method may be performed in any order, and may be carried out, e.g., by the processor(s), as described elsewhere herein. Method 1100 begins with step 1104, which involves operating a sensor for detecting neuromuscular activity associated with a non-lip region of a head of an individual. In step 1106, subvocalization signals indicative of the neuromuscular activity are received in an absence of perceptible vocalization. These signals are processed in step 1108 to determine a plurality of subvocalized words, including a first subvocalized word and a second subvocalized word. Steps 1110 and 1112 involve determining a first voice characteristic from a first set of the subvocalization signals associated with the first subvocalized word, and a second voice characteristic from a second set of subvocalization signals associated with the second subvocalized word, respectively. Step 1114 involves synthesizing audible output of the plurality of subvocalized words, where each word is audibly outputted in a manner reflecting its associated voice characteristic. As discussed elsewhere herein, this method allows for the production of speech from subvocal signals while preserving different voice characteristics for different words. The method may be applied in various contexts, such as in translation services where subvocalized speech is detected, translated, and then synthesized with appropriate voice characteristics in the target language, or in gaming applications where players can silently communicate with in-game characters using subvocalized speech that reflects their character's emotional state.

[0219] By way of non-limiting example, reference is made to Figure 12A, which is an exemplary block diagram of a processing system for synthesizing subvocalized speech, consistent with some disclosed embodiments. The system includes various input data received from various sources or sensors, including nerve mapping data 1202, skin / muscle deformation data 1204, optical sensor data 1206, and potentially other data 1208 associated with detected speech or subvocalizations. A processor 1218 is configured to process the input data and determine subvocalized word(s) 1210, as well as analyze the input data or additional input data to determine an emotional condition 1212, a physical condition 1214, and context 1216 associated with the subvocalized speech. The system further includes an auxiliary data input interface 1220 for receiving additional data, a user input interface 1222 for user interaction and for receiving user input (e.g., user preferences or custom information), and storage 1224 for storing data, instructions, or reference information. As discussed elsewhere herein, the processor 1218 integrates information from the various input sources and determined conditions to synthesize audible output that reflects not only the content of the subvocalized words but also the associated voice characteristics and contextual information. ThisAttorney Docket No. 16198.0057-00304 system may be used in various applications, such as in medical settings where patients can communicate silently while the system takes into account their physical condition, or in customer service environments where representatives may subvocalize responses that are then synthesized with appropriate emotional tones based on the detected emotional conditions, physical conditions, or context of the interaction.

[0220] By way of non-limiting example, reference is made to Figure 12B, which is an exemplary block diagram of a speech detection system 1225 consistent with some disclosed embodiments. The system includes a wearable and / or non-wearable device 1230 connected to a sensor array 1250, which may comprise various sensors configured to detect signals associated with speech, subvocalization, emotion, physical condition, voice characteristics, and other detectable parameters. System 1225 comprises a processor 1255 connected to various components, including an auxiliary data input interface 1240, a user input interface 1245, memory 1265, and an output 1260. As discussed elsewhere herein, processor 1255 receives input from sensor array 1250 and processes this data along with information from auxiliary data input interface 1240 and user input interface 1245 to detect and process speech or subvocalized speech, as well as voice characteristics thereof. This system may be applied in diverse scenarios, such as in automotive interfaces where drivers may silently control vehicle functions using subvocalized commands, or in accessibility technologies where individuals with speech disorders may communicate using subvocalized speech that is synthesized with natural - sounding voice characteristics.

[0221] Some disclosed embodiments involve a wearable system for speech detection. The term “wearable system,” as used herein, refers to a computing or sensing apparatus configured to be worn on the body of an individual. Examples of a wearable system include smart glasses, head -mounted display (HMD), smartwatch, bone -conduction headset, neckband sensor, ear-worn devices (e.g., earbuds with embedded sensors), and facial patches or stickers embedded with microelectromechanical systems (MEMS) sensors. In particular, the wearable system may be positioned such that one or more components are placed in proximity to the individual’s head for speech detection. The phrase “speech detection” broadly refers to identifying, recognizing, or otherwise determining the presence or characteristics of speech -related activity. Exemplary processes of speech detection may involve computational techniques such as, optical flow analysis for detecting subtle facial movements associated with speech articulation; computer vision algorithms such as convolutional neural networks (CNNs) for facial muscle tracking; electromyography (EMG) signal processing to detect neuromuscular activity in the jaw, throat, or facial muscles; machine learning classifiers (e.g., support vector machines or recurrent neural networks) trained on multimodal sensor data to distinguish speech -related patterns from noise or non-speech activity.

[0222] In some examples, speech detection processes may involve analysis of light -reflection signals, visual signals, physiological signals, and / or biomechanical signals. In some cases, speechAttorney Docket No. 16198.0057-00304 detection may include identifying subvocalized or silent speech by measuring neuromuscular activity that occurs without vocalization. For example, the wearable system may detect the articulation of words during subvocalization by monitoring skin micromovements in the facial region corresponding to muscle engagement associated with speech formation. In other cases, speech detection may include identifying vocalized speech by measuring audio signals or processing any other signals acquired by the at least one sensor. Consistent with the present disclosure, the wearable system may further include a speaker. A speaker is a device or component that outputs audio. The speaker may be configured so that its operation does not interfere with the operation of the at least one sensor used from speech detection. For example, a wearable system may limit the speaker from generating an audio output in frequencies that cause vibratory interference with operation of the at least one sensor.

[0223] Some disclosed embodiments involve a head-mountable housing. The term “head- mountable housing” refers to any structure or enclosure designed for connection to or location on a human head, such as in a manner configured to be worn by a user. The head-mountable housing may be configured to contain or support one or more electronic components or sensors. In one example, the head-mountable housing is configured for association with a pair of glasses. In another example, the head-mountable housing is associated with an earbud. In yet other examples, the head-mountable housing may be adapted to be worn around a neck or may include a skin patch. The head -mountable housing may have a cross-section that is button-shaped, P-shaped, square, rectangular, rounded rectangular, or any other regular or irregular shape capable of being worn by a user. Such a stmcture may permit the head-mountable housing to be worn on, in, or around a body part associated with a head of the user (e.g., on the ear, in the ear, around the neck). The head-mountable housing may be made of plastic, metal, composite, a combination of two or more of plastic, metal and composite, or other suitable material. Wearable housings 110 in Fig. 1 and Fig. 2A are non-limiting examples of a head -mountable housing, consistent with the present disclosure.

[0224] In some embodiments, the head -mountable housing is configured for mounting on an ear of the wearer. The phrase “configured for mounting” refers to a structural or mechanical arrangement of the housing that enables attachment to a specific anatomical region of the wearer’s body, in this case, a wearer’s ear. This configuration may include contours, clips, hooks, or other features that conform to the shape of the anatomical region of the wearer’s body. As disclosed herein, example mounting configurations on an ear of the wearer may include: (1) in-the-ear (ITE), inserted into the ear canal and held by the ear’s shape; (2) behind -the-ear (BTE), seated behind the ear with a tube to the ear canal; (3) over-the-ear (OTE), resting on the ear and supported by a headband (e.g., headphones). By way of a non-limiting example, in Fig. 13, wearable system 1300 may include head-mountable housing 1302.

[0225] Some disclosed embodiments involve at least one sensor. The term “sensor” broadly refers to any device, element, or system capable of measuring or detecting one or more properties andAttorney Docket No. 16198.0057-00304 generating an output based on the measured or detected properties. In the context of this disclosure, the at least one sensor is configured to measure signals indicative of physical or bioelectric activity associated with vocalization or subvocalization. The signals may be indicative of muscle activity, nerve activity, facial movements, vocal tract dynamics, or other physiological or biomechanical activity. By way of example, as described throughout the disclosure, the at least one sensor may include a light detector configured to detect light reflections from a facial region to determine (e.g., via speckle analysis) muscle recruitment. By way of another example, the at least one sensor may include a non-optical detector, such as a pressure sensor, configured to detect mechanical displacements of the skin. By way of another example, the at least one sensor may include a detector (e.g., a piezoelectric sensor) adapted to detect skin micromovements or vibrations corresponding to neuromuscular engagement. Specifically, the detector may be configured to detect ear canal distortion caused by muscle or nerve activity when the user intends to speak. By way of another example, the at least one sensor may include a displacement-sensing element configured to measure skin movement associated with facial muscle activity during subvocalization. Such facial muscle activity may include contraction of the orbicularis oris, buccinator, mentalis, or any other muscles. By way of another example, the at least one sensor may include an electrode configured to detect electrical signals from muscles innervated by facial nerves, or to measure the propagation speed and amplitude of impulses associated with facial neural activity. Such facial neural activity may include cranial nerve VII activation, cranial nerve XII activity associated with tongue movement during speech preparation, cranial nerve V activity related to jaw positioning, and reflexive or anticipatory neural firing in response to intended vocalization or articulation.

[0226] Some disclosed embodiments involve at least one sensor integrated with the head- mountable housing and configured for direction towards a wearer’s head. The term “integrated with” refers to being physically or wirelessly connected or linked to another part. The at least one sensor may be “integrated” with the head-mountable housing in that it may be incorporated within the head- mountable housing, may extend from the housing, or may be pairable with electronics in the housing. Examples of integration include: a sensor embedded within the housing shell; a sensor mounted on an arm extending from the housing; a flexible sensor laminated onto the inner surface of the housing; or a wireless sensor that communicates with a processing unit embedded in the housing. The term “configured for” refers to an arrangement or design of a component that enables it to perform a specified function or achieve a particular result. This arrangement or design may be achieved through physical stmcture, software programming, electrical circuitry, mechanical design, or any combination thereof. The term does not require that the function is actively being performed at all times, but rather that the component can perform the function when appropriately activated or used. The phrase “configured for direction towards a wearer’s head” broadly refers to an arrangement in which the sensor is positioned, oriented, or otherwise adapted to detect signals, properties, or activity originatingAttorney Docket No. 16198.0057-00304 from the head or a region thereof when the individual wears the system. This may include, for example, alignment for optical, acoustic, mechanical, electrical, or other sensing modalities. In one example, when a sensor is integrated with earphones, the sensor may be oriented towards the ear canal to detect vibrations, sound, or physiological signals. In another example, when a sensor is integrated with glasses, the sensor may be oriented towards the wearer’s face to detect muscle movements, skin deformations, or other facial activity.

[0227] By way of a non-limiting example, in Fig. 13, wearable system 1300 may include at least one sensor 1304 integrated with head -mountable housing 1302 and configured for direction towards a wearer’s head 1306. Arrow 1308 depicts the direction of an area being monitored by at least one sensor 1304, i.e., the region of the wearer’s head from which the sensor is configured to detect signals or activity. In the illustrated example, when wearable system 1300 is worn, at least one sensor 1304 may be directed to a non-lip facial region associated with specific muscles, such as the zygomaticus muscle or the risorius muscle.

[0228] In some disclosed embodiments, the at least one sensor is configured to detect neuromuscular activity. “Detecting” refers to the process of discovering, identifying, or determining the existence, presence, or characteristics of one or more signals, events, conditions, or phenomena using one or more sensing modalities. This may involve acquiring data from a sensor and analyzing it to recognize patterns, thresholds, or changes that are indicative of a target state or activity. The detection may be performed in real time or retrospectively. The term “neuromuscular activity” broadly refers to any electrical, mechanical, or physiological manifestation that occurs via nerves or when nerves stimulate muscles, whether voluntary or involuntary. By way of example, the neuromuscular activity may include or involve minute contractions, tension changes, micro - movements in muscle fibers, and / or electrical impulses transmitted along associated nerves. Specifically, neuromuscular activity may involve subtle movements of non-lip facial skin or nerve signals associated with speech articulation, even in the absence of audible sound. In the context of this disclosure, the at least one sensor is configured to measure neuromuscular activity or associated signals. The detection of neuromuscular activity may involve a processing device for identifying patterns or events linked to muscle activation, such as changes in the position or tension of facial skin. In some cases, the processing device may detect facial skin movements by analyzing light reflections. In other cases, detection may include sensing electrical potentials from cranial or facial nerves, measuring skin surface vibrations with a piezoelectric element, monitoring pressure variations from muscle contractions, measuring deformations of the ear canal resulting from subvocalization, capturing changes in tissue impedance resulting from nerve -induced activation, and any other process associated with the corresponding type of sensor.

[0229] By way of a non-limiting example, in Fig. 13, at least one sensor 1304 is configured to detect neuromuscular activity 1310. In the illustrated example, neuromuscular activity 1310 includesAttorney Docket No. 16198.0057-00304 micromovements of facial skin in a non -lip facial region. Neuromuscular activity 1310 may be associated with a silent question 1312 from a user of wearable system 1300, such as “where is my next meeting?” posed to a virtual assistant associated with wearable system 1300.

[0230] Some disclosed embodiments involve a speaker integrated with the head -mountable housing for outputting audio. As discussed previously, a speaker is a device that outputs audio. It may include, for example, all manners of electronic devices that convert electrical or analog signals into sound waves. For example, a speaker may include a driver or transducer, an enclosure, and an amplifier. The speaker may receive electrical signals and convert them into sound waves, which may then be emitted in a manner enabling hearing (e.g., by projecting sound). For example, the speaker may be incorporated into a loudspeaker, earbuds, audio headphones, a wearable hearing aid device, a bone conduction headphone, or any other device capable of converting an electrical audio signal into a corresponding sound. The speaker may be “integrated” with the head -mountable housing in a manner similar to that described with respect to integration of the at least one sensor with the head- mountable housing. For example, the speaker may be incorporated within the head -mountable housing, or attached to or mounted on an appropriate portion of the head -mountable housing. In one implementation, the speaker may be housed within the internal stmcture of the head -mountable housing or connected to the head -mountable housing via an appropriate stmcture. In this implementation, the speaker may be integrated with the head -mountable housing by being connected to one or more components included in the housing via a wired connection. Alternatively, the speaker may be physically separate from the head-mountable housing and may be integrated with the head- mountable housing by being connected to one or more components included in the housing via a wireless connection. For example, the wearable system may transmit audio data wirelessly (e.g., via Bluetooth® or other short-range communication) to a remote audio device such as a standalone earbud, which is not physically attached to or embedded in the head -mountable housing but still functions to output audio. As used herein, “outputting audio” refers to the process of generating, transmitting, emitting, or otherwise providing sound signals through one or more audio output components, such as speakers, transducers, or audio ports. This may include, for example, playing another party’s voice during a two-way communication session, outputting pre-recorded sounds, generating notification tones, or any other audible content intended for perception by a user of the wearable system. Outputting audio may occur in real time or be triggered by specific events, conditions, or user interactions.

[0231] By way of a non-limiting example, in Fig. 13, wearable system 1300 may include speaker1314 integrated with the head -mountable housing 1302 for generating audio output.

[0232] Some disclosed embodiments involve at least one filter for limiting the speaker from generating an audio output in at least one frequency range . “Generating” refers to causing any type of electronic device to initiate an action for creating, producing, originating, or making something. AnAttorney Docket No. 16198.0057-00304“audio output” refers to sound signals produced by a device, such as a speaker or transducer, intended to be heard by a user. The term “frequency range” refers to a span or interval of frequencies of a waveform signal. The frequency range may be defined by a lower bound and an upper bound (e.g., 300 Hz to 400 Hz), and may pertain to acoustic or other types of signals. The phrase “generating audio output in at least one frequency range” refers to a process by which an audio output component (e.g., the speaker) produces sound signals that fall within a defined segment of the frequency spectmm. The audio output generated by the wearable system may be continuous or intermittent, and may include speech, music, alerts, or synthetic signals.

[0233] By way of a non-limiting example, Fig. 14 illustrates an audio output associated with a frequency range between 0-4.5 kHz. Specifically, Fig. 14 illustrates example representations of an audio output 1400 that may be generated by the wearable system, consistent with some disclosed embodiments. The top graph 1402 shows a waveform of audio output 1400 overtime, where the vertical axis represents amplitude and the horizontal axis represents time. As shown, audio output 1400 includes multiple clusters of higher amplitude oscillations separated by intervals of lower amplitude that correspond to spoken portions and pauses, respectively. The bottom graph 1404 shows a spectrogram of audio output 1400, where the vertical axis represents frequencies and the horizontal axis represents time. The brightness intensity represents signal eneigy at each frequency, with brighter regions indicating higher energy.

[0234] Consistent with the present disclosure, the system may be configured to selectively generate audio in one or more frequency ranges depending on operational context or user interaction. For example, the wearable system may selectively generate output audio in specific frequency ranges to avoid overlapping with sensor operation frequencies using at least one filter for limiting the speaker from generating an audio output in at least one frequency range . The term “filter” broadly refers to any device, circuit, algorithm, or process configured to selectively allow, attenuate, block, or modify certain frequency components of a signal while leaving other frequency components substantially unaffected. A filter may operate in the analog domain, the digital domain, or a combination thereof, and may be implemented using electrical, mechanical, optical, or computational means. Examples of filters for audio signals include low-pass filters that attenuate high-frequency components, high-pass filters that block low-frequency noise, band-pass filters that isolate a specific frequency band, and notch filters that suppress narrow frequency ranges. In this context, “limiting” means reducing, suppressing, or constraining the generation or transmission of certain signal components to prevent undesired effects. Accordingly, the phrase “limiting the speaker from generating an audio output in at least one frequency range” refers to reducing the amplitude, power, or presence of one or more frequency components in the signal driving the speaker to a level below a threshold, thereby reducing or eliminating the generation of audio output in the at least one frequency range. In some disclosedAttorney Docket No. 16198.0057-00304 embodiments, the threshold may be predetermined and the at least one frequency range includes one or more frequencies between 1 and 300 Hz or a plurality of frequencies under 200 Hz.

[0235] By way of a non-limiting example, Fig. 15A illustrates audio output when at least one filter (e.g., filter 1316) limits the speaker (e.g., speaker 1314) from generating audio output in a first frequency range (e.g., between approximately 2.6 kHz and 2.8 kHz) and in a second frequency range (e.g., between approximately 0.8 kHz and 1 kHz). The first horizontal blacked -out region 1500 is associated with the first frequency range and the second horizontal blacked -out region 1502 is associated with the second frequency range. The first and second regions represent frequencies in which audio output 1404 has been attenuated or suppressed. For example, the first and second regions correspond to portions of the signal spectrum where the filter has reduced the amplitude or presence of audio components below a predetermined threshold, thereby minimizing or eliminating sound generation in those ranges.

[0236] Some disclosed embodiments involve limiting the speaker from generating an audio output in at least one frequency range that causes vibratory interference with operation of the at least one sensor. “Causing” refers to initiating, inducing, or contributing to the occurrence of an effect, condition, or outcome, either directly or indirectly. The term “vibratory interference” refers to the presence or transmission of unwanted mechanical vibrations, pressure fluctuations, or acoustic energy that may disrupt the normal operation, stability, or accuracy of nearby components or systems. Such interference may result from physical coupling, airborne transmission, or structural resonance, and may affect performance, introduce noise, or lead to degradation overtime. For example, the vibratory interference may be caused by sound waves generated by a speaker of the wearable system. The phrase “causing vibratory interference with the operation of a sensor” refers to the introduction of unwanted movements into the sensor’s measurement environment in a manner that distorts, masks, or otherwise disrupts its ability to accurately detect or measure a desired signal. This interference may occur in real time, during active sensing, or may persist overtime, due to residual vibrations, signal contamination, or mechanical fatigue. For example, vibratory interference may arise when acoustic energy from the speaker couples through the device housing, through the wearer’s body, or directly through the surrounding air. Such interference may result in the sensor misinterpreting ambient vibrations as neuromuscular activity, registering false signals, or failing to detect subtle physiological changes. Additionally, vibratory interference may damage sensitive components of the sensor, particularly if the audio output includes frequencies near the sensor’s mechanical or electrical resonance frequency. This may lead to excessive oscillation, component stress, or long-term degradation in performance. Accordingly, “limiting the speaker from generating an audio output in at least one frequency range that causes vibratory interference with operation of the at least one sensor” means to reduce, suppress, or eliminate the generation of sound signals caused by the speaker within specific frequency range known to adversely interfere with the sensing of neuromuscularAttorney Docket No. 16198.0057-00304 activity for speech detection. This may be achieved by attenuating the amplitude of the audio signal within the identified frequency range, blocking signal components entirely, or dynamically adjusting the speaker’s output based on contextual factors such as sensor activity or environmental conditions. Consistent with the present disclosure, the extent of vibratory interference may depend on factors such as the distance and the physical coupling between the speaker and the sensor, the geometry and material composition of the head-mountable housing, and the amplitude and frequency content of the audio output.

[0237] In a first example, when the at least one sensor includes an optical displacement sensor used to detect micromovement using speckle analysis, audio output in the range of 120 Hz to 350 Hz may excite mechanical resonances of the sensor’s optical mounting or lens assembly, causing measurable jitter in the optical path and degrading displacement accuracy. In a second example, when the at least one sensor includes a piezoelectric vibration sensor used to detect skin or bone conduction signals, audio output in the range of 1 .5 kHz to 4 kHz may couple directly into the piezoelectric element via the housing stmcture, exciting the sensor’s own resonance modes and causing persistent ringing that may mask actual skin or bone motion signals. In addition, the sound waves may cause minute skin vibrations corresponding to mechanical oscillations in nearby tissue that may cause false positives. In a third example, when the at least one sensor includes a capacitive proximity sensor configured to detect facial micromovements, audio output in the range of approximately 500 Hz to 1 .2 kHz may introduce pressure fluctuations in the surrounding air that modulate the dielectric environment near the sensor surface. A person skilled in the art would recognize that a sensor combination of differing types of sensors may be associated with multiple frequency ranges that cause vibratory interference with operation of the at least one. For example, when the at least one the sensor includes both an optical displacement sensor and a piezoelectric vibration sensor, two differing frequency ranges may cause vibratory interference with operation of the at least one sensor. Therefore, in this example, the at least one filter may be configured to limit the speaker from generating an audio output the two differing frequency ranges.

[0238] By way of a non-limiting example, in Fig. 13, wearable system 1300 may include at least one filter 1316 for limiting the speaker from generating an audio output in at least one frequency range that causes vibratory interference with operation of the at least one sensor.

[0239] In some embodiments, limiting the speaker from generating audio output in the at least one frequency range includes preventing audio output in the at least one frequency range. “Preventing” refers to actively inhibiting, blocking, or substantially eliminating the occurrence of a specified action or output, such that the intended result is substantially avoided or does not occur. The phrase “preventing audio output in the at least one frequency range” refers to substantially removing or blocking those components from the signal driving the speaker such that the speaker produces negligible or no acoustic energy within the specified frequency range. For example, in some cases,Attorney Docket No. 16198.0057-00304 preventing audio output in the at least one frequency range may be implemented by applying a digital notch filter to the audio signal prior to amplification, thereby removing energy within a narrow frequency band known to cause interference. Additionally, preventing audio output in the at least one frequency range may be implemented by using a hardware-based band-stop filter in the speaker’s driver circuitry to suppress signal components within a predefined frequency range.

[0240] In some embodiments, limiting the speaker from generating audio output in the at least one frequency range includes generating a modified audio output with reduced intensity in the at least one frequency range. The phrase “modified audio output” refers to an audio signal that has been altered in amplitude, spectral content, phase, or other characteristics relative to an unmodified version of the same signal. For example, in a modified audio one or more frequency components may be attenuated, reshaped, or eliminated. The term “reduced intensity” refers to a decrease in the amplitude, power, or energy of a signal component, such that its effect or perceptibility is diminished relative to its original level. In the context of the present disclosure, a modified audio output may be generated by attenuating specific frequency components within the at least one frequency range known to cause vibratory interference with operation of the at least one sensor. For example, if an optical displacement sensor exhibits sensitivity to structural vibrations in the range of 150 Hz to 300 Hz, the audio output may be digitally processed to reduce the amplitude of any components within this range by at least 5 dB, 10 dB , 20 dB, thereby lowering the mechanical excitation transmitted through the head-mountable housing to a level insufficient to disrupt sensor operation. Similarly, in the case of a piezoelectric vibration sensor susceptible to resonance -induced artifacts between 2 kHz and 3 kHz, applying targeted attenuation within this range can prevent the speaker’s output from exciting the sensor’s resonance modes, thus preserving measurement accuracy. In other embodiments, preventing audio output in the at least one frequency range and or limiting the speaker’s output involves noise cancellation where a counter audio signal reduces one or more unwanted frequencies.

[0241] In some embodiments, the filter may be an analog band-stop filter configured to create a high impedance for the at least one frequency range. The term “analog band-stop filter” refers to an electrical network of passive and / or active components configured to attenuate signals within a specified band of frequencies while allowing frequencies outside that band to pass with minimal attenuation. An analog band-stop filter may be implemented using inductors, capacitors, resistors, operational amplifiers, or combinations thereof. The phrase “create a high impedance for the at least one frequency range” refers to configuring the analog band-stop filter such that the electrical impedance presented to signals within the specified frequency range is significantly increased, thereby impeding the flow of current associated with those frequencies. This elevated impedance results in reduced signal transmission through the filter at those frequencies, effectively attenuating or blocking them from reaching the speaker. For example, a parallel resonant circuit composed of an inductor andAttorney Docket No. 16198.0057-00304 capacitor may be tuned to resonate at a target frequency range, presenting a high impedance path and thereby suppressing signal components within that range.

[0242] Consistent with the present disclosure, the analog band -stop filter may include at least one capacitor. For example, the analog band -stop filter may include at least one capacitor with a value no greater than 10 microfarads, (pF), no greater than 15 pF, no greater than, 20 pF or no greater than 50 pF. The term “capacitor” refers to an electrical component that stores energy in an electric field between a pair of conductive plates separated by a dielectric material. Capacitors in an analog band- stop filter form part of a frequency -selective network that attenuates signals within a specified range while allowing frequencies outside that range to pass substantially unaffected. The impedance of a capacitor is inversely proportional to frequency, enabling it to block lower-frequency components and pass higher-frequency components when combined with resistors and / or inductors in the filter circuit. In a band-stop configuration, capacitors are arranged with resistive and / or inductive elements to shunt orblock signal eneigy in the stop band, thereby reducing the amplitude of targeted frequencies. The capacitance value, together with associated resistance and / or inductance, determines the cutoff points of the stop band according to known filter design equations, allowing precise targeting of frequencies that may cause vibratory interference with a sensor. By selecting appropriate capacitor values, the analog band-stop filter can be tuned to attenuate specific frequency range that would otherwise mechanically excite components of the sensor or its mounting structure.

[0243] By way of a non-limiting example, in Fig. 13, at least one filter 1316 may be analog bandstop filter including at least one capacitor with a value no greater than 10 pF .

[0244] In other embodiments, the filter may be a digital band-stop filter configured for implementation by at least one digital signal processor. The term “digital band-stop filter” refers to a computations that attenuate signal components within a specified frequency range while allowing other frequencies to pass. A digital band-stop filter may be implemented using finite impulse response (FIR) or infinite impulse response (IIR) techniques, and may be applied in real time or as part of a pre-processing stage. The at least one digital signal processor may be integrated into the wearable system or implemented in a remote processing device in wired or wireless communication with the wearable system. Specifically, in some cases, the at least one digital signal processor may transform an audio signal to a frequency domain, modify the transformed signal by applying a filter function in the frequency domain, transform back the modified signal into a time domain, and output a modified version of the audio signal. In one example, the at least one digital signal processor may transform an audio signal from the time domain to the frequency domain using a mathematical operation such as a fast Fourier transform (FFT). The time-domain audio signal represents amplitude as a function of time, while the frequency -domain representation expresses the same signal as a set of discrete frequency components, each with an associated magnitude and phase. Once in the frequency domain, the processor may apply a filter function — such as a band-stop function — that selectively attenuates orAttorney Docket No. 16198.0057-00304 removes components within the at least one frequency range that causes vibratory interference with operation of the at least one sensor. The processor may then transform the modified frequency -domain data back into the time domain using an inverse fast Fourier transform (IFFT), producing a waveform in which the targeted frequency components have been reduced or eliminated. The resulting time- domain signal, now constituting a modified audio output, may be transmitted to the speaker, thereby limiting or preventing the generation of acoustic eneigy in the problematic frequency range while preserving the remaining portions of the audio signal.

[0245] By way of a non-limiting example, in Fig. 13, wearable system 1300 may include at least one processor 1318 for implementing a digital band-stop filter to limit the speaker. In some embodiments, the at least one processor 1318 may form part of at least one filter 1316 integrated within the wearable system 1300, such that the processor executes instructions or algorithms to attenuate frequency components within a designated stop band of the audio signal prior to output by speaker 1314. In other embodiments, the implementation of the digital band-stop filter may be performed by a processor that is separate from the wearable system 1300, such as a processor located in a remote computing device in wireless communication with the wearable system. In such cases, the remote processor may receive an audio signal, apply the digital band -stop filter to generate a modified version of the audio signal, and transmit the modified signal to the wearable system for output.

[0246] In some embodiments, the system further includes an equalizer configured to modify frequencies not affected by the filter to maintain desired audio characteristics. The term “equalizer” broadly refers to a device, circuit, or code that adjusts the amplitude of selected frequency components of an audio signal to achieve a taiget tonal balance, frequency response, or sound quality. An equalizer may operate in the analog domain, the digital domain, or a combination thereof, and may include fixed-band, graphic, or parametric configurations. In operation, the equalizer modifies frequencies by amplifying (boosting) or attenuating (reducing) selected frequency bands, thereby shaping the spectral content of the output signal. In the context of the present disclosure, the equalizer may be used to adjust unaffected frequency ranges (i.e., frequencies other than within at least one frequency range) so that the overall perceived audio quality remains consistent despite the removal or attenuation of certain frequencies by the filter. The phrase “modify frequencies not affected by the filter to maintain desired audio characteristics” refers to selectively adjusting the amplitude or spectral balance of frequency components outside the filtered range to compensate for filtered frequencies in the at least one frequency range and to maintain the intended acoustic profile of the audio output. For example, if a band-stop filter removes mid-range frequencies to avoid sensor interference, the equalizer may boost adjacent low or high frequencies to preserve clarity or fullness in the perceived sound.Attorney Docket No. 16198.0057-00304

[0247] By way of a non-limiting example, in Fig. 13, wearable system 1300 may include an equalizer 1320 for modifying frequencies not affected by the filter to maintain desired audio characteristics.

[0248] In some cases, the equalizer is configured to add between 0.1 to 3 dB to some frequencies to compensate for a lack of audio output in the at least one frequency range . The phrase “compensate for a lack of audio output in the at least one frequency range” means to restore perceived sound balance when certain frequency bands are suppressed or lost. For example, adding between 0. 1 dB and 3 dB by the equalizer increases the amplitude of those unaffected frequencies by a small, controlled amount so that the overall loudness and tonal balance of the audio signal is preserved. For example, if a band-stop filter removes a portion of midrange frequencies between 500 Hz and 800 Hz, the equalizer may boost nearby unaffected bands — such as 400 Hz to 475 Hz and / or 1 kHz to 1 .2 kHz by 1-2 dB to offset the perceived dip in sound eneigy. This compensation helps maintain a smooth and natural audio experience for the user while still preventing generation of frequencies that would cause vibratory interference with the at least one sensor.

[0249] By way of a non -limiting example, Fig. 15B illustrates an example of a modified audio output 1510 after equalizer (e.g., equalizer 1320) modified frequencies not affected by the filter, i.e., frequencies not included in first region 1500 and in the second region 1502 associated with the at least one frequency range that causes vibratory interference with operation of the at least one sensor. Modified audio output 1510 may maintain at least some of the audio characteristics of audio output 1400.

[0250] In one implementation, the at least one sensor includes a light source and a light detector configured to detect reflections of light projected from the light source, and wherein the reflections are indicative of non -lip facial skin micromovements associated with the neuromuscular activity. The terms “light source,” “light detector,” “light reflections,” and “non -lip facial skin micromovement” are described throughout the present disclosure and illustrated in, for example, Figs. 5A and 5B. Consistent with the present disclosure vibrations originating from the speaker can interfere with the operation of at least one of the light source or the light detector by introducing unintended relative motion between optical components, altering the angle or position of the projected light beam, or causing jitter in the detected light signal. Such vibration -induced perturbations may modulate the intensity, phase, or spatial distribution of the reflected light in a manner unrelated to the actual micromovement of the skin, thereby introducing noise or false readings into the measurement data acquired by the at least one sensor. In some embodiments, the light source is located closer to the light detector than to the speaker. This refers to the speaker being further from the light detector than the light sensor. By way of example, a wearable system may include a head -mounted frame in which the light source and light detector are mounted adjacent to one another on a common support arm extending from a temple portion of the frame, while the speaker is positioned farther away — such asAttorney Docket No. 16198.0057-00304 in or adjacent to the ear. This arrangement can reduce mechanical coupling of vibrations from the speaker to the optical components. For example, as shown in Fig. 1, the light source and light detector may be housed within optical sensing unit 116, while the speaker is positioned within output unit 114, the two units being separated by arm 118.

[0251] In some embodiments, the light source is configured to project coherent light at a wavelength between about 650 nm and 1150 nm towards a non-lip region of the wearer’s face. Coherent light refers to light waves having a fixed phase relationship, for example, light that produced by a laser diode. For example, a coherent light source emitting at 850 nm may be used for speckle pattern analysis of facial skin micromovements. In some embodiments, the light source is configured to project noncoherent light outside a visible spectrum towards a non-lip region of the wearer’s face. Noncoherent light refers to light waves having random phase relationships, such as that produced by light-emitting diodes (LEDs). The phrase “non-lip region of the wearer’s face” broadly refers to any facial area other than the lips, including, for example, the cheeks, temples, jawline, or area around the eyes. Targeting a non-lip region can allow detection of subtle micromovements indicative of neuromuscular activity.

[0252] Some embodiments include at least one processor configured to alter operation of at least one of the light source or the light detector when the speaker generates audio output to thereby maintain measurement accuracy in a presence of vibrations. The phrase “alter operation of at least one of the light source or the light detector” refers to modifying how either the light source, the light detector, or both function. By way of example, this may occur by adjusting their timing, intensity, sensitivity, or activation pattern — in response to specific conditions (e.g., when the speaker generates audio output). The phrase “maintain measurement accuracy” refers to preserving the integrity, reliability, and resolution of sensor measurements such that detected signals remain representative of actual neuromuscular activity rather than artifacts caused by vibration. Consistent with the present disclosure, in addition to limiting the speaker from generating an audio output in at least one frequency range that causes vibratory interference, the wearable system may alter the operation of the at least one sensor to maintain its measurement accuracy. In some cases, altering the operation involves increasing light intensity of a light source to improve the signal -to-noise ratio in the presence of vibratory noise. For example, if audio output at 200 Hz induces mechanical jitter in the light source, increasing the light source’s drive current to temporarily raise optical power by 25% may help ensure that reflected signals remain detectable above vibration -induced variations. Other examples of altering the light source to maintain measurement accuracy may include adjusting the beam focus to optimize the illuminated spot size, changing the modulation frequency of the light source to a range less affected by the vibratory interference, or switching to a different wavelength that is less sensitive to vibration-induced scattering. Additionally, altering the operation of the light detector may include reducing detector exposure times. For example, shortening the detector’s integration period from 10Attorney Docket No. 16198.0057-00304 ms to 5 ms during periods of audio output may prevent blur in the captured speckle pattern. Other examples of altering the light detector to maintain measurement accuracy may include dynamically adjusting the detector’s gain to compensate for increased light intensity, implementing synchronous detection matched to the modulation frequency of the light source, or applying real-time digital filtering to suppress frequency components associated with the vibratory interference.

[0253] Fig. 16 is a flowchart of example process 1600 for speech detection, consistent with some disclosed embodiments. In some embodiments, process 1600 may be performed by at least one processor associated with the wearable system (e.g., processing device 400 in Fig. 4) to perform operations or functions described herein. In some embodiments, some aspects of process 1600 may be implemented as software (e.g., program codes or instructions) that are stored in a memory (e.g., memory device 402) or a non-transitory computer readable medium. In some embodiments, some aspects of process 1600 may be implemented as hardware (e.g., a specific purpose circuit). In some embodiments, process 1600 may be implemented as a combination of software and hardware.

[0254] Process 1600 may include a step 1602 of operating at least one sensor to detect neuromuscular activity. The at least one sensor may be integrated with a head -mountable housing and configured for direction towards a wearer’s head. This step enables a system (e.g., wearable system 1300) to monitor subtle electrical or muscular signals generated by the user’s facial movements, speech articulation, or other neuromuscular actions. The sensor’s placement ensures optimal proximity to the source of activity to improve signal fidelity.

[0255] Process 1600 may include a step 1604 of operating a speaker for outputting audio. The speaker may or may not be integrated with the head -mountable housing. This step allows the system to deliver audio feedback, prompts, or media content to the user. Whether embedded in the housing or external, the speaker functions as the primary output channel for auditory information.

[0256] Process 1600 may include a step 1606 limiting the speaker from generating an audio output in at least one frequency range that causes vibratory interference with operation of the at least one sensor. This step involves preventing the speaker from emitting sounds that could mechanically or electromagnetically interfere with the sensor’s ability to detect neuromuscular signals. The at least one frequency range known to cause vibratory interference is attenuated or removed from the audio output. In some cases, the wearable system may dynamically adjust the filtered range based on sensor feedback or environmental conditions. In some cases, to preserve audio quality, an equalizer may compensate for the attenuation of certain frequencies by boosting adjacent unaffected frequencies, thereby maintaining desired audio characteristics.

[0257] Some disclosed embodiments involve performing operations for resolving detected speech ambiguities. The term “ambiguity” refers to uncertainty or inexactness in meaning. In speech, ambiguity may arise from unclear articulation or foreign accents that alter pronunciation and intonation, particularly when the listener is unfamiliar with the speaker’s dialect. Speech ambiguityAttorney Docket No. 16198.0057-00304 may also result from speech impairments such as stuttering, slurred speech, or atypical articulation patterns. Technical and environmental factors may further cause speech ambiguity, including poor call quality (e.g., low signal strength, dropped connections, or compression artifacts) and background noise (e.g., wind, traffic, or overlapping conversations). For example, during a mobile call with intermittent connectivity, portions of speech may not be captured by the microphone. Similarly, if a person speaks in a crowded or noisy environment, individual words may be masked and not clearly perceived by the listener.

[0258] The phrase “resolving detected speech ambiguities” refers to the process of addressing and eliminating (or attempting to eliminate) uncertainties that arise when detected speech could have more than one possible interpretation. For example, if the system detects a vocalization that could correspond to two different words, resolving involves applying additional analysis and generating an output that clarifies the ambiguity. By way of a non -limiting example, as described in detail below, the detected speech ambiguities may be resolved by the generation of a hybrid output that includes the user’s voice for unambiguous portions and a synthesized voice for ambiguous portions.

[0259] Some disclosed embodiments involve receiving audio signals representing a plurality of words vocalized by an individual. The term “receiving” may include retrieving, acquiring, or otherwise gaining access to data, such as reading data from memory and / or receiving data from a computing device via a wired or wireless communications channel. The term “audio signals” refers to representations of sound waves, in the form of analog or digital signals, that convey audible information such as speech, music, or environmental sounds. These signals may be transmitted through physical media (e.g., air, wires, or fiber optics) or wirelessly, and can be processed, recorded, amplified, or reproduced by electronic systems. Audio signals may include any representation of sound, typically using either a changing level of electrical voltage for analog signals, or a series of binary numbers for digital signals. Examples of audio signals include waveforms, frequencies, amplitudes, decibels, bits, and pressure levels. The act of receiving audio signals may occur via a microphone, sensor, or other audio input device associated with, integrated into, or connected to the disclosed system. These devices capture the sound waves produced by the individual and convert them into the audio signals, which are then processed by a processing device. For example, the disclosed system may receive audio signals when the individual vocalized words by detecting the sound through its built-in microphone. The term “vocalized” means that sounds are audibly produced by the individua...

Claims

Attorney Docket No. 16198.0057-00304CLAIMSWhat is claimed is:1 . A non-transitory computer readable medium containing instructions that when executed by at least one processor cause the at least one processor to perform operations for synthesizing voice from subvocalized speech, the operations comprising: operating at least one sensor for detecting neuromuscular activity associated with a non -lip region of a head of an individual; receiving, from the at least one sensor, subvocalization signals indicative of the neuromuscular activity associated with the non-lip region, wherein the subvocalization signals are captured in an absence of perceptible vocalization; processing the received subvocalization signals to determine a plurality of subvocalized words including a first subvocalized word and a second subvocalized word; determining a first voice characteristic from a first set of the subvocalization signals associated with the first subvocalized word; determining a second voice characteristic from a second set of subvocalization signals associated with the second subvocalized word, the second voice characteristic differing from the first voice characteristic; and synthesizing audible output of the plurality of subvocalized words, wherein the first subvocalized word is audibly outputted in a manner reflecting the first voice characteristic and the second subvocalized word is audibly outputted in a manner reflecting the second voice characteristic.

2. The non-transitory computer readable medium of claim 1, wherein each of the first voice characteristic and the second voice characteristic includes at least one of intonation, pitch, pace, intensity, prosody, articulation, or volume, and wherein the first voice characteristic and the second voice characteristic differ from each other by at least one of the intonation, pitch, pace, intensity, prosody, articulation, or volume.

3. The non-transitory computer readable medium of claim 1, wherein the at least one sensor includes a light source and a light detector configured to detect reflections of light projected from the light source, and wherein the reflections are indicative of non-lip facial skin micromovements associated with the neuromuscular activity.

4. The non-transitory computer readable medium of claim 1, wherein the at least one sensor includes a sensor configured to detect mechanical deformation caused by muscle engagement, and wherein the detected deformation is indicative of facial muscle activity in the non-lip region associated with the neuromuscular activity.Attorney Docket No. 16198.0057-003045. The non-transitory computer readable medium of claim 1, wherein the at least one sensor includes an electrode configured to detect electrical signals transmitted through cranial nerves, and wherein the detected electrical signals are indicative of facial nerve activity in the non-lip region associated with the neuromuscular activity.

6. The non-transitory computer readable medium of claim 1, wherein the operations further include analyzing the first set of the subvocalization signals to determine a sequence of non-lip neuromuscular actions indicative of the first subvocalized word, and determining the first voice characteristic based on a value of at least one parameter associated with the sequence of non-lip neuromuscular activities.

7. The non-transitory computer readable medium of claim 6, wherein the at least one parameter associated with the sequence of non-lip neuromuscular actions includes at least one of: an amplitude associated with facial micromovements, a frequency of facial micromovements, timing of facial micromovements, an amplitude of muscle contraction, a frequency of muscle contraction, timing of muscle contraction, an amplitude of electrical signals transmitted through facial nerves, a frequency of electrical signals transmitted through facial nerves, or a timing of electrical signals transmitted through cranial nerves.

8. The non-transitory computer readable medium of claim 6, wherein the operations further include determining that the first voice characteristic has a first manifestation when the value of the at least one parameter is greater than a threshold and a second manifestation when the value of the at least one parameter is less than the threshold.

9. The non-transitory computer readable medium of claim 1, wherein the operations further include determining an emotional condition based on the neuromuscular activity, and wherein determining the first voice characteristic is further based on the determined emotional condition.

10. The non-transitory computer readable medium of claim 9, wherein the emotional condition includes at least one of happiness, fear, excitement, embarrassment, sadness, anger, surprise, disgust, or contempt.11 . The non-transitory computer readable medium of claim 9, wherein the operations further include determining a change in the emotional condition, and modifying the second voice characteristic based on the determined change in the emotional condition.Attorney Docket No. 16198.0057-0030412. The non -transitory computer readable medium of claim 1, wherein the operations further include determining a physical condition based on the neuromuscular activity and wherein determining the first voice characteristic is further based on the determined physical condition.

13. The non-transitory computer readable medium of claim 12, wherein the physical condition includes at least one of pain, fatigue, anxiety, stress, or respiratory abnormality.

14. The non-transitory computer readable medium of claim 12, wherein the operations further include determining a change in the physical condition, and modifying the second voice characteristic based on the change in the physical condition.

15. The non-transitory computer readable medium of claim 1, wherein the operations further include receiving auxiliary data, using the auxiliary data to determine a context for the plurality of sub vocalized words, and determining the first voice characteristic is further based on the determined context.

16. The non-transitory computer readable medium of claim 15, wherein the auxiliary data includes at least one of interaction data, location data, or environmental data; and wherein the determined context includes at least one of a location, an environmental condition, or an identity of a person with whom the individual is interacting.

17. The non-transitory computer readable medium of claim 1, wherein the operations further include receiving input from the individual and personalizing a manner in which the first subvocalized word is audibly outputted based on the received input.

18. The non-transitory computer readable medium of claim 1, wherein the operations further include storing the first voice characteristic associated with the first subvocalized word and storing the second voice characteristic associated with the second sub vocalized word as reference data for use in future synthesization.

19. A method for synthesizing voice from subvocalized speech, comprising: operating at least one sensor for detecting neuromuscular activity associated with a non -lip region of a head of an individual; receiving, from the at least one sensor, subvocalization signals indicative of the neuromuscular activity associated with the non-lip region, wherein the subvocalization signals are captured in an absence of perceptible vocalization;Attorney Docket No. 16198.0057-00304 processing the received subvocalization signals to determine a plurality of subvocalized words including a first subvocalized word and a second subvocalized word; determining a first voice characteristic from a first set of the subvocalization signals associated with the first subvocalized word; determining a second voice characteristic from a second set of subvocalization signals associated with the second subvocalized word, the second voice characteristic differing from the first voice characteristic; and synthesizing audible output of the plurality of subvocalized words, wherein the first subvocalized word is audibly outputted in a manner reflecting the first voice characteristic and the second subvocalized word is audibly outputted in a manner reflecting the second voice characteristic.

20. A system for synthesizing voice from subvocalized speech, the system comprising: at least one sensor configured to detect neuromuscular activity associated with a non -lip region of a head of an individual; and at least one processor configured to: receive, from the at least one sensor, subvocalization signals indicative of the neuromuscular activity, wherein the subvocalization signals are captured in an absence of perceptible vocalization; process the received subvocalization signals to determine a plurality of subvocalized words including a first subvocalized word and a second subvocalized word; determine a first voice characteristic from a first set of the subvocalization signals associated with the first subvocalized word; determine a second voice characteristic from a second set of subvocalization signals associated with the second subvocalized word, the second voice characteristic differing from the first voice characteristic; and synthesize audible output of the plurality of subvocalized words, wherein the first subvocalized word is audibly outputted in a manner reflecting the first voice characteristic and the second subvocalized word is audibly outputted in a manner reflecting the second voice characteristic.

21. A wearable system for speech detection, the wearable system comprising: a head-mountable housing; at least one sensor integrated with the head -mountable housing and configured for direction towards a wearer’s head, wherein the at least one sensor is configured to detect neuromuscular activity; a speaker integrated with the head -mountable housing for outputting audio; andAttorney Docket No. 16198.0057-00304 at least one filter for limiting the speaker from generating an audio output in at least one frequency range that causes vibratory interference with operation of the at least one sensor.22 .The wearable system of claim 21, wherein the at least one frequency range includes one or more frequencies between 1 and 300 Hz.

23. The wearable system of claim 21, wherein the at least one frequency range includes a plurality of frequencies under 200 Hz.

24. The wearable system of claim 21, wherein the head -mountable housing is configured for mounting on an ear of the wearer.

25. The wearable system of claim 21, wherein the filter is an analog band -stop filter configured to create a high impedance for the at least one frequency range.

26. The wearable system of claim 25, wherein the analog band-stop filter includes at least one capacitor with a value no greater than 10 microfarads (pF).

27. The wearable system of claim 21, wherein the filter is a digital band -stop filter configured for implementation by at least one digital signal processor.

28. The wearable system of claim 27, wherein the at least one digital signal processor is configured to transform an audio signal to a frequency domain, modify the transformed signal by applying a filter function in the frequency domain, transform back the modified signal into a time domain, and output a modified version of the audio signal.

29. The wearable system of claim 21, wherein the filter is configured to limit the speaker from generating audio output in the at least one frequency range by preventing audio output in the at least one frequency range.

30. The wearable system of claim 21, wherein the filter is configured to limit the speaker from generating audio output in the at least one frequency range by generating a modified audio output with reduced intensity in the at least one frequency range.

31. The wearable system of claim 21, wherein the system further includes an equalizer configured to modify frequencies not affected by the filter to maintain desired audio characteristics.Attorney Docket No. 16198.0057-0030432. The wearable system of claim 31 , wherein the equalizer is configured to add between 0. 1 to 3 dB to some frequencies to compensate for a lack of audio output in the at least one frequency range .

33. The wearable system of claim 21, wherein the at least one sensor includes a light source and a light detector configured to detect reflections of light projected from the light source, and wherein the reflections are indicative of non-lip facial skin micromovements associated with the neuromuscular activity.

34. The wearable system of claim 33, wherein the light source is located closer to the light detector than to the speaker.

35. The wearable system of claim 33, wherein the light source is configured to project coherent light at a wavelength between about 650 nm and 1150 nm towards a non-lip region of the wearer’s head.

36. The wearable system of claim 33, wherein the light source is configured to project noncoherent light outside a visible spectmm towards a non-lip region of the wearer’s head.

37. The wearable system of claim 33, further including at least one processor configured to alter operation of at least one of the light source or the light detector when the speaker generates audio output to thereby maintain measurement accuracy in a presence of vibrations.

38. The wearable system of claim 37, wherein altering the operation of at least one of the light source or the light detector includes increasing light intensity or reducing detector exposure times.

39. A method for speech detection, comprising: operating at least one sensor to detect neuromuscular activity, wherein the at least one sensor is integrated with a head -mountable housing and configured for direction towards a wearer’s head; operating a speaker for outputting audio; and limiting the speaker from generating an audio output in at least one frequency range that causes vibratory interference with operation of the at least one sensor.

40. Anon-transitory computer readable medium containing instructions that when executed by at least one processor cause the at least one processor to perform operations for speech detection, the operations comprising:Attorney Docket No. 16198.0057-00304 operating at least one sensor to detect neuromuscular activity, wherein the at least one sensor is integrated with a head -mountable housing and configured for direction towards wearer’s head; operating a speaker for outputting audio; and limiting the speaker from generating an audio output in at least one frequency range that causes vibratory interference with operation of the at least one sensor.

41. A non -transitory computer readable medium containing instructions that when executed by at least one processor cause the at least one processor to perform operations for resolving detected speech ambiguities, the operations comprising: receiving audio signals representing a plurality of words vocalized by an individual; determining an ambiguity in the audio signals; during receiving of the audio signals, operating at least one sensor directed towards a non -lip region of a head of the individual; receiving, from the at least one sensor, non -audio signals indicative of neuromuscular activity associated with the non-lip region; analyzing the non-audio signals to resolve the ambiguity through an identification of at least one phoneme corresponding to the ambiguity; and generating a hybrid output of the plurality of vocalized words, wherein the hybrid output includes a first portion derived from the audio signals and a second portion derived from the non- audio signals, the second portion including a representation of the at least one phoneme.

42. The non -transitory computer readable medium of claim 41, wherein the first portion includes an original voice of the individual and the second portion includes a synthesized voice.

43. The non -transitory computer readable medium of claim 42, wherein the synthesized voice simulates the voice of the individual.

44. The non-transitory computer readable medium of claim 42, wherein the operations further include analyzing the audio signals to determine the original voice of the individual, and determining the synthesized voice based on the original voice of the individual.

45. The non-transitory computer readable medium of claim 41, wherein generating the hybrid output of the plurality of vocalized words includes indistinguishably joining the first portion including some of the plurality of words and the second portion including the representation of the at least one phoneme.Attorney Docket No. 16198.0057-0030446. The non -transitory computer readable medium of claim 41, wherein the at least one phoneme includes a single phoneme in a word.

47. The non -transitory computer readable medium of claim 41, wherein the at least one phoneme includes a plurality of phonemes making up one or more words.

48. The non-transitory computer readable medium of claim 41, wherein the hybrid output is generated for a live phone call between the individual and at least one other individual, and wherein the operations include imposing a delay of less than 0.25 seconds for generating the hybrid output.

49. The non-transitory computer readable medium of claim 46, wherein the operations further include, prior to generating the hybrid output, receiving consent from the individual to generate the hybrid output.

50. The non-transitory computer readable medium of claim 41, wherein the operations further include analyzing the non -audio signals to detect neuromuscular activity that occurred prior to an onset of an ambiguous vocalization, and using the neuromuscular activity to identify the at least one phoneme corresponding to the ambiguity.

51. The non-transitory computer readable medium of claim 41 , wherein the operations further include using audio processing on the audio signals to identify the plurality of words and using contextual analysis on the plurality of words to identify the at least one phoneme that corresponds to the ambiguity.

52. The non-transitory computer readable medium of claim 41, wherein the operations further include accessing a personal record of a vocabulary of the individual, and analyzing the personal record to identify the at least one phoneme most likely to correspond to the ambiguity.

53. The non-transitory computer readable medium of claim 41, wherein the operations further include analyzing the audio signals and the non-audio signals to determine a cause of the ambiguity in the audio signals.

54. The non-transitory computer readable medium of claim 53, wherein when the cause of the ambiguity is associated with an environmental condition, the operations further include initiating at least one action to reduce future ambiguities.Attorney Docket No. 16198.0057-0030455. The non -transitory computer readable medium of claim 54, wherein the at least one action includes performing noise reduction, adjusting microphone settings, suggesting microphone positioning, or recommending moving to a quieter environment.

56. The non-transitory computer readable medium of claim 53, wherein when the cause of the ambiguity is associated with a condition of the individual, the operations further include receiving input from the individual and generating the hybrid output based on the received input.

57. The non-transitory computer readable medium of claim 56, wherein when the condition of the individual is continuous, the input is received before receipt of the audio signals.

58. The non-transitory computer readable medium of claim 53, wherein when the cause of the ambiguity is due to use of a secondary language and the plurality of words are generally spoken in a primary language, the operations further include translating the at least one phoneme to the primary language.

59. A method for resolving detected speech ambiguities, the method comprising: receiving audio signals representing a plurality of words vocalized by an individual; determining an ambiguity in the audio signals; during receiving of the audio signals, operating at least one sensor directed towards a non -lip region of a head of the individual; receiving, from the at least one sensor, non -audio signals indicative of neuromuscular activity associated with the non-lip region; analyzing the non-audio signals to resolve the ambiguity through an identification of at least one phoneme corresponding to the ambiguity; and generating a hybrid output of the plurality of vocalized words, wherein the hybrid output includes a first portion derived from the audio signals and a second portion derived from the non- audio signals, the second portion including a representation of the at least one phoneme.

60. A system for resolving detected speech ambiguities, the system comprising: at least one processor configured: receive audio signals representing a plurality of words vocalized by an individual; determine an ambiguity in the audio signals; during receiving of the audio signals, operate at least one sensor directed towards a non-lip region of a head of the individual;Attorney Docket No. 16198.0057-00304 receive, from the at least one sensor, non-audio signals indicative of neuromuscular activity associated with the non-lip region; analyze the non-audio signals to resolve the ambiguity through an identification of at least one phoneme corresponding to the ambiguity; and generate a hybrid output of the plurality of vocalized words, wherein the hybrid output includes a first portion derived from the audio signals and a second portion derived from the non-audio signals, the second portion including a representation of the at least one phoneme.