Optical sensing of facial movements

EP4747864A1Pending Publication Date: 2026-05-27Q (CUE) LTD

Patent Information

Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
Q (CUE) LTD
Filing Date
2024-07-18
Publication Date
2026-05-27

AI Technical Summary

Technical Problem

Existing technologies face challenges in accurately detecting fine movements of the skin for speech sensing, particularly in noisy environments and without vocalization, which affects privacy and communication efficiency.

Method used

The development of optical sensing devices that use interferometric, confocal chromatic, and ellipsometric sensors to detect changes in coherent light, broadband light, and polarized light reflected from the skin, respectively, to generate a speech output without direct contact.

Benefits of technology

These devices enable accurate detection of subvocalization and silent speech, improving privacy and communication efficiency by converting skin movements into speech outputs, even in noisy environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2024056967_23012025_PF_FP_ABST
    Figure IB2024056967_23012025_PF_FP_ABST
Patent Text Reader

Abstract

A sensing device (20, 60) configured to fit on a head of a user (24) includes an optical sensing head (28, 68) held by the device in a location in proximity to a face of the user and includes an emitter (70) configured to direct coherent light toward multiple locations on a body surface of the user and an interferometric sensor (76) configured to sense changes in a phase of the coherent light that is reflected from the multiple locations on the body surface. Processing circuitry (36) is configured to apply the sensed changes in the phase in generating a speech output. Other sensing modalities are also disclosed.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] OPTICAL SENSING OF FACIAL MOVEMENTS

[0002] CROSS-REFERENCE TO RELATED APPLICATION

[0003] This application claims the benefit ofU.S. Provisional Patent Application 63 / 514,154, filed July 18, 2023, which is incorporated herein by reference.

[0004] FIELD OF THE INVENTION

[0005] The present invention relates generally to physiological sensing, and particularly to methods and apparatus for sensing human speech.

[0006] BACKGROUND

[0007] The process of speech activates nerves and muscles in the chest, neck, and face. Thus, for example, electromyography (EMG) has been used to capture muscle impulses for purposes of speech sensing.

[0008] Secondary speckle patterns have been used for monitoring movement of skin on the human body. Secondary speckle typically occurs in diffuse reflections of a laser beam from a rough surface, such as the skin. By tracking both temporal and amplitude changes of secondary speckle produced by reflection from human skin when illuminated by a laser beam, investigators have measured blood pulse pressure and other vital signs. For example, U.S. Patent 10,398,314 describes a method for monitoring conditions of a subject’s body using image data that is indicative of a sequence of speckle patterns generated by the body.

[0009] Secondary speckle patterns on the skin can also be used in detecting speech. For example, PCT International Publication WO 2023 / 012527, whose disclosure is incorporated herein by reference, describes a sensing device, which includes a bracket configured to fit an ear of a user of the device. An optical sensing head is held by the bracket in a location in proximity to a face of the user and configured to sense light reflected from the face and to output a signal in response to the detected light. Processing circuitry processes the signal to generate a speech output.

[0010] As another example, PCT International Publication WO 2023 / 012527, whose disclosure is incorporated herein by reference, describes a method for generating speech that includes uploading a reference set of features that were extracted from sensed movements of one or more target regions of skin on faces of one or more reference human subjects in response to words articulated by the subjects and without contacting the one or more target regions. A test set of features is extracted from the sensed movements of at least one of the target regions of skin on a face of a test subject in response to words articulated silently by the test subject and without contacting the one or more target regions. The extracted test set of features is compared to the reference set of features, and based on the comparison, a speech output is generated, including the articulated words of the test subject.

[0011] Self-mixing interferometry has also been proposed as a means for detecting skin movement. For example, U.S. Patent 11,473,898 describes wearable devices that use self-mixing interferometry signals of a self-mixing interferometry sensor to recognize user inputs. The user inputs may include voiced commands or silent gesture commands. The devices may be wearable on the user's head, with the self-mixing interferometry sensor configured to direct a beam of light toward a location on the user's head. Skin deformations or vibrations at the location may be caused by the user’s speech or the user’s silent gestures and recognized using the self-mixing interferometry signal. The self-mixing interferometry signals may be used for bioauthentication and / or audio conditioning of received sound or voice inputs to a microphone.

[0012] SUMMARY

[0013] Embodiments of the present invention that are described hereinbelow provide improved devices and methods for detection of fine movements of the skin.

[0014] There is therefore provided, in accordance with an embodiment of the invention, a sensing device configured to fit on a head of a user. The device includes an optical sensing head held by the device in a location in proximity to a face of the user and including an emitter configured to direct coherent light toward multiple locations on a body surface of the user and an interferometric sensor configured to sense changes in a phase of the coherent light that is reflected from the multiple locations on the body surface. Processing circuitry is configured to apply the sensed changes in the phase in generating a speech output.

[0015] In a disclosed embodiment, the interferometric sensor includes a speckle interferometer.

[0016] In some embodiments, the emitter and the sensor direct the coherent light toward the body surface and receive the coherent light from the body surface in a monostatic configuration. Alternatively, the emitter and the sensor direct the coherent light toward the body surface and receive the coherent light from the body surface in a bistatic configuration. In a disclosed embodiment, the device includes a waveguide coupled to transmit a reference beam from the emitter to the sensor.

[0017] There is also provided, in accordance with an embodiment of the invention, a sensing device configured to fit on a head of a user. The device includes an optical sensing head held by the device in a location in proximity to a face of the user and including an emitter configured to direct broadband light toward a body surface of the user and a confocal chromatic sensor configured to sense changes in a wavelength of the light that is reflected from the body surface and focused by the confocal chromatic sensor. Processing circuitry is configured to apply the sensed changes in the wavelength in generating a speech output.

[0018] In some embodiments, the confocal chromatic sensor is configured to sense the changes in the wavelength of the light that is reflected from multiple locations on the body surface. In a disclosed embodiment, the confocal chromatic sensor includes an array of optical fibers and confocal optics configured to image the body surface confocally onto the optical fibers, such that the fibers capture the light reflected at respective confocal wavelengths from corresponding points on the body surface. A spectrometer is coupled to the optical fibers and configured to sense the changes in the respective confocal wavelengths.

[0019] There is additionally provided, in accordance with an embodiment of the invention, a sensing device configured to fit on a head of a user. The device includes an optical sensing head held by the device in a location in proximity to a face of the user and including an emitter configured to direct polarized light toward a body surface of the user and an ellipsometric sensor configured to sense changes in a polarization of the light that is reflected from the body surface. Processing circuitry is configured to apply the sensed changes in the polarization in generating a speech output.

[0020] In a disclosed embodiment, the ellipsometric sensor is configured to sense the changes in the polarization of the light that is reflected from multiple locations on the body surface.

[0021] In some embodiments, the body surface from which the optical sensing head receives the reflected light includes an area of the face of the user. Alternatively or additionally, the body surface from which the optical sensing head receives the reflected light includes an area of a neck of the user. Further additionally or alternatively, the body surface from which the optical sensing head receives the reflected light is in an ear canal of the user.

[0022] In another embodiment, the device includes a pressure sensor configured to fit in an ear of the user of the device and to output a signal in response to pressure changes associated with subvocalization by the user, wherein the processing circuitry is configured to apply the signal in generating the speech output.

[0023] There is further provided, in accordance with an embodiment of the invention, a sensing device, which includes a pressure sensor configured to fit in an ear of a user of the device and to output a signal in response to pressure changes associated with subvocalization by the user. Processing circuitry is configured to process the signal to generate a speech output. In a disclosed embodiment, the pressure sensor is configured to output the signal in response to the pressure changes due to movements of a tongue of the user. Additionally or alternatively, the device includes an earphone configured to fit in the ear of the user, wherein the pressure sensor is integrated with the earphone.

[0024] There is moreover provided, in accordance with an embodiment of the invention, a method for sensing, which includes mounting an optical sensing head in a location in proximity to a face of a user and directing coherent light from the sensing head toward multiple locations on a body surface of the user. Changes in a phase of the coherent light that is reflected from the multiple locations on the body surface to the optical sensing head are sensed interferometrically, and a speech output is generated responsively to the sensed changes.

[0025] There is furthermore provided, in accordance with an embodiment of the invention, a method for sensing, which includes mounting an optical sensing head in a location in proximity to a face of a user and directing broadband light from the sensing head toward a body surface of the user. Changes in a wavelength of the light that is reflected from the body surface are sensed using a confocal chromatic sensor in the optical sensing head, and a speech output is generated responsively to the sensed changes.

[0026] There is also provided, in accordance with an embodiment of the invention, a method for sensing, which includes mounting an optical sensing head in a location in proximity to a face of a user and directing polarized light from the sensing head toward a body surface of the user. Changes in a polarization of the light that is reflected from the body surface are sensed using an ellipsometric sensor in the optical sensing head, and a speech output is generated responsively to the sensed changes.

[0027] There is additionally provided, in accordance with an embodiment of the invention, a method for sensing, which includes inserting a pressure sensor into an ear of a user. Using the pressure sensor, pressure changes associated with subvocalization by the user are sensed and processed to generate a speech output.

[0028] The present invention will be more fully understood from the following detailed description of the embodiments thereof, taken together with the drawings in which:

[0029] BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Fig. 1 is a schematic pictorial illustration of a system for speech sensing, in accordance with an embodiment of the invention; Fig. 2 is a schematic pictorial illustration of a speech sensing device, in accordance with an embodiment of the invention;

[0031] Fig. 3 is a schematic side view of an interferometric sensor, in accordance with an embodiment of the invention;

[0032] Fig. 4 is a schematic frontal view of an interferometric sensor mounted on a spectacle frame, in accordance with an embodiment of the invention;

[0033] Fig. 5 is a schematic side view of a confocal chromatic sensor, in accordance with an embodiment of the invention; and

[0034] Fig. 6 is a schematic side view of an ellipsometric sensor, in accordance with an embodiment of the invention.

[0035] DETAILED DESCRIPTION OF EMBODIMENTS

[0036] The widespread use of mobile telephones in public spaces creates a cacophony of noise and often raises privacy concerns, since conversations are easily overheard by passersby. At the same time, when one of the parties in a telephone conversation is in a noisy location, the other party or parties may have difficulty in understanding what they are hearing due to background noise. Text communications provide a solution to these problems, but text input to a mobile telephone is slow and interferes with the users’ ability to see where they are going.

[0037] The above-mentioned PCT publications address these problems by using optical sensing of secondary laser speckle patterns cast on the face of a subject to detect minute movements of the skin surface and thus reconstruct the sequence of words articulated by the subject. By sensing fine movements of the skin (indicative of activation of subcutaneous nerves and muscles), occurring in response to words articulated by the subject with or without vocalization, a speech output is generated. The term “speech output” refers to modalities that are typically conveyed by speech, such as text, voice, identity, and emotions.

[0038] Further techniques of this sort are described in PCT International Publication WO 2024 / 018400, whose disclosure is also incorporated herein by reference. This publication describes systems for detecting and utilizing facial skin micromovements. In some non-limiting embodiments, the detection of the facial skin micromovements occurs using a speech detection system that may include a wearable housing, a light source (either a coherent light source or a noncoherent light source), a light detector, and at least one processor. One or more processors may be configured to analyze light reflections received from a facial region to determine the facial skin micromovements, and extract meaning from the determined facial skin micromovements. Examples of meaning that may be extracted from the determined facial skin micromovements may include words spoken by the individual (either silently spoken or vocally spoken), an identification of the individual, an emotional state of the individual, a heart rate of the individual, a respiration rate of the individual, or any other biometric, emotion, or speech-related indicator.

[0039] Using the systems and methods described in WO 2024 / 018400 and the other PCT publications cited above, facial skin micromovements may be detected during subvocalization. The term “during subvocalization” refers to any speech-related activity that takes place without utterance, before utterance, or preceding an imperceptible utterance. In one embodiment, the speech-related activity may include silent speech (i.e., when air flow from the lungs is absent but the facial muscles articulate the desired sounds). In another embodiment, the speech-related activity may include speaking soundlessly (i.e., when some air flows from the lungs, but words are articulated in a manner that is not perceptible using an audio sensor). In yet another embodiment, the speech-related activity may include prevocalization muscle recruitments (i.e., subvocalization that occurs prior to an onset of vocalization). In some cases, the prevocalization facial skin micromovements may be triggered by voluntary muscle recruitments that occur when certain craniofacial muscles start to vocalize words. In other cases, the prevocalization facial skin micromovements may be triggered by involuntary facial muscle recruitments that the individual makes when certain craniofacial muscles prepare to vocalize words.

[0040] Embodiments of the present invention that are described herein may similarly be applied in sensing fine movements of the skin during subvocalization, irrespective of lip movements by the speaker, by sensing light reflected from areas of the user’s body surface, such as the face, neck, or ear canal (including but not limited to the eardrum).

[0041] The term “light,” as used in the present description and in the claims, refers to electromagnetic radiation, which may be coherent or incoherent, in any or all of the infrared, visible, and ultraviolet ranges. In alternative embodiments, movements of the skin surface may be sensed using radiation in other spectral ranges, such as microwaves.

[0042] Fig. 1 is a schematic pictorial illustration of a system 18 for speech sensing, in accordance with an embodiment of the invention. System 18 is based on a sensing device 20, in which a bracket, in the form of an ear clip 22, fits over the ear of a user 24 of the device. An earphone 26 attached to ear clip 22 fits into the user’s ear. An optical sensing head 28 is connected by an arm 30 to ear clip 22 and thus is held in a location in proximity to the user’s face. In the pictured embodiment, device 20 has the form and appearance of a clip-on headphone, with optical sensing head 28 in place of (or in addition to) the microphone. Optical sensing head 28 directs one or more beams of light toward different, respective locations on the face of user 24. In the pictured embodiment, these beams create an array of spots 32 extending over an area 34 of the face (and specifically over the user’s cheek). In the present embodiment, optical sensing head 28 does not contact the user’s skin at all, but rather is held at a certain distance from the skin surface.

[0043] Optical sensing head 28 senses the light that is reflected from spots 32 the face and outputs a signal in response to the reflected light. Device 20 may sense and process the signals due to all of spots 32 or of only a certain subset of spots 32. In the various embodiments that are described hereinbelow, sensing head 28 implements a number of different sensing modalities:

[0044] • Interferometry, for example using miniature Michelson interferometers.

[0045] • Speckle interferometry.

[0046] • Confocal chromatic sensing.

[0047] • Ellipsometry.

[0048] • Pressure sensing.

[0049] Details of the structure and operation of optical sensing head 28 in these sensing modalities, including projection, detection, and processing of the detected radiation, are described below. These modalities may be implemented individually or in any suitable combination and may be used instead of or in conjunction with the modalities described in the above-mentioned PCT publications.

[0050] In alternative embodiments (not shown in the figures), an optical sensing head may direct light toward and sense light reflected from spots on other areas of the body, such as the neck or the ear canal. Specifically, fine motions in the ear canal, such as motions of the eardrum, can be indicative of contraction of facial muscles, as well as pressure changes due to movements of the mouth, and sensing these fine motions can be helpful in improving the fidelity of detection of subvocal speech.

[0051] Processing circuitry in system 18 processes the signal that is output by optical sensing head 28 to generate a speech output. As noted earlier, the processing circuitry is capable of sensing movements of the skin of user 22 and generating the speech output even without vocalization of the speech or utterance of any other sounds by user 22. The functions of the processing circuitry in system 18 may be carried out entirely within device 20, or they may alternatively be distributed between device 20 and an external processor, such as a processor in a smartphone 36 running suitable application software. Smartphone 36 may also access a server 38 over a data network, such as the Internet, in order to upload data and download software updates, for example.

[0052] In an alternative embodiment, sensing head 28 comprises optical components, such as lens arrays and / or metasurfaces, which optically process the light that is reflected from spots 32 to extract features of the speckles, for example using optical convolution techniques. The extracted features can be used as an input to an electronic image analysis or video analysis system.

[0053] In the pictured embodiment, device 20 also comprises a pressure sensor 35, which is connected to ear clip 22 and inserted into the user’s ear. For example, pressure sensor 35 may comprise a piezoelectric sensor or strain gauge, which may be integrated with earphone 26. Pressure sensor 35 senses pressure changes due to muscle movements (including micromovements) associated with subvocalization, including movements of the user’s tongue, for example. The sensed pressure changes can be used in conjunction with the signals output by optical sensing head 28 in generating the speech output from sensing device 20.

[0054] Fig. 2 is a schematic pictorial illustration of a speech sensing device 60, in accordance with another embodiment of the invention. In this embodiment, ear clip 22 is integrated with or otherwise attached to a spectacle frame 62. Additionally or alternatively, device 60 includes one or more additional optical sensing heads 68, similar to optical sensing head 28, for sensing skin movements in other areas of the user’s body surface. Because sensing heads 68 are located directly in front of the user’s cheeks, the optical axes of these sensing heads can be incident on the user’s skin at angles close to the normal, as opposed to the higher angles of incidence that are typical of optical sensing head 28. For this reason, the location of sensing heads 68 can be advantageous in the different sensing modalities listed above. These additional optical sensing heads 68 may be used together with or instead of optical sensing head 28.

[0055] In the embodiment of Fig 2, nasal electrodes 64 and temporal electrodes 66 are optionally attached to frame 62 and contact the user’s skin surface for the purpose of body surface electromyogram (sEMG) sensing. The sEMG measurement can be used in conjunction with the signals output by optical sensing heads 28 and / or 68 in generating a speech output Additionally or alternatively, a pressure sensor may be used to sense tongue movements as in the previous embodiment.

[0056] Fig. 3 is a schematic side view of sensing head 28, containing an interferometric sensor, in accordance with an embodiment of the invention. This sort of sensor can also be used, mutatis mutandis, in sensing head 68. In this interferometric sensor, a radiation source 70, such as an infrared laser diode, directs coherent light via a mirror 71 and a beamsplitter 72 toward area 34 of the subject’s face. Beamsplitter 72 splits off part of the illumination beam to serve as a reference beam, for example by reflection from a mirror 74 as in a classical Michelson interferometer. The light that is reflected from area 34 mixes with the reference beam to form an interference pattern in an imaging system, for example on a suitable image sensor 76. Small displacements of the subject’s skin surface cause shifts in the phase of the reflected light, which give rise to changes in the interference pattern. These changes are decoded to generate a speech output. The sensitivity of detection depends on the wavelength and the sampling rate, which should be in the multi-kilohertz range.

[0057] Although Fig. 3 shows only a single source of coherent radiation, in practice an array of sources and detectors can be used to sense local skin displacement at multiple locations over one or more areas of interest on the face, as illustrated in Fig. 1, or on another body surface. Furthermore, the sensor may use other interferometric configurations that are known in the art, such as a Mach-Zehnder configuration.

[0058] Fig. 4 is a schematic frontal view of sensing head 68, comprising an interferometric sensor mounted on spectacle frame 62, in accordance with another embodiment of the invention. Unlike the monostatic configuration of the sensor of Fig. 3, the present embodiment has a bistatic configuration, with a receiver 76 (containing a sensor with imaging optics) displaced transversely relative to a laser source 70. The reference beam in this case is transmitted from the source to the receiver via a different path, for example through a waveguide, such as an optical fiber 78.

[0059] Although only a single receiver is shown in Fig. 4, multiple receivers may alternatively be arrayed along spectacle frame 62. The beam output by a single laser source can be split to illuminate the skin in the fields of view of all the receivers and to provide respective reference beams to all the receivers.

[0060] In another embodiment, one or more of sensing heads 28 and / or 68 are configured as speckle interferometers. In speckle interferometry, the speckled image of an object under coherent irradiation, such as area 34 of the subject’s skin, is made to interfere with a reference beam. In the present embodiment, the reference beam is produced by the same laser that illuminates the skin. Any displacement of the skin surface then results in changes in the intensity distribution in the speckle pattern The speckle pattern on a small area of the skin is captured by a high-speed camera, and the camera images are processed to track the changes in the speckles. The use of coherent (interferometric) speckle sensing enhances the sensitivity of detection of small movements of the skin surface. Fig. 5 is a schematic side view of sensing head 28 incorporating a confocal chromatic sensor, in accordance with another embodiment of the invention. The confocal chromatic sensor uses the confocal principle and chromatic aberration to measure distances. The confocal principle allows the optical system to detect only a focused image. The objective optics of the sensor intentionally have a strong chromatic aberration. Thus, the focus of the optical system depends on the wavelength of the image, and the confocal principle chooses only a single wavelength to be in focus at each point in the image. A hyperspectral camera or other suitable spectrometer measures the wavelength of the sensed points on the skin surface captured by the sensor. By spectroscopically measuring changes in the wavelength, the sensor is able to sense local movements of the skin, which give rise to small changes in the distance from the sensor to each point on the skin.

[0061] In the pictured embodiment, a light source 80 projects broadband (white) light toward multiple points on area 34 of the skin surface. Objective optics 82 image the skin surface confocally onto an array of optical fibers 84, such that each fiber captures the confocal wavelength from a corresponding point on the skin surface. Fibers 84 output the captured light to a spectrometer 86, which measures the wavelength (and changes in the wavelength) at each point.

[0062] Fig. 6 is a schematic side view of sensing head 28 containing an ellipsometric sensor, in accordance with yet another embodiment of the invention. This sensor measures the changes in polarization of light that is reflected from the surface of the skin in response to movement of the skin surface. Because the skin is a Lambertian reflector, the present sensor is able to measure ellipsometric changes of diffuse reflections from the skin, with an angle of reflection that is not necessarily equal to the angle of incidence.

[0063] In the pictured embodiment, a light source 90 outputs a beam of polarized light toward area 34 of the skin surface. The light source may be coherent or broadband. If the light source is not inherently polarized, a polarizer 92 selects the desired linear polarization and, optionally, a compensator 94, such as a suitable quarter-wave plate, converts the polarization to elliptical. The light reflected from the skin is (optionally) recompensated back to linear polarization, and a polarization analyzer 98 selects the linear polarization component that is to be measured by a detector 100. The detector may sense the polarization of reflections from a single point, or it may comprise an array of detector elements, such as an image sensor, with suitable optics for imaging the skin surface onto the array. The intensity of the detected light, and hence the detector output signal, is indicative of the polarization rotation at each point that is sampled in this manner. When a detector array, such as an image sensor, is used in the ellipsometric sensor of Fig.

[0064] 6, the sensor is able to sense skin movement simultaneously at multiple points across a wide area of the face. This sort of sensor is capable of sampling skin movement at rates of up to 100 frames / sec, with a spatial resolution of 10 pm or less. The techniques that are described above may be used individually or in any combination, as well as in conjunction with other methods of sensing, image processing, feature extraction, and analysis that are described in the above-mentioned PCT publications. Machine learning techniques may be applied in fusing the data provided by multiple different types of sensors.

[0065] It will thus be appreciated that the embodiments described above are cited by way of example, and that the present invention is not limited to what has been particularly shown and described hereinabove. Rather, the scope of the present invention includes both combinations and subcombinations of the various features described hereinabove, as well as variations and modifications thereof which would occur to persons skilled in the art upon reading the foregoing description and which are not disclosed in the prior art.

Claims

CLAIMS1. A sensing device configured to fit on a head of a user, comprising: an optical sensing head held by the device in a location in proximity to a face of the user and comprising an emitter configured to direct coherent light toward multiple locations on a body surface of the user and an interferometric sensor configured to sense changes in a phase of the coherent light that is reflected from the multiple locations on the body surface; and processing circuitry configured to apply the sensed changes in the phase in generating a speech output.

2. The device according to claim 1, wherein the interferometric sensor comprises a speckle interferometer.

3. The device according to claim 1, wherein the emitter and the sensor direct the coherent light toward the body surface and receive the coherent light from the body surface in a monostatic configuration.

4. The device according to claim 1, wherein the emitter and the sensor direct the coherent light toward the body surface and receive the coherent light from the body surface in a bistatic configuration.

5. The device according to claim 4, and comprising a waveguide coupled to transmit a reference beam from the emitter to the sensor.

6. A sensing device configured to fit on a head of a user, comprising: an optical sensing head held by the device in a location in proximity to a face of the user and comprising an emitter configured to direct broadband light toward a body surface of the user and a confocal chromatic sensor configured to sense changes in a wavelength of the light that is reflected from the body surface and focused by the confocal chromatic sensor; and processing circuitry configured to apply the sensed changes in the wavelength in generating a speech output.

7. The device according to claim 6, wherein the confocal chromatic sensor is configured to sense the changes in the wavelength of the light that is reflected from multiple locations on the body surface.

8. The device according to claim 7, wherein the confocal chromatic sensor comprises: an array of optical fibers;confocal optics configured to image the body surface confocally onto the optical fibers, such that the fibers capture the light reflected at respective confocal wavelengths from corresponding points on the body surface; and a spectrometer coupled to the optical fibers and configured to sense the changes in the respective confocal wavelengths.

9. A sensing device configured to fit on a head of a user, comprising: an optical sensing head held by the device in a location in proximity to a face of the user and comprising an emitter configured to direct polarized light toward a body surface of the user and an ellipsometric sensor configured to sense changes in a polarization of the light that is reflected from the body surface; and processing circuitry configured to apply the sensed changes in the polarization in generating a speech output.

10. The device according to claim 9, wherein the ellipsometric sensor is configured to sense the changes in the polarization of the light that is reflected from multiple locations on the body surface.

11. The device according to any of claims 1-10, wherein the body surface from which the optical sensing head receives the reflected light comprises an area of the face of the user.

12. The device according to any of claims 1-10, wherein the body surface from which the optical sensing head receives the reflected light comprises an area of a neck of the user.

13. The device according to any of claims 1-10, wherein the body surface from which the optical sensing head receives the reflected light is in an ear canal of the user.

14. The device according to any of claims 1-10, and comprising: a pressure sensor configured to fit in an ear of the user of the device and to output a signal in response to pressure changes associated with subvocalization by the user, wherein the processing circuitry is configured to apply the signal in generating the speech output.

15. A sensing device, comprising: a pressure sensor configured to fit in an ear of a user of the device and to output a signal in response to pressure changes associated with subvocalization by the user; and processing circuitry configured to process the signal to generate a speech output.

16. The device according to claim 15, wherein the pressure sensor is configured to output the signal in response to the pressure changes due to movements of a tongue of the user.

17. The device according to claim 15 or 16, and comprising an earphone configured to fit in the ear of the user, wherein the pressure sensor is integrated with the earphone.

18. A method for sensing, comprising: mounting an optical sensing head in a location in proximity to a face of a user; directing coherent light from the sensing head toward multiple locations on a body surface of the user; interferometrically sensing changes in a phase of the coherent light that is reflected from the multiple locations on the body surface to the optical sensing head; and generating a speech output responsively to the sensed changes.

19. The method according to claim 18, wherein interferometrically sensing the changes comprises applying speckle interferometry to laser speckles generated on the body surface by the coherent light.

20. The method according to claim 18, wherein interferometrically sensing the changes comprises directing the coherent light toward the body surface and receiving the coherent light from the body surface in a monostatic configuration.

21. The method according to claim 18, wherein interferometrically sensing the changes comprises directing the coherent light toward the body surface and receiving the coherent light from the body surface in a bistatic configuration.

22. The method according to claim 21, and comprising transmitting a reference beam via a waveguide from an emitter of the coherent light to a receiver of the reflected coherent light.

23. A method for sensing, comprising: mounting an optical sensing head in a location in proximity to a face of a user; directing broadband light from the sensing head toward a body surface of the user; sensing changes in a wavelength of the light that is reflected from the body surface using a confocal chromatic sensor in the optical sensing head; and generating a speech output responsively to the sensed changes.

24. The method according to claim 23, wherein sensing the changes comprises applying the confocal chromatic sensor to sense the changes in the wavelength of the light that is reflected from multiple locations on the body surface.

25. The method according to claim 24, wherein the confocal chromatic sensor comprises: an array of optical fibers; confocal optics configured to image the body surface confocally onto the optical fibers, such that the fibers capture the light reflected at respective confocal wavelengths from corresponding points on the body surface; and a spectrometer coupled to the optical fibers and configured to sense the changes in the respective confocal wavelengths.

26. A method for sensing, comprising: mounting an optical sensing head in a location in proximity to a face of a user; directing polarized light from the sensing head toward a body surface of the user; sensing changes in a polarization of the light that is reflected from the body surface using an ellipsometric sensor in the optical sensing head; and generating a speech output responsively to the sensed changes.

27. The method according to claim 26, wherein sensing the changes comprises applying the ellipsometric sensor to sense the changes in the polarization of the light that is reflected from multiple locations on the body surface.

28. The method according to any of claims 18-27, wherein the body surface from which the optical sensing head receives the reflected light comprises an area of the face of the user.

29. The method according to any of claims 18-27, wherein the body surface from which the optical sensing head receives the reflected light comprises an area of a neck of the user.

30. The method according to any of claims 18-27, wherein the body surface from which the optical sensing head receives the reflected light is in an ear canal of the user.

31. The method according to any of claims 18-27, and comprising: inserting a pressure sensor into in an ear of the user; and sensing, using the pressure sensor, pressure changes associated with subvocalization by the user, wherein generating the speech output comprises applying the sensed pressure changes in producing the speech output.

32. A method for sensing, comprising: inserting a pressure sensor into an ear of a user; sensing, using the pressure sensor, pressure changes associated with subvocalization by the user; and processing the sensed pressure changes to generate a speech output.

33. The method according to claim 32, wherein sensing the pressure changes comprises sensing the pressure changes due to movements of a tongue of the user.

34. The device according to claim 32 or 33, wherein the pressure sensor is integrated with an earphone for insertion into the ear.