Data glasses with expression detection function
By using laser feedback interferometer sensors and VCSEL technology, the space and power consumption issues of existing facial expression detection systems are solved, and efficient and miniaturized facial expression detection is achieved.
Patent Information
- Application Number
- CN202480011733.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-02-10
- Filing Date
- 2024-01-16
- Publication Date
- 2025-09-19
AI Technical Summary
Existing camera-based facial expression detection systems have large structural space, high power consumption, and are subject to significant light exposure limitations.
A laser feedback interferometer (LFI) sensor is used to detect facial area information through laser radiation emission and reflection, and optical feedback interferometry is used to evaluate the reflected laser radiation. VCSEL and ViP technologies are combined to achieve miniaturization and low energy consumption.
It achieves robust detection of facial expressions, adapts to changes in head geometry of different users, and reduces the size and power consumption of the device.
Smart Images

Figure CN120677407A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a device, in particular a pair of data glasses, which are designed to be worn by a user of the device in a predetermined wearing position on the user's body, in particular on the user's head, and to a method for operating such a device.
[0002] Furthermore, the present disclosure relates to a communication system comprising such a device and a communication method. Background Art
[0003] Camera-based systems for detecting faces or facial areas are known from the prior art, wherein faces are detected using a camera and the detected image data are evaluated using image processing methods, for example based on neural networks, in order to identify so-called facial reference points in the detected image data and to derive information therefrom, for example information about the facial expression.
[0004] Such camera-based systems require a corresponding amount of space and have a relatively high power consumption. Furthermore, the operation of the cameras is restricted by the increased light exposure.
[0005] The object of the present disclosure is to provide a device of the type mentioned at the outset which overcomes the aforementioned disadvantages. Summary of the Invention
[0006] An embodiment relates to a device, in particular a pair of data glasses, which are configured to be worn by a user of the device on the user's body, in particular on the user's head, in a prescribed wearing state, wherein the device includes at least one laser feedback interferometer (LFI) sensor having at least one laser light source, in particular a laser diode, wherein the LFI sensor is arranged and configured to emit laser radiation into a reference area and detect a component of the reflection of the laser radiation at the device, the reference area being located in a first area of the face of the device user outside the eyes, and wherein the device is configured to derive information about the reference area based on the component of the reflection of the laser radiation, and wherein the device is configured to provide the information about the reference area for inserting a virtual target object, in particular an avatar representing the device user, in particular to another device.
[0007] According to the present disclosure, the LFI sensor is configured to emit laser radiation into a reference region. This is understood to mean that the LFI sensor illuminates the reference region at least partially, almost completely, or completely. Illumination of the reference region may occur, for example, by illuminating the surface of the reference region or by illuminating a plurality of discrete reference points within the reference region.
[0008] Whether by illuminating the surface of the reference area or by illuminating multiple points within the reference area, the device, in particular the detection of the reflected laser radiation and / or the derived information, can be made robust to movement of the glasses. Furthermore, the detection and deduction are more robust to different head geometries of different users.
[0009] A laser feedback interferometer (LFI) sensor is a sensor that is configured to emit laser radiation by means of a laser light source and to detect reflected laser radiation or a variable related thereto.
[0010] In a particularly advantageous manner, the evaluation of the backscattered and / or reflected radiation is performed based on optical feedback interferometry. The measurement principle of this method is preferably based on a method also known as self-mixing interferometry (SMI). Here, laser radiation is reflected at an object, for example at a reference area, and scattered or reflected back into the laser cavity generating the laser light. The returning light then interferes with the light beam generated in the laser cavity, i.e., primarily with the corresponding standing wave in the laser cavity, resulting in changes in the optical and / or electrical properties of the laser. This typically leads to intensity fluctuations in the laser output power. By analyzing these changes, information about the object, for example, the reference area where the laser radiation was reflected or scattered, can be obtained.
[0011] To facilitate understanding, the principle is first explained based on a single point, where the laser beam is scattered and / or reflected. If twice the distance between the LFI sensor and the object (e.g., the point where the radiation is scattered and / or reflected) is an integer multiple of the laser radiation wavelength, the scattered or reflected radiation is in phase with the radiation at the LFI sensor. This results in positive / constructive interference, which lowers the laser threshold and slightly increases the laser power. When the distance is slightly greater than an integer multiple, the two radiation waves experience a phase difference and negative interference occurs, reducing the laser output power. If the distance between the LFI sensor and the object (e.g., a reference point where the radiation is scattered and / or reflected) changes at a constant speed, the laser power fluctuates between a maximum value, which results in constructive interference, and a minimum value, which results in destructive interference. The resulting oscillation is a function of the speed of the object (e.g., the reference point) and the laser wavelength.
[0012] For example, the velocity of an object can be determined by analyzing the amplitude in the frequency domain. In a simple, exemplary case, when a reference point moves at a constant velocity relative to the LFI sensor and the LFI sensor's laser light source is unmodulated—that is, the wavelength and frequency of the laser light do not change over time—the peak frequency (also known as the center frequency) in the LFI sensor's amplitude / frequency spectrum is directly related to the velocity component in the direction of the beam.
[0013] In another simple example, when a reference point moves at a constant velocity relative to the LFI sensor, the peak frequency shifts upward or downward, or, correspondingly, left or right in the LFI sensor's amplitude / frequency spectrum. The direction of the shift depends on the modulation ramp used to operate the laser (rising or falling) and the direction of the velocity vector of the reference point relative to the LFI sensor (toward or away from the LFI sensor). Based on the distance between the peak frequencies, the direction and absolute value of the reference point's motion can be determined.
[0014] A similar effect occurs when an object, such as a reference point from which laser radiation is scattered and / or reflected, moves parallel to the emitted laser beam. Due to the Doppler effect, the backscattered laser light undergoes a frequency shift. At low velocities, this can be approximated as a phase shift of the backscattered laser light in the laser cavity, leading to positive and negative interference oscillations similar to the aforementioned effect, i.e., the formation of a beat frequency f. b (Also called beat frequency). Distance-dependent beat frequency f b It can be obtained by fast Fourier transform (FFT).
[0015] When irradiating a reference region, a superposition of multiple frequencies occurs in the components of the laser radiation reflected from the reference region. In principle, this can be viewed as reflections at multiple points within the reference region. This superposition results in what is known as a spectral distribution in the range spectrum. Similarly, there is a velocity spectrum that contains the Doppler frequency f d If different points within the reference area move at different speeds, superposition can also be detected in the spectrum. The Doppler frequency f in the spectrum d The superposition of is referred to as the velocity spectrum in the following text.
[0016] In this way, information about the reference area can be derived based on the range and velocity spectra of the reflected components of the laser radiation. This information about the reference area includes, for example, the positions of discrete reference points in the reference area, such as their absolute positions or their relative positions relative to the LFI sensor. Alternatively or additionally, changes in the positions of the reference points and / or the speed of these changes and / or changes in the speed can also be derived. Furthermore, the topology of the reference area relative to the LFI sensor and / or changes in the topology or the speed of these changes can be derived. "Topology" is understood to mean the position and / or arrangement of the reference area relative to the LFI sensor in terms of its geometry.
[0017] By means of information about the reference area, the shape of a certain area of the user's face can be detected and thereby expressions and / or changes in the shape of the area of the user's face and thereby movements and changes in the user's expressions can be detected.
[0018] In a preferred embodiment of the method or device according to the invention, a surface emitter is used as the laser diode. Surface emitters, also known as VCSELs (vertical cavity surface emitting lasers), have various advantages over edge emitters. Firstly, VCSELs require only a very small space, in particular the sensor structure space is less than 200x200 μm, so that such laser radiation generating units are particularly suitable for miniaturized applications. In addition, VCSELs are relatively cheap and have lower energy consumption compared to conventional edge emitters. For the measurement principle on which the method according to the invention is based and for the use of VCSELs in miniaturized applications, reference is made to the publication by Pruijmboom et al. "VCSEL-based miniature laser-Doppler inteferometer" (Proc. of SPIE Vol. 6908, 69080I-1-7).
[0019] In the case of vertical-cavity surface-emitting lasers (VCSELs), the mirror structure is constructed as a distributed Bragg reflector (DBR). On one side of the laser cavity, the DBR reflector has a transmission of approximately 1%, so the laser radiation can be coupled into free space.
[0020] In a particularly preferred embodiment of the method according to the invention, a surface emitter unit with an integrated photodiode or possibly a plurality of photodiodes is used, which is also referred to as a ViP (VCSEL, Vertical-Cavity-Surface-Emitting-Laser, integrated photodiode). The integrated photodiode allows immediate analysis of backscattered or reflected laser light that interferes with standing waves in the laser cavity. When manufacturing the corresponding surface emitter unit, the photodiode can be integrated directly into the production process of the laser diode (e.g., produced as a semiconductor component) during semiconductor processing.
[0021] In the case of ViP, the photodiode is located on the other side of the laser resonator, so it does not interfere with the free-space coupling. A unique feature of ViP is that the photodiode is directly integrated into the laser's lower Bragg reflector. Therefore, the size is primarily determined by the lens used, making it possible to achieve a laser / photodiode unit with dimensions of less than 2 x 2 mm. As a result, ViP can be integrated virtually invisibly into, for example, data glasses.
[0022] According to one embodiment, the LFI sensor includes at least one optical element configured to widen the laser beam emitted by the laser light source along at least one line. For example, the reference area extends along a line, so that the LFI sensor illuminates the reference area by widening the laser radiation at least along this line. A line is understood to be a line extending along a straight line. Alternatively, the line can also be a line that is bent or curved once or multiple times. Widening can also occur in two or more directions. For example, an illumination area of virtually any geometric shape can be generated. For example, the illumination area can include a rectangular or approximately rectangular shape, a circle, an ellipse, or any other shape. The optical element for widening the laser beam is, for example, a lens, in particular a cylindrical lens, or a diffractive optical element (DOE) or a holographic optical element (HOE).
[0023] According to another embodiment, illumination of the reference area is understood to mean that the LFI sensor at least partially illuminates the reference area, such that the laser beam emitted by the laser light source of the LFI sensor is split into at least two, in particular multiple, discrete partial beams, thereby illuminating at least two discrete reference points within the reference area. According to one embodiment, the LFI sensor is configured to illuminate at least two discrete reference points within the reference area. At each of these reference points, light is scattered and / or reflected, so that the scattered and / or reflected components of the partial beams are detected by the LFI sensor, resulting in a superposition in the distance spectrum and / or velocity spectrum. In principle, multiple discrete reference points within the reference area can be scanned in this manner. For example, a diffractive optical element (DOE) or a holographic optical element (HOE) can be provided to split the laser beam into discrete partial beams. Alternatively, the separation can also be performed by scanning the reference points during the scanning process. As a scanner, for example, a microscanner comprising a movable single mirror (also called MEMS), a surface light modulator (also called SLM) comprising a mirror matrix, a reflective system based on LCoS (Liquid Crystal on Silicon) technology, or an optical phase shifter can be used.
[0024] According to one embodiment, the device is configured such that a discrete reference point can be illuminated according to at least one first illumination pattern and at least one second illumination pattern, and can be switched between the first and second illumination patterns. For example, a diffractive optical element (DOE) or a holographic optical element (HOE) can be switched between different states using electronically controllable liquid crystals. Using a scanner, different illumination patterns can also be generated through corresponding manipulation.
[0025] According to an advantageous embodiment, the device comprises at least two or more laser feedback interferometers (LFI) sensors. In this case, for example, two or more LFI sensors can be provided, which are arranged and configured to emit laser radiation towards two or more reference areas in the first area of the face and / or in other areas of the face.
[0026] The first area and / or at least one other area is an area in the user's face, such as the cheek area, in particular in the right or left half of the face, or the eyebrow area, in particular in the right or left half of the face, or the nose area, or the chin area, or the mouth area, or the eyelid area, in particular in the right or left half of the face.
[0027] The apparatus may include, for example, at least one control device or a plurality of control devices for controlling at least one LFI sensor or a plurality of LFI sensors. Control may include, for example, switching the LFI sensor and / or its laser light source on and off. For example, the laser light source of an LFI sensor, particularly an LFS sensor, may be manipulated to emit laser radiation having a modulated frequency.
[0028] The device comprises, for example, at least one computing device for deriving information about the reference region or the reference region based on a component of the reflection of the laser radiation. A plurality of computing devices can also be provided.
[0029] The device is configured to provide information about the reference region for inserting a virtual target object, particularly an avatar representing a user of the device, particularly to another device. The device itself may also be configured to insert a virtual target object. The device may also be configured to predefine an expression for the virtual target object based on the reference region or the information about the reference region. For example, the virtual target object may be a virtual representation of the user's face.
[0030] The insertion of virtual objects can be accomplished using various techniques. For example, the virtual objects can be projected, for example, using a laser, onto a display or onto the lenses of the device, particularly data glasses. It is also conceivable to project the virtual objects directly into the field of view of the device user, onto the retina of the user's eye.
[0031] According to one embodiment, the device can be configured to provide the reference region or information about the reference region to another device. For example, the device comprises a suitable communication interface for this purpose.
[0032] According to one embodiment, the device is configured to receive data from another device, particularly data glasses, of another user. The data received may be received, for example, via a suitable communication interface. The data to be received includes information about at least one reference region within at least one first region of the other user's face. The device is configured to insert a virtual target object, particularly an avatar representing the other user, based on the data from the other device. Insertion may be performed, for example, by projecting the virtual target object, for example, onto a display or device lens, or directly onto the retina of the device user's eye. The information about the reference region includes information about the positions of discrete reference points within the reference region and / or information about their changes and / or information about the topology of the reference region and / or information about their changes. Based on this information, a shape and thus an expression of a certain region of the user's face can be detected, and / or changes in the shape of the region of the user's face and thus motion and changes in the user's expression within the region of the user's face can be detected. The device is configured to predetermine an expression for the virtual target object based on the data from the other device. The virtual target object may be, for example, a virtual representation of the other user's face.
[0033] According to one embodiment, the device is configured such that predefining an expression of the virtual target object includes modulating at least one spline in the first region of the virtual target object based on information about at least one reference region of the first region. For example, the position of an anchor point can be modulated based on the position of a discrete reference point in the reference region, and the spline can be modulated based on the anchor point. The anchor points serve, for example, as nodes of the spline. Alternatively, the position of the anchor point can be determined or modulated based on a distance spectrum and / or velocity spectrum determined based on components of reflected laser radiation detected by means of at least one LFI sensor, or the spline can be determined or modulated directly.
[0034] Other embodiments relate to a method for operating a device according to an embodiment. The method comprises the following steps:
[0035] emitting laser radiation into at least one reference area in a first area of a face of a user of the device and detecting a reflected component of the laser radiation;
[0036] deriving information about the reference area based on the components of the reflection of the laser radiation,
[0037] Information about the reference area is provided for inserting a virtual target object, in particular an avatar representing a user of the device, in particular to another device.
[0038] The laser radiation is emitted at a modulation frequency. The information about the reference region is derived as described above based on the distance spectrum and / or velocity spectrum ascertained from the reflected components of the laser radiation detected by means of the at least one LFI sensor.
[0039] According to one embodiment, the method includes: receiving data from another device of another user, in particular data glasses, wherein the data includes information about at least one reference area in a first area of the face of the other user, and inserting a virtual target object, in particular an avatar representing the other user, on the device based on the data of the other device, and presetting an expression of the target object based on the data of the other device.
[0040] According to one embodiment, predefining the expression of the avatar includes modulating at least one spline in at least one first area of the virtual target object based on information about a reference area of the first area.
[0041] According to one embodiment, it is provided that the respective reference region is assigned or associated with at least one spline of the virtual target object, and the respective spline is modulated based on information about the respective reference region.
[0042] According to one embodiment, a range spectrum and / or velocity spectrum determined based on components of reflected laser radiation detected by at least one LFI sensor is provided as input data to at least one trained neural network, and information about a reference region or a spline of a virtual target object assigned to the corresponding reference region is derived from the range spectrum and / or velocity spectrum by the trained neural network, and the spline of the virtual target object is modulated based on the derived information. Alternatively, at least one spectrum is detected by at least two LFI sensors, each of which is assigned to a reference point in a region of the user's face, and the LFI sensor spectra are arranged one above the other and / or side by side in a matrix along a frequency axis and / or a time axis as input data to the trained neural network, and information about the corresponding reference region is derived from the matrixed spectra by the trained neural network.
[0043] According to one specific embodiment, the method includes a training phase for adapting a trained neural network to a user of the device, wherein the neural network is adapted to the user using a camera, wherein image data recorded with the camera are used as labeled training data for adapting the neural network.
[0044] Further embodiments relate to a communication system comprising at least one first device and at least one further device, wherein the first device and the further device are constructed according to any one of claims 1 to 6, and wherein the first and the further device are configured to perform the method according to any one of claims 7 to 12.
[0045] Further embodiments relate to a method for communication between at least two users via a communication network comprising a communication system according to claim 13, wherein a first user is connected to the communication network via a first terminal device and a second user via a second terminal device, wherein the first terminal device and the second terminal device are each a device according to any one of claims 1 to 6, and wherein at least the first user is inserted into the second terminal device of the second user as a virtual target object, in particular an avatar representing the first user.
[0046] Further advantages are apparent from the description and drawings. Exemplary embodiments of the present invention are illustrated in the drawings and explained in detail in the following description. Identical reference numerals in different figures denote identical or at least functionally comparable elements. When describing a corresponding figure, reference may also be made to elements in other figures. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] They are shown in schematic form:
[0048] Figure 1 An embodiment showing a device in a prescribed wearing state on a user's head is shown;
[0049] Figure 2 Show the basis Figure 1 a fragment of the device;
[0050] Figure 3 shows a fragment of a device according to another embodiment;
[0051] Figure 4 shows an exemplary amplitude spectrum / frequency domain spectrum of an LFI sensor when the laser light source of the LFI sensor is moved without being modulated;
[0052] Figure 5 shows an exemplary amplitude spectrum / frequency domain spectrum of an LFI sensor when moving with its laser light source modulated;
[0053] Figure 6 a) shows an example of a ramp-shaped modulation of the current of a laser light source for operating an LFI sensor;
[0054] Figure 6 b) Example graph showing the output power of a laser light source; and the corresponding frequency shift
[0055] Figure 7 shows the different components of a device,
[0056] Figure 8 shows the functional principle of an LFI sensor for deriving information from reflected laser radiation;
[0057] Figure 9 Show the basis Figure 1 a device and a virtual object inserted, for example, by means of the device or by means of another device;
[0058] Figure 10 Show Figure 8 fragments;
[0059] Figure 11 An exemplary embodiment is shown with schematically illustrated neural network input data;
[0060] Figure 12 Another embodiment of a schematically illustrated neural network input data is shown;
[0061] Figure 13 A communication system for performing a communication method with at least two users. DETAILED DESCRIPTION
[0062] Figure 1 A device 10, in particular a pair of data glasses, is shown, which is worn by a user 12 of the device 10 in a prescribed wearing position shown by way of example on the head of the user 12. According to the embodiment shown, the data glasses 10 comprise a frame 14 containing two spectacle lenses 15 and two temples 16.
[0063] The device 10 comprises at least one laser feedback interferometer, LFI, having at least one laser light source, and a sensor 18. In the example shown, a plurality of LFI sensors 18 are shown, four in total.
[0064] The LFI sensor 18 is arranged and configured to emit laser radiation into a reference area 20, which is located in an area 22, 24 of the face of the user 12 of the device 10 outside the eyes, and to detect a reflected component of the laser radiation at the device 10. The laser radiation is schematically represented by a dashed line.
[0065] In the example, one LFI sensor 18 is provided on each half of the face and is configured to emit a laser beam toward a first area 22 of the face of the user 12, in the example, the eyebrow area. In the example, another LFI sensor 18 is provided on each half of the face and is configured to emit a laser beam toward another area 24 of the face of the user 12, in the example, the cheek area. Exemplary areas include the cheek area, particularly in the right or left half of the face, the eyebrow area, particularly in the right or left half of the face, the nose area, the chin area, the mouth area, or the eyelid area, particularly in the right or left half of the face.
[0066] The laser feedback interferometer (LFI) sensor 18 is a sensor configured to emit laser radiation by means of a laser light source and to detect reflected laser radiation or a variable related thereto.
[0067] In a particularly advantageous manner, the evaluation of backscattered and / or reflected radiation is performed based on optical feedback interferometry. The measurement principle of this method is preferably based on a method also known as self-mixing interferometry (SMI). Here, a laser beam is reflected at an object and scattered or reflected back into the laser cavity that generates the laser light. The returned light then interferes with the beam generated in the laser cavity, i.e., primarily with the corresponding standing wave in the laser cavity, resulting in changes in the optical and / or electrical properties of the laser. This typically results in intensity fluctuations in the laser output power. By analyzing these changes, information about the object on which the laser beam was reflected or scattered can be obtained.
[0068] To facilitate understanding, the principle is first explained based on a single reference point, for example a discrete reference point located in a reference area.
[0069] If twice the distance between the LFI sensor 18 and the object, such as a reference point in the reference area 20 where radiation is scattered and / or reflected, is an integer multiple of the laser radiation wavelength, the scattered or reflected radiation is in phase with the radiation at the LFI sensor. This results in positive / constructive interference, which lowers the laser threshold and slightly increases the laser power. When the distance is slightly greater than an integer multiple, the two radiation waves experience a phase difference and negative interference occurs, reducing the laser output power. If the distance between the LFI sensor 18 and the object, such as a reference point where radiation is scattered and / or reflected, changes at a constant speed, the laser power fluctuates between a maximum value, which results in constructive interference, and a minimum value, which results in destructive interference. The resulting oscillation is a function of the speed of the object, such as the reference point 20, and the laser wavelength.
[0070] For example, the velocity of an object can be determined by analyzing the amplitude in the frequency domain. In a simple, exemplary case, when a reference point moves at a constant velocity relative to the LFI sensor and the LFI sensor's laser light source is unmodulated—that is, the wavelength and frequency of the laser light do not change over time—the peak frequency (also called the center frequency) in the LFI sensor's amplitude spectrum / frequency domain spectrum is directly related to the velocity component in the direction of the beam. Figure 4 An exemplary amplitude / frequency domain spectrum of the LFI sensor 18 is shown when a reference point moves at a constant speed relative to the LFI sensor and the laser light source of the LFI sensor is not modulated. Figure 4 The amplitude A is shown as a function of frequency f. The peak frequency (also called center frequency) f1 is directly related to the velocity component in the direction of the beam.
[0071] Alternatively, the current of the laser light source operating the LFI sensor 18 can be ramp-modulated, thereby modulating the wavelength of the laser radiation. When the distance between the LFI sensor 18 and the object, such as a reference point where the laser radiation is scattered and / or reflected, is fixed, this also changes the number of wavelengths that "fit" into the optical path, resulting in the aforementioned oscillation-temporal interference pattern. If the optical radiation power is now detected, for example, using a photodiode, the intensity variation of the backscattered laser power can be inferred from the amplitude variation of the radiation power. By analyzing the number of oscillations, for example by counting zero crossings or maximum values, or by calculating the Fourier spectrum using an FFT and analyzing the amplitude in the frequency domain, the number of oscillations, i.e., the number of constructive and destructive interference passes, can also be determined. Thus, given a known laser wavelength, the distance between the LFI sensor 18 and the object, such as the reference point where the laser radiation is scattered and / or reflected, can be determined.
[0072] Figure 5 shows an example amplitude spectrum for a modulated operation. If there were no object movement, the spectrum would look like Figure 4 shown. Figure 4 The spectrum shown is also called the range spectrum. If the object reflecting the laser radiation moves at a constant speed, the spectrum will be as follows Figure 5 shown. Figure 5 The spectrum with the Doppler frequencies superimposed is shown. The superposition of Doppler frequencies is also called the velocity spectrum. In this case, the distance between the LFI sensor and the object can be determined from the peak frequencies f1, f1'. If the object moves further, the peak frequencies f1, f1' will shift upwards or downwards, see Figure 4 The offset direction depends on the modulation ramp used to operate the laser (rising or falling) and the direction of the object's velocity vector relative to the LFI sensor (towards or away from the LFI sensor). Figure 5 Two spectra of a falling and rising modulated ramp (left and right) are shown. Based on the distance a between the peak frequencies, the direction and absolute value of the object's motion can be determined.
[0073] A similar effect occurs when an object, such as a reference point from which laser radiation is scattered and / or reflected, moves parallel to the emitted laser beam. Due to the Doppler effect, the backscattered laser light undergoes a frequency shift. At low velocities, this can be approximated as a phase shift of the backscattered laser light in the laser cavity, leading to positive and negative interference oscillations similar to the aforementioned effect, i.e., the formation of a beat frequency f. b , also called beat frequency. b It is proportional to the speed of the moving object, for example the reference point where the laser radiation is scattered and / or reflected. The speed of light c0, the angle a between the laser beam and the motion vector, and the excitation laser frequency f0 are known. The beat frequency (also called beat frequency) fb , using the formula f b =2v / c0*f0 cos(a) determines the velocity of the object.
[0074] Distance-dependent beat frequency f b The triangular modulated output power of the laser can be obtained by FFT, see Figure 8 .
[0075] Doppler frequency f d The triangular modulated output power of the laser can also be obtained by FFT, see Figure 8 .
[0076] Time signal, measured photocurrent I p (t) is segmented according to the modulating signal (here a triangle ramp). In the example, these segments are called segments T up and Section T down , see Figure 8 c).
[0077] Then, the time series data I p (T) corresponding segment T up and Section T down Apply FFT, see Figure 8 d).
[0078] f b and f d The calculation is performed, for example, according to the following mathematical relationship:
[0079] and Among them, through and
[0080] The velocity v of the reference point 20 can be determined T Or the distance L from the reference point 20 to the LFI sensor 18 ext .
[0081] When irradiating the reference region 20, a superposition of multiple frequencies occurs in the components of the reflected laser radiation from the reference region 20. In principle, this can be viewed as reflections at multiple points within the reference region 20. This superposition results in what is known as a spectral distribution in the range spectrum. Similarly, there is a velocity spectrum that contains the Doppler frequency f d If different points within the reference region 20 move at different speeds, superposition can also be detected in the spectrum. Doppler frequency f in the spectrum (Spektrum) d The superposition of is referred to as the velocity spectrum in the following text.
[0082] In this way, information about reference area 20 can be derived based on the distance and velocity spectra of the reflected components of the laser radiation. This information about reference area 20 includes, for example, the positions of discrete reference points in reference area 20, such as the absolute positions of the reference points or their relative positions relative to LFI sensor 18. Alternatively or additionally, changes in the positions of the reference points and / or the speed of these changes and / or changes in speed can also be derived. Furthermore, the topology of reference area 20 relative to LFI sensor 18 and / or changes in the topology or the speed of these changes can be derived. Topology is understood to mean the geometric position and / or arrangement of reference area 20 relative to LFI sensor 18.
[0083] By means of information about the reference area 20, a shape of a certain area of the user's face and thereby an expression and / or a change in the shape of the area of the user's face and thereby a movement in the area of the user's face and a change in the user's expression can be detected.
[0084] In a preferred embodiment of the method or device according to the invention, a surface emitter is used as the laser diode. Surface emitters, also known as VCSELs (vertical cavity surface emitting lasers), have various advantages over edge emitters. Firstly, VCSELs require only a very small space, in particular the sensor structure space is less than 200 x 200 μm, so that such laser radiation generating units are particularly suitable for miniaturized applications. In addition, VCSELs are relatively cheap and have lower energy consumption compared to conventional edge emitters. For the measurement principle on which the method according to the invention is based, and for the use of VCSELs in miniaturized applications, reference is made to the publication by Pruijmboom et al. "VCSEL-based miniature laser-Doppler inteferometer" (Proc. of SPIE Vol. 6908, 69080I-1-7).
[0085] In the case of vertical-cavity surface-emitting lasers (VCSELs), the mirror structure is constructed as a distributed Bragg reflector (DBR). On one side of the laser cavity, the DBR reflector has a transmission of approximately 1%, so the laser radiation can be coupled into free space.
[0086] In a particularly preferred embodiment of the method according to the invention, a surface emitter unit with an integrated photodiode or possibly a plurality of photodiodes is used, which is also referred to as a ViP (VCSEL, Vertical Cavity Surface Emitting Laser, Integrated Photodiode). The integrated photodiode allows immediate analysis of backscattered or reflected laser light that interferes with standing waves in the laser cavity. When manufacturing the corresponding surface emitter unit, the photodiode can be integrated directly during the production of the laser diode (e.g., as a semiconductor component) during semiconductor processing.
[0087] In the case of ViP, the photodiode is located on the other side of the laser resonator, so it does not interfere with the free-space coupling. A unique feature of ViP is that the photodiode is directly integrated into the laser's lower Bragg reflector. Therefore, the size is primarily determined by the lens used, making it possible to achieve a laser / photodiode unit with dimensions of less than 2 x 2 mm. As a result, ViP can be integrated virtually invisibly into, for example, data glasses.
[0088] Figure 2 show Figure 1 Detailed view of the Figure 1 fragment.
[0089] according to Figure 2 , the LFI sensor 18-1 is arranged on the upper part 14a of the eyeglass frame 14 or is integrated into the eyeglass frame 14. In the example, the LFI sensor 18-1 is arranged, aligned and constructed such that the laser radiation is emitted into a reference area 20, which is located in a first area 22 of the face of the user 12 outside the eyes, i.e., in the area of the eyebrows.
[0090] As described above, information about the reference area 20 can then be derived based on the reflected laser radiation detected by means of the LFI sensor 18-1. Based on the derived information, movements in this area of the user's face can be detected and thus the user's expression can be inferred. For example, the movement and / or position of the eyebrows can be detected. Movements or changes in the position of the eyebrows are caused, for example, by frowning, pulling the forehead up or down, blinking, squinting, closing or partially closing the eyelids, or movements of the upper facial muscles. Thus, based on the detected movement and / or position of the eyebrows, one of the aforementioned expressions and / or possibly other facial expressions can be inferred. This can also be done with the help of a neural network. Reference will be made later to Figure 10 and Figure 11 Provide explanation.
[0091] according to Figure 2Another LFI sensor 18-2 is arranged on the lower portion 14a of the eyeglass frame 14 or is integrated into the eyeglass frame. In the example, the LFI sensors 18-2, 18-5, 18-6 are arranged, aligned, and configured such that laser radiation is emitted into a reference area 20 located in a second area 24 of the face of the user 12 outside the eyes, i.e., in the cheek area.
[0092] As described above, information about reference region 20 can then be derived based on the reflected laser radiation detected by LFI sensor 18-2. Based on this derived information, movement in this region of the user's face can be detected and, thus, the user's facial expression can be inferred. For example, movement in the cheek region can be detected. Movement or position changes in the cheek region can be caused, for example, by smiling, squinting, closing the eyes, pursing the lips, speaking, wrinkling the nose, and other movements of the lower facial muscles. Thus, based on the detected movement and / or position in the cheek region, one of the aforementioned facial expressions and / or possibly other facial expressions can be inferred. This can also be performed using a neural network.
[0093] According to the present disclosure, the LFI sensor 18 is configured to emit laser radiation into the reference area 20. This is to be understood as meaning that the LFI sensor 18 illuminates the reference area 20 at least partially, almost completely, or completely. In order to be able to illuminate the reference area 20, it is provided, for example, that the LFI sensor 18 includes at least one optical element, which is configured to widen the laser beam emitted by the laser light source of the LFI sensor 18 at least along one line, see Figure 2 . In the example, the reference area extends along a line 21, so that the LFI sensor illuminates the reference area by widening the laser radiation at least along this line 21. A line is understood to be a line that extends along a straight line. Alternatively, the line can also be a line that is bent or curved once or multiple times. The widening can also take place in two or more directions. For example, an irradiation area of almost any geometric shape can be produced. For example, the irradiation area can include a rectangular or approximately rectangular shape, a circle, an ellipse or any other shape. The optical element for widening the laser beam is, for example, a lens, in particular a cylindrical lens, or a diffractive optical element (DOE for short), or a holographic optical element (HOE for short).
[0094] According to another embodiment, the illumination of the reference region 20 is understood to mean that, see Figure 3, the LFI sensor 18 at least partially illuminates the reference area 20, specifically in such a way that the laser beam emitted by the laser light source of the LFI sensor 18 is split into at least two, in particular a plurality of, discrete partial beams, thereby illuminating at least two discrete reference points in the reference area. At each such reference point, the light is scattered and / or reflected, so that the scattered and / or reflected components of the partial beam are detected by the LFI sensor, resulting in a superposition in the distance spectrum and / or the velocity spectrum. For example, in order to split the laser beam into discrete partial beams, a diffraction optical element (DOE for short) or a holographic optical element (HOE for short) can be provided. Alternatively, the separation can also be performed by scanning the reference points during the scanning process. As a scanner, for example, a microscanner comprising a movable single mirror (also called MEMS), or a surface light modulator (also called SLM) comprising a mirror matrix, or a reflective system based on LCoS (Liquid Crystal on Silicon) technology, or an optical phase shifter can be used. In Figure 3 As shown in FIG. 1 , for example, a laser beam emitted by the LFI sensor 18 - 1 is split into three discrete partial beams, wherein the partial beams illuminate three discrete reference points 20 - 1 , 20 - 2 , 20 - 3 within the reference area 20 .
[0095] It can also be provided that the discrete reference points can be illuminated according to at least one first illumination pattern and at least one second illumination pattern. For example, the LFI sensor 18 can be operated in such a way that it is possible to switch between the first illumination pattern and the second illumination pattern. For example, a diffractive optical element (DOE) or a holographic optical element (HOE) can be switched between different states by means of electronically controllable liquid crystals. Different illumination patterns can also be generated by corresponding control using a scanner. Figure 3 For example, it is shown that, by means of the LFI sensor 18-2, according to the first illumination pattern B1 (in Figure 3 The three reference points are illuminated according to the second illumination pattern B2 (marked with circles in Figure 3 The three reference points are irradiated (marked with a cross in FIG). Irradiation according to the irradiation patterns B1 and B2 does not have to be performed simultaneously.
[0096] The LFI sensors 18 are, for example, almost completely or completely recessed into the frame 14 of the device 10. The optics of the LFI sensors (not shown in detail) can be applied to the frame from the outside or recessed into the frame together with the LFI sensors 18. For convenient installation, it can be provided that the respective LFI sensors 18 are arranged on a common circuit board arrangement 26, such as a printed circuit board, a flexible printed circuit board, or a flexible line. In the example, the LFI sensors are connected to central electronics 28 via corresponding connections, such as a flexible line, a printed circuit board, or another line. In the example, the central electronics is arranged in one leg 16 of the device 10. Other arrangements within the device 10 are also possible. The central electronics 28 is or includes, for example, a computing device and / or a control device.
[0097] In the example, signal processing is performed in the central electronics 28. The signal processing includes, for example, controlling and / or activating the LFI sensors 18 and / or the respective LFI sensors 18, 18-1 to 18-4 and / or generating drive signals for the LFI sensors 18 and / or generating modulation signals for the LFI sensors 18 and / or reading interference signals detected by the respective LFI sensors 18 and / or converting the interference signals of the respective LFI sensors 18, for example by means of an FFT, deriving information about the reference region 20 and, if necessary, determining information for modulation anchor points and / or splines.
[0098] exist Figure 7 1 shows a central electronic device 28 and an LFI sensor 18 in the form of a VIP connected thereto. In this example, the LFI sensor 18 comprises an optical device 30, such as a collimating lens. Figure 6 As shown, central electronics 28 includes components assigned to a digital domain 32 and components assigned to an analog domain 34 .
[0099] In the example, the digital domain 34 is implemented as an application-specific integrated circuit 36 (ASIC). Other implementations are also conceivable. In the example, the circuit 36 includes a D / A (digital / analog) converter 38, an A / D (analog / digital) converter 40, as well as an exemplary segmentation component 42, an exemplary FFT component 44 for converting the interference signal, and a component 46 for deriving information about the reference region from the detected laser radiation.
[0100] In the example, this is simplified to deriving the distance d(t) between the LFI sensor 18 and the reference point 20 - 1 and the speed v(t) at which the reference point moves.
[0101] The analog domain 34 includes, in this example, a driver 48 for the laser light source of the LFI sensor 18 and a driver 48 for amplifying the photocurrent I of the photodiode of the LFI sensor 18 . p (t) Photodiode amplifier 50.
[0102] Functional principle reference of segmentation component 42 and FFT component Figure 7 (New) explanation.
[0103] Figure 7 a) exemplarily shows an LFI sensor 18 and a reference point 20. In the example, the LFI sensor 18 comprises an integrated photodiode.
[0104] In the example, the current I of the laser light source operating the LFI sensor 18 is modulated in a ramp-shaped manner. The photocurrent I is detected by means of the LFI sensor. p (t), see Figure 7 b).
[0105] The segmentation component 42 modulates the time signal (measured photocurrent I) according to the triangular ramp of the modulation signal. p (t)) is divided into segments, which are called segment T in the example. up and Section T down , see Figure 7 c).
[0106] Then, the time series data I p (T) corresponding segment T up and Section T down Apply FFT, see Figure 7 d). In the example, this step is performed by the FFT component 44.
[0107] f b and f d The calculation is performed, for example, according to the following mathematical relationship:
[0108] and Among them, through and
[0109] The velocity v of the reference point 20-1 can be determined T Or the distance L from the reference point 20 - 1 to the LFI sensor 18 ext .
[0110] Signal processing can also be decentralized. For example, multiple computing devices and / or control devices can be provided. For example, the corresponding LFI sensor 18 can include its own computing device and / or its own control device. This allows the corresponding signal processing, such as manipulation and / or modulation and / or conversion, to be performed near the sensor.
[0111] The components of the digital domain 32 and the components of the analog domain 34 can be implemented either in the central electronics or separately in the respective LFI sensors.
[0112] The device 10 is configured to provide information about the reference area 20 for inserting a virtual target object 12', in particular an avatar representing the user 12 of the device 10. Figure 7 In an embodiment of the present invention, the device is exemplarily configured to insert a virtual target object 12'. The device 10 can also be configured to predefine an expression of the virtual target object 12' based on information about the reference area 20 or the reference area 20. In this example, the virtual target object 12' is a virtual representation of the face of the user 12.
[0113] The virtual target object 12' can be inserted using various techniques. For example, the virtual target object 12' can be projected, for example, using a laser, onto a display or onto the lens of the device 10, particularly data glasses. It is also conceivable to project the virtual target object 12' directly into the field of view of the user of the device 10, onto the retina of the user's 12 eye.
[0114] In the example, the virtual target object 12' includes a plurality of "virtual" anchor points 20'. The respective anchor point 20'-1 of the virtual target object 12' can, but need not, be associated with or linked to a reference point 20-1 on the face of the user 12, for example. However, one anchor point 20' can also be associated with a plurality of reference points 20. One or more anchor points 20' can also be associated with a reference area 20. For example, if a movement is detected at a reference point or reference area 20 by the LFI sensor 18, the movement can be reproduced in the virtual target object 12' by modulating the respective anchor point 20' or multiple anchor points 20'. Modulating an anchor point 20' is, for example, to be understood as a change in the position of the anchor point 20'. Alternatively, the virtual target object 12' can also include at least one or more splines 52. The splines are, for example, assigned to or linked to the reference area 20.
[0115] Figure 10 Exemplary display Figure 9 Information about the reference region 20 is derived by means of the sensor 18 - 1 , for example by determining the distances d1 ( t ), d2 ( t ) and d3 ( t ) to the reference points 20 - 1 , 20 - 2 , 20 - 3 based on the reflected laser radiation.
[0116] For example, this information is used to modulate anchor points 20 ′- 1 , 20 ′- 2 , and 20 ′- 3 . To determine the positions of anchor points 20 ′- 1 , 20 ′- 2 , and 20 ′- 3 , for example, a geometric model can be used. For example, the positions of anchor points 20 ′- 1 , 20 ′- 2 , and 20 ′- 3 in virtual target object 12 ′ relative to the position of device 10 can be determined based on d1(t), d2(t), and d3(t) and the known positions of the corresponding LFI sensors 18 relative to device 10 .
[0117] In the example, a spline 52 can be calculated based on the determined positions of the anchor points 20 ′- 1 , 20 ′- 2 , and 20 ′- 3 , so as to modulate the eyebrows of the virtual target object 12 ′ in a targeted manner, for example. The anchor points are, for example, nodes of the spline 52 .
[0118] Alternatively, the position of the anchor point may also be determined from the range spectrum and the speed spectrum determined by the LFI sensor 18 or the spline may be determined directly.
[0119] In the example, this is done by using a trained neural network 54. The range spectrum and the speed spectrum ascertained by the LFI sensor 18 are supplied as input data to the neural network.
[0120] For one LFI sensor 18, the corresponding determined distance spectrum and velocity spectrum are provided as input data to the neural network. If multiple sensors are used, the distance spectrum is advantageously arranged in a matrix along the frequency axis f and the velocity spectrum is provided along the time axis t as input data, see for example Figure 11 The resolution of the time axis t corresponds to the modulation period of the triangular modulation. With the aid of the trained neural network 54, information about the reference region or about one or more anchor points and / or splines of the virtual target object assigned to the respective reference region is derived from the matrix-arranged spectra S1, S2, S3.
[0121] In this example, the neural network is a CNN (Convolutional Neural Network) model. In this example, three CNN layers 56 are shown. Based on the measured distance and velocity spectra arranged in a matrix, features are extracted using the CNN model and then input to an optional Fully Connected layer 58. Information about reference points, anchor points of virtual objects, and / or splines can be derived from the extracted features via a regressor 60.
[0122] Based on the derived information, the expression of the target object can then be modulated, for example by modulating the anchor points and / or directly modulating the splines.
[0123] Using the neural network 54 , the feature extraction can be pre-trained on a correspondingly large dataset suitable for this purpose, so that the detection is robust to different facial shapes and positions of the glasses and / or illumination points of the sensor.
[0124] According to one embodiment, the method includes a training phase for adapting the trained neural network 54 to the user 12 of the device 10. The neural network 54 is adapted to the user using a camera, with image data captured by the camera being used as labeled training data for adapting the neural network. The training phase is based on, for example, few-shot learning. For example, the user can specifically predefine different expressions, which are captured by the camera and are then available as labeled training data.
[0125] Figure 12 Alternative permutations of the neural network input data are shown.
[0126] Figure 13 A communication system 70 is shown for performing a communication method with at least two users 12-1 and 12-2. Each user wears a device 10-1, 10-2, in the example in the form of data glasses. The two users 12-1, 12-2 can be located in different locations. The devices 10-1, 10-2 each include a communication interface (not shown in detail) that allows data to be exchanged between the devices 10-1, 10-2. The two devices 10-1, 10-2 are configured according to the aforementioned embodiments and are used to derive information about a reference area 20 in the face of each user 12-1, 12-2 using an LFI sensor 18. Each device 10-1, 10-2 provides this information to the other device 10-1, 10-2 via data transmission. Each device 10-1, 10-2 receives data from the other device 10-1, 10-2. Based on the received data, each device 10-1, 10-2 inserts a virtual target object 12'-1, 12'-2 representing the other user 12-1, 12-2. Each device is configured to predefine the expression of the virtual target object 12'-1, 12'-2 based on the data of the other device 10-1, 10-2.
Claims
1. A device (10, 10-1, 10-2), in particular data glasses, which is designed to be worn by a user of the device in a predetermined wearing position on the user's body, in particular on the user's head, wherein: The device (10) comprises at least one laser feedback interferometer sensor (18), i.e., an LFI sensor, having at least one laser light source, in particular a laser diode, wherein the LFI sensor (18) is arranged and configured to emit laser radiation into a reference area (20) in a first area of the face of the user (12) of the device (10) outside the eyes and to detect a component of the reflection of the laser radiation at the device (10), and wherein the device (10) is configured to derive information about a reference area (20) based on the reflected component of the laser radiation, And wherein the device (10) is designed to provide information about the reference area (20) for inserting a virtual target object (12'), in particular an avatar representing the user of the device (10), in particular to another device (10).
2. The device (10) according to claim 1, wherein The LFI sensor (18) includes at least one optical element configured to widen a laser beam emitted by a laser light source at least along one line.
3. The device (10) according to claim 1 or 2, wherein The LFI sensor (18) is designed such that the laser beam emitted by the laser light source is split into at least two, in particular a plurality of, discrete partial beams, so that at least two discrete reference points within a reference area (20) are illuminated.
4. The device (10) according to claim 3, wherein The apparatus (10) is configured such that discrete reference points can be illuminated according to at least one first illumination pattern and at least one second illumination pattern, and can be switched between the first illumination pattern and the second illumination pattern.
5. The device (10) according to any one of claims 1 to 4, wherein The device (10) is configured to receive data from another device (10) of another user, the other device being in particular data glasses, wherein the data comprises information about at least one reference area (20) in at least one first area (22) of the face of the other user, and wherein the device (10) is configured to insert a virtual target object (12'), in particular an avatar representing the other user, based on the data from the other device (10), and wherein the device (10) is configured to predefine an expression of the virtual target object (12') based on the data from the other device (10).
6. The device (10) according to claim 5, wherein The device (10) is configured such that predetermining the expression of the virtual target object (12') includes modulating at least one spline (20') in a first area of the virtual target object (12') based on information about at least one reference area (20) of the first area (22).
7. A method for operating a device (10) according to any one of claims 1 to 6, comprising the following steps: emitting laser radiation into at least one reference area (20) in a first area (22) of a face of a user (12) of the device (10) and detecting a reflected component of the laser radiation; deriving information about a reference area (20) based on a component of the reflection of the laser radiation, Information about the reference area (20) is provided for inserting a virtual target object (12'), in particular an avatar representing a user of the device, in particular provided to another device (10).
8. The method according to claim 7, wherein: The method comprises: receiving data from another device (10) of another user, the other device being in particular data glasses, wherein the data comprises information about at least one reference area (20) in a first area (22) of the face of the other user (12); inserting a virtual target object (12'), in particular an avatar representing the other user, on the device based on the data of the other device (10); and presetting an expression of the target object (12') based on the data of the other device.
9. The method according to any one of claims 7 or 8, wherein Predetermining the expression of the avatar comprises modulating at least one spline (20') in at least one first area of the virtual target object (12') based on information about the reference area (20') of the first area.
10. The method according to claim 9, wherein: A corresponding reference area (20) is assigned to at least one spline (20') of the virtual target object (12'), or to at least one spline (20') of the virtual target object (12'). The method is to associate at least one spline of an image with the image and modulate the corresponding spline (20') based on information about the corresponding reference region (20).
11. The method according to any one of claims 7 to 10, wherein A distance spectrum and / or a velocity spectrum determined based on a component of the reflection of the laser radiation detected by means of at least one LFI sensor (18) is provided as input data to at least one trained neural network (54), and information about the reference area or information about a spline of the virtual target object (12') assigned to the corresponding reference area (20) is derived from the distance spectrum and / or the velocity spectrum by means of the trained neural network (54), and the spline of the virtual target object (12') is modulated based on the derived information.
12. The method according to claim 11, wherein The method comprises a training phase for adapting a trained neural network (54) to a user of the device (10), wherein the neural network (54) is adapted to the user using a camera, wherein image data recorded with the camera are used as labeled training data for adapting the neural network.
13. A communication system (70), comprising at least one first device (10-1) and at least one further device (10-2), wherein: The first device and the further device (10-1, 10-2) are configured according to any one of claims 1 to 6, and wherein the first device and the further device (10-1, 10-2) are configured to implement the method according to any one of claims 7 to 12.
14. A method of communication between at least two users (12-1, 12-2) via a communication network comprising a communication system (70) according to claim 13, wherein: The first user (12-1) is connected to the communication network (70) via a first terminal device (10-1) and the second user (12-2) is connected via a second terminal device (10-2), wherein the first terminal device (10-1) and the second terminal device (10-2) are respectively devices (10) according to any one of claims 1 to 6, and wherein at least the first user (12-1) is inserted into the second terminal device (10-2) of the second user (12-2) as a virtual target object (12'-1), in particular as an avatar representing the first user (12-1).