Facial expression detection
By using inertial measurement units and machine learning algorithms on portable devices, the accuracy and adaptability problems of the portable facial expression detection system are solved, and efficient facial expression detection and device control are achieved.
Patent Information
- Application Number
- CN202080020530.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-03-11
- Filing Date
- 2020-02-24
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2040-02-24
AI Technical Summary
Existing facial expression detection systems are difficult to achieve highly free head movement on portable devices, while ensuring that facial expressions are not abrupt and accurate.
Inertial measurement units (IMUs), such as ear-mounted devices, are used to combine machine learning algorithms, including neural networks and hidden Markov models, to determine facial expression information through IMU data and control the functions of electronic devices.
It realizes efficient and accurate detection of facial expressions on portable devices, provides convenient user input and feedback, adapts to different environmental conditions, and reduces signal and noise interference.
Smart Images

Figure CN113557490B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to facial expression detection. Some relate to using information from an inertial measurement unit of at least one wearable device for facial expression detection. Background Art
[0002] Facial expressions provide powerful and necessary non-verbal signals for social interaction. Facial expressions convey clues about human emotions, empathy, and feelings. Systems that can accurately detect facial expressions open up new markets for useful products and services.
[0003] It is difficult to design such a system to allow a high degree of freedom of head movement, allow changing surrounding conditions, and ensure that the system is not obtrusive when it is portable. Summary of the Invention
[0004] According to various but not necessarily all embodiments, there is provided an apparatus including components for: receiving information from at least one inertial measurement unit configured to be worn on a user's head; and at least partially causing determination of facial expression information at least based on the received information.
[0005] In some but not necessarily all examples, the at least one inertial measurement unit includes a gyroscope.
[0006] In some but not necessarily all examples, the inertial measurement unit is configured to be part of an earable device.
[0007] In some but not necessarily all examples, the facial expression information is determined based on the information and machine learning.
[0008] In some but not necessarily all examples, the machine learning includes a machine learning algorithm that includes a neural network or a hidden Markov model.
[0009] In some but not necessarily all examples, the machine learning algorithm includes one or more convolutional layers and one or more long short-term memory layers.
[0010] In some but not necessarily all examples, the apparatus includes components for at least partially causing control of an electronic device function based on the facial expression information.
[0011] In some but not necessarily all examples, controlling the electronic device function includes controlling output of feedback information through an output device based on the facial expression information.
[0012] In some but not necessarily all examples, the feedback information includes a recommended change to a task.
[0013] In some but not necessarily all examples, the feedback information includes recommended changes in how the task is performed.
[0014] In some but not necessarily all examples, controlling the functions of an electronic device includes interpreting facial expression information as an input command made by a user and causing the functions of the electronic device to be controlled in accordance with the input command.
[0015] According to various but not necessarily all embodiments, a handheld portable electronic device including a device is provided.
[0016] According to various but not necessarily all embodiments, a system including a device and an inertial measurement unit is provided.
[0017] According to various but not necessarily all embodiments, a method includes: receiving information from at least one inertial measurement unit configured to be worn on a user's head; and at least partially causing facial expression information to be determined based at least in part on the received information.
[0018] According to various but not necessarily all embodiments, a computer program, when run on a computer, performs: causing information to be received from at least one inertial measurement unit configured to be worn on a user's head; and at least partially causing facial expression information to be determined based at least in part on the received information.
[0019] According to various but not necessarily all embodiments, an example as claimed in the appended claims is provided. Description of the Drawings
[0020] Some example embodiments will now be described with reference to the drawings, in which:
[0021] Figure 1 An example of a method is illustrated;
[0022] Figure 2A An example of an ear-worn device is illustrated and Figure 2B An example of components of the ear-worn device is illustrated;
[0023] Figure 3 A facial expression illustrating six action units is illustrated;
[0024] Figure 4 Time histories of inertial measurement unit data for six action units are illustrated;
[0025] Figure 5A An example of a hidden Markov model algorithm is illustrated, Figure 5B An example of a convolutional neural network algorithm is illustrated, and Figure 5CIllustrates an example of an improved convolutional neural network algorithm;
[0026] Figure 6 Illustrates an example of a facial expression information server;
[0027] Figure 7A Illustrates examples of devices, apparatuses, and systems, and Figure 7B Illustrates an example of a computer-readable storage medium. Detailed implementation manners
[0028] Figure 1 Illustrates an example of method 100, including: at block 110, receiving information from at least one inertial measurement unit (IMU) (such as Figure 2B the IMU 204 shown in) configured to be worn on a user's head; and at block 120, at least partially causing determination of facial expression information based at least in part on the received information. Optional block 130 includes at least partially causing control of a human-machine interface function based on the facial expression information.
[0029] As described herein, measurements from the IMU 204 can be related to facial expressions. The IMU 204 is small and inexpensive. The IMU 204 can also be discrete because the sensor does not need to be in continuous contact with the user's skin to measure the inertial effects of moving facial muscles on the surface of the skin. For the same reason, there is no need for implantation or other invasive procedures to install the IMU 204.
[0030] First, various example implementations of block 110 will be described in detail. To receive useful IMU information, the IMU 204 is first worn.
[0031] The location on the user's head where the IMU 204 is worn. For the purposes of this disclosure, this location is any location on the human head that is moved in a detectable manner by the IMU 204 in accordance with the contraction and / or relaxation of facial muscles. Such locations include locations on the head and can also include locations in the upper neck region that are otherwise anatomically classified as part of the neck.
[0032] In some but not necessarily all examples, more than one IMU 204 is worn. Wearing multiple IMU 204s can include wearing more than one IMU 204 at a first location. Wearing multiple IMU 204s can include wearing IMU 204s that provide different sensing modalities. For example, different sensing modalities can include gyroscopes and accelerometers. Wearing multiple IMU 204s can include wearing one IMU 204 per axis for up to three measurement axes. Thus, three accelerometer IMUs can be configured to provide the function of a three-axis accelerometer, and three gyroscope IMUs can be configured to provide the function of a three-axis gyroscope.
[0033] Wearing multiple IMUs can include wearing the IMU 204 at different positions on the user's head. In some examples, the different positions can be on the left and right sides of the head. The positions can be on symmetrically opposite sides of the head. This provides better discrimination between symmetric and asymmetric facial expressions (e.g., a smile vs. a half-smile). In other embodiments, the distribution of positions can target different facial muscles and may or may not involve symmetric IMU positioning.
[0034] The following describes example attributes of a wearable device for positioning the (multiple) IMU 204 at the (multiple) desired positions.
[0035] The wearable device including the IMU 204 can be configured to be worn in a reusable manner. A reusable manner means that the wearable device can be removed and then re-worn without irreversible damage upon removal. The wearable device can be worn outside the user's body so that no implantation is required.
[0036] The IMU 204 can be provided or embedded on the wearable device. The IMU 204 can be positioned with reference to the wearable device so as not to touch or continuously touch the user's skin during use when worn, for increased comfort.
[0037] The wearable device can provide a wearable accessory function. An accessory as described herein means a wearable device that provides at least aesthetic and / or non-medical functions. Examples of wearable accessories include ear-worn devices (or audible devices), virtual reality headsets, glasses, clothing, jewelry, and hair accessories. An ear-worn device is a wearable accessory that can be worn inside or on the ear. An audible device is defined herein as an ear-worn device having an audio speaker.
[0038] Examples of further functions of wearable accessories include, but are not limited to, providing a human-machine interface (input and / or output), noise cancellation, positioning additional sensors for other uses, etc. Some wearable accessories can even include additional medical / non-accessory functions, such as corrective / tinted glasses lenses, positioning health monitoring sensors.
[0039] The wearable device can be configured not to be single-use. For example, the wearable device can be configured for a friction and / or bias fit. This avoids the need for single-use adhesives, etc. However, in alternative implementations, the wearable device is configured for single-use operation, such as the wearable device can include an adhesive patch.
[0040] Figure 2A and Figure 2BIllustrated is an example implementation of a wearable device 200 including an ear-worn device 201. The ear-worn device 201 has the advantage of being convenient compared to, for example, wearing specific clothing or unnecessary glasses. Another advantage is that the ear-worn device 201 is positioned close to several facial muscles highly related to common facial expressions, and the ear-worn device 201 can provide additional functions such as a headphone function or locate other sensors. The correlation will be discussed later.
[0041] Figure 2A Two ear-worn devices 201 are shown, one for the left ear and one for the right ear respectively. In other examples, only one ear-worn device 201 is provided for use with only one ear.
[0042] Figure 2A An internal view of the ear-worn device is shown in Figure 2B The ear-worn device 201 includes a human-machine interface that includes at least an audio speaker 210 for audio output so that the functions of the audible device are provided. The illustrated ear-worn device 201 includes at least one IMU 204. In an example implementation, the ear-worn device 201 includes a three-axis gyroscope and a three-axis accelerometer.
[0043] The illustrated ear-worn 201 (or other wearable device) includes circuitry 206 for operating the (multiple) IMUs 204. The circuitry 206 can operate the audio speaker 210. The circuitry 206 can be powered by a power source (not shown). If needed, an interface such as a wire or an antenna (not shown) can provide a communication link between at least the IMU 204 and an external device.
[0044] Figure 2A and Figure 2B The ear-worn device 201 is an in-ear ear-worn device 201 for embedding in the auricle. The in-ear ear-worn device 201 can be configured to be embedded close to the ear canal. The in-ear ear-worn device 201 can be configured to be embedded in the outer ear or the outer ear cavity. One advantage is that there is a strong correlation between the movement of the facial muscles forming common facial expressions and the deformation or movement of the part of the ear in contact with the ear-worn device 201. This correlated movement can be exploited by positioning the IMU 204 within the ear-worn device 201 because the IMU output depends on the movement or deformation of the ear. Thus, compared to other wearable devices, the ear-worn device 201, such as the in-ear ear-worn device 201, reduces the amount of data processing required to isolate meaningful signals from signal noise. Other wearable devices can operate when positioned at various head positions specified herein and form part of this disclosure. However, the ear-worn device 201 provides a favorable compromise between correlation (required data processing) and obtrusiveness to the wearer 400 (the user wearing the IMU 204).
[0045] The ear-worn device 201 can be configured to maintain a predetermined orientation of the IMU 204 relative to the user to ensure that clean data is obtained. In Figure 2A and Figure 2B the example of, the ear-worn device 201 includes an element 208 configured to engage with the intertragic notch of the user's ear. The element 208 can include a sleeve for a wire, which is configured to increase the effective stiffness of the wire and reduce bending fatigue. If the ear-worn device 201 is wireless, the element 208 can include an internal antenna for wireless communication. In other examples, the element 208 has no other purpose than to engage with the intertragic notch to position the ear-worn device 201 in a predetermined orientation.
[0046] It should be understood that Figure 2A the ear-worn device 201 of is one of many possible alternative wearable devices that can include the IMU 204.
[0047] As described above, the information provided by the IMU 204 is received as block 110 of method 100. This information can be received at a device that is part of the same wearable device that includes the IMU 204, or at a device that is remote from the IMU 204 via a communication link. The information can be received directly from the sensor in its raw form as an analog signal. Alternatively, the information can be received in digital form and / or can have been pre-processed, such as to filter noise.
[0048] Once the information has been received from the IMU 204, block 110 is complete and method 100 proceeds to block 120. At block 120, method 100 includes at least partially causing the determination of facial expression information based at least on the received information. The facial expression information is determined by processing the received IMU information and optionally additional information.
[0049] The determination of the facial expression information can be performed locally in the circuitry 206 using locally available processing resources, or caused to occur remotely, such as on a remote server that utilizes improved processing resources.
[0050] The determined facial expression information indicates which one of a plurality of different facial expressions is indicated by the received information. Thus, determining the facial expression information can include determining which one of a plurality of different facial expressions is indicated by the received information. The selected one facial expression defines the facial expression information. The plurality of different facial expressions can correspond to specific user-defined or machine-defined tags or categories.
[0051] Based on the received IMU information, the determined facial expression information can distinguish different upper facial expressions and / or different lower facial expressions. The upper facial expression can be associated with at least the eyebrows and / or eyes. The lower facial expression can be associated with at least the mouth. Multiple different facial expressions can indicate different upper facial expressions and / or different lower facial expressions. In some examples, both a change in the upper facial expression and a change in the lower facial expression can change the determined facial expression information. This improves the accuracy of emotion capture. In a non-limiting example, different facial expression information can be determined for a smile with symmetric eyebrows compared to a smile with raised eyebrows.
[0052] An experiment describing facial expressions will be referred to below using the "Action Unit" (AU) codes specified by the Facial Action Coding System (FACS) developed by P. Ekman and W. Friesen (Facial Action Coding System: A Technique for the Measurement of Facial Movement. Consulting Psychologists Press, Palo Alto, 1978). Figure 3 Facial expressions corresponding to AU2 (Outer Brow Raiser), AU4 (Brow Lowerer), AU6 (Cheek Raiser), AU12 (Lip Corner Puller), AU15 (Lip Corner Depressor), and AU18 (Lip Puckerer) are shown.
[0053] Figure 4 The time history of the IMU data for the in-ear ear-worn device 201 of FIG. 2 is illustrated, which was collected as the wearer 400 adopted Figure 3 each of the six action units illustrated. The wearer 400 did not perform other facial activities, such as talking or eating, while adopting the AUs. The plotted IMU data includes triaxial accelerometer data and triaxial gyroscope data.
[0054] Figure 4 The results of Figure 4 show a correlation between the IMU data and the AUs. Some correlations are stronger than others. In Figure 4 the study, but not necessarily for all facial expressions or wearable devices, stronger correlations can be found in the data from the gyroscope compared to the accelerometer. Stronger correlations can be found in the x-axis and y-axis data of the gyroscope compared to the z-axis data of the gyroscope. In this example, the x-axis is approximately in the forward direction, the y-axis is approximately in the upward direction, and the z-axis is approximately in the lateral direction.
[0055] Significantly, Figure 4The shape of the time course varies between AUs, enabling different AUs to be distinguished. This shows that it is possible to determine which of the multiple facial expressions is indicated by the information received from at least one IMU 204 in block 120. Only one IMU 204 may be provided, although using multiple IMU 204s as shown in the figure improves accuracy.
[0056] As shown in the figure, the gyroscope provides clear signals for eyebrow movement, cheek movement, and lip corner movement. The accelerometer captures clear signals for lip corner movement.
[0057] The in-ear ear-worn device IMU 204 can provide clear signals for both upper face AUs and lower face AUs. Table 1 shows the AUs for which the clearest signals were found:
[0058] Table 1: AUs for which the clearest signals were found by the ear-worn device IMU 204 and their muscle bases.
[0059]
[0060] The list of the above AUs that can be detected via the ear-worn device IMU 204 is not exhaustive and is preferably stated to include any AU that involves the facial muscles of Table 1 above, either individually or in combination. Facial expressions with higher impulse responses can be detected more easily than those with lower ones. Additional AUs and muscle dependencies can be detected using more sensitive IMU204s and / or improved data processing methods. For Figure 4 The IMU 204 for the experiment was an inexpensive MPU6500 model.
[0061] The manner in which facial expression information can be recognized is described.
[0062] It is possible to accurately determine facial expression information without machine learning. In a simple implementation, predetermined thresholds can be defined for the instantaneous values and / or time derivatives of data from one or more IMU 204s. When the defined threshold(s) is / are exceeded, it is determined which of the multiple facial expressions is indicated by the received information.
[0063] In various but not necessarily all examples of the present disclosure, the determination in block 120 depends on both information and machine learning. Machine learning can improve reliability. Various machine learning algorithms for determining facial expression information are described.
[0064] The machine learning algorithm can be a supervised machine learning algorithm. The supervised machine learning algorithm can perform classification to determine which predefined category of facial expression is indicated by the information. This enables the class labels to be predefined using terms for facial expressions such as "smile", "frown", etc. to improve user recognition.
[0065] The class labels do not necessarily correspond to individual AUs, but can correspond to facial expression classes that are best described as combinations of AUs. For example, a smile includes a combination of AU6 and AU12. A combination of AU4 and AU15 represents a frown. The algorithm can use at least the following class labels: smile; frown; none.
[0066] The machine learning algorithm can alternatively be an unsupervised machine learning algorithm. The unsupervised machine learning algorithm can perform clustering without class labels or training. This avoids the training burden, which otherwise might be important for interpreting the overall variations in facial geometry and IMU wear location (e.g., different orientations of different users' ears).
[0067] The machine learning algorithm can include a convolutional neural network (CNN) and / or a recurrent neural network (RNN) and / or a temporal convolutional network. A Hidden Markov Model (HMM) can be used for lower latency, although with sufficient training, CNNs and RNNs can achieve higher accuracy than HMMs. In an example implementation, the machine learning algorithm includes one or more convolutional layers, one or more long short-term memory (LSTM) layers, and an attention mechanism to improve the accuracy of a basic CNN to enable real-time facial expression information tracking with sufficiently low processing requirements. In another alternative example, the machine learning algorithm can include a deep neural network (DNN) for maximum accuracy at the cost of greater processing requirements.
[0068] Examples of supervised machine learning algorithms that have been experimented with are described below, and their F1 scores are provided to rank their performance relative to each other. The performance is evaluated based on the ability to detect the following three facial expressions: smile (AU6+AU12); frown (AU4+AU15); and none. The experiment included nine participants performing the above facial expressions, and each participant repeated them 20 times. For Figure 4 the device as described above.
[0069] Figure 5A An example structure of an HMM-based learning scheme for the experiment is shown. The HMM has low latency, effectively characterizes sequential data with an embedded structure (=AU), and is robust to variable input sizes.
[0070] The HMM pipeline extracts a list of 8-dimensional vectors (3-axis acceleration, 3-axis gyroscope signals, acceleration magnitude, and gyroscope magnitude) during a period when the wearer 400 adopts a facial expression. The HMM algorithm is trained using the Baum-Welch algorithm. The HMM is configured using a 12-hidden state left-right model with Gaussian emissions. The log-likelihood of the observation sequence for each class is determined using the forward algorithm. The facial expression model with the maximum log-likelihood is selected as the final result to represent the facial expression information of box 120.
[0071] From the HMM experiment, the average F1 score is 0.88, which implies that the HMM classifier can capture the intermittent and micro-muscle movements during facial expressions. For example, most of the time, smiling is correctly detected (F1 score = 0.96). The F1 score for frowning is 0.89. The F1 score for a neutral expression is 0.79.
[0072] Figure 5B Shows an example structure of a CNN-based learning scheme for the Figure 5A same experimental data. The CNN consists of a chain of four temporal convolutional layers "Conv1", "Conv2", "Conv3", "Conv4", and a pooling layer before the top fully-connected layer and softmax group. Each convolutional layer includes 64 filters (nf = 64). "Conv1" includes 3 kernels (kernel = 3). The other convolutional layers each include 5 kernels. "Conv1" has a stride of 2, "Conv2" and "Conv3" have a stride of 1, and "Conv4" has a stride of 3. The data of the "global average" layer has a data size of (1, 64) and the "Dense" layer has a data size of (1, 3). T is the window size.
[0073] Figure 5B The average F1 score of the CNN is 0.54, which is significantly above random chance and can be improved through further training to exceed the capabilities of the HMM. It should be understood that in use, the values of nf, kernel, stride, T, data size, number of layers, and any other configurable attributes of the CNN may vary based on the implementation.
[0074] Figure 5CAn example structure of an improved CNN-based learning scheme, referred to herein as "ConvAttention", is shown. The key features of ConvAttention are the adoption of LSTM (a special type of RNN) and an attention mechanism to better highlight the kinematic features of IMU signals made by facial expressions. LSTM is used to exploit the temporal patterns of AUs because LSTM is designed to utilize the temporal dependencies within the data. The attention mechanism is adopted because it can enable the recurrent network to reduce false positives from noise by targeting the regions of interest where the facial expressions in the data actually change and assigning higher weights to the regions of interest. Figure 5C Two convolutional layers Conv1 (nf = 64, kernel = 5, stride = 1) and Conv2 (nf = 64, kernel = 5, stride = 3) are shown, followed by an LSTM layer that returns the attention weights for each time point. The probabilities are multiplied by the feature vectors from the convolutional layers and averaged to produce a single feature vector. Subsequently, the feature vector is non-linearly transformed into class likelihoods through a fully connected layer.
[0075] The average F1 score of ConvAttention is 0.79, which is significantly above random chance and can be improved beyond the capabilities of HMMs through further training. It should be understood that configurable attributes may vary based on the implementation.
[0076] Once the facial expression information has been determined, box 120 is completed. Subsequently, the facial expression information can be used for various purposes.
[0077] Figure 6 A potential architecture of a facial expression information server 500 is shown, which can provide facial expression information to a requesting client 514. For example, the client 514 can be a client software application. The server 500 can reside in software implemented in one or more controllers and / or can reside in hardware. The server 500 executes Figure 1 method 100 for the client 514.
[0078] An example implementation of the server 500 is described below.
[0079] The server includes a sensor agent 502, which is configured to receive information from at least one IMU 204 and which executes box 110. In some but not necessarily all examples, information from additional sensors of different modalities can be received by the sensor agent 502 for synthesis and use in box 120 of method 100. Additional sensors (not shown) that can detect facial expressions include, but are not limited to:
[0080] - Proximity sensors on glasses;
[0081] - Force sensors on the ear - worn device 201;
[0082] - Bending sensors on the ear - worn device 201;
[0083] - Capacitive sensors on the wires close to the ear - worn device 201 and
[0084] - Electromyography sensors on the ear - worn device 201.
[0085] The proximity sensor can be configured to detect the distance between the glasses and the corresponding local positioning on the face. When the muscles around the eyes and nose (such as the orbicularis oculi, frontalis, levator labii superioris, nasalis) are tense (such as contempt, disgust, sadness), the face around the eyes and nose may bulge, and thus change the distance between the local positioning on the face and the corresponding proximity sensor.
[0086] The force sensor can be configured to detect the pressure on the force sensor through the deformation of the ear. The deformation of the ear can be caused by the tension of the upper auricle and the zygomatic major muscle, which is related to fear, anger, and surprise.
[0087] The bending sensor can be configured to detect the bending of the wires of the ear - worn device 201 if the ear - worn device 201 is wired (such as a headphone cable hanging down from the ear). When the masseter, zygomaticus, and buccinator muscles are tense (happy), the bulging face pushes the wires and causes some bending. An example of a compact bending sensor for detecting small bends in the wires is a nanosensor that includes a torsional optomechanical resonator and a waveguide such as an optical fiber cable for detecting torsion (bending).
[0088] The capacitive sensor can be configured to detect a change in capacitance at the wires of the ear - worn device 201. The capacitive sensor can be provided in the wires. When the head moves or the facial expression changes, the face can touch the wires, causing a change in the capacitance of the wires at a certain position along the wires. Happiness (smile) can be detected using the capacitive sensor.
[0089] The sensor agent 502 is configured to provide the received information to an optional noise filter 504. The noise filter 504 can include a high - pass filter, a low - pass filter, a band - pass filter, an independent component analysis filter, or a spatio - temporal filter such as a discrete wavelet transform filter. In an example implementation, the noise filter 504 includes a low - pass filter.
[0090] Subsequently, the filtered information is passed to the facial expression detector 510, and the facial expression detector can execute block 120.
[0091] The optional resource manager 508 adjusts the sampling rate and monitoring interval of the IMU 204, for example, according to resource availability and / or requests from client applications.
[0092] An optional Application Programming Interface (API) 512 is provided that enables a client 514, such as a software application or other requester, to request facial expression information.
[0093] In some but not necessarily all examples, the API 512 may support multiple request types, such as 1) continuous query, 2) on-the-spot query, and / or 3) historical query. A continuous query may cause the server 500 to continuously or periodically monitor the user's facial expressions and provide a final result at a given time. An on-the-spot query may cause the server 500 to return the most recent facial expression made by the user. A historical query may cause the server 500 to return a list of past facial expressions within a specified time range of the request.
[0094] An optional database (DB) 506 maintains facial expression information and / or raw IMU data, for example in response to a historical query.
[0095] Method 100 may terminate or loop back after block 120 is completed. After block 120, the determined facial expression information may be stored in a memory. For the client-server model described above, the determined facial expression information may be stored in the database 506 and / or provided to the requesting client 514. Thus, in the client-server model, method 100 may include receiving a request for facial expression information from the client 514. The request may indicate one of the above request types. Method 100 may include providing facial expression information in response. The provided information may conform to the request type.
[0096] Once the client 514 has received the facial expression information, the client may subsequently control an electronic device function based on the facial expression information. Thus, the server 500 or other device that executes method 100 may be summarized as being capable of causing, at least in part (e.g., via the client), an electronic device function to be controlled based on the facial expression information. Accordingly, an optional block 130 of method 100 is provided that includes causing, at least in part, an electronic device function to be controlled based on the facial expression information.
[0097] Some example use cases are provided below for how an application may control an electronic device function based on facial expression information. The example use cases represent situations where a user may desire or at least accept wearing a wearable device that includes an IMU 204. They also represent situations where it may be undesirable or impractical to ensure that the wearer 400 is within the field of view of a camera used to track facial expressions.
[0098] In some but not necessarily all examples, controlling the functions of an electronic device includes controlling an actuator. Examples of actuators that can be controlled include, but are not limited to: environmental control actuators (e.g., thermostats); navigation actuators (e.g., CCTV pan / zoom, steering); or medical device actuators.
[0099] In some but not necessarily all examples, controlling the functions of an electronic device includes controlling a human-machine interface (HMI) function. Controlling the HMI function can include interpreting facial expression information as a command input by the user and causing the functions of the electronic device to be controlled according to the input command. This enables the user to deliberately modify their facial expression to provide user input for controlling the electronic device. Additionally or alternatively, controlling the HMI function can include controlling the output of feedback information according to the facial expression information through an output device. This enables information dependent on the wearer's facial expression to be fed back to the wearer user or a different user.
[0100] In an example where at least one output function is controlled, the output function can be a user output function provided by one or more of the following output devices for user output: a display; a printer; a haptic feedback unit; an audio speaker (e.g., 210); or an odor synthesizer.
[0101] In an example where the output device is a display, the information displayed by the display according to the facial expression information can include text, images, or any other suitable graphical content. The displayed information can indicate to the user of the client application 514 the current facial expression information or current emotional state information associated with the wearer. The user of the client application 514 can be the monitored user (IMU wearer 400) or another user.
[0102] The facial expression information displayed as mentioned above can simply provide an indication of the facial expression (e.g., smile, frown). The emotional state information displayed as described above can provide an indication of the emotion determined to be associated with the facial expression (e.g., smile = happy, frown = sad / confused). Additional processing can be performed to determine the emotional state information from the facial expression information. This is because the emotional state is not necessarily indicated by an immediate facial expression but can be apparent from the temporal history of the facial expressions. The emotional state information indicating fatigue can be related to frequent expressions within the category associated with negative emotions (e.g., anger, disgust, contempt). Therefore, the emotional state information can be determined according to the temporal history of the facial expression information.
[0103] In an example where at least one output function is controlled, the control of block 130 can include controlling the output of feedback information by the output device according to the facial expression information. The feedback information can indicate the current emotional state of the wearer 400.
[0104] Feedback information can include recommended changes to a task. This is advantageous for use cases where wearer fatigue or harmful emotions may affect the wearer's ability to perform a task. In various examples, the wearer 400 can be an employee. The employee may be performing a safety-critical task, such as driving a vehicle, manufacturing or assembling safety-critical components, handling hazardous chemicals, or working in a nuclear power plant, etc.
[0105] For employee monitoring, the method 100 can include receiving a request for facial expression information. The request can include a continuous query as described above, or an on-site query or a historical query. The request can come from the client application 514, which can be an employer-side client application or an employee-side client application. The request can be triggered by the determination of the task that the wearer 400 is currently performing, such as the determination that the user has started working or a shift. Techniques can be used to make the determination, such as: tracking the wearer's location using a location sensor; determining whether the wearable device has been put on by the user; receiving information from a calendar application; or receiving user input indicating that the user is performing a task. Figure 1 The method 100 can be executed in response to the request.
[0106] The method 100 can additionally include deciding whether to output feedback information recommending a change to the task from the determined current task based on the determined facial expression information. The decision can be made in Figure 6 the client application 514 or the server 500.
[0107] The decision can be based on the emotion state information as described above. If the determined emotion state information has a first attribute or value (e.g., emotion category), the decision can be to loop the method 100 back (continuous query) or terminate (on-site query, historical query), and / or output information indicating the emotion state. If the determined emotion state information has a second attribute or value (e.g., a different emotion category), the decision can be to output feedback information recommending a change to the task. Using the above "fatigue" example, the first attribute / value may not be associated with fatigue and the second attribute / value may be associated with fatigue. In other examples, the decision can be based on facial expression information without determining emotion state information.
[0108] The recommended change to the task can include recommending to temporarily or permanently stop the task, such as taking a break or stopping. If the wearable device 200 including the IMU 204 also includes an audio speaker 210, the feedback information can be output to the audio speaker 210. This is convenient because the wearer does not need to be near an external audio speaker and does not need to wear an audio speaker device separately. The feedback information can be configured to be output at a headphone volume level so that other nearby users are not alerted to the feedback information. However, it should be understood that the feedback information can be provided to any suitable output device.
[0109] In response to a recommended change to a task, the employer may direct the employee to take a break or stop, or the wearer 400 may decide on their own to take a break or stop.
[0110] The recommended change to a task does not have to recommend taking a break or stopping the task in all instances. For example, the user may be working through a sequence of tasks (such as a hobby, cooking, watching TV), and the recommendation may be based on the emotional state as to when to change tasks. In a fitness monitoring use case, the recommended change to a task may be to start or stop exercising.
[0111] According to the above use case, the recommendation is to change the task. However, in additional or alternative use cases, the feedback information may include a recommended change to how the task is performed, without necessarily changing the task. Except for giving different feedback, the steps involved may be the same as or different from the above use case for recommending a change to a task.
[0112] Examples of changing how a task is performed include optimizing facial expressions during a communication task. Facial expressions are a very important form of non-verbal communication, arguably as important as the words chosen by the wearer 400. If the wearer's facial expression contradicts the image they are trying to convey, the feedback will improve the user's communication ability.
[0113] The user may wish to optimize their facial expressions during high-pressure face-to-face communication tasks (such as job interviews, sales interactions, doctor-patient interactions, meetings, funerals). A vision-based emotion tracking system using a camera may not be an available option because personal devices with cameras may need to be left in a pocket. This makes the wearable IMU method desirable. In other implementations, the communication may be video communication. For communication tasks, constantly turning to a personal device for on-site or historical queries may be impolite, so the ability to make continuous queries is beneficial for communication tasks.
[0114] Detecting that the user is performing a communication task may be as described above (such as location tracking, when worn, calendar information, or manual input). The decision as to whether to recommend a change in how the task is performed may use the above methods (such as based on emotional state information or just facial expression information).
[0115] The recommended change to how a task is performed is not necessarily limited to communication tasks. For example, the recommended change to how a task is performed may include an increase or decrease in the intensity of the task (such as exercise intensity, vehicle driving speed / acceleration, or other fatiguing tasks). In an employee monitoring example, if the emotional state does not improve, a change in task intensity may be recommended first before recommending a break.
[0116] Further examples will now be described, where the HMI function controlled by block 130 of method 100 includes an input function. For example, client application 514 may interpret facial expression information as an input command by the user, and may cause a device function to be controlled based on the input command.
[0117] The input command may include at least one of the following: selection of an option provided by a user interface; navigation within the user interface; insertion of an object (such as an emoji, text, and / or image); changing the device power state (on, off, sleep); activating or deactivating a peripheral device or subsystem, etc.
[0118] A useful example of using facial expressions for input is when a device with an input HMI is not easily accessible. For example, if a user is driving or in a meeting, they may be prohibited by law or deterred by etiquette from using a personal device such as a mobile phone. If the personal device has a camera, the personal device may even be put away, which prevents the use of vision-based emotion tracking. In such a situation, the use of the wearable device IMU 204 is advantageous.
[0119] The input command may control a hands-free device function. The hands-free function includes one or more of the following: accepting and / or rejecting an incoming request to start a communication session (such as an incoming voice / video call request); terminating a communication session (such as hanging up); replying to text-based communication (such as using SMS or an instant messaging application); changing the user status on an application (such as busy, idle); listening to voicemail; changing device settings (such as loud, mute, airplane); canceling or postponing notifications (such as alerts, incoming text-based communication), etc.
[0120] In some examples, the hands-free function may be used for a virtual assistant service. The hands-free function may be used to give instructions or to respond to queries from a virtual assistant service. The interface for the virtual assistant service may be provided by a device such as the ear-worn device 201 that lacks a touch-based human-machine interface for interacting with the virtual assistant service and / or the graphical user interface.
[0121] In some but not necessarily all examples, when the facial expression information is associated with a first facial expression, the input command is a first input command, and when the facial expression information is associated with a second (different) facial expression, the input command is a second (different) input command. For example, a first facial expression such as a smile can initiate a reply or confirmation function (e.g., sending a confirmation of a missed call, confirming an alert), and a second facial expression such as a frown can initiate a cancellation function (e.g., canceling a missed call notification, delaying an alert). If the facial expression does not fall into either of the above two or cannot be determined, no facial expression-related action can be performed. In other examples, only one type of facial expression can be recognized, such as smiling or not smiling, or more than two recognizable facial expressions can provide more than two or three results.
[0122] The methods described herein can be performed by a device 602 such as Figure 7A the device 602 shown in. The device 602 can be provided in the wearable device 200 together with the IMU 204, or can be provided in a device 601 separate from the device including the IMU 204. The device 601 can include an output device 612. The output device 612 can perform the functions of one or more of the previously disclosed output devices. In other implementations, the output device 612 can be provided separately from the device 601.
[0123] Thus, in one example, a device 601 including the device 602 and the IMU 204 is provided, and in another example, a system 600 including the device 602 and a separate IMU 204 is provided, which are coupled wired or wirelessly. The system 600 can optionally include an output device 612.
[0124] Figure 7A The device 601 of can optionally include:
[0125] One or more cameras (not shown), such as one or more front cameras and / or one or more rear cameras;
[0126] A user interface (not shown), such as a touch screen, buttons, sliders, or other known underlying technologies;
[0127] An input / output communication device (not shown) configured to send and / or receive the data / information described herein, such as an antenna or a wired interface.
[0128] Figure 7A The device 601 of can be the personal device mentioned herein. The device 601 can be configured to provide the electronic device functions mentioned herein. The device 601 can be a hand-held portable electronic device 601. The hand-held portable electronic device 601 can be a smart phone, a tablet computer, or a laptop computer.
[0129] Figure 7A An example of the controller 604 is illustrated. The implementation of the controller 604 can be as a controller circuit system. The controller 604 can be implemented solely in hardware, have certain aspects in software including separate firmware, or can be a combination of hardware and software (including firmware).
[0130] As Figure 7A illustrated, the controller 604 can be implemented using instructions that enable hardware functions, for example, by using executable instructions of a computer program 610 that can be stored on a computer-readable storage medium (disk, memory, etc.) in a general-purpose or special-purpose processor 606 for execution by such a processor 606.
[0131] The processor 606 is configured to read from and write to the memory 608. The processor 606 may also include an output interface through which the processor 606 outputs data and / or commands and an input interface through which data and / or commands are input into the processor 606.
[0132] The memory 608 stores a computer program 610 that includes computer program instructions (computer program code), which controls the operation of the device 602 when the computer program 610 is loaded into the processor 606. The computer program instructions of the computer program 610 provide the logic and routines that enable the device to perform Figure 1 the illustrated method 100. The processor 606 can load and execute the computer program 610 by reading the memory 608.
[0133] Thus, the device 602 includes:
[0134] at least one processor 606; and
[0135] at least one memory 608 that includes computer program code,
[0136] the at least one memory 608 and the computer program code are configured to, together with the at least one processor 606, cause the device 602 to at least perform:
[0137] receive information from at least one inertial measurement unit configured to be worn on a user's head;
[0138] at least partially cause determination of facial expression information at least based on the received information; and
[0139] at least partially cause control of a human-machine interface function based on the facial expression information.
[0140] As Figure 7BAs illustrated, computer program 610 can reach device 602 via any suitable delivery mechanism 614. The delivery mechanism 614 can be, for example, a machine-readable medium, a computer-readable medium, a non-transitory computer-readable storage medium, a computer program product, a memory device, a recording medium (such as a compact disc read-only memory (CD-ROM) or a digital versatile disc (DVD) or solid-state memory, an article of manufacture including or tangibly embodying computer program 610). The delivery mechanism can be a signal configured to reliably transmit computer program 610. Device 602 can propagate or transmit computer program 610 as a computer data signal.
[0141] Computer program instructions for causing a device to at least perform the following operations or for at least performing the following operations: causing information to be received from at least one inertial measurement unit configured to be worn on a user's head; at least partially causing facial expression information to be determined at least based on the received information; and at least partially causing a human-machine interface function to be controlled based on the facial expression information.
[0142] The computer program instructions can be included in a computer program, a non-transitory computer-readable medium, a computer program product, a machine-readable medium. In some but not necessarily all examples, the computer program instructions can be distributed over more than one computer program.
[0143] Although memory 608 is illustrated as a single component / circuit system, it can be implemented as one or more separate component / circuit systems, some or all of which can be integrated / removable and / or can provide permanent / semi-permanent / dynamic / cache storage.
[0144] Although processor 606 is illustrated as a single component / circuit system, it can be implemented as one or more separate component / circuit systems, some or all of which can be integrated / removable. Processor 606 can be a single-core or multi-core processor.
[0145] References to "computer-readable storage medium", "computer program product", "tangibly embodied computer program", etc. or "controller", "computer", "processor", etc. should be understood to include not only computers having different architectures such as single / multi-processor architectures and sequential (von Neumann) / parallel architectures, but also dedicated circuits such as field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), signal processing devices, and other processing circuit systems. References to computer programs, instructions, code, etc. should be understood to include software for programmable processors or firmware, such as, for example, programmable content of a hardware device, whether instructions for a processor or configuration settings for a fixed function device, a gate array, or a programmable logic device, etc.
[0146] As used in this application, the term "circuitry" may refer to one or more or all of the following:
[0147] (a) Only hardware circuitry implementations (such as implemented in only analog and / or digital circuitry) and
[0148] (b) Combinations of hardware circuitry and software, such as (where applicable):
[0149] (i) Combinations of analog and / or digital hardware circuitry with software / firmware and
[0150] (ii) Any part of a hardware processor with software (including digital signal processors), software, and memories working together to cause a device (such as a mobile phone or server) to perform various functions and
[0151] (c) (Multiple) hardware circuitry and / or (multiple) processors (such as (multiple) microprocessors or a part of (multiple) microprocessors) require software (such as firmware) for operation, but the software may be absent when it is not needed for operation.
[0152] This definition of circuitry applies to all uses of the term in this application, including in any claims. As a further example, as used in this application, the term circuitry also encompasses implementations of only hardware circuitry or processors and their (or their) accompanying software and / or firmware. The term circuitry also encompasses (for example and if applicable to a particular claim element) a baseband integrated circuit for a mobile device or a server, a cellular network device, or a similar integrated circuit in other computing or network devices.
[0153] Figure 1 And the blocks illustrated in FIG. 5 may represent steps in a method and / or fragments of code in a computer program 610. The specification of a particular order of the blocks does not necessarily imply a required or preferred order for the blocks, and the order and arrangement of the blocks may vary. Additionally, it is possible that some blocks are omitted.
[0154] The technical effect of method 100 is an improved physiological sensor. This is because facial expressions convey information about the physiology of the human making the expression and can cause a physiological response in humans who can see the facial expression. The sensor is improved at least because, unlike other physiological sensors, the inertial measurement unit does not need to be in continuous direct contact with the user's skin and is small, light, and inexpensive for use in wearable device accessories.
[0155] The technical impact of the IMU 204 in the ear-worn device 201 is that the IMU 204 can enable services additional to the facial expression information service. In some but not necessarily all examples, the devices and methods described herein can be configured to determine head pose information from ear-worn IMU information and provide the head pose information to an application. The application can include virtual reality functionality, augmented reality functionality, or mixed reality functionality configured to control the rendered gaze direction based on the head pose information. Another potential application can include an attention alarm function that can provide an alarm when the head droops (e.g., during driving). In further examples, the audio attributes of the audio rendered by the audio speaker 210 of the ear-worn device can be controlled based on the ear-worn device IMU information.
[0156] In further examples, the devices and methods described herein can be configured to determine location information from IMU information using, for example, dead reckoning. The location information can indicate the current location of the wearer and / or the navigation path of the wearer. The application can include a map function, a guidance-giving function, and / or a tracking function for tracking the wearer (e.g., an employee).
[0157] Where structural features have been described, they can be replaced by one or more functions that perform the functions of the structural features, whether the functions or those functions are described explicitly or implicitly.
[0158] The capture of data can include only a temporary record, or it can include a permanent record or it can also include both a temporary record and a permanent record. A temporary record implies a temporary recording of data. For example, this can occur during sensing or image capture, occur in dynamic memory, occur in buffers such as circular buffers, registers, caches, or similar buffers. A permanent record implies that the data is in the form of an addressable data structure, retrievable from an addressable storage space, and thus can be stored and retrieved until deleted or overwritten, although long-term storage may or may not occur. The use of the term "capture" with respect to an image relates to the temporary or permanent recording of the data of the image.
[0159] Systems, apparatuses, methods, and computer programs can use machine learning, which can include statistical learning. Machine learning is a field of computer science that gives a computer the ability to learn without being explicitly programmed. A computer learns from experience E with respect to some class of tasks T and performance measure P if the performance of the computer on task T, as measured by P, improves with experience E. A computer can typically learn from previous training data to make predictions about future data. Machine learning includes fully or partially supervised learning and fully or partially unsupervised learning. It can enable discrete outputs (e.g., classification, clustering) and continuous outputs (e.g., regression). Machine learning can be implemented, for example, using different means such as cost function minimization, artificial neural networks, support vector machines, and Bayesian networks. For example, cost function minimization can be used for linear and polynomial regression and K-means clustering. Artificial neural networks, e.g., having one or more hidden layers, model complex relationships between input vectors and output vectors. Support vector machines can be used for supervised learning. A Bayesian network is a directed acyclic graph that represents the conditional independence of multiple random variables.
[0160] As used in this document, the term "comprising" has an inclusive non-exclusive meaning. That is, any reference to X that comprises Y indicates that X can include only one Y or can include more than one Y. If it is intended to use "comprising" with an exclusive meaning, then it will be referred to in the context as "comprising only one..." or by using "consisting of".
[0161] In this specification, various examples have been referred to. A description of an example feature or function indicates that those features or functions exist in that example. The use of the terms "example" or "for example" or "able to" or "can" in the text means that, whether explicitly stated or not, such a feature or function exists at least in the described example, whether described as an example or not, and they may but do not necessarily exist in some or all other examples. Thus, "example", "for example", "able to", or "can" refer to a particular instance in a class of examples. The attributes of an instance can be attributes of only that instance or attributes of the class or a subclass of the class that includes some but not all instances of the class. Thus, features described with reference to one example rather than another are implicitly disclosed and can, where possible, be used as part of a working combination in that other example, but do not necessarily have to be used in other examples.
[0162] In this specification, control that at least partially causes the functionality of an electronic device, which can include directly controlling an input device and / or an output device and / or an actuator, or providing data to a requesting client to cause the client to control an input device and / or an output device and / or an actuator.
[0163] Although the embodiments have been described with reference to various examples in the foregoing paragraphs, it should be understood that modifications may be made to a given example without departing from the scope of the claims.
[0164] The features described in the foregoing description may be used in combinations other than those explicitly described above.
[0165] Although functions have been described with reference to certain features, those functions may be performed by other features, whether or not described.
[0166] Although features have been described with reference to certain embodiments, those features may also be present in other embodiments, whether or not described.
[0167] As used in this document, the term "a" or "the" has an inclusive but not exclusive meaning. That is, any reference to X that includes a / the Y indicates that X may include only one Y or may include more than one Y, unless the context clearly indicates the contrary. If an exclusive meaning of "a" or "the" is intended, it will be made clear in the context. In some cases, the use of "at least one" or "one or more" may be used to emphasize the inclusive meaning, but the absence of these terms should not be taken as inferring an exclusive meaning.
[0168] The presence of a feature (or combination of features) in a claim is a reference to that feature or (combination of features) itself and is also a reference to features (equivalent features) that achieve substantially the same technical effect. Equivalent features include, for example, variants and features that achieve substantially the same result in substantially the same way. Equivalent features include, for example, features that perform substantially the same function and achieve substantially the same result in substantially the same way.
[0169] In this specification, various examples have been cited that use adjectives or adjective phrases to describe example features. Such a description of a characteristic of an example indicates that the characteristic exists exactly as described in some examples and exists substantially as described in other examples.
[0170] While efforts have been made in the foregoing specification to draw attention to those features regarded as important, it should be understood that the applicant may seek protection via the claims for any patentable feature or combination of features mentioned and / or shown in the foregoing and / or in the drawings, whether or not emphasis has been placed thereon.
Claims
1. An apparatus for examining one or more facial expressions, comprising: at least one processor; and at least one memory including computer program code, wherein the at least one memory and the computer program code are configured to, together with the at least one processor, cause the apparatus to at least: receive information from an inertial measurement unit positioned within an in-ear ear-worn device configured to be embedded in the auricle of a user's head such that the output of the inertial measurement unit depends on the movement or deformation of the part of the ear in contact with the in-ear ear-worn device; and at least partially cause determination of facial expression information based at least on a correlation between the received information and the movement of facial muscles forming a facial expression and the movement or deformation of the part of the ear in contact with the in-ear ear-worn device.
2. The apparatus according to claim 1, wherein the in-ear ear-worn device includes an element configured to engage with the intertragal notch of the ear to maintain a predetermined orientation of the inertial measurement unit relative to the user.
3. The apparatus according to claim 1, wherein the at least one memory and the computer program code are configured to, together with the at least one processor, further cause the apparatus to: receive information from a second inertial measurement unit configured to be worn on a symmetrically opposite side of the user's head, wherein determination of the facial expression information further depends on the information from the second inertial measurement unit to enable discrimination between symmetric and asymmetric facial expressions.
4. The apparatus according to claim 1, wherein the facial expression information is determined based on the information and machine learning.
5. The apparatus according to claim 4, wherein the machine learning includes a machine learning algorithm, and the machine learning algorithm includes a neural network or a hidden Markov model.
6. The apparatus according to claim 5, wherein the machine learning algorithm includes one or more convolutional layers and one or more long short-term memory layers.
7. The apparatus according to claim 1, wherein the at least one memory and the computer program code are configured to, together with the at least one processor, further cause the apparatus to: at least partially cause a component for controlling an electronic device function based on the facial expression information.
8. The apparatus according to claim 7, wherein controlling the electronic device function includes controlling the output of feedback information through an output device based on the facial expression information.
9. The apparatus according to claim 8, wherein the feedback information includes a recommended change to a task.
10. The apparatus according to claim 8, wherein the feedback information includes a recommended change in how a task is performed.
11. The apparatus according to claim 7, wherein controlling the electronic device function includes interpreting the facial expression information as an input command made by the user and causing the electronic device function to be controlled based on the input command.
12. A system for examining one or more facial expressions, comprising: An inertial measurement unit, the inertial measurement unit being positioned within an in-ear ear-worn device configured to be embedded in the auricle of a user's head such that the output of the inertial measurement unit depends on the movement or deformation of the portion of the ear in contact with the in-ear ear-worn device; and a device, the device comprising: at least one processor; and at least one memory including computer program code, the at least one memory and the computer program code being configured to, together with the at least one processor, cause the device to at least: receive information from the inertial measurement unit; and at least partially determine facial expression information at least based on a correlation between the received information and the movement of facial muscles forming a facial expression and the movement or deformation of the portion of the ear in contact with the in-ear ear-worn device.
13. A method for examining one or more facial expressions, comprising: receiving information from an inertial measurement unit positioned within an in-ear ear-worn device configured to be embedded in the auricle of a user's head such that the output of the inertial measurement unit depends on the movement or deformation of the portion of the ear in contact with the in-ear ear-worn device; and at least partially determining facial expression information at least based on a correlation between the received information and the movement of facial muscles forming a facial expression and the movement or deformation of the portion of the ear in contact with the in-ear ear-worn device.
14. The method according to claim 13, wherein the in-ear ear-worn device comprises an element configured to engage with the intertragal notch of the ear to maintain a predetermined orientation of the inertial measurement unit relative to the user.
15. The method according to claim 13, wherein the facial expression information is determined based on the information and machine learning.
16. The method according to claim 13, further comprising controlling an electronic device function based on the facial expression information.