Systems, methods, and computer programs for detecting and confirming user emotions
By using signals representing user head movement and physiological parameters, input and processing into machine learning models, the problem of determining user emotions without session data is solved, and efficient emotion recognition with minimal hardware is achieved.
Patent Information
- Application Number
- CN202380072947.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-21
- Filing Date
- 2023-10-20
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art is difficult to reliably determine user sentiment without session data using minimal hardware.
The user's emotions are determined and confirmed by providing the machine learning model with a first signal representing the user's head movement and processing a second signal representing the user's physiological parameters. The system can be embedded in a single head-mounted device, using only limited hardware.
A method of accurately determining user emotions without the wearer's conscious interaction is realized, and only limited hardware is used, improving the reliability and efficiency of emotion recognition.
Smart Images

Figure CN120077345A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of affective computing.
[0002] More specifically, the present invention relates to a system, a corresponding method, and a corresponding computer program for determining a user's mood. Background Art
[0003] Emotion recognition is the process of identifying human emotions. Using technology for emotion recognition is a relatively new research area, which is said to contribute to the emergence of the so-called Internet of Emotions or Feelings.
[0004] So far, most work has been done in automatically recognizing conversational data (including recognizing facial expressions from video, spoken expressions from audio, or written expressions from text). Generally, this kind of recognition works best if it uses multiple modalities in context.
[0005] Research has also shown that the human visual system is highly sensitive to biological motion and can extract information about an individual's emotions, intentions, personality traits, kinematic style, and biological attributes when the biological motion pattern of any individual is presented. Thus, models for automatically analyzing biological motion patterns and determining the associated emotions have been developed in recent years. These models use, both in training and in production, datasets representing the global motion of the body, such as datasets obtained using a full-body motion sensing input device. By analyzing the dynamics of the full-body feature points corresponding to the body joints in a laboratory environment, emotions during a subject's walking activity have been successfully detected. However, recording these dynamics requires specific equipment and stable methods, which are limiting factors for the widespread use of biological motion pattern analysis.
[0006] Therefore, there is a need for methods and products that allow reliable determination of a user's mood even when no conversational data is available.
[0007] Any relevant data should further be obtainable using a minimum of hardware, preferably only using the mobile device provided to the user. Summary of the Invention
[0008] The present invention has been made in view of the above problems.
[0009] According to one aspect of the proposed technology, a system is provided that is configured to:
[0010] - Provide a first signal representing at least the movement of a user's head to a machine learning model, thereby obtaining an output representing the mood of the user, the machine learning model having been pre-trained using a model signal database representing the movement of the head of a model user and associated with at least one mood of the model user, and
[0011] - Process a second signal representing a physiological parameter of the user to confirm the user's mood.
[0012] The proposed technology allows determining the user's mood without any conscious interaction from the wearer and with only limited hardware. In fact, the proposed technology does not require a large number of available signals representing the movements of multiple body landmarks. Instead, in terms of the wearer's movement, it is sufficient to provide a first signal representing the movement of the head. In some embodiments, the required hardware can even be embedded in a single head-mounted device. The combination of the specific nature of the first and second signals with the specific method of first determining the mood based on a machine learning analysis of the first signal and then confirming the mood based on the processing of the second signal has surprisingly allowed obtaining better results than other envisioned methods.
[0013] Optionally, the system further includes a wearable device that includes sensors configured to sense at least the movement of the user's head and the user's physiological parameters when the wearable device is worn by the user. It is more convenient for the wearer when all the required sensors are incorporated in a single head-mounted device rather than multiple devices.
[0014] Optionally, the model signal database represents only the movement of the head of the model user. This allows minimizing the storage space required to execute the proposed method.
[0015] Optionally, the model signal database represents the movement of the head of the model user more extensively compared to the movement of the body of the model user. This allows providing different versions of the mood prediction service to a group of wearers according to each wearer's device. When the wearer is equipped with a single motion sensor embedded in a head-mounted device, a simpler version of the service is based only on head movement. A more complex version can provide further adjustments, for example, taking into account hand movement when the wearer is further equipped with a hand-held or wrist-mounted motion sensor. In order to still achieve good results with the simpler version, it is recommended that the model signal database represents the movement of the head of the model user extensively enough, especially in comparison to the movements of other body parts of the model user.
[0016] Optionally, the system further includes a control interface configured to generate a control signal based on a condition related to the detected mood or the confirmed mood.
[0017] For example, a smart glasses system can be provided to a user to be worn throughout their daily activities, including, for example, while walking in a crowd. Such a smart glasses system can include various sensors, among which there is an IMU that can be configured to obtain a first signal, and a camera system and a sound recording system that can be configured to obtain pictures, videos, and / or audio recordings. The first signal can be obtained at the current moment and provided in real time to a machine learning model, which in turn indicates the detected emotion. Then, tests can be automatically run using the detected emotion, such as checking a given criterion. An example of such a criterion can be to check whether the detected emotion is considered positive by, for example, corresponding to a positive valence and / or a high arousal intensity of the emotion. In this example, when the detected emotion is considered positive, the camera can be automatically activated to take a picture, and the picture is tagged according to the detected emotion. Further, during the user's daily activities, he / she can further actively use the camera, for example, to take pictures, and can tag these pictures according to their own perceived emotional state. These tags can be associated with the detected emotion and provided as tagged training data to the machine learning model. Then, instructions can be generated to identify, among a set of tagged pictures, the picture with the most positive emotion score based on the tags of the pictures. It may later be further proposed to print, visualize, or share these identified pictures as an album with the user's friends and / or family members.
[0018] Optionally, the machine learning model can be configured to output the detected emotion together with a trust level associated with the detected emotion, and the pictures can be further tagged according to the trust level. When considering the above example, the trust level threshold can be used as a criterion for filtering the tagged pictures. This allows, for example, to ignore any pictures with tags representing a detected emotion associated with an insufficient trust level.
[0019] Optionally, the system further includes a wearable device having an activatable and / or controllable optical function, such as a refractive function and / or a transmissive function, and the wearable device is configured to activate and / or control the optical function upon receiving a control signal.
[0020] For example, an electrochromic wearable device that can switch between different possible hues can be provided to the user. The selection of hue appearance and comfort can be pre-programmed. With the user interface application of the electrochromic wearable device (which can be available, for example, on a mobile device), the user can trigger one or another hue management algorithm proposed regarding his / her perceived needs in light environment management. Separately, a machine learning model can also run and be trained on the user interface application to detect the user's mood based on multiple signals obtained from sensors on the device worn by the wearer, such sensors including, for example, an IMU, a heart rate sensor, and a respiration sensor. The detected mood can be timestamped and stored. Further, each hue change manually triggered by the user or automatically triggered by the hue management algorithm can also be identified as a timestamped event and stored. Associating the detected mood with events having a timestamp essentially the same as the timestamp itself or by other known means allows both the machine learning model that outputs the detected mood and the hue management algorithm to be improved by a co-learning method. The advantage of this co-learning method is to avoid facing inappropriate hues in specific social / environmental mood situations. For example, in principle, a user may wish to wear glasses with dark hues outdoors in summer. But when having an intimate social interaction such as a face-to-face discussion with another person, the user may still prefer a more transparent hue so as to better express his / her mood with his / her eyes. This allows the other person to perceive and detect the mood through eye contact to increase the engagement in the relationship. Further, the head dynamics caused by the postures and movements of such social interactions can help distinguish between different scenarios and allow the detection of mood. Another advantage of the co-learning method is that the hue change can be used as user feedback. This includes any hue change manually triggered by the user, as well as any hue change automatically triggered by the hue management algorithm without causing a negative reaction from the user, a negative reaction such as manually reverting the hue change, or expressing perceivable discomfort through changes in postures and / or physiological parameters.
[0021] Optionally, the system further includes a feedback module, such as a visual, audio, and / or tactile feedback module, configured to provide feedback to the user upon receiving a control signal.
[0022] Optionally, using user feedback to evaluate the currently detected mood through the interface can help further train the machine learning model. For example, such feedback can be voluntarily provided by the user by selecting predefined options on the interface and / or as an answer to an automatic notification. The feedback is user input that can, for example, confirm the detected mood or better represent the user's current mood of self-assessment. The machine learning model can use this feedback from the user to evaluate its performance as a valid trust level for the detected mood represented by its output.
[0023] Optionally, the system further includes a messaging module configured to send a message to the user upon receipt of a control signal.
[0024] Optionally, when the detected user emotion will conflict with the social emotion, the system will send a message to the user.
[0025] Optionally, the model signal further includes model context data, and the first signal further includes context data.
[0026] Optionally, the system further includes a context data module configured to obtain context data, and the context data module includes sensors and / or communication interfaces.
[0027] The context data may include the time of day, location, activity data, etc. Optionally, the context data includes an indication that the user is wearing a wearable device, and the determined and / or confirmed emotion represents the comfort associated with wearing the wearable device (e.g., as a score on a rating scale).
[0028] Optionally, the model signal further includes model user profile data associated with the model user, and the first signal further includes user profile data associated with the user.
[0029] Optionally, the system further includes a user profile module configured to obtain user profile data, and the user profile module includes sensors and / or communication interfaces.
[0030] The model user profile data and the user profile data may include, for example, gender, weight, height, age, user identifier, etc.
[0031] Optionally, the system further includes: a feed module configured to feed a user profile with the determined or confirmed emotion; and / or a tagging module configured to tag context data based on the determined or confirmed emotion.
[0032] For example, when the detected emotion meets a predetermined criterion (e.g., is a positively expressed emotion), the camera can be turned on to capture an image of the user. In the example, the camera is embedded in the front of the spectacle frame worn by the user. Further, before capturing the image, the user can optionally be requested to face a mirror. In this example, image processing of the captured image can allow for the identification of the user's facial expression and / or pose, which in turn can be used to confirm the detected emotion. Alternatively, a picture of the user can be taken by a friend and shared on a social network. The picture of the user and the detected emotion can be tagged based on, for example, metadata as referring to the same event. Cross-comparing the detected emotion with the image processing result of the picture allows for the confirmation of the detected emotion.
[0033] According to another aspect of the proposed technology, a method for determining a user's emotion is provided, the method comprising:
[0034] - providing at least a first signal representing the movement of the user's head as input to a machine learning model, to obtain an output representing the emotion of the user, the machine learning model having been pre-trained using a database of model signals representing the movement of the head of a model user and associated with the emotion of the model user, and
[0035] - processing a second signal representing the physiological parameters of the user to confirm the emotion of the user.
[0036] According to another aspect of the proposed technology, a computer program is provided, the computer program comprising one or more sequences of stored instructions that are accessible by a processing unit and that, when executed by the processing unit, cause the processing unit to perform the above method. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] To more fully understand the description provided herein and its advantages, reference is now made to the following brief description taken in conjunction with the accompanying drawings and detailed description, in which like reference numerals represent like parts.
[0038] Figure 1 and Figure 2 Schematically shows a machine learning model training algorithm according to an embodiment.
[0039] Figure 3 and Figure 4 Schematically shows a trained machine learning model according to an embodiment.
[0040] Figure 5 Schematically shows a possible use of the detected and / or confirmed emotion of a user according to an embodiment. DETAILED DESCRIPTION
[0041] The proposed technology aims to determine the emotion of a user.
[0042] Emotions can first be detected based on a signal representing head movement being input into a machine learning model. Then, additional information can be used to confirm the output of the machine model, and the additional information includes signals representing the physiological parameters of the user.
[0043] The proposed technology allows for reliably determining a user's emotion using a minimal combination of sensors. For example, a signal representing head movement can be obtained from an inertial measurement unit, which can be part of a head-mounted device such as glasses, goggles, a head-mounted unit, headphones, etc., while a signal representing the user's physiological parameters can be obtained from a physiological sensor, which can be part of the same device or from a different wearable device such as a wristband.
[0044] The proposed technology can be implemented by running a computer program on a processing unit, and the computer program includes instructions stored on a non-transitory storage medium. When executed by the processing unit, these instructions cause the processing unit to perform the method according to the proposed technology. The processing unit can further be operably connected to a variety of devices via one or more wired or wireless communication interfaces.
[0045] For example, the devices operably connected to the processing unit can include one or more sensors from the above sensor combination. This enables the processing unit to perform various functions, including controlling these sensors and / or processing the data sensed by these sensors over time. For example, the devices operably connected to the processing unit can include one or more terminal devices provided to the user, and such terminal devices include, for example, one or more user interfaces such as a graphical UI, a tactile UI, etc. Such user interfaces enable the processing unit to perform additional functions, including obtaining feedback from the user and / or providing information to the user. More generally, any device operably connected to the processing unit is selected and provided by a person of ordinary skill in the art according to the functions to be achieved by the above computer program.
[0046] The proposed technology can find many applications. Since determining the user's mood at a given moment and monitoring this mood in daily life is beneficial in various fields. Determining the mood (e.g., discomfort of the lens wearer) is valuable for decision-making in clinical practice (e.g., tele-optometry) or for sales purposes (e.g., digital retail). Determining the audience mood of media content streamed in a real environment (e.g., smart home) or a virtual environment (e.g., metaverse) can help determine whether the desired effect has been achieved on the audience. In the field of education, the learning sequence can be extended or shortened based on the evolution of the learner's attention level or fatigue over time. Social robots can use the judgment of the patient's emotional state to adjust their interaction with the patient. For example, such social robots may help support the emotional development of autistic children. Monitoring the mood of vehicle occupants can be used to trigger specific safety measures based on preset conditions (e.g., when the driver is detected to be angry), for example, through V2X communication.
[0047] Now definitions are proposed for various terms and expressions used in this article.
[0048] "Wearer" refers to the user whose mood is to be determined. The wearer wears a head-mounted device that includes at least one or more sensors for sensing a first signal.
[0049] "First signal" is an analog or numerical signal representing the translational and / or rotational movement of the wearer's head. An inertial measurement unit (IMU) is a known combination of sensors that allows the measurement of movement by recording the three-dimensional position and orientation, translational velocity and rotational velocity, translational acceleration and rotational acceleration with timestamps. The IMU can be easily integrated into the frames of various mounting devices or wearable devices and has low power consumption compared to other types of sensors (e.g., cameras). Possible mounting devices include IMU clips, smart ears, true wireless earphones, smartphones held pressed against the head, AR / VR head-mounted devices, smart glasses, smart frames, etc. Then, the first signal can be, for example, a signal output by an IMU embedded in the head-mounted device or wearable device.
[0050] "The second signal" is an analog or numerical signal representing a user's physiological parameter. Well-known methods and biomedical sensors allow for obtaining various physiological parameters, including, for example, skin conductance response, skin temperature, analytes from sweat, or heart rate variability. Depending on the nature of the physiological parameter at hand or according to the context, suitable sensors can be arranged in the user's environment or embedded in various wearable devices. The second signal can be output, for example, by an image processing module that receives an image of the user as input. In particular, the input image can at least show the user's face. The input image can originate from a camera, which is an exemplary suitable sensor. The second signal can represent the user's physiological condition, including, for example, blushing, which can be associated with a list of emotions or possible emotions. The second signal can also represent the physiological manifestation of an emotional state or the physiological response to a stimulus. This applies in particular to various facial expressions, including smiling. Facial expressions can be extracted by processing an image showing the user's face and can be associated with a list of emotions or possible emotions. In an embodiment, a single head-mounted device can be embedded with all the required sensors to obtain both the first signal and the second signal.
[0051] The first signal and the second signal are not necessarily different signals originating from different sensors, but may relate to different aspects combined in a single sensed signal. This is the case, for example, when an IMU is placed on or near the skin of the wearer at a location suitable for detecting a complex motion representing the combined effects of head motion, heart rate motion, and respiratory motion.
[0052] A machine learning model is a model that has been trained to associate one or more inputs, including at least the first signal, with a corresponding output variable representing the emotion of the wearer. Training the machine learning model is performed using labeled training data, which in this case are signals that each represent the motion of the model wearer's head and are each associated with the emotion of the model wearer. The first signal can be recorded periodically and form a time series. Based on this time series, the machine learning model can determine the output variable representing the wearer's emotion on various time scales: instantaneous (up to a few seconds), short-term (up to a few minutes), or long-term (hours, days, weeks, or longer). The machine learning model can further be configured to determine a latent variable representing the persistent inner emotional state of the user based on this time series. This latent variable can be used to determine whether the output variable corresponds to a steady, stable emotional state or a transient emotional state.
[0053] In the field of affective computing, appropriate metrics can be used to classify emotions.
[0054] One possibility related to supervised training is to define a finite list of possible emotion categories (e.g., anger, fear, neutral, sadness, surprise, disgust, happiness) to label the training data according to these categories and train a machine learning model to infer the correct category when the training data is input. When labeling the training data, emotions can be classified, for example, according to very simple head movement marker points as follows.
[0055]
[0056] One possibility for automatically labeling training data without human supervision is to configure a clustering model (e.g., a neural network) to perform an unsupervised clustering mechanism. Then, a corresponding identifier corresponding to the category can be assigned to each cluster. This identifier can have a numerical value. This identifier can also be valence. Valence (also known as pleasantness) refers to an exemplary classification metric that refers to the inherent attractiveness or aversion to an event, object, or situation. Valence can be used as a way to characterize emotions. Thus, for example, the valence of anger and fear is negative, while the valence of joy is positive. Numbers can be assigned to valence, and the possible values of these numbers are used to label the first model signal with as many categories as possible. These values can be further output by the machine learning model and associated with the first signal. The machine learning model can be, for example, a binary classifier between a first category and a second category, where the first category reorganizes all positive valence emotions and the second category reorganizes all negative valence emotions. Classical metrics that may be associated with the binary classifier can include accuracy, sensitivity, and specificity. Of course, regardless of the number of emotion categories considered, the machine learning model can also be a fuzzy classifier or a probability classifier. The machine learning model can process the variables and internal emotional states of the wearer according to various possible topologies. Thus, the machine learning model can be a linear classifier or a non-linear Euclidean space classifier.
[0057] A technical problem is how to describe the wearer's emotion based on the movement when only detecting the wearer's movement from the wearer's head. In response to this technical problem, the inventors have determined that the head movement considered in this way can reflect the control and execution of body movement and the emotions caused by social or physical environment interactions. The head movement further provides clues about the wearer that can be de-correlated with the wearer's emotion. For example, in some motor activities such as walking activities, gender-specific differences have been analyzed. This analysis reveals that the dynamic part of the whole body movement contains information about gender. Since, as already mentioned, the head movement reflects this dynamic, knowing the wearer's gender helps to interpret their head movement and correctly identify their emotion based on the interpreted head movement of the wearer.
[0058] Now refer to as Figure 1A simple example of a machine learning model training algorithm. Fundamentally, a machine learning model aims to identify specific types of patterns. The model is trained on a set of data to provide an algorithm that can be used to reason about and learn from this data. This is the so-called "training phase". Once the model is trained, it can be used to reason about data it has not seen before and make predictions about this data. This is the so-called "production phase". As Figure 1 The purpose of the training algorithm shown is to optimize the model to reconstruct the labeled training data. The two main types of machine learning models are supervised learning models and unsupervised learning models. Supervised learning involves learning a function that maps inputs to outputs based on example input-output pairs. In contrast, unsupervised learning is used to draw inferences and discover patterns from input data without reference to labeled results. As Figure 1 The training algorithm shown can be either a supervised or an unsupervised training algorithm without distinction.
[0059] Obtain (2) a model signal. The model signal represents at least the translational and / or rotational movement of the head of the model wearer and may further represent the movement of other body parts of the model wearer. The model wearer can be an actual wearer, in which case the model signal is sensed. The model wearer can also represent an actual wearer or a group of actual wearers, in which case the model signal is either fully generated or sampled from a database of sensed first signals corresponding to the (multiple) actual wearers. The model signal can represent the actin activity of the model wearer, such as walking activity. The model signal can be acquired when such activity is detected and within an appropriate time window corresponding to all or part of the duration of such activity. The duration of the time window can depend on context data, such as the environment of the model wearer or the nature of the activity, or can be set to a predefined value independent of the context.
[0060] The model signal is further associated with a variable representing the emotion of the model wearer by any known means.
[0061] Provide (4) the model signal to a machine learning model, which then outputs (6) an output variable representing a prediction of the wearer's emotion.
[0062] Further obtain (8) a variable associated with the model signal and compare (10) this variable with the output variable. The result of the comparison is provided as feedback (12) to the machine learning model. The feedback received after each prediction indicates whether the prediction is correct or incorrect.
[0063] Repeat the whole process using different model signals that jointly form a training data set. Between iterations, adjust the internal logic of the machine learning model based on aggregated feedback, with the aim of maximizing the number of correct predictions using the training data set.
[0064] Now refer to a more complex example of a machine learning model training algorithm as Figure 2 shown.
[0065] In this more complex example, for each model signal provided to the machine learning model for making predictions, additional data related to that model signal is also provided to the machine learning model. The resulting multimodal and / or temporal analysis performed by the machine learning model generally allows for better predictions.
[0066] For example, a set of model signals of a given model wearer can be obtained during an extended period of time in which the given model wearer performs various activities such as working, walking, running, sleeping, watching television, driving, etc. The given model wearer can be further equipped with a device or interface that is adapted to sense or more generally obtain context data for each model signal. Data is said to be context - related if it refers to elements that can contextualize and assist in the interpretation of the first signal. Context elements can be events, states, situations, etc. that are permanently or temporarily applied to the model wearer and / or their environment when the model signal is sensed. The context data can indicate one or more types of context elements, including for example relevant physiological parameters of the model wearer, the model wearer's sociodemographic information (such as gender), the model wearer's facial expression, the noise level near the model wearer, the light level near the model wearer, the geographical location data of the model wearer, a timestamp, the activities or tasks experienced by the model wearer, etc. Then, the (14) context data can be obtained and provided to the machine learning model. The machine learning model can then use the corresponding context data to jointly analyze each model signal of the model wearer when the model signal is sensed. This allows for establishing correlations. For example, some head movements may be related to specific context elements (such as a specific activity or group of similar activities). Emotions can further be related to specific context elements. Such correlations assist in the training of the machine learning model and specifically in improving the prediction accuracy of the model wearer's emotions.
[0067] When training a dataset, the model signals representing the head movements of a given model wearer can be further labeled with the identifier of the given model wearer. Further obtaining (16) these identifiers and providing them to the machine learning model allows the model to establish a correlation specific to that model wearer or group of model wearers between the head movement patterns of one model wearer or a group of model wearers and their emotions. Thus, after the training phase, the machine learning model can perform different analyses on the first signal, and thus output different output variables according to whether the first signal represents the head movement of the first wearer or the head movement of the second wearer.
[0068] The model signals representing the head movements of the same model wearer can be further organized and provided to the machine learning model as a set of time series related to different aspects of the head movement. In other words, when considering the model signals corresponding to a given time point obtained (2) and provided (4) to the machine learning model, one or more historical model signals corresponding to one or more earlier time points can be further obtained (18) and provided to the machine learning model. This enables the machine learning model to not only interpret instantaneous indications, but also a series of consecutive indications of the model wearer's head movement up to a given time point. Then, the machine learning model can reliably identify various types of data based on the temporal correlation and possible multimodal correlation between consecutive model signals related to consecutive time points up to a given time point. Furthermore, the machine learning model can then associate the data types thus identified with the corresponding emotions and output the output variables accordingly. Examples of data types that can be identified are, for example, the head movement patterns of the model wearer, such as the head lifting for a moment and then dropping again. Another example of a data type that can be identified is the estimation or confirmation of context elements (such as the detected vibrations characteristic of a model wearer cycling). Yet another example of a data type that can be identified is the estimation of the overall body dynamic acceleration as a measure reflecting the energy consumption of the model wearer. In fact, as already mentioned, analyzing head movements can provide an indication of the overall body movement.
[0069] The inventors have developed a proof of concept for extrapolating emotions based solely on head movements, and will now describe it. Code showing the movement of 15 body landmarks of a model walker has been retrieved, such as that of BiomotionLab on its website https: / / www.biomotionlab.ca / html5-bml-walker / Provided above. For each of the 15 body landmark points, according to the retrieved codes, corresponding motion data were regenerated for different full-scale emotions respectively labeled with "sad", "happy", and "neutral". Then, a database was established to associate the motion data with the labeled emotions, and the database was split into a training dataset and a test dataset. A logistic regression model was trained using the training database, and the trained model was applied to the test dataset. An identification score represented as the AUC value (AUC stands for "area under the curve") was determined. For the "happy" and "sad" emotions, the identification score was 0.65 when using only head data, and 0.67 when using head + hand data. For the neutral emotion, the identification score was higher, reaching 0.73.
[0070] Now refer to Figure 3 and Figure 4 , which depict examples of the possible use of a machine learning model in the production phase. The machine learning model was previously trained as depicted, for example, in Figure 1 or Figure 2 .
[0071] First, a first signal is obtained (20) and provided (22) to the machine learning model. The machine learning model can extract, for example, one or more of the following data from the first signal: head acceleration, velocity, frequency-amplitude analysis, entropy, jerk, and analyze the extracted data to identify patterns that were also found when analyzing the training data, thereby outputting (24) an output variable.
[0072] Various additional data can be further obtained and provided to the machine learning model to benefit from the correlations identified in the training phase and better distinguish between similar emotions. For example, data fusion from multiple sensors can enhance reliability and eliminate ambiguity.
[0073] Obtaining (26) and providing a second signal to the model can allow the model to extract relevant data from the second signal, identify patterns in the multimodal data extracted from the first and second signals, compare the identified patterns with the patterns that were also found when analyzing the training data, thereby outputting (24) an output variable. The principle of multimodal analysis for enhancing the results also applies when obtaining (32) and providing context data to the model, or when obtaining (34) and providing a wearer identifier to the model.
[0074] The wearer identifier can be combined with user or wearer profile elements and processed by a machine learning model to customize emotion determination. For example, the kinematics and dynamics of head movement may vary between individuals and be unrelated to their emotions. For example, let's consider that a downward head movement may be associated with negative emotions such as disgust or sadness. When neither any wearer identifier nor any wearer profile element associated with the first signal is provided to the machine learning model, the machine learning model is likely to simply set one or more thresholds, for example, in terms of the amplitude or angular velocity of the downward angle, and when the amplitude or angular velocity of the downward angle exceeds the threshold, it will result in an output variable representing such negative emotions. In contrast, when a wearer identifier or a wearer profile element associated with the first signal is provided to the machine learning model, it is further possible to set different values of the above (multiple) thresholds for different wearers.
[0075] For example, assuming a walking activity, the estimated stance swing and walking speed obtained from additional data can be combined with the first signal and allow the machine learning model to decorrelate a part of the head dynamics included in the first signal that is completely or mainly caused by the walking activity from another part of the head dynamics included in the first signal that is completely or mainly caused by the wearer's emotions.
[0076] Additional data provided by, for example, sensors can further help identify joint events such as climbing stairs or mixing with a fast-moving crowd, which may help detect the wearer's emotional transitions.
[0077] Additional data related to the context (such as the wearer's environment or activity), daily time, previous values of the output variable, etc. can be further used to preprocess or filter an incoming series of first signals. For example, one or more criteria that the additional data may meet or not meet can be set, and different actions can be triggered based on whether these criteria are met. One possible action is to provide the first signal to the machine learning model only when one or more criteria are met, and conversely, not to provide the first signal to the machine learning model when one or more criteria are not met. Another possible action is to modify (e.g., mark) the first signal based on whether a certain criterion is not met before providing the first signal to the machine learning model. Considering an example of a criterion based on daily time, the first signals can be preprocessed by adding a mark indicating day or night to all of them, and then the marked first signals are provided to the machine learning model.
[0078] The output variable can be stored for further use by the machine learning model. In particular, an emotion kinematic pattern may have been identified during training data, such that the output variable representing the wearer's emotion at a given time point can also be a cue for:
[0079] - Determine the wearer's emotions in the near future, or
[0080] - Determine whether the wearer's emotions are stable within a set amount of time, or
[0081] - Detect an event that causes a short-term or long-term emotional shift in the wearer. In other words, the event is related to an output variable representing the wearer's apparent or current emotions, or to a latent variable representing the wearer's persistent emotions.
[0082] Then, the output variable and / or the latent variable can be associated with various parameters, where there is an indication of whether the represented emotion is steady / stable or transitional. When the represented emotion is transitional, a kinematic indication of the transition can be further provided. When the represented emotion is stable, an indication of the duration associated with the represented emotion and / or an indication of the initial time and the final time can be further provided. Such an indication allows for enhanced downstream management of the wearer's emotions.
[0083] Another example of the further use of the stored output variable is combined with a request for feedback from the wearer on the emotion represented by the stored output variable. The request can be a pop-up window or an audio signal or a tactile signal, etc., that prompts the wearer to interact with the human-machine interface. Then, the feedback can be obtained (36) in the form of a signal generated by such an interaction or the absence of such an interaction before a set timer expires. The feedback can confirm or deny the emotion represented by the stored output variable. The feedback can be associated with the stored output variable and provided to a machine learning model to help further train the model. In particular, providing wearer feedback associated with a wearer identifier allows the machine training model to continue learning from personal data.
[0084] In addition to the output variable output (24) by the machine learning model, additional data is further obtained (26, 32, 34, 36) and processed (28) to confirm (30) the wearer's emotions. The additional data includes at least a second signal and may further include context data, a wearer identifier, or feedback on the confirmed emotion.
[0085] A possible way to process the second signal (28) will now be described by way of example. For example, the second signal can be output by a pupillometer. A look-up table with a list of entries can be provided, each entry mapping a list of allowed values of the output variable to a corresponding range of pupil size values. In this example, processing the second signal can include identifying the range of pupil size to which the wearer's pupil size belongs from the second signal, and then retrieving the entry in the look-up table corresponding to the identified range. Next, it can be checked whether the value of the output variable can be found in the list of allowed values. Then, if the result of the check is positive, the wearer's emotion can be confirmed, or if the result of the check is negative, the wearer's emotion can be denied. Of course, the same principle can also be applied to process the second signal input by any other sensor or to process any other type of additional data. For example, if a second signal is received from sensors such as a heart rate monitor and / or a respiratory rate sensor, the entries in the look-up table may be related to ranges of heart rate values and / or respiratory rate values.
[0086] Alternatively, a machine learning model can conceptually be divided into a first internal model and a second internal model. The first internal model receives at least the first signal and outputs (24) an output variable, and the second internal model receives and processes (28) the second signal to output a processing result. Then, the output variable and the output processing result can be processed together to confirm (30) the emotion represented by the output variable.
[0087] Now referring to Figure 5 , which provides a non-exhaustive list of possible uses of the detected and confirmed emotions of the user.
[0088] After the emotion is confirmed (30), the output variable can be processed (38) to generate (40) a control signal, which may contain instructions for various hardware and / or software modules.
[0089] Then, a messaging service that receives the control signal can send (42) appropriate messages at an appropriate time, for example, via an application and / or a device (e.g., via a worn device such as a mixed reality headset). These messages can be specifically targeted at the wearer or at other users. These messages can convey positive reinforcement of the detected and confirmed positive emotions and can also convey an alert, a warning, or a suggestion when a negative emotion is detected and confirmed.
[0090] A Wearer Vision Profile Manager can create and feed (44) a real-time vision ergonomic wearer experience profile that includes a set of preferences that are automatically adjusted in real-time for the wearer based on output variables. The adjustment can be further based on, for example, context data including ambient light level. The adjustment can be further based on latent variables. When operating a device with a visual interface, such as a computer, tablet, smartphone, active optical element placed in front of one or both eyes of the wearer, a head-mounted device adapted to display visual elements to the wearer, etc., the vision profile can be automatically applied.
[0091] Timestamped data can be provided to a data annotator. The control signal can include instructions for the data annotator to tag (46) any timestamped data that corresponds to a time associated with an output variable representing a detected emotional event of the wearer. The data annotator can access and tag several types of timestamped data, including situation-related timestamped data (such as collected pictures, environmental features, vision ergonomic specifications, vision tasks, emotional tasks, and actin tasks), user-related timestamped physical data (such as geographical location, head movement, lighting, weather, crowding), timestamped data streams (such as music, phone calls, messages, application notifications), etc.
[0092] An Active Optical Function Controller can control various optical functions, including activating filters, displaying images in augmented reality and / or virtual reality, activating or adjusting the transmission function to achieve, for example, a desired hue and / or color intensity, controlling the refractive function to achieve, for example, magnification or focusing in a specific gaze direction and / or for a specific gaze distance, etc. The control signal can include instructions for the Active Optical Function Controller to control (48) the optical functions in a differentiated manner based at least on the output variables and optionally further based on latent variables, context data, wearer profile data, etc. Thus, for example, the wearer's lifestyle, behavior, and emotions can be linked, and the optical functions can be customized accordingly for the wearer. Further, an emotion and the co-occurrence of a specific element in the wearer's field of view can be detected and confirmed, and thus the active refractive function can be controlled at a specific gaze direction and / or specific gaze distance to form a region of interest corresponding to the position of the specific element in the wearer's field of view. Further, an emotion can be detected and confirmed and simultaneous context data related to the wearer's environment (such as a sudden change in ambient light intensity) can be obtained, and thus a change in the transmission function can be triggered to compensate for the sudden change.
[0093] The feedback module may be configured to provide (50) feedback to the user regarding their environment, contextual data, tasks, long-term emotions, etc. Possible feedback examples may correspond to automatically taking a picture using a camera embedded in the head mounted device or recording a sound using an embedded microphone whenever the output variable indicates, for example, the wearer's happiness and the latent variable further indicates that the emotion does not change over a sufficient amount of time. The taken picture and / or recorded sound may then be stored for later retrieval and provision to the user.
[0094] All of the above modules can be activated and modulated in a differentiated manner for each possible emotion of the wearer that can be represented by an output variable. Further differentiation can be based on whether the represented emotion is stable or transient.
[0095] Now let's briefly describe a few use cases.
[0096] In the first use case, the wearer is considered to be focused on the task, and the regional transmission function is activated by default to dim the peripheral light in order to avoid disturbing the wearer and help the wearer focus on the task. Throughout the task, the first signal is obtained periodically (for example, every few seconds or every minute). The first signal can be filtered to meet specific measurement conditions, that is, if a given first signal meets the specific measurement conditions, the given first signal is provided to the machine learning model, otherwise the given first signal is ignored. As a preprocessing behavior, an exemplary purpose of such filtering can be to remove outliers. For each provided first signal, the machine learning model outputs a corresponding output variable, which represents the wearer's emotion obtained at least from the analysis of the provided first signal. The output variables are collectively interpreted as monitoring the wearer's emotional state over time, and more specifically checking whether the wearer is healthy or bored. When boredom is detected, a possible action may be to automatically suggest the wearer to pause the task through an appropriate message. Another possible action may be to interrupt the dimming of the peripheral light to further encourage the wearer to rest.
[0097] In a second use case, a consumer tries on glasses either physically or virtually. The consumer may verbally express a subjective emotion while trying on the frames. The objective emotion may be further detected by providing a first signal to a machine learning model and outputting an output variable. The objective emotion may be further confirmed by processing a second signal. The confirmed objective emotion may be compared to the expressed subjective emotion to identify whether it is consistent or inconsistent. In the case of inconsistency, such as when the subjective emotion is positive and the objective emotion is not positive, an alert or suggestion to try on another frame may be generated despite the positive subjective emotion.
Claims
1. A system configured to: - Provide (22) a first signal representing at least the movement of the user's head to a machine learning model to obtain (24) an output representing the detected emotion of the user, the machine learning model having been pre-trained using a model signal database representing at least the movement of the head of at least one model user and associated with at least one emotion of the at least one model user, and - Process (28) a second signal representing the physiological parameters of the user to obtain (30) the confirmed emotion of the user.
2. The system according to claim 1, wherein, the system further includes a wearable device, the wearable device includes a sensor configured to sense (20, 26) at least the movement of the user's head and the physiological parameters of the user when the wearable device is worn by the user.
3. The system according to claim 1 or 2, wherein: - The model signal database represents only the movement of the head of the model user, or - Compared with the movement of the body of the model user, the model signal database represents the movement of the head of the model user more extensively.
4. The system according to any one of claims 1 to 3, wherein, the system further includes a control interface configured to generate (40) a control signal based on the detected emotion of the user or the confirmed emotion of the user.
5. The system according to claim 4, wherein, the system further includes a wearable device having an activatable and / or controllable optical function, the wearable device being configured to activate and / or control (48) the optical function according to the control signal.
6. The system according to claim 4 or 5, wherein, the system further includes a feedback module configured to receive (36) feedback from the user on the detected emotion of the user or the confirmed emotion of the user.
7. The system according to any one of claims 4 to 6, wherein, the system further includes a messaging module configured to send (42) a message to the user according to the control signal.
8. The system according to any one of claims 1 to 7, wherein, the model signal further includes model context data, and the first signal further includes context data.
9. The system according to claim 8, wherein: - The context data includes an indication that the user wears the wearable device, and - The determined emotion of the user and / or the confirmed emotion of the user represents the comfort associated with wearing the wearable device.
10. The system according to claim 8 or 9, wherein, the system further includes a context data module configured to obtain (32) the context data, the context data module including a sensor and / or a communication interface.
11. The system according to any one of claims 1 to 10, wherein, The model signal further includes model user profile data associated with the model user, and the first signal further includes user profile data associated with the user.
12. The system according to claim 11, wherein, the system further includes a user profile module configured to obtain (34) the user profile data, and the user profile module includes sensors and / or communication interfaces.
13. The system according to any one of claims 1 to 12, the system further comprises: - a feed module configured to feed (44) the user profile with the determined or confirmed emotion, and / or - a tagging module configured to tag (46) context data based on the determined or confirmed emotion.
14. A method for determining a user's emotion, the method comprises: - providing at least a first signal representing the movement of the user's head to a machine learning model as an input, thereby obtaining (24) an output representing the detected emotion of the user, the machine learning model having been pre-trained using a model signal database representing the movement of the head of at least one model user and associated with the emotion of the at least one model user, and - processing (28) a second signal representing the physiological parameters of the user to obtain (30) the confirmed emotion of the user.
15. A computer program comprising one or more sequences of stored instructions that are accessible by a processing unit and, when executed by the processing unit, cause the processing unit to perform the method according to claim 14.