System, method and computer program for detecting and verifying user emotions

A wearable device using head and physiological sensors with a machine learning model effectively identifies emotions, overcoming the limitations of existing technologies by integrating head movement and physiological data for reliable emotion recognition.

JP2025535401APending Publication Date: 2025-10-24ESSILOR INTERNATIONAL(COMPAGNIE GENERALE D OPTIQUE)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025522696
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-10-21
Filing Date
2023-10-20
Publication Date
2025-10-24

AI Technical Summary

Technical Problem

Existing emotion recognition technologies require specialized equipment and methodologies, limiting their widespread adoption, and there is a need for reliable emotion identification without speech data using minimal hardware.

Method used

A system that uses a machine learning model trained on head movements and physiological parameters, integrated into a single wearable device, to identify emotions without requiring conscious interaction, with optional additional sensors for hand movements.

Benefits of technology

Provides reliable emotion identification using minimal hardware, achieving better results than multiple signal approaches by combining head movement analysis with physiological data, enabling applications in various fields such as clinical practice, sales, and smart environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025535401000001_ABST
    Figure 2025535401000001_ABST
Patent Text Reader

Abstract

The present disclosure relates to a system configured to provide a first signal representing at least a movement of a user's head to a machine learning model and thereby obtain an output representing an emotion of the user, the machine learning model having been pre-trained using a database of model signals representing the head movements of a model user and associated with at least an emotion of the model user, and to process a second signal representing a physiological parameter of the user to ascertain the emotion of the user. The present disclosure further relates to a corresponding method and a corresponding computer program.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to the field of affective computing.

[0002] More particularly, the present invention relates to a system for identifying a user's emotion, a corresponding method and a corresponding computer program. [Background technology]

[0003] Emotion recognition is the process of identifying human emotions. The use of technology to perform emotion recognition is a relatively emerging field of research and is said to contribute to the emergence of the so-called affective or emotional Internet.

[0004] To date, most research has focused on automatic recognition of speech data, such as facial expressions from video, words from audio, and written expressions from text. In general, this recognition works best when multiple modalities are used in the context.

[0005] Research has shown that the human visual system is highly sensitive to biological motion, and that when presented with an individual's biological motion patterns, information about that individual's emotions, intentions, personality traits, movement style, and biological attributes can be extracted. In recent years, models have been developed to automatically analyze biological motion patterns and identify associated emotions. These models use datasets representing the entire body's movements, such as those obtained using full-body motion sensing input devices, for both training and production. Analysis of the dynamics of full-body landmarks corresponding to body joints in a laboratory setting has been successful in detecting emotions during a subject's walking activity. However, recording these dynamics requires specialized equipment and robust methodologies, limiting the widespread adoption of biological motion pattern analysis. Summary of the Invention [Problem to be solved by the invention]

[0006] Therefore, there is a need for methods and products that allow for reliable identification of a user's emotions even when speech data is not available.

[0007] Any relevant data should also be retrievable using minimal hardware, preferably using only the mobile device provided to the user. [Means for solving the problem]

[0008] The present invention has been made in view of the above problems.

[0009] According to an aspect of the proposed technique, there is provided a system comprising: - providing a first signal representative of at least a movement of the user's head to a machine learning model, and thereby obtaining an output representative of the user's emotion, the machine learning model having been pre-trained using a database of model signals representative of the head movements of the model user and associated with at least an emotion of the model user; - processing a second signal representative of a physiological parameter of the user to obtain an emotion of the user; A system is provided that is configured to:

[0010] The proposed technique allows for identifying a user's emotions using limited hardware without requiring any conscious interaction from the wearer. Indeed, the proposed technique does not require multiple signals representing the movements of multiple body markers to be available. Rather, it is only necessary to provide a first signal representing head movement as far as the wearer's movements are concerned. In some embodiments, the required hardware can even be embedded in a single head-worn device. The specificity of the first and second signals, combined with a specific approach that first identifies emotions from machine learning-based analysis of the first signal and then confirms emotions from processing the second signal, surprisingly allows for better results than other possible approaches.

[0011] Optionally, the system further comprises a wearable device comprising sensors configured to sense at least head movements of the user and physiological parameters of the user when the wearable device is worn by the user, it being more convenient for the wearer if all necessary sensors are grouped together in a single head-worn device rather than in multiple devices.

[0012] Optionally, the database of model signals represents only the head movements of the model user, which allows minimizing the storage space required to implement the proposed method.

[0013] Optionally, the database of model signals represents the model user's head movements to a greater extent than the model user's body movements. This makes it possible to provide an emotion prediction service to a group of wearers with different versions depending on the equipment of each wearer. A simpler version of the service is based only on head movements when the wearer is equipped with a single motion sensor integrated into a head-worn device. A more complex version can provide further adjustments, for example taking into account hand movements, when the wearer is additionally equipped with hand- or wrist-worn motion sensors. To still achieve good results with the simpler version, it is recommended that the database of model signals represents the model user's head movements sufficiently widely, especially compared to the movements of other body parts of the model user.

[0014] Optionally, the system further comprises a control interface configured to generate a control signal based on the detected or confirmed emotional state.

[0015] For example, a smart eyewear system may be provided to a user to wear throughout daily activities, including, for example, walking through a crowd. Such a smart eyewear system may include various sensors, particularly an IMU that may be configured to acquire a first signal, and a camera and audio recording system that may be configured to acquire photographs, videos, and / or audio recordings. The first signal may be acquired at a current time and thus provided in real time to a machine learning model indicating a detected emotion. A test, such as checking a predetermined criterion, may then be automatically performed using the detected emotion. An example of such a criterion may be checking whether the detected emotion is considered positive, for example, by corresponding to a positive emotional valence and / or a high arousal intensity of the emotion. In this example, if the detected emotion is considered positive, the camera may automatically activate to take a photo that is labeled according to the detected emotion. Furthermore, during the user's daily activities, the user may actively take photos, for example, using the camera and label these photos according to the user's own perceived emotional state. These labels may be associated with the detected emotions and provided to the machine learning model as labeled training data. Instructions may then be generated to identify photos from the set of labeled photos based on those labels that have the most positive emotional ratings, which may then be further suggested to the user for printing, visualization, or as an album to share with friends and / or family.

[0016] Optionally, the machine learning model may be configured to output, along with the detected emotion, a confidence level associated with the detected emotion, and the photo may be further labeled according to the confidence level. Considering the above example, the confidence level threshold may be used as a criterion for filtering the labeled photos. This makes it possible, for example, to ignore any photos having a label representing a detected emotion associated with an insufficient confidence level.

[0017] Optionally, the system further includes a wearable device having an actuatable and / or controllable optical function, such as a refractive function and / or a transmissive function, the wearable device configured to actuate and / or control the optical function upon receiving a control signal.

[0018] For example, a user may be provided with an electrochromic wearable that can switch between different possible shades. The appearance and comfort of the shades may be pre-programmed. For example, a user interface application for the electrochromic wearable, which may be available on a mobile device, may allow the user to trigger one or another shade management algorithm suggested based on the user's perceived needs for light environment management. Separately, a machine learning model may also be run on the user interface application and trained to detect the user's emotions based on multiple signals obtained from sensors on the device worn by the wearer, including, for example, an IMU, a heart rate sensor, and a respiration sensor. The detected emotions may be time-stamped and stored. Furthermore, any shade changes triggered manually by the user or automatically by the shade management algorithm may also be identified as events that are time-stamped and stored. By associating detected emotions with events that have approximately the same timestamp, based on the timestamp itself or by other known means, a co-learning approach may improve both the machine learning model that outputs the detected emotions and the shade management algorithm. An advantage of this co-learning approach is that it avoids encountering inappropriate shades in certain social / environmental emotional situations. For example, a user may generally desire to wear dark-tinted eyewear outdoors in daylight saving time situations. However, the user may also desire more pronounced tints when engaged in close social interactions, such as face-to-face discussions, with others to better express the user's emotions through their eyes. This allows the other party to increase their level of engagement in the relationship by perceiving and detecting emotions through eye contact. Furthermore, head dynamics induced by postures and movements in such social interactions may help distinguish between different scenarios and enable emotion detection. Another advantage of the co-learning approach is that tint changes can be used as user feedback.This includes any shade change that is manually triggered by the user, and any shade change that is automatically triggered by a shade management algorithm and does not provoke a negative reaction by the user, such as manually reversing the shade change or expressing perceptible discomfort due to changes in posture and / or physiological parameters.

[0019] Optionally, the system further comprises a feedback module, for example a visual, audio and / or haptic feedback module, configured to provide feedback to the user upon receiving the control signal.

[0020] Optionally, using user feedback through the interface to evaluate the currently detected emotion may help further train the machine learning model. For example, such feedback may be provided voluntarily by the user by selecting a predefined option on the interface and / or in response to an automatic notification. The feedback is, for example, user input that may confirm the detected emotion or may better represent the user's self-assessed current emotion. The machine learning model may use this feedback from the user to evaluate its performance as an asserted level of confidence in the detected emotion represented by its output.

[0021] Optionally, the system further comprises a messaging module configured to send a message to a user upon receiving the control signal.

[0022] Optionally, the system sends a message to the user when the detected user sentiment is detected to be inconsistent with the social sentiment.

[0023] Optionally, the model signal further comprises model context data, and the first signal further comprises context data.

[0024] Optionally, the system further comprises a context data module configured to obtain the context data, the context data module comprising a sensor and / or a communication interface.

[0025] The context data may include time, location, activity data, etc. Optionally, the context data includes an indication of a user wearing the wearable device, and the identified and / or confirmed emotion represents a degree of comfort associated with wearing the wearable device, for example as a rating on a rating scale.

[0026] Optionally, the model signal further includes model user profile data associated with the model user, and the first signal further includes user profile data associated with the user.

[0027] Optionally, the system further comprises a user profile module configured to obtain user profile data, the user profile module comprising a sensor and / or a communication interface.

[0028] The model user profile data and the user profile data may include, for example, gender, weight, height, age, a user identifier, and the like.

[0029] Optionally, the system further comprises a providing module configured to provide a user profile with the identified or confirmed emotion and / or a tagging module configured to tag the context data based on the identified or confirmed emotion.

[0030] For example, if the detected emotion meets a predetermined criterion, such as being a positively indicated emotion, a camera may be activated to capture an image of the user. In one example, the camera is embedded in the front of eyeglass frames worn by the user. Furthermore, the user may optionally be required to face a mirror before the image is captured. In this example, image processing of the captured image may make it possible to identify the user's facial expression and / or posture, which may therefore be used to confirm the detected emotion. Alternatively, a photo of the user may be taken by a friend and shared on a social network. The user's photo and the detected emotion may be tagged as referring to the same event, for example, based on metadata. Multiplying the detected emotion with the results of image processing makes it possible to confirm the detected emotion.

[0031] According to another aspect of the proposed technique, there is provided a method for identifying a user's emotion, the method comprising: - providing at least a first signal representative of a user's head movement as an input to a machine learning model, and thereby obtaining an output representative of the user's emotion, the machine learning model having been pre-trained using a database of model signals representative of a model user's head movement and associated with the model user's emotion; - processing the second signal representative of the physiological parameter of the user to obtain a confirmed emotion of the user; A method is provided that includes:

[0032] According to another aspect of the proposed technique, there is provided a computer program accessible to a processing unit and comprising one or more stored instruction sequences which, when executed by the processing unit, cause the processing unit to perform the above method.

[0033] For a more detailed understanding of the description provided herein and its advantages, reference is now made to the following brief description taken in conjunction with the accompanying drawings and detailed description, wherein like reference numerals represent like parts. [Brief explanation of the drawings]

[0034] [Figure 1] FIG. 1 illustrates a schematic diagram of a machine learning model training algorithm according to an embodiment. [Figure 2] FIG. 1 illustrates a schematic diagram of a machine learning model training algorithm according to an embodiment. [Figure 3] FIG. 1 illustrates a schematic diagram of a trained machine learning model according to an embodiment. [Figure 4] FIG. 1 illustrates a schematic diagram of a trained machine learning model according to an embodiment. [Figure 5] 10A-10C illustrate schematically possible uses for detected and / or confirmed emotions of a user according to an embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0035] The proposed technique aims to identify the user's emotions.

[0036] Emotions may be detected first based on a machine learning model to which signals representing head movements are input. The output of the machine model may then be confirmed using additional information, including signals representing physiological parameters of the user.

[0037] The proposed technique allows to reliably identify a user's emotions using a minimal combination of sensors: for example, signals representing head movements may be obtained from an inertial measurement unit, which may be part of a head-worn device such as glasses, goggles, a headset, earphones, etc., while signals representing the user's physiological parameters may be obtained from physiological sensors, which may be part of the same device, or from a different worn device such as a wristband.

[0038] The proposed technique may be implemented by executing a computer program including instructions stored on a non-transitory storage medium on a processing unit. When executed by the processing unit, the instructions cause the processing unit to perform a method according to the proposed technique. The processing unit may further be operably connected to various devices by one or more wired or wireless communication interfaces.

[0039] For example, a device operably connected to a processing unit may include one or more sensors from a combination of the above sensors, enabling the processing unit to perform various functions, including controlling these sensors and / or processing data sensed by these sensors over time. For example, a device operably connected to a processing unit may include one or more end devices provided to a user, including, for example, one or more user interfaces, such as a graphical UI, a haptic UI, etc. Such user interfaces may enable the processing unit to perform additional functions, including obtaining feedback from the user and / or providing information to the user. More generally, any device operably connected to a processing unit may be selected and provided by one skilled in the art depending on which functions are to be achieved by the computer program described above.

[0040] The proposed technique may find many applications, as both identifying a user's emotions at a given moment and monitoring such emotions in daily life can be beneficial in various fields. Identifying a lens wearer's emotions, such as discomfort, can be useful for decision-making in clinical practice (e.g., remote vision testing) or for sales purposes (e.g., digital retail). Identifying viewers' emotions of media content streamed in a real environment (e.g., a smart house) or a virtual environment (e.g., the metaverse) can help determine whether a desired effect is achieved on the viewer. In the education field, learning sequences can be extended or shortened depending on the learner's assessed concentration level or fatigue progression over time. Determining a patient's emotional state can be used by a social robot to adapt its interactions with the patient. For example, such a social robot can help support the emotional development of children with autism. Monitoring the emotions of vehicle occupants can be used, for example, through V2X communication, to activate specific safety measures based on preset conditions, such as detecting that the driver is angry.

[0041] Definitions are now proposed for various terms and expressions used herein.

[0042] "Wearer" refers to a user whose emotions are to be identified. The wearer wears a head-worn device that includes at least one sensor for detecting a first signal.

[0043] The "first signal" is an analog or numerical signal representing the translational and / or rotational movement of the wearer's head. An inertial measurement unit (IMU) is a combination of known sensors that allows for measuring movement by recording time-stamped three-dimensional position and orientation, translational and rotational velocity, and translational and rotational acceleration. IMUs can be easily integrated into various mounting devices or frames of wearables and have low power consumption compared to other types of sensors, such as cameras. Possible mounting devices include IMU clips, smart ears, true wireless earphones, smartphones that are held pressed against the head, AR / VR headsets, smart glasses, smart frames, etc. The first signal can then be a signal output by an IMU embedded in, for example, a head-mounted device or wearable.

[0044] The "second signal" is an analog or numerical signal representing a physiological parameter of the user. Known methods and biosensors allow for the acquisition of a wide range of physiological parameters, including, for example, galvanic skin response, skin temperature, sweat analytes, or heart rate variability. Suitable sensors can be placed in the user's environment or embedded in various wearable devices, depending on the nature and context of the physiological parameter at hand. The second signal can be output, for example, by an image processing module that receives an image of the user as input. In particular, the input image can show at least the user's face. The input image can originate from a camera, as an exemplary suitable sensor. The second signal can be correlated to an emotion or a list of possible emotions and can represent the user's physiological state, including, for example, facial flushing. The second signal can also represent physiological signs of an emotional state or a physiological response to a stimulus. This applies in particular to various facial expressions, including smiling. Facial expressions can be extracted by processing an image showing the user's face and correlated to an emotion or a list of possible emotions. In an embodiment, a single head-worn device may incorporate all the sensors necessary to acquire both the first and second signals.

[0045] The first and second signals are not necessarily separate signals originating from different sensors, but may relate to different aspects that are combined into a single sensed signal, such as when the IMU is placed on or near the wearer's skin in a position suitable to detect complex movements representing the combined effects of head movements, heartbeat movements, and breathing movements.

[0046] The machine learning model is a model trained to associate one or more inputs, including at least a first signal, with a corresponding output variable representing the wearer's emotion. The machine learning model is trained using labeled training data, in this case signals each representing a model wearer's head movement and each associated with the model wearer's emotion. The first signal may be recorded periodically to form a time series. Based on this time series, the machine learning model may identify output variables representing the wearer's emotion over various time scales: instantaneous (up to a few seconds), short-term (up to a few minutes), or long-term (hours, days, weeks, or more). The machine learning model may be further configured to identify a latent variable representing the user's enduring internal emotional state based on this time series. This latent variable may be used to determine whether the output variable corresponds to a steady, stable emotional state or a transient emotional state.

[0047] In the field of affective computing, emotions can be classified using appropriate metrics.

[0048] A possibility for supervised training is to define a limited list of possible emotion categories (e.g., anger, fear, neutral, sadness, surprise, disgust, happiness), label the training data according to these categories, and train a machine learning model to infer the correct category when fed with the training data. When labeling the training data, emotions can be classified as a function of very simple markers, for example, head movement, as follows: Anger / fear: Head moves up Neutral / Sad ⇒ No movement / Head moves down Surprise: Head moves up Disgust ⇒ Head moves downwards Happiness ⇒ No movement / head moves up.

[0049] A possibility for automatically labeling training data without human supervision is to construct a clustering model, e.g., a neural network, to perform an unsupervised clustering mechanism. Each cluster can then be assigned a corresponding identifier corresponding to the category. The identifier can have a numerical value. The identifier can also be an emotional valence. Emotional valence, also known as hedonic tone, is an exemplary classification metric that refers to the intrinsic attractiveness or aversion of an event, object, or situation. Emotional valence can be used as a way to characterize emotions, for example, anger and fear are negatively valenced, while joy is positively valenced. The emotional valence can be assigned a number of possible values ​​that are used to label the first model signal with a desired number of categories. These values ​​can be further output by the machine learning model and associated with the first signal. The machine learning model can be, for example, a binary classifier between a first category that regroups all positively valenced emotions and a second category that regroups all negatively valenced emotions. Classic metrics that can be associated with a binary classifier can include accuracy, sensitivity, and specificity. Of course, regardless of the number of emotional categories considered, the machine learning model can also be a fuzzy or probabilistic classifier. The machine learning model can address variables and internal emotional states of the wearer according to a variety of possible topologies. Thus, the machine learning model can be a linear or nonlinear Euclidean space classifier.

[0050] A technical challenge is how to interpret the wearer's emotions from this movement when it is detected solely from the head. In response to this technical challenge, the inventors have identified that head movement can therefore reflect emotions evoked by the control and execution of body movements and interactions with the social or physical environment. Head movement further provides the wearer with clues that may be uncorrelated with the wearer's emotions. For example, gender differences have been analyzed for some motor activities, such as walking. This analysis has revealed that the dynamic portion of whole-body movement contains information about gender. As mentioned above, because head movement reflects such dynamics, knowing the wearer's gender is helpful in interpreting the wearer's head movement and correctly identifying the wearer's emotions from the interpreted head movement.

[0051] We now refer to a simple example of a machine learning model training algorithm, as shown in Figure 1. Essentially, a machine learning model aims to recognize certain types of patterns. The model is trained on a set of data, thus providing an algorithm that can be used to infer and learn from that data. This is the so-called "training phase." Once the model is trained, it can be used to infer and make predictions about previously unseen data. This is the so-called "production phase." The goal of a training algorithm, such as that shown in Figure 1, is to optimize the model to reconstruct labeled training data. The two main types of machine learning models are supervised and unsupervised. Supervised learning involves learning a function that maps inputs to outputs based on example input-output pairs. Unsupervised learning, on the other hand, is used to draw inferences and discover patterns from input data without reference to labeled outcomes. Training algorithms, such as those shown in Figure 1, can be supervised or unsupervised.

[0052] A model signal is acquired (2). The model signal represents at least translational and / or rotational movement of the model wearer's head and may further represent movement of other body parts of the model wearer. The model wearer may be a real wearer, in which case the model signal is sensed. The model wearer may also represent a real wearer or a group of real wearers, in which case the model signal is either entirely generated or sampled from a database of sensed first signals corresponding to the real wearers. The model signal may represent a motor activity of the model wearer, such as walking activity. The model signal may be acquired within an appropriate time window corresponding to the detection of such activity and the duration of all or part of such activity. The duration of the time window may depend on contextual data, such as the model wearer's environment or the nature of the activity, or may be set to a predefined value that is not related to the context.

[0053] The model signal is further related, by any known means, to variables representative of the model wearer's emotions.

[0054] The model signal is provided to a machine learning model (4), which then outputs an output variable representing a prediction of the wearer's emotion (6).

[0055] Variables associated with the model signal are also obtained 8 and compared 10 with the output variables. The results of the comparison are provided to the machine learning model as feedback 12. The feedback received after each prediction indicates whether the prediction was correct or incorrect.

[0056] The entire process is repeated with different model signals that collectively form the training dataset. Between iterations, the internal logic of the machine learning model is adapted based on the aggregate feedback with the goal of maximizing the number of correct predictions made by the training dataset.

[0057] We now turn to a more complex example of a machine learning model training algorithm, as shown in Figure 2.

[0058] In this more complex example, for each model signal provided to the machine learning model to perform a prediction, additional data about the model signal is also provided to the machine learning model, and the resulting multimodal and / or temporal analysis performed by the machine learning model generally enables better predictions.

[0059] For example, a set of model signals for a given model wearer may be acquired over an extended period of time while the given model wearer engages in various activities, such as working, walking, running, sleeping, watching television, driving, etc. The given model wearer may further include a device or interface adapted to sense, or more generally, acquire, context data for each model signal. Data is said to be contextual if it refers to elements that may contextualize a first signal and facilitate its interpretation. A context element may be an event, condition, situation, etc. that applies either permanently or temporarily to the model wearer and / or its environment while the model signal is sensed. The context data may indicate one or more types of context elements, including, for example, relevant physiological parameters of the model wearer, socio-demographic information of the model wearer, such as gender, facial expression of the model wearer, noise levels near the model wearer, light levels near the model wearer, geolocation data of the model wearer, timestamps, activities or tasks experienced by the model wearer, etc. The context data may then be acquired (14) and provided to the machine learning model. The machine learning model can then jointly analyze each model signal for the model wearer with the corresponding context data at the time the model signal is sensed. This allows correlations to be built. For example, some head movements may be correlated with specific contextual elements, such as specific activities or groups of similar activities. Emotions may further be correlated with specific contextual elements. Such correlations contribute to training the machine learning model, particularly to improving the accuracy of predicting the emotions of the model wearer.

[0060] In the training data set, model signals representing head movements of a given model wearer may be further tagged with an identifier for the given model wearer. Further, obtaining (16) these identifiers and providing them to the machine learning model enables the model to build correlations between emotions and head movement patterns specific to a model wearer or group of model wearers. As a result, after the training phase, the machine learning model may analyze a first signal differently and, as a result, output different output variables depending on whether the first signal represents head movements of the first wearer or the second wearer.

[0061] The model signals representing the head movements of the same model wearer can be further organized as a set of time series for different aspects of head movement and provided to the machine learning model. In other words, when considering a model signal corresponding to a given time point acquired (2) and provided to the machine learning model (4), one or more historical model signals corresponding to one or more previous time points can be further acquired (18) and provided to the machine learning model. This allows the machine learning model to interpret a series of continuous indicators of the model wearer's head movements up to the given time point, rather than just momentary indicators. The machine learning model can then reliably identify various types of data from the temporal, and possibly multimodal, correlations between the successive model signals for successive time points up to the given time point. Accordingly, the machine learning model can then associate the identified types of data with corresponding emotions and output output variables accordingly. An example of the type of data that can be identified is the model wearer's head movement pattern, such as a brief moment of moving the head up, followed by moving the head down again. Another example of the type of data that can be identified is the estimation or confirmation of a contextual element, such as detected vibrations characteristic of the model wearer riding a bicycle. Yet another example of the type of data that can be identified is an estimate of the overall dynamic body acceleration as a measure reflecting the energy expenditure of the model wearer. Indeed, as already mentioned, analyzing head movement can provide an indication regarding body movement as a whole.

[0062] A proof-of-concept for extrapolating emotions based solely on head movement was developed and described here. Code showing the movements of 15 body markers of a model walker, such as that provided by the BiomotionLab website (https: / / www.biomotionlab.ca / html5-bml-walker / ), was retrieved. From the retrieved code, corresponding motion data for each of the 15 body markers was regenerated for different full-scale emotions, labeled "sad," "happy," and "neutral," respectively. A database was then constructed associating the motion data with the labeled emotions, and the database was split into a training dataset and a test dataset. A logistic regression model was trained using the training database, and the trained model was applied to the test dataset. A recognition score, expressed as an AUC value (AUC stands for "area under the curve"), was determined. For the emotions "happy" and "sad," the recognition score was 0.65 for head data alone and 0.67 for head and hands. For the neutral emotion, a high recognition score of 0.73 was achieved.

[0063] Reference is now made to Figures 3 and 4, which illustrate examples of possible uses in the production stage of previously trained machine learning models, for example as shown in Figure 1 or Figure 2.

[0064] A first signal is first acquired (20) and provided to a machine learning model (22). The machine learning model may, for example, extract one or more of the following data from the first signal: head acceleration, velocity, frequency-amplitude analysis, entropy, and jerk, analyze the extracted data to identify patterns also found in the analysis of the training data, and thereby output output variables (24).

[0065] Various additional data can be further acquired and provided to the machine learning model to benefit from correlations identified in the training phase and to better distinguish between similar emotions. Increased reliability and disambiguation can result, for example, from data fusion from multiple sensors.

[0066] A second signal may be acquired 26 and provided to the model, enabling the model to extract relevant data from the second signal, identify patterns in the multimodal data extracted from the first and second signals, compare the identified patterns to patterns also found in the analysis of the training data, and result in output variables 24. The same principles of multimodal analysis for improved results apply when contextual data is acquired 32 and provided to the model or when a wearer identifier is acquired 34 and provided to the model.

[0067] The wearer identifier may be combined with a user or wearer profile element and processed by a machine learning model to customize emotion identification. For example, the kinematics and dynamics of head movements may vary between individuals regardless of emotion. For example, a downward head movement may be associated with a negative emotion, such as disgust or sadness. If the machine learning model is not provided with any wearer identifier or any wearer profile element associated with the first signal, the machine learning model may simply set one or more thresholds, for example, regarding the amplitude of the downward angle or angular velocity, which, when exceeded, may lead to outputting an output variable representing such a negative emotion. Conversely, if the machine learning model is provided with a wearer identifier or a wearer profile element associated with the first signal, a further possibility is that different values ​​of the above thresholds are set for different wearers.

[0068] For example, assuming walking activity, estimates of postural sway and walking speed derived from the additional data may be combined with the first signal, allowing the machine learning model to de-correlate a portion of the head dynamics contained in the first signal that is attributable solely or predominantly to walking activity from another portion of the head dynamics contained in the first signal that is attributable solely or predominantly to the wearer's emotions.

[0069] For example, the additional data provided by the sensors may further help identify joint events such as climbing stairs or mingling with a fast-moving crowd, which may contribute to detecting emotional shifts in the wearer.

[0070] Additional data related to the context, such as the environment or the wearer's activity, the time of day, previous values ​​of the output variables, etc., can be further used to preprocess or filter the input series of first signals. For example, it is possible to set one or more criteria that may or may not be satisfied by the additional data, and trigger different actions depending on whether the criteria are satisfied. A possible action is to provide the first signal to the machine learning model only if one or more criteria are satisfied, or conversely, not to provide the first signal to the machine learning model if one or more criteria are not satisfied. Another possible action is to modify, for example tag, the first signal before providing it to the machine learning model depending on whether a particular criterion is not satisfied. Considering the example of criteria based on time of day, it is possible to preprocess all first signals by tagging them with a tag indicating daytime or nighttime before providing the tagged first signals to the machine learning model.

[0071] The output variables may be stored for further use by the machine learning model. In particular, emotional-kinematic patterns may have been identified in the training data, so that the output variables representing the wearer's emotions at a given time may also be stored. - To determine the wearer's emotions in the near future, or - To determine whether the wearer's emotions are stable for a set amount of time, or - To detect events that cause emotional transitions in the wearer, either short-term or long-term (in other words, this relates to output variables representing the user's overt or current emotions or to latent variables representing the user's enduring emotions). This could be a hint.

[0072] The output variables and / or latent variables may then be associated with various parameters that indicate whether the expressed emotion is stationary / stable or transitioning. If the expressed emotion is transitioning, an indication of the kinematics of the transition may further be provided. If the expressed emotion is stable, an indication of the duration and / or initial and final times associated with the expressed emotion may further be provided. Such indications may enable enhanced downstream management of the wearer's emotions.

[0073] Another example of a further use of the stored output variables is in combination with requesting feedback from the wearer regarding the emotion represented by the stored output variables. The request may be a pop-up window or an audio or haptic signal, etc., prompting the wearer to interact with the human-machine interface. Feedback may then be obtained in the form of a signal resulting from such interaction or the absence of such interaction before the expiration of a set timer (36). The feedback may confirm or discuss the emotion represented by the stored output variables. The feedback may be associated with the stored output variables and provided to the machine learning model to contribute to further training of the model. In particular, providing wearer feedback associated with a wearer identifier allows the machine learning model to continue learning from individual data.

[0074] In addition to the output variables being output by the machine learning model (24), additional data is further acquired (26, 32, 34, 36) and processed (28) to ascertain the wearer's emotion (30). The additional data includes at least the second signal and may further include contextual data, a wearer identifier, or feedback regarding the ascertained emotion.

[0075] Here, a possible method (28) for processing the second signal will be described by way of example. For example, the second signal may be output by a pupillometer. The lookup table may include a list of entries, each mapping a list of allowable values ​​for the output variable to a corresponding range of pupil size values. In this example, processing the second signal may involve identifying from the second signal the range of pupil sizes to which the wearer's pupil size belongs and then searching the lookup table for an entry corresponding to the identified range. Next, it may be checked whether the value of the output variable is found within the list of allowable values. The wearer's emotions may then be confirmed if the check results in a positive result, or discussed if the check results in a negative result. Naturally, the same principles may be applied to process second signals input by any other sensor or to process any other type of additional data. For example, if the second signal is received from a sensor such as a heart rate monitor and / or a respiration rate sensor, the lookup table entries may relate to ranges of heart rate and / or respiration values.

[0076] Alternatively, the machine learning model may be conceptually separated into at least a first internal model that receives a first signal and outputs output variables (24) and a second internal model that receives a second signal, processes it (28), and outputs a processing result. The output variables and the output processing result may then be processed together to ascertain the emotion represented by the output variables (30).

[0077] Reference is now made to FIG. 5, which provides a non-exhaustive list of possible uses for detected and confirmed user emotions.

[0078] After the emotion is ascertained (30), the output variables may be processed (38) to generate control signals (40), which may include instructions for various hardware and / or software modules.

[0079] The messaging service receiving the control signal can then, at the appropriate time, send an appropriate message (42), e.g., via an application and / or device, e.g., via a wearable device such as a mixed reality headset. The message can be intended for the wearer or another user. The message can convey positive reinforcement of detected and confirmed positive emotions, and can also convey alerts, warnings, or suggestions when negative emotions are detected and confirmed.

[0080] The wearer visual profile manager may create and deliver a visual live ergonomic wearer experience profile (44) that includes a set of preferences that are automatically adjusted to the wearer in real time based on the output variables. The adjustments may be further based on contextual data including, for example, ambient light levels. The adjustments may be further based on latent variables. The visual profile may be automatically applied when operating a device having a visual interface, such as a computer, tablet, smartphone, active optical element placed in front of one or both of the wearer's eyes, head-mounted device adapted to display visual elements to the wearer, etc.

[0081] The data labeler can be provided with the time-stamped data. The control signal can include instructions for the data labeler to label (46) any time-stamped data with a timestamp corresponding to a time associated with an output variable representing a detected emotional event of the wearer. Several types of time-stamped data can be accessed and labeled by the data labeler, including collected photographs, time-stamped data about situations such as environmental features, visual ergonomic specifications, visual, emotional, and motor tasks, time-stamped physical data about the user such as geolocation, head movement, lighting, weather, crowding, time-stamped data flows such as music, phone calls, messages, application notifications, etc.

[0082] The active optical function controller may control various optical functions, including activating filters, displaying images in augmented and / or virtual reality, activating or adjusting transmission functions, e.g., to achieve a desired hue and / or color intensity, controlling refraction functions, e.g., to achieve magnification or focus for a particular gaze direction and / or a particular gaze distance, etc. The control signal may further include instructions for the active optical function controller to control (48) the optical functions in a differentiated manner in response to at least the output variables, and optionally based on latent variables, contextual data, wearer profile data, etc. As a result, it is possible, for example, to bridge the wearer's lifestyle, behavior, and emotions and customize the optical functions for the wearer accordingly. Furthermore, it is possible to detect and confirm the wearer's emotions and the simultaneous occurrence of specific elements in the field of view, thereby controlling the active refraction function at specific gaze directions and / or specific gaze distances that form a region of interest corresponding to the location of the specific elements in the wearer's field of view. Furthermore, emotions can be detected and confirmed, and concurrent contextual data about the wearer's environment, such as a sudden change in ambient light intensity, can be obtained, thereby triggering a change in the transmittance function to compensate for this sudden change.

[0083] The feedback module may be configured to provide 50 feedback to the user about the user's environment, contextual data, task, long-term emotions, etc. An example of possible feedback may correspond to automatically taking a photo using a camera embedded in the head-worn device or recording sound using an embedded microphone whenever the output variable represents, for example, the wearer's happiness and the latent variable further represents the constancy of the emotion over sufficient time. The taken photo and / or recorded audio may then be stored for later retrieval and provision to the user.

[0084] All the above modules can be activated and modulated in a different way for each possible emotion of the wearer that can be represented by the output variables. A further distinction can be based on whether the expressed emotion is stable or transient.

[0085] Here we briefly describe some use cases.

[0086] In a first use case, the wearer is considered to be focused on a task, and the zone transmission function is activated to dim ambient light by default to avoid disturbing the wearer and promote focus. Throughout the task, a first signal is acquired periodically, for example, every few seconds or every minute. The first signal may be filtered to meet specific measurement conditions. If the specific measurement conditions are met for the predetermined first signal, the predetermined first signal is provided to a machine learning model; otherwise, the predetermined first signal is ignored. An exemplary purpose of such filtering as a preprocessing action may be to eliminate outliers. For each provided first signal, the machine learning model outputs a corresponding output variable representing the wearer's emotion, as derived at least from an analysis of the provided first signal. The output variables are collectively interpreted to monitor the wearer's emotional state over time and, more specifically, to check the wearer's wellness or boredom. If boredom is detected, a possible action may be to automatically suggest to the wearer to pause their work with an appropriate message. Another possible action may be to discontinue dimming the ambient light to further encourage the wearer to take a break.

[0087] In a second use case, a consumer tries on eyeglasses, either physically or virtually. While trying on eyeglass frames, the consumer may verbally express a subjective feeling. The objective feeling may be further detected by a machine learning model that is provided with a first signal and outputs an output variable. The objective feeling may be further confirmed by processing a second signal. The confirmed objective feeling and the expressed subjective feeling may be compared to identify a match or a mismatch. In the case of a mismatch, for example, if the subjective feeling is positive but the objective feeling is not, a warning or recommendation to try on another eyeglass frame may be generated despite the positive subjective feeling.

Claims

1. - providing (22) a first signal representative of at least a movement of a user's head to a machine learning model, and thereby obtaining (24) an output representative of a detected emotion of said user, said machine learning model having been pre-trained using a database of model signals representative of head movements of at least one model user and associated with at least an emotion of said at least one model user; - processing (28) a second signal representative of a physiological parameter of said user to obtain (30) a confirmed emotion of said user; 2. A system configured to:

2. 2. The system of claim 1, further comprising a wearable device including a sensor configured to detect (20, 26) at least the movement of the head of the user and the physiological parameters of the user when the wearable device is worn by the user.

3. - the database of model signals represents movements of only the head of the model user, or A system according to claim 1 or 2, wherein the database of model signals represents the head movements of the model user to a greater extent than the body movements of the model user.

4. The system of any one of claims 1 to 3, further comprising a control interface configured to generate (40) a control signal based on the detected emotion of the user or the confirmed emotion of the user.

5. 5. The system of claim 4, further comprising a wearable device having an actuatable and / or controllable optical function, the wearable device configured to actuate and / or control (48) the optical function according to the control signal.

6. The system of claim 4 or 5, further comprising a feedback module configured to receive (36) feedback from the user of the detected emotion of the user or the confirmed emotion of the user.

7. The system of any one of claims 4 to 6, further comprising a messaging module configured to send (42) a message to the user according to the control signal.

8. The system of any one of claims 1 to 7, wherein the model signal further comprises model context data and the first signal further comprises context data.

9. the context data includes indicators of the user wearing the wearable device; and The system of claim 8, wherein the identified emotion of the user and / or the confirmed emotion of the user represents a degree of comfort associated with wearing the wearable device.

10. The system of claim 8 or 9, further comprising a context data module configured to acquire (32) the context data, the context data module comprising a sensor and / or a communication interface.

11. 11. The system of claim 1, wherein the model signal further comprises model user profile data associated with the model user, and the first signal further comprises user profile data associated with the user.

12. The system of claim 11 , further comprising a user profile module configured to acquire (34) the user profile data, the user profile module comprising a sensor and / or a communication interface.

13. a provisioning module configured to provide (44) a user profile with identified or confirmed emotions, and / or a tagging module configured to tag (46) the context data based on the identified or confirmed emotion; The system of any one of claims 1 to 12, further comprising:

14. 1. A method for identifying a user's emotion, comprising: - providing (22) at least a first signal representative of the user's head movements as input to a machine learning model, and thereby obtaining (24) an output representative of the detected emotion of the user, the machine learning model having been pre-trained using a database of model signals representative of the head movements of at least one model user and associated with the emotions of the at least one model user; - processing (28) a second signal representative of a physiological parameter of said user to obtain (30) a confirmed emotion of said user; A method comprising:

15. 15. A computer program accessible to a processing unit and comprising one or more stored sequences of instructions that, when executed by said processing unit, cause said processing unit to perform the method of claim 14.