Method and device for controlling an on-board vehicle system using standardized sets of action units
By normalizing facial emotion predictions using standardized action units, the method improves accuracy and safety in vehicle systems, addressing challenges related to variations in facial expressions and lighting conditions.
Patent Information
- Authority / Receiving Office
- FR · FR
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-18
- Publication Date
- 2026-03-20
AI Technical Summary
Existing facial emotion recognition systems for vehicle drivers face challenges in accuracy due to variations in facial position, lighting, occlusions, and differences in emotional expressions among individuals, leading to reduced performance and compromised driving safety.
A method that normalizes facial emotion predictions by using standardized sets of action units, obtained by averaging and normalizing features from driver images, allowing for improved facial emotion recognition without requiring labeled data for each driver.
Enhances the accuracy of facial emotion recognition in vehicle systems, improving driver assistance and autonomous vehicle operations, and ultimately enhancing road safety.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Title of the invention: Method and device for controlling an embedded vehicle system using standardized sets of action units. Technical field
[0001] The invention relates to methods and devices for controlling an embedded system of a vehicle based on a prediction of facial emotions of a driver of the vehicle and more particularly on the control of this embedded system when the prediction of facial emotions is obtained as output of a facial emotion prediction model trained when a set of action units is presented as input to said model. Technological background
[0002] It is known that systems can predict facial emotions from images of people's faces. For example, it is known that systems are configured to predict emotions such as anger, disgust, fear, joy, sadness, surprise, neutrality, as well as personal characteristics such as sex, age, facial features, head position and orientation, and gaze direction.
[0003] Some of these facial emotion prediction systems include a facial emotion prediction module, such as a neural network, which is trained on images of people's faces during a learning phase. During the interference phase, facial images of people are acquired, and the facial emotions of these people are predicted as output from the trained facial emotion prediction model when these images are presented as input to this model.
[0004] However, the practical deployment of such facial emotion prediction systems remains a difficult task, particularly when it comes to predicting the facial emotions of vehicle drivers from images of their faces captured while they are driving. Indeed, captured images of a driver's face can be noisy due to variations in facial position, lighting inside the vehicle, facial occlusions, and differences in the expression of emotions among people of different ages, genders, races, or cultures. Noise in images captured inside vehicles reduces the accuracy of predicting driver facial emotions, and improving this accuracy remains a challenge.
[0005] Facial emotion recognition (FER) can be used by human-vehicle communication systems, as the driver's facial expression is one of the key characteristics that allows for the assessment of the driver's current psychological state. In turn, the driver's psychological state determines the cognitive load that the various driver assistance systems can impose on them. Indeed, if the cognitive load is high, it can lead to a decrease in driving performance and compromise driving safety. One of the main problems in the field of facial emotion recognition is the dependence of the expression of a facial emotion on a specific person, in that different people can express the same emotion in different ways.If we consider that the facial emotion recognition system is trained from different users, but that facial emotion recognition must be performed for other users (which generally happens in practice), the accuracy of facial emotion recognition could be reduced.
[0006] This is one of the major challenges of learning a model for facial emotion recognition / classification / prediction.
[0007] Several solutions have been proposed to address this problem. Some of these solutions involve not processing the image of a person's face but using pixels from the image as features on which facial emotion recognition is based. Other solutions extract landmarks from the face and use different distances between these landmarks as features on which facial emotion recognition is based. Still others extract action units from the image of a person's face and use them to create a set of features on which facial emotion recognition is based.
[0008] However, these solutions provide predictions of facial emotions which still need improvement when these facial emotions are predicted from noisy face images due to variations in the position of the driver's face in the images, lighting inside the passenger compartment, facial occlusions, differences in the expression of emotions of people of different ages, sexes, races or cultures. Summary of the present invention
[0009] One object of the present invention is to solve at least one of the problems of the technological background described above.
[0010] An object of the present invention is, for example, to improve the accuracy of predicting facial emotions of a vehicle driver.
[0011] Another object of the present invention is to control an embedded system of a vehicle based on the prediction of facial emotions of the vehicle driver.
[0012] Another object of the present invention is to improve the operation of the vehicle's embedded systems, in particular those related to driver assistance or autonomous vehicle operation.
[0013] Another object of the present invention is to improve road safety.
[0014] According to a first aspect, the present invention relates to a method of controlling an embedded system of a vehicle from a prediction of facial emotions of a driver of the vehicle obtained at the output of a trained facial emotion prediction model when a set of action units, representative of a facial emotion of the driver of the vehicle, is presented at the input of said model, each action unit representing a contraction or relaxation of facial muscles of the driver of the vehicle resulting in movements of parts of the face of the driver of the vehicle; said process being implemented by at least one computer embedded in said vehicle, said process comprising the following steps: - obtaining a first set of action units by averaging second sets of action units, each second set of action units being representative of a facial emotion of the vehicle driver and each second set of action units being obtained from features extracted from at least one image of the vehicle driver's face; - obtaining, from a memory, a third set of action units representative of a neutral facial emotion of the vehicle driver; - obtaining a normalized set of action units by normalizing the first average set of action units from the third set of action units; - obtaining the prediction of the vehicle driver's facial emotions from the trained facial emotion prediction model when the normalized set of action units is presented as input to said model; and - control of the vehicle's on-board system based on the prediction of facial emotions obtained.
[0015] The method implements a preprocessing of features extracted from images of the vehicle driver's face by normalizing these features to make them invariant with respect to the driver's personality. This method is advantageous because it does not require labeled data for each driver of the vehicle whose facial emotions must be recognized, as is the case with reinforcement learning methods used to adjust pre-trained facial emotion prediction models.
[0016] The method improves the performance of facial emotion recognition systems by including this method during the preprocessing phase of image features of these facial emotion recognition / prediction systems.
[0017] When a new driver of a vehicle first uses these facial emotion recognition (prediction) systems, they are asked to look at a camera for a certain period of time (for example, 10 seconds) while displaying a neutral expression. A third set of action units is then obtained from the images captured by this camera. This third set of action units is then stored and used to normalize any initial set of action units for that driver that will be obtained subsequently for the purpose of recognizing facial emotions during future journeys.
[0018] The method increases the accuracy of facial emotion prediction through the use of standardized sets of action units. Considering that a driver's neutral emotion face can be interpreted as an undisturbed face, i.e., as a "zero" emotion state for the driver, then a non-neutral emotion face for that driver can be interpreted as a disturbed face. The standardized sets of action units quantify deviations (divergences) between a first set of action units, representative, on average, of a typical facial emotion of the vehicle's driver, and a third set of action units, representative of a neutral facial emotion for the vehicle's driver. These standardized sets of action units make it possible, on the one hand, to preserve the internal characteristics of the driver's face and, on the other hand, to distinguish a non-neutral facial emotion from a neutral facial emotion for that driver.
[0019] According to a particular and non-limiting embodiment of the present invention, the method further comprises a step of obtaining the third neutral set of action units by averaging fourth sets of action units, each fourth set of action units being representative of a neutral facial emotion of the driver of the vehicle and each fourth set of action units being obtained from features extracted from at least one image of the face of the driver of the vehicle, and a step of storing in said memory the third set of action units of the driver of the vehicle.
[0020] According to a particular and non-limiting embodiment of the present invention, the facial emotion prediction model is trained during a learning phase comprising the following steps: - obtaining fifth sets of action units representative of different facial emotions of different faces of vehicle drivers, each fifth set of action units being obtained by averaging sixth sets of action units, each sixth set of action units being representative of a facial emotion of a vehicle driver and each sixth set of action units being obtained from features extracted from at least one image of the face of said vehicle driver; - obtaining seventh sets of action units representative of neutral facial emotions of different vehicle driver faces, each seventh set of learning action units being representative of a neutral facial emotion of a vehicle driver; - obtaining standardized sets of learning action units, each standardized set of action units being obtained by normalizing each fifth set of action units corresponding to a vehicle driver from the seventh set of action units corresponding to said vehicle driver; - learning the prediction model by comparing the output sets of action units present at the output of said prediction model when the normalized sets of action units are presented at the input of said prediction model.
[0021] According to a particular and non-limiting embodiment of the present invention, the prediction model comprises a trained neural network.
[0022] According to a particular and non-limiting embodiment of the present invention, the standardized set of action units is obtained by dividing the action units of the first set of action units by the action units of the third set of action units.
[0023] According to a particular and non-limiting embodiment of the present invention, the standardized set of action units is obtained from different values of the action units of the first set of action units and the action units of the third set of action units.
[0024] According to a second aspect, the present invention relates to a control device for an embedded system of a vehicle based on a prediction of facial emotions of a driver of the vehicle, the device comprising a memory associated with a processor configured for the implementation of the steps of the method according to the first aspect of the present invention.
[0025] According to a third aspect, the present invention relates to a vehicle, for example of the automobile type, comprising a device as described above according to the second aspect of the present invention.
[0026] According to a fourth aspect, the present invention relates to a computer program which includes instructions adapted for carrying out the steps of the process according to the first aspect of the present invention, in particular when the computer program is executed by at least one processor.
[0027] Such a computer program may use any programming language, and be in the form of source code, object code, or an intermediate code between source code and object code, such as in a partially compiled form, or in any other desirable form.
[0028] According to a fifth aspect, the present invention relates to a computer-readable recording medium on which is recorded a computer program comprising instructions for carrying out the steps of the process according to the first aspect of the present invention.
[0029] On the one hand, the recording medium can be any entity or device capable of storing the program. For example, the medium can include a storage means, such as a ROM, a CD-ROM or a microelectronic circuit-type ROM, or a magnetic recording means or a hard disk drive.
[0030] On the other hand, this recording medium can also be a transmissible medium such as an electrical or optical signal, such a signal being able to be transmitted via an electrical or optical cable, by conventional or radio frequency, by self-directing laser beam, or by other means. The computer program according to the present invention can, in particular, be downloaded from an Internet-type network.
[0031] Alternatively, the recording medium may be an integrated circuit in which the computer program is incorporated, the integrated circuit being adapted to execute or to be used in the execution of the process in question. Brief description of the figures
[0032] Other features and advantages of the present invention will become apparent from the description of the particular and non-limiting embodiments of the present invention below, with reference to the attached Figures 1 to 5, in which:
[0033] [Fig-1] schematically illustrates part of a vehicle passenger compartment, according to a a particular embodiment of the present invention.
[0034] [Fig.2] schematically illustrates a diagram of the steps of a method for controlling an on-board vehicle system from a prediction of facial emotions of a vehicle driver, according to particular and non-limiting examples of the present invention.
[0035] [Fig.3] schematically illustrates a diagram of the steps of a process for obtaining a first set of action units 201, according to a particular and non-limiting embodiment of the present invention.
[0036] [Fig.4] schematically illustrates a diagram of the steps of a method for learning the facial emotion prediction model, according to a particular and non-limiting embodiment of the present invention.
[0037] [Fig.5] schematically illustrates a device configured for the control of an on-board vehicle system based on a prediction of facial emotions of a vehicle driver, according to a particular and non-limiting embodiment of the present invention. Description of examples of achievements
[0038] A method and a device for controlling an on-board system of a vehicle based on a prediction of facial emotions of a driver of the vehicle will now be described in what follows with joint reference to Figures 1 to 5. The same elements are identified with the same reference signs throughout the description that follows.
[0039] The terms "first," "second" (or "firsts," "seconds"), etc., are used in this document by arbitrary convention to allow for the identification and distinction of different elements (such as operations, means, etc.) implemented in the embodiments described below. Such elements may be distinct or correspond to a single element, depending on the embodiment.
[0040] Fig. 1 schematically illustrates part of the passenger compartment of a vehicle 10, according to a particular and non-limiting embodiment of the present invention.
[0041] Vehicle 10 corresponds, for example, to a vehicle with an internal combustion engine, with electric motor(s), or even a hybrid vehicle with an internal combustion engine and one or more electric motors. Vehicle 10 thus corresponds, for example, to a land vehicle comprising a passenger compartment, for example a car, a truck, a bus.
[0042] The vehicle 10 advantageously carries one or more cameras 13 configured for the acquisition of image data of the passenger compartment of the vehicle 10, in particular for the acquisition of images of the face of the driver of the vehicle 10 and of any occupants of the vehicle 10.
[0043] The camera 13 is, for example, arranged in the passenger compartment of the vehicle 10 at the level of the interior rearview mirror. Such a camera 13 has a field of vision corresponding to the front of the passenger compartment including the front seats and possibly one or more rear seats.
[0044] According to one embodiment, the vehicle 10 further comprises another camera 13 (not shown) arranged on the dashboard, for example in a space behind the steering wheel. Such a camera 13 can also be configured for acquiring images of the face of the driver of the vehicle 10. Such a camera corresponds, for example, to the camera of a driver attention monitoring system, known as a DMS (Driver Monitoring System).
[0045] Camera 13 includes, for example, the following elements: - a photosensitive sensor corresponding for example to a matrix of photoreceptors associated for example with a Bayer filter; - an optical assembly arranged in front of the sensor with respect to the scene to be acquired by the sensor, the optical assembly comprising, for example, an arrangement of one or more lenses; and - optionally one or more computers associated with memory and configured for processing images acquired by the sensor.
[0046] The image data received or obtained from each camera, in particular camera 13, are thus representative of one or more images of the interior of the vehicle 10 and in particular of the face of a driver of the vehicle 10. This image data corresponds for example to data representative of a matrix of pixels, color values being for example associated with each pixel, for example according to one or more color channels; for example, the pixel data are coded in the form of RGB values (from the English "Red, Green, Blue" or "Red, Green, Blue" in French).
[0047] The vehicle 10 may also include a display system comprising one or more computers controlling one or more display devices belonging to the display system. The display system includes, for example, a touchscreen 12 and a computer configured to control the display of content from a graphical HMI on the touchscreen 12, for example, integrated into the dashboard 11.
[0048] The computers controlling the display system (such as the computer of the infotainment system, known as the IVI computer (from the English "In-Vehicle Infotainment" or in French "Infodivertissement étoilé")), the various components of the vehicle 10 and a set of embedded driver assistance systems of the AD AS type (from the English "Advanced Driver-Assistance System") form, for example, a multiplexed architecture for the implementation of various services useful for the proper functioning of the vehicle.Computers communicate and exchange data with each other via one or more computer buses, for example a CAN (Controller Area Network), CAN FD (Controller Area Network Flexible Data-Rate), FlexRay (according to ISO 17458), LIN (Local Interconnect Network), or Ethernet (according to ISO / IEC 802-3) type communication bus.
[0049] A method for controlling an embedded system of vehicle 10 based on a prediction of facial emotions of a driver of vehicle 10 is implemented by one or more devices embedded in vehicle 10, for example by one or more embedded computers of vehicle 10, i.e. by one or more processors of the computer(s) in association with one or more memories, by For example, a memory of the computer(s). Examples of the implementation of such a process are described opposite figures 2 to 4 below.
[0050] Figure 2 schematically illustrates a diagram of the steps in the control process. of an embedded vehicle system based on a prediction of facial emotions of a vehicle driver, according to particular and non-limiting embodiments of the present invention.
[0051] The method in [Fig. 2] uses a trained facial emotion prediction model.
[0052] In the field of facial emotion prediction (classification), and particularly in the field of training such models, it is known to define facial emotion classes. Typically, a facial emotion prediction model is trained on training data to provide output data that identify one or more facial emotion classes when input data is presented to these models. During the inference phase of these trained models, data is presented as input to these trained prediction models, which then provide output data that identify one or more facial emotion classes.
[0053] In a step 21, a first set of action units 201 is obtained by averaging second sets of action units. Each second set of action units is representative of a facial emotion of the vehicle driver and each second set of action units is obtained from features extracted from at least one image of the vehicle driver's face.
[0054] Fig. 3 schematically illustrates a diagram of the steps of a process for obtaining (step 21) a first set of action units 201, according to a particular and non-limiting embodiment of the present invention.
[0055] In a step 30, a continuous video stream 301 is captured by one or more cameras 13 of the vehicle 10. This video stream is sampled according to a determined sampling frequency to obtain a set of consecutive images 302 of the driver's face at each cycle of the sampling frequency, for example every second.
[0056] In a step 31, the first set of action units 201 is obtained as output from a facial action coding system (J. Yang et al., Facial Expression Recognition Based on Facial Action Unit, October 2019, Conference: Tenth International Green and Sustainable Computing Conference (IGSC)).
[0057] This facial action coding system, one implementation of which is known as OpenFace (https: / / github.com / TadasBaltrusaitis / OpenFace / wiki / Action-Units), allows the facial movements of a person to be described by action units that make it possible to taxonomite human facial movements according to their appearance on the face. The movements of individual facial muscles Facial expressions are encoded by the system based on slight, instantaneous changes in facial appearance. The system can encode almost any anatomically possible facial expression by breaking it down into specific facial action units that produced the expression. The system allows for an objective description of facial expressions. Essentially, in substep 311 of step 31, each image in the consecutive image set 302 is processed to detect the face of the vehicle driver 10 and to identify landmarks. In substep 312, these detected landmarks are used to identify or calculate image feature values, called action units, for each image in the consecutive image set 302. The action unit values measure specific facial muscle movements for each image in the consecutive image set 302.The values of the action units change according to the facial emotions present on the face of the driver of vehicle 10. The use of action units makes it possible to take only the most informative units of the driver's face when classifying facial emotions and to minimize the number of features in the facial emotion classification models. A set of action units 201 is then formed from the action units obtained for each image of the set of consecutive images 302. In a substep 313, the first set of action units 201, i.e., action unit values, is then obtained by aggregating the action unit values calculated for the images of the set of consecutive images.
[0058] For example, said at least one image of the face of the driver of the vehicle 10 can be received by the computer(s) implementing the control method of [Fig. 2] from the camera 13 as the video stream is acquired by the camera 13. The method of [Fig. 2] can then calculate the first set of action units in real time. In an alternative embodiment, the first set of action units can be received from a memory of the vehicle 10 in which this first set of action units is calculated and temporarily stored after acquisition of said at least one image of the face of the driver of the vehicle 10 in order to be processed as explained below.
[0059] In a step 22 of the control process ([Fig.2]), a third set of action units 221 representing a neutral emotion of the vehicle driver is obtained from a memory.
[0060] According to a particular and non-limiting embodiment of the present invention, the third neutral set of action units 221 is obtained by averaging fourth sets of action units, each fourth set of action units being representative of a neutral facial emotion of the vehicle driver and each fourth set of action units being obtained from extracted features of at least one image of the face of the driver of vehicle 10. The third set of action units 231 is stored in memory.
[0061] In a step 23, a normalized set of action units 231 is obtained by normalizing the first set of action units 201 from the third set of action units 221.
[0062] According to a particular and non-limiting embodiment of the present invention, the standardized set of action units 231 is obtained by dividing the action units of the first set of action units 201 by the action units of the third set of action units 221.
[0063] According to a particular and non-limiting embodiment of the present invention, the standardized set of action units 231 is obtained from different values of the action units of the first set of action units 201 and the action units of the third set of action units 231.
[0064] The normalization of the first set of action units by the third set of action units is not limited to the above embodiment examples but extends to any means of obtaining a normalized set of action units which, on the one hand, preserves the internal features of the driver's face and which, on the other hand, distinguishes a non-neutral facial emotion from a neutral facial emotion.
[0065] In a step 24, the prediction of facial emotions of the vehicle driver is obtained from the trained facial emotion prediction model when the normalized set of action units 231 is presented as input to said model.
[0066] In a step 25 of the control process, the vehicle's on-board system 10 is controlled according to the prediction of facial emotions obtained.
[0067] The use of driver facial emotion prediction can be used to control various vehicle on-board systems 10.
[0068] For example, a speed control system can be configured to adapt the speed setting based on the prediction of the driver's facial expressions. A setting set by the driver can thus be capped, for example, if the prediction of the facial expressions of the driver of vehicle 10 indicates that the driver is angry.
[0069] The present invention is not limited to this example of control nor to cruise control but extends to any type of control of any type of vehicle on-board system (ADAS and others).
[0070] Figure 4 schematically illustrates a diagram of the steps of a method for learning the facial emotion prediction model, according to a particular and non-limiting embodiment of the present invention.
[0071] In a step 400, output data 402 can be generated as output of the facial emotion prediction model to be trained when training data 401 are presented as input to the facial emotion prediction model to be trained.
[0072] The input data 401 correspond to normalised sets of learning action units, each normalised set of action units being obtained by normalising a fifth set of action units corresponding to a vehicle driver from a seventh set of action units corresponding to said vehicle driver.
[0073] The fifth sets of action units are representative of different facial emotions of different vehicle driver faces, each fifth set of action units being obtained by averaging sixth sets of action units, each sixth set of action units being representative of a facial emotion of a vehicle driver and each sixth set of action units being obtained from features extracted from at least one image of the face of said vehicle driver.
[0074] The seventh sets of action units are representative of neutral facial emotions of different vehicle driver faces, each seventh set of learning action units being representative of a neutral facial emotion of a vehicle driver.
[0075] According to a particular and non-limiting embodiment of the present invention, each standardized set of learning action units can be obtained by dividing the action units of a fifth set of action units by the action units of a seventh set of action units.
[0076] According to a particular and non-limiting embodiment of the present invention, each standardized set of learning action units is obtained from different values of the action units of a fifth set of action units and the action units of a seventh set of action units.
[0077] The output data 402 are representative of the emotion prediction of vehicle drivers which are provided as output of the facial emotion prediction model when the input data 401 are presented as input to the facial emotion prediction model.
[0078] The ground reality data associates the 401 input data with predictions of facial emotions that are experimentally defined and annotated by an operator.
[0079] In a step 410, internal parameters of the facial emotion prediction model to be trained can be adjusted to minimize a loss function quantifying the differences between the output data 402 and the real-world data 411, i.e., the expected facial emotion predictions when the data Training 401 are presented as input to the facial emotion prediction model.
[0080]
[0081]
[0082]
[0083]
[0084]
[0085]
[0086]
[0087]
[0088]
[0089]
[0090]
[0091] According to a particular and non-limiting embodiment of the present invention, the facial emotion prediction model may include a trained deep neural network comprising a set of artificial neuron layers. For example, each artificial neuron is a perceptron, that is, a linear classifier generally comprising several inputs and a single output, and characterized by an activation function, weights (or synaptic coefficients) and a bias (or threshold). For example, a perceptron with n inputs (x^ ..., x„) and a single output o is defined by nweights (Mq, and a bias (or threshold) #: Read o = j\z> = \ 1 1 0 otherwise The output o then results from applying the Heaviside function to the postsynaptic potential z given by: In - 6 with a non-linear activation function H(x) given for example by: Vr eP, € € H(x) - 0 if x < 0 1 if x > 1 The internal parameters of the neural network are then connection weights and biases for the set of perceptrons. The present invention is not limited to this definition of perceptron nor to the use of other basic elements forming a layer of the neural network. It is also not limited by the number of perceptrons (or other basic elements) used per layer, nor by the number of layers. The internal parameters of the neural network are usually weights and biases, regardless of the basic elements of the neural network layers. During the learning phase, the deep neural network provides output data when training data is presented to its input, and the deep neural network is trained when it provides output data that corresponds to the real-world data. More specifically, the learning phase of a neural network is iterative. At each iteration, output data is obtained from the neural network (step 400) and the internal parameters of the deep neural network are adjusted (step 410) by minimizing a loss function which can be of the maximum likelihood type, i.e. so that the output data is as close as possible to the real-world data.
[0092] According to a particular and non-limiting embodiment of the present invention, during the learning phase, the internal parameters of the deep neural network can be optimized by backpropagation and gradient descent based, for example, on the method described by Robbins and Monro (Robbins, H. and S. Monro (1951). “A Stochastic Approximation Method.” In The Annals of Mathematical Statistics 22.3, pp. 400-407).
[0093] Fig. 5 schematically illustrates a device 3 configured for the control of an on-board system of a vehicle from a prediction of facial emotions of a driver of the vehicle implemented in the form of a neural network, according to examples of embodiments of the invention.
[0094] Device 3 advantageously corresponds to a data processing device embedded in a vehicle, for example a computer.
[0095] Device 3 is, for example, configured to implement the steps of the processes described opposite Figures 1 to 4. Examples of such a device 3 include, but are not limited to, embedded electronic equipment such as a vehicle's on-board computer or an electronic control unit such as an ECU (Electronic Control Unit). The elements of device 3, individually or in combination, may be integrated into a single integrated circuit, into several integrated circuits, and / or into discrete components. Device 3 may be implemented in the form of electronic circuits or software (or computer) modules, or a combination of electronic circuits and software modules.
[0096] The device 5 comprises one (or more) processor(s) 50 configured to execute instructions for carrying out the steps of the process and / or for executing instructions from the software embedded in the device 5. The processor 50 may include integrated memory, an input / output interface, and various circuits known to those skilled in the art. The device 5 further comprises at least one memory 51, corresponding, for example, to volatile and / or non-volatile memory, and / or includes a memory storage device that may include volatile and / or non-volatile memory, such as EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic disk, or optical disk.
[0097] The computer code of the embedded software(s) including the instructions to be loaded and executed by the processor is for example stored on memory 51.
[0098] According to various particular and non-limiting embodiments, the device 5 is coupled in communication with other similar devices or systems and / or with communication devices, for example a TCU (Telematic Control Unit), for example via a communication bus or through dedicated input / output ports.
[0099] According to a particular and non-limiting embodiment, the device 5 includes a block 52 of interface elements for communicating with external devices. The interface elements of the block 52 include one or more of the following interfaces: - radio frequency RF interface, for example of the Wi-Fi® type (according to IEEE 802.11), for example in the 2.4 or 5 GHz frequency bands, or of the Bluetooth® type (according to IEEE 802.15.1), in the 2.4 GHz frequency band, or of the Sigfox type using UBN (Ultra Narrow Band) radio technology, or LoRa in the 868 MHz frequency band, LTE (Long-Term Evolution), LTE-Advanced; - USB interface (from the English "Universal Serial Bus" or "Universal Serial Bus" in French); - HDMI interface (from the English "High Definition Multimedia Interface", or "High Definition Multimedia Interface" in French); - LIN interface (from the English "Local Interconnect Network", or in French "Réseau interconnecté local").
[0100] According to another particular and non-limiting embodiment, the device 5 includes a communication interface 53 which enables communication with other devices (such as other computers in the embedded system) via a communication channel 54. The communication interface 53 corresponds, for example, to a transmitter configured to transmit and receive information and / or data via the communication channel 54. The communication interface 53 corresponds, for example, to a wired network of the CAN (Controller Area Network), CAN FD (Controller Area Network Flexible Data-Rate), FlexRay (standardized by ISO 17458) or Ethernet (standardized by ISO / IEC 802-3) type.
[0101] Internal parameters (weights) of the pre-trained neural network can for example be received via the communication interface 53.
[0102] According to a particular and non-limiting embodiment, the device 5 can provide output signals to one or more external devices, such as a display screen 55, touch or not, one or more speakers 56 and / or other peripherals 57 via output interfaces 58, 59 and 60 respectively. According to a variant, one or more of the external devices is integrated into the device 5.
[0103] Of course, the present invention is not limited to the embodiments described above but extends to a method of controlling an on-board system of a vehicle based on a prediction of the facial emotions of a vehicle driver which would include secondary steps without falling outside the scope of the present invention. The same would apply to a device configured for implementing such a method.
[0104] The present invention also relates to a vehicle, for example an automobile or more generally an autonomous land-powered vehicle, comprising the device 5 of [Fig.5].
Claims
Demands
1. Method of controlling an embedded system of a vehicle from a prediction of facial emotions of a driver of the vehicle obtained at the output of a trained facial emotion prediction model when a set of action units, representative of a facial emotion of the driver of the vehicle, is presented at the input of said model, each action unit representing a contraction or relaxation of facial muscles of the driver of the vehicle resulting in movements of parts of the face of the driver of the vehicle;said method being implemented by at least one computer embedded in said vehicle, said method comprising the following steps: - obtaining (21) a first set of action units by averaging second sets of action units, each second set of action units being representative of a facial emotion of the driver of the vehicle and each second set of action units being obtained from features extracted from at least one image of the face of the driver of the vehicle; - obtaining (22), from a memory, a third set of action units representative of a neutral facial emotion of the driver of the vehicle; - obtaining (23) a normalized set of action units by normalizing the first average set of action units from the third set of action units;- obtaining (24) the prediction of facial emotions of the vehicle driver from the trained facial emotion prediction model when the normalized set of action units is presented as input to said model; and - controlling (25) the vehicle's on-board system based on the obtained facial emotion prediction.
2. A method according to claim 1, further comprising a step of obtaining the third neutral set of action units by averaging a fourth set of action units, each fourth set of action units being representative of a neutral facial emotion of the vehicle driver and each fourth set of action units being obtained from features extracted from at least one image of the vehicle driver's face, and a memorization step in said memory of the third set of action units of the vehicle driver.
3. A method according to claim 1, wherein the facial emotion prediction model is trained during a learning phase comprising the following steps: - obtaining fifth sets of action units representative of different facial emotions of different faces of vehicle drivers, each fifth set of action units being obtained by averaging sixth sets of action units, each sixth set of action units being representative of a facial emotion of a vehicle driver and each sixth set of action units being obtained from features extracted from at least one image of the face of said vehicle driver;- obtaining a seventh set of action units representing neutral facial emotions of different vehicle drivers' faces, each seventh set of learning action units being representative of a neutral facial emotion of a vehicle driver; - obtaining normalized sets of learning action units, each normalized set of action units being obtained by normalizing each fifth set of action units corresponding to a vehicle driver from the seventh set of action units corresponding to said vehicle driver; - learning the prediction model by comparing the output sets of action units present at the output of said prediction model when the normalized sets of action units are presented as input to said prediction model.
4. A method according to claim 1, wherein the prediction model comprises a trained neural network.
5. A method according to any one of the preceding claims, wherein the standardized set of action units is obtained by dividing the action units of the first set of action units by the action units of the third set of action units.
6. A method according to any one of the preceding claims, wherein the standardized set of action units is obtained from different values of the action units of the first set of action units and the action units of the third set of action units.
7. A computer program comprising instructions for carrying out the method according to any one of the preceding claims, when such instructions are executed by at least one processor.
8. A computer-readable recording medium on which is recorded a computer program comprising instructions for carrying out the steps of the process according to any one of claims 1 to 6
9. 1 d O. Device (5) for controlling an on-board system of a vehicle from a prediction of facial emotions of a driver of the vehicle, said device (5) comprising a memory (51) associated with at least one processor (50) configured for the implementation of at least one step of the method according to any one of claims 1 to 6.
10. Vehicle comprising the device (5) according to claim 9.
Citation Information
Patent Citations
Expression recognition method and system based on facial expression coding system and electronic equipment
CN112016368A
Image Processing System for Extracting a Behavioral Profile from Images of an Individual Specific to an Event
US20210326586A1