Method and device for controlling an on-board system of a vehicle based on a prediction of facial emotions of a vehicle driver

The method improves facial emotion prediction in vehicle drivers by using progressive refinement through multiple classification models, addressing noise issues and enhancing the accuracy of emotion interpretation for better system control and safety.

FR3166358A1Pending Publication Date: 2026-03-20STELLANTIS AUTO SAS +1
1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
FR · FR
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-17
Publication Date
2026-03-20

AI Technical Summary

Technical Problem

Existing facial emotion prediction systems for vehicle drivers face challenges in accuracy due to noise in captured images caused by variations in face position, lighting, and emotional expressions, which affects the performance of driver assistance and autonomous vehicle operations, and compromises road safety.

Method used

A method using progressive refinement of facial emotion prediction through multiple classification models, each trained to optimize the classification of facial emotions at different levels, including binary, probability, and identification models, to improve accuracy by filtering and discriminating facial muscle movements.

Benefits of technology

Enhances the accuracy of facial emotion prediction in vehicle drivers, enabling better control of embedded systems and improving road safety by accurately interpreting driver emotions.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The present invention relates to a method and device for controlling an embedded vehicle system based on a prediction of the facial emotions of the vehicle's driver. The method implements a progressive prediction of the driver's facial emotions based on facial emotion classification models trained on various training datasets dedicated to detecting certain facial emotions associated with facial emotion classes within a group of facial emotion classes. Figure 2 (for the abstract)
Need to check novelty before this filing date? Find Prior Art

Description

Title of the invention: Method and device for controlling an embedded system of a vehicle based on a prediction of facial emotions of a vehicle driver technical field

[0001] The invention relates to methods and devices for information processing, and more particularly to methods and devices for controlling an embedded system of a vehicle based on a prediction of the facial emotions of a vehicle driver. Technological background

[0002] Systems for predicting facial emotions from images of people's faces are known. For example, systems configured to predict emotions such as anger, disgust, fear, joy, sadness, surprise, neutrality, excitement, gaze direction, head orientation, as well as personal characteristics such as sex and age, facial features, or even the position and orientation of the head and gaze direction of these people are known.

[0003] Some of these systems are based on learning deep neural networks. These systems usually include a trained facial emotion prediction module based on captured facial images of people. The prediction model is typically based on a neural network that is trained during a learning phase. During the interference phase, facial images of people are acquired, and the emotions of these people are predicted from the trained facial emotion prediction model when these images are presented as input to this model.

[0004] However, the practical deployment of such facial emotion prediction systems remains a difficult task, particularly when working with images of drivers' faces captured while they are driving. Indeed, the captured facial images can be noisy due to variations in the driver's face position, lighting inside the vehicle, facial occlusions, and differences in the expression of emotions among people of different ages, genders, races, or cultures. Noise in images captured inside vehicles reduces the accuracy of driver facial emotion prediction, and improving this accuracy remains a challenge. Summary of the present invention

[0005] One object of the present invention is to solve at least one of the problems of the technological background described above.

[0006] An object of the present invention is, for example, to improve the accuracy of predicting facial emotions of a vehicle driver.

[0007] Another object of the present invention is to control an embedded system of a vehicle based on the prediction of facial emotions of the vehicle's driver.

[0008] Another object of the present invention is to improve the operation of the vehicle's embedded systems, in particular those related to driver assistance or autonomous operation of the vehicle.

[0009] Another object of the present invention is to improve road safety.

[0010] According to a first aspect, the present invention relates to a method for controlling an embedded system of a vehicle based on a prediction of the facial emotions of a driver of the vehicle, said method being implemented by at least one computer embedded in said vehicle, said method comprising the following steps: - obtaining a sample of image features from at least one image of the driver's face, said sample of image features comprising action units representing contractions or relaxations of facial muscles of the driver of the vehicle resulting in movements of parts of the driver's face; - obtaining binary data from a first trained facial emotion classification model to associate, with the sample of image features obtained, the binary data whose value indicates whether the sample of image features obtained represents a facial emotion associated with a first class of facial emotion or whether the sample of image features obtained represents a facial emotion from a group of facial emotion classes; - if the value of the binary data indicates that the sample of image features obtained represents the facial emotion associated with the first class of facial emotion then a first prediction of facial emotions of the vehicle driver is given by the facial emotion associated with the first class of facial emotion; - if the binary data value indicates that the obtained image feature sample represents a facial emotion from the facial emotion class group, a probability value vector is obtained from a second facial emotion classification model trained to associate the probability value vector with the obtained image feature sample, each probability value quantifying the veracity of the obtained image feature sample to represent a facial emotion associated with a facial emotion class from a subgroup of facial emotion classes determined from the emotion class group facial, the dimension of the probability value vector being defined by a number of subgroups of facial emotion classes; - selection, based on the probability value vector obtained, of a subgroup of facial emotion from among the subgroups of facial emotion determined from the group of facial emotion classes; - if the selected facial emotion class subgroup includes only one facial emotion class, called the second facial emotion class, then a second prediction of the vehicle driver's facial emotions is given by a facial emotion associated with the second facial emotion class; - otherwise, obtaining an identification data for a third class of facial emotion from a third facial emotion classification model trained to associate with the sample of image features obtained the identification data for the third class of facial emotion whose value indicates that the sample of image features obtained represents a facial emotion associated with a class of facial emotion, called the third class of facial emotion, from the selected subgroup of classes of facial emotion classes, a third prediction of facial emotions of the vehicle driver then being the facial emotion associated with the third class of facial emotion; - obtaining a prediction of the facial emotions of the vehicle driver from the first prediction of facial emotions of the vehicle driver if it exists, from the second prediction of facial emotions of the vehicle driver if it exists, and from the third prediction of facial emotions of the vehicle driver if it exists, and - control of said vehicle's on-board system based on the prediction of the vehicle driver's facial emotions obtained.

[0011] The method predicts the facial emotions of a vehicle driver from a sample of image features of the driver's face. The prediction of these facial emotions differs from the prior art by using a progressive refinement of the prediction, which improves its accuracy compared to the prior art. The progressive prediction of facial emotions takes place in three steps: in the first step, it is determined whether the driver's facial emotion can be predicted by the facial emotion associated with the first class of facial emotions, for example, by a so-called "neutral" facial emotion, or whether the driver's facial emotion belongs to a group of facial emotion classes that are described, for example, as "non-neutral."If it is determined that the facial emotion of the vehicle driver is predicted by the facial emotion associated with the first class of facial emotion (emotion "neutral" according to the example) then a first prediction of facial emotions of the driver. The driver's facial emotion is predicted by the facial emotion associated with the first facial emotion class. If not, in a second step, it is determined whether the driver's facial emotion can be predicted by a facial emotion associated with a facial emotion class from a subgroup of facial emotion classes selected from the group of facial emotion classes. If this subgroup of facial emotion classes contains only one facial emotion class, then a second prediction of the driver's facial emotion is given by a facial emotion associated with the second facial emotion class. Otherwise, a third prediction of the driver's facial emotion is determined to be the facial emotion associated with one of the facial emotion classes from the subgroup of facial emotion classes.The facial emotion of the vehicle driver is then predicted from the first, second, and / or third prediction of facial emotions of the vehicle driver.

[0012] The method improves the prediction of the facial emotions of the vehicle driver because it takes into account the natural a priori differences between distinct facial emotions and groups of facial emotions, and constructs a prediction procedure based on these differences.

[0013] The method uses facial emotion classification models at each of the three steps of the process. Each classification model used in a step is trained so that the classification of the sample of image features is optimal at that step. For example, the first classification model is trained only to distinguish feature samples that represent a first class of facial emotion from those that represent other classes of facial emotions within a group of facial emotion classes. The second model is trained only to determine which subgroup(s) of facial emotion class(es) are represented by the image feature samples presented as input to the second model.The third model is trained for each subgroup of facial emotions only so that it can distinguish image feature samples that represent one facial emotion from that subgroup of facial emotions from those that represent any other facial emotion from that subgroup of facial emotions. This process thus improves the prediction of a driver's facial emotions.

[0014] According to a particular and non-limiting embodiment of the present invention, the image feature sample obtained is filtered before obtaining the binary data to limit the number of action units of the image feature sample obtained to those representing contractions or relaxations of facial muscles of the vehicle driver that are discriminating to determine if the sample The image characteristics obtained represent or do not represent the first class of facial emotion.

[0015] According to a particular and non-limiting embodiment of the present invention, the image feature sample obtained is filtered before obtaining the probability value vector to limit the number of action units of the image feature sample obtained to those representing contractions or relaxations of facial muscles of the vehicle driver which are discriminating to determine whether the image feature sample represents one of the facial emotions of the facial emotion class group.

[0016] According to a particular and non-limiting embodiment of the present invention, the image feature sample obtained is filtered, before obtaining the identification data of the third facial emotion class, to limit the number of action units to those representing contractions or relaxations of facial muscles of the vehicle driver that are discriminating to distinguish the facial emotions associated with the facial emotion classes of the selected subgroup of facial emotion classes.

[0017] According to a particular and non-limiting embodiment of the present invention, the first facial emotion classification model is a binary classification model and the third facial emotion classification model of a subgroup of facial emotion classes is a multi-value classification model.

[0018] According to a particular and non-limiting embodiment of the present invention, the first, second and third classification models each comprise a trained neural network.

[0019] According to a second aspect, the present invention relates to a control device for an embedded system of a vehicle based on a prediction of facial emotions of a driver of the vehicle, the device comprising a memory associated with a processor configured for the implementation of the steps of the method according to the first aspect of the present invention.

[0020] According to a third aspect, the present invention relates to a vehicle, for example of the automobile type, comprising a device as described above according to the second aspect of the present invention.

[0021] According to a fourth aspect, the present invention relates to a computer program which includes instructions adapted for carrying out the steps of the process according to the first aspect of the present invention, in particular when the computer program is executed by at least one processor.

[0022] Such a computer program can use any programming language, and be in the form of source code, object code, or a intermediate code between source code and object code, such as in a partially compiled form, or in any other desirable form.

[0023] According to a fifth aspect, the present invention relates to a computer-readable recording medium on which is recorded a computer program comprising instructions for carrying out the steps of the process according to the first aspect of the present invention.

[0024] On the one hand, the recording medium can be any entity or device capable of storing the program. For example, the medium can include a storage means, such as a ROM, a CD-ROM or a microelectronic circuit-type ROM, or a magnetic recording means or a hard disk drive.

[0025] On the other hand, this recording medium can also be a transmissible medium such as an electrical or optical signal, such a signal being able to be transmitted via an electrical or optical cable, by conventional or radio frequency, by self-directing laser beam, or by other means. The computer program according to the present invention can, in particular, be downloaded from an Internet-type network.

[0026] Alternatively, the recording medium may be an integrated circuit in which the computer program is incorporated, the integrated circuit being adapted to execute or to be used in the execution of the process in question. Brief description of the figures

[0027] Other features and advantages of the present invention will become apparent from the description of the particular and non-limiting embodiments of the present invention below, with reference to the attached Figures 1 to 6, in which:

[0028] [Fig-1] schematically illustrates part of a vehicle passenger compartment, according to a example of a particular embodiment of the present invention;

[0029] [Fig.2] schematically illustrates a method of controlling an on-board vehicle system from a prediction of facial emotions of a vehicle driver, according to particular and non-limiting examples of embodiments of the present invention;

[0030] [Fig.3] schematically illustrates obtaining a sample of image characteristics, according to a particular and non-limiting embodiment of the present invention;

[0031] [Fig.4] schematically illustrates the control method of [Fig.2], according to other particular and non-limiting embodiments of the present invention;

[0032] [Fig.5] schematically illustrates a learning phase of a facial emotion classification model, according to a particular and non-limiting embodiment of the present invention;

[0033] [Fig.6] schematically illustrates a device configured for the control of an on-board vehicle system based on a prediction of facial emotions of a vehicle driver, according to a particular and non-limiting embodiment of the present invention. Description of examples of achievements

[0034] A method and a device for controlling an on-board system of a vehicle based on a prediction of facial emotions of a driver of the vehicle will now be described in what follows with joint reference to Figures 1 to 6. The same elements are identified with the same reference signs throughout the description that follows.

[0035] The terms "first," "second" (or "firsts," "seconds"), etc., are used in this document by arbitrary convention to allow for the identification and distinction of different elements (such as operations, means, etc.) implemented in the embodiments described below. Such elements may be distinct or correspond to a single element, depending on the embodiment.

[0036] Fig. 1 schematically illustrates part of the passenger compartment of a vehicle 10, according to a particular and non-limiting embodiment of the present invention.

[0037] Vehicle 10 corresponds, for example, to a vehicle with an internal combustion engine, with electric motor(s), or even a hybrid vehicle with an internal combustion engine and one or more electric motors. Vehicle 10 thus corresponds, for example, to a land vehicle comprising a passenger compartment, for example a car, a truck, a bus.

[0038] The vehicle 10 advantageously carries one or more cameras 13 configured for the acquisition of image data of the passenger compartment of the vehicle 10, in particular for the acquisition of images of the face of the driver of the vehicle 10 and of any occupants of the vehicle 10.

[0039] The camera 13 is, for example, arranged in the passenger compartment of the vehicle 10 at the level of the interior rearview mirror. Such a camera 13 has a field of vision corresponding to the front of the passenger compartment including the front seats and possibly one or more rear seats.

[0040] According to one embodiment, the vehicle 10 further comprises another camera 13 (not shown) arranged on the dashboard, for example in a space behind the steering wheel. Such a camera 13 can also be configured for acquiring images of the face of the driver of the vehicle 10. Such a camera corresponds, for example, to the camera of a driver attention monitoring system, known as a DMS (Driver Monitoring System).

[0041] Camera 13 includes, for example, the following elements: - a photosensitive sensor corresponding for example to a matrix of photoreceptors associated for example with a Bayer filter; - an optical assembly arranged in front of the sensor with respect to the scene to be acquired by the sensor, the optical assembly comprising, for example, an arrangement of one or more lenses; and - optionally one or more computers associated with memory and configured for processing images acquired by the sensor.

[0042] The image data received or obtained from each camera, in particular camera 13, are thus representative of one or more images of the interior of the vehicle 10, this image data corresponding for example to data representative of a pixel matrix, color values ​​being for example associated with each pixel, for example according to one or more color channels; for example, the pixel data are coded in the form of RGB values ​​(from the English "Red, Green, Blue" or "Rouge, vert, bleu" in French).

[0043] The vehicle 10 may also include a display system comprising one or more computers controlling one or more display devices belonging to the display system. The display system includes, for example, a touchscreen 12 and a computer configured to control the display of content from a graphical HMI on the touchscreen 12, for example, integrated into the dashboard 11.

[0044] The computers controlling the display system (such as the computer of the infotainment system, known as the IVI computer (from the English "In-Vehicle Infotainment" or in French "Infodivertissement étoilé")), the various components of the vehicle 10 and a set of embedded driver assistance systems of the AD AS type (from the English "Advanced Driver-Assistance System") form, for example, a multiplexed architecture for the implementation of various services useful for the proper functioning of the vehicle.Computers communicate and exchange data with each other via one or more computer buses, for example a CAN (Controller Area Network), CAN FD (Controller Area Network Flexible Data-Rate), FlexRay (according to ISO 17458), LIN (Local Interconnect Network), or Ethernet (according to ISO / IEC 802-3) type communication bus.

[0045] A method for controlling an on-board system of vehicle 10 based on a prediction of facial emotions of a driver of vehicle 10 is implemented by one or more devices on-board in vehicle 10, for example by one or more on-board computers of vehicle 10, i.e. by one or more processors of the computer(s) in association with one or more memories, for example a memory of the computer(s). Examples of implementation of such a method are described opposite Figures 2 to 4 below.

[0046] Figure 2 schematically illustrates this method of controlling an on-board vehicle system based on a prediction of facial emotions of a vehicle driver, according to particular and non-limiting examples of the present invention.

[0047] The method in [Fig.2] uses trained facial emotion classification models.

[0048] In the field of facial emotion classification, and particularly in the field of model training, it is known to define facial emotion classes. Typically, a facial emotion classification model is trained on training data to provide output data that identify one or more facial emotion classes when input data is presented to these models. During the inference phase of these trained models, data is presented as input to these trained models, which then provide output data that identify one or more facial emotion classes.

[0049] In a step 21 of the control process, a sample of image features 201 is obtained from at least one image of a set of consecutive images of the face of the driver of the vehicle 10, said sample of image features comprising action units representing contractions or relaxations of facial muscles of the driver of the vehicle 10 resulting in movements of parts of the face of the driver of the vehicle 10.

[0050] Fig. 3 schematically illustrates the obtaining (step 21) of a sample of image features 201, according to a particular and non-limiting embodiment of the present invention.

[0051] In a step 30, a continuous video stream 301 is captured by one or more cameras 13 of the vehicle 10. This video stream is sampled according to a determined sampling frequency to obtain a set of consecutive images 302 of the driver's face at each cycle of the sampling frequency, for example every second.

[0052] In a step 31, the image feature sample 201 is obtained as output from a Facial Action Coding System (FACS) representing facial emotions, such as the system known as OpenFace (https: / / github.com / TadasBaltrusaitis / OpenFace / wiki / Action-Units). A facial action coding system is a system for describing facial movements using action units that allow human facial movements to be taxonomicated according to their appearance on the face. The movements of Individual facial muscles are coded by the system based on slight, instantaneous changes in facial appearance. The system can code almost any anatomically possible facial expression by breaking it down into specific facial action units that produced the expression. The system allows for an objective description of facial expressions. Essentially, in substep 311 of step 31, each image in the consecutive image set 302 is processed to detect the face of the vehicle driver 10 and to identify landmarks. In substep 312, these detected landmarks are used to identify or calculate image feature values, called action units, for each image in the consecutive image set 302. The action unit values ​​measure specific facial muscle movements for each image in the consecutive image set 302.The values ​​of the action units change according to the facial emotions present on the face of the driver of vehicle 10. The use of action units makes it possible to take only the most informative units of the driver's face when classifying facial emotions and to minimize the number of features in the facial emotion classification models. The image feature sample 201 is then formed from the action units obtained for each image of the set of consecutive images 302. In a substep 313, the image feature sample 201, i.e., the action unit values, is then obtained by aggregating the action unit values ​​calculated for the images of the set of consecutive images.

[0053] For example, each image feature sample 201 can be received directly by the computer(s) implementing the control method of [Fig.2] from the camera 13 and as the video stream is acquired by the camera 13 or the image feature sample 201 can be received from a vehicle memory 10 in which the images are temporarily stored after acquisition in order to be processed as explained below.

[0054] In a step 22 of the control process ([Fig.2]), a binary data 221 is obtained from a first facial emotion classification model trained to associate, with the sample of image features obtained 201, the binary data 221 whose value indicates whether the sample of image features obtained 201 represents a facial emotion associated with a first class of facial emotion identified by a reference 222 or whether the sample of image features obtained 201 represents a facial emotion from a group of facial emotion classes.

[0055] For example, the first facial emotion class 222 can be associated with a so-called "neutral" facial emotion, and the facial emotion classes group can include several facial emotion classes associated with so-called "non" facial emotions neutral” such as, for example, “anger”, “calm”, “disgust”, “fear”, “joy”, “sadness” or “surprise”.

[0056] If the value of the binary data 221 indicates that the image feature sample obtained 201 represents the facial emotion associated with the first facial emotion class 222 then a first prediction of facial emotions of the vehicle driver is given by the facial emotion associated with the first facial emotion class 222.

[0057] If the value of the binary data 221 indicates that the obtained image feature sample 201 represents a facial emotion from the facial emotion class group, then, in a step 23, a probability value vector 231 is obtained from a second facial emotion classification model trained to associate the probability value vector 231 with the obtained image feature sample 201. Each probability value quantifies the veracity of the obtained image feature sample 201 in representing a facial emotion associated with a facial emotion class from a subgroup of facial emotion classes identified by a reference 241i (i = 1 to N, where N is an integer) and determined from the facial emotion class group. The dimension of the probability value vector is defined by a number N of subgroups of facial emotion classes.

[0058] In a step 24, a subgroup of facial emotion 241i is selected, according to the vector of probability values ​​obtained, from among the subgroups of facial emotion determined from the group of facial emotion classes selected.

[0059] According to a particular and non-limiting embodiment of the present invention, the selected facial emotion subgroup 241i corresponds to the facial emotion subgroup that corresponds to the highest probability value of the obtained probability value vector 231.

[0060] If the selected facial emotion class subgroup 241i (i=1 to N) comprises only one facial emotion class, called the second facial emotion class 241u, then a second prediction of facial emotions of the vehicle driver is given by a facial emotion associated with the second facial emotion class 241^.

[0061] Alternatively, in a step 25; (i=l to N), an identification data piece for a third facial emotion class 25lij(j=l to M, M an integer value) is obtained from a third facial emotion classification model trained to associate with the obtained image feature sample 201 the identification data piece for the third facial emotion class 25lij whose value indicates that the obtained image feature sample 201 represents a facial emotion associated with a facial emotion class, called the third facial emotion class 251^, of the selected subgroup of facial emotion classes 241i, a third prediction of facial emotions of the vehicle driver being then the facial emotion associated with the third class of facial emotion 251ij.

[0062] According to one example, each subgroup of facial emotion classes consists of facial emotion classes that are associated with facial emotions that are most often confused with each other.

[0063] The subgroups of facial emotion classes can be determined experimentally, for example by analyzing a facial emotion recognition confusion matrix that indicates percentages of facial emotion recognition relative to real-life situations. For example, assuming that a group of facial emotion classes includes facial emotion classes such as "anger," "calm," "disgust," "fear," "joy," "neutral," "sadness," and "surprise," seven subgroups of facial emotion classes can be determined.For example, one subgroup of facial emotion classes might include two classes of facial emotions associated with the facial emotions "fear" and "sadness," another might include two classes of facial emotions associated with "anger" and "fear," another might include two classes of facial emotions associated with "neutral" and "sadness," and yet another might include two classes of facial emotions associated with "disgust" and "sadness." Other subgroups of facial emotion classes might contain only a single class. For example, one subgroup of facial emotion classes might include a class of facial emotions associated with the facial emotion "calm," another might include a class of facial emotions associated with the facial emotion "joy," and yet another might include a class of facial emotions associated with the facial emotion "surprise."It can be noted that a subgroup of facial emotion classes may include the first facial emotion class (associated, for example, with the "neutral" facial emotion). This may seem to contradict step 22, which determines the image feature samples that represent a facial emotion associated with the first facial emotion class. However, including the first facial emotion class (associated, for example, with the "neutral" facial emotion) in at least one subgroup of facial emotion classes compensates for imperfections in the first classification model, which might not detect in step 22 that the image feature sample 201 is not representative of the "neutral" facial emotion.

[0064] The present invention is not limited to examples of determining subgroups of facial emotion classes, nor to their number or to the definition of these facial emotion classes, but extends to any type of facial emotion class, to any type of subgroups of emotion classes and to any number of facial emotion classes composing each of these emotion class subgroups.

[0065] In a step 26, a prediction of the facial emotions of the vehicle driver is obtained from the first prediction of facial emotions of the vehicle driver (222) if it exists, from the second prediction of facial emotions of the vehicle driver (241u) if it exists and from the third prediction of facial emotions of the vehicle driver (25 Lj) if it exists.

[0066] According to one exemplary and non-limiting embodiment of the present invention, the method provides a prediction of a single facial emotion for each feature sample 201.

[0067] According to a particular and non-limiting embodiment of the present invention, the first (222), second (241^), and third (25hj) predictions of the vehicle driver's facial emotion can each be associated with a probability value that the image feature sample 201 represents an emotion associated with the first, second, and third facial emotion classes, respectively. For example, the prediction of the vehicle driver's facial emotions (step 26) can then be obtained solely from the first, second, and / or third facial emotion prediction if the probability values ​​associated with the first, second, and / or third facial emotion prediction are greater than a threshold value. The method can then provide a prediction of one, two, or three facial emotions of the vehicle driver with associated probability values.

[0068] In a step 27, at least one on-board system of the vehicle 10 is controlled according to the prediction of the facial emotions of the driver of the vehicle 10 obtained.

[0069] The use of driver facial emotion prediction can be used to control various vehicle on-board systems 10.

[0070] For example, a speed control system can be configured to adapt the speed setting based on the prediction of the driver's facial expressions. A setting set by the driver can thus be capped, for example, if the prediction of the facial expressions of the driver of vehicle 10 indicates that the driver is angry.

[0071] The present invention is not limited to this example of control or to cruise control but extends to any type of control of any type of vehicle on-board system.

[0072] Fig. 4 schematically illustrates a method of controlling Fig. 2, according to other particular and non-limiting embodiments of the present invention.

[0073] According to one example, in a step 215, the image feature sample 201 is filtered before obtaining the binary data 221 (step 22) to limit the number of action units in the image feature sample 201 to those representing contractions or relaxations of facial muscles of the vehicle driver 10 which are discriminating to determine whether the image feature sample 201 represents the first class of facial emotion or not.

[0074] According to one example, in a step 235, the image feature sample 201 is filtered before obtaining the probability value vector 231 (step 23) to limit the number of action units of the image feature sample 201 to those representing contractions or relaxations of facial muscles of the vehicle driver that are discriminating to determine whether the image feature sample 201 represents one of the facial emotions of the facial emotion class group.

[0075] According to a variant of step 235, the image features of the image feature sample 201, optionally filtered, are further scaled.

[0076] According to one example, in each step 252; (i=l to N), the sample of image features 201 is filtered, before obtaining the identification data of the third facial emotion class 251ij, to limit the number of action units to those representing contractions or relaxations of facial muscles of the vehicle driver that are discriminating to distinguish the facial emotions associated with the facial emotion classes of the selected subgroup of facial emotion classes 241i

[0077] According to a particular and non-limiting embodiment of the present invention, the first facial emotion classification model is a binary classification model and the third facial emotion classification model of a subgroup of facial emotion classes is a multi-value classification model.

[0078] The first, second and each third facial emotion classification model is trained during a learning phase.

[0079] Figure 5 schematically illustrates a learning phase of a facial emotion classification model, according to a particular and non-limiting embodiment of the present invention.

[0080] In a step 500, output data 502 can be generated as output of the facial emotion classification model to be trained when training data 501 are presented as input to the facial emotion classification model to be trained.

[0081] In a step 510, internal parameters of the facial emotion classification model to be trained can be adjusted to minimize a loss function quantifying the differences between the output data 502 and the field reality data 511, i.e., the expected output data when the training data 501 are presented as input to the facial emotion classification model. The field reality data are generally determined experimentally and annotated by an operator.

[0082]

[0083]

[0084]

[0085]

[0086]

[0087]

[0088]

[0089]

[0090]

[0091]

[0092] According to a particular and non-limiting embodiment of the present invention, each facial emotion classification model (first, second and third) may include a trained deep neural network that comprises a set of artificial neuron layers. For example, each artificial neuron is a perceptron, that is to say, a classifier. linear generally comprising several inputs and a single output and characterized by an activation function, weights (or synaptic coefficients) and a bias (or threshold). For example, a perceptron with n inputs (x^ ...,x„) and a single output o is defined by nweights (Mq, and a bias (or threshold) 0: o = f(z) = 1 if L^iXi > 0 0 otherwise The output o then results from applying the Heaviside function to the postsynaptic potential z given by: z = 6 with a non-linear activation function H(x) given for example by: f Û SI X < 0 If X > 1 The internal parameters of the neural network are then connection weights and biases for the set of perceptrons. The present invention is not limited to this definition of a perceptron, nor to the use of other basic elements forming a layer of the neural network. It is also not limited to the number of perceptrons (or other basic elements) used per layer, nor to the number of layers. The internal parameters of the neural network are usually weights and biases, regardless of the basic elements of the neural network layers. During the learning phase, the deep neural network provides output data when training data is presented to its input, and the deep neural network is trained when it provides output data that corresponds to the real-world data. More specifically, the learning phase of a neural network is iterative. At each iteration, output data is obtained from the neural network (step 500) and the internal parameters of the deep neural network are adjusted (step 510) by minimizing a loss function which can be of the maximum likelihood type, i.e. so that the output data is as close as possible to the real-world data.

[0093] According to a particular and non-limiting embodiment of the present invention, during the learning phase, the internal parameters of the deep neural network can be optimized by backpropagation and gradient descent based, for example, on the method described by Robbins and Monro (Robbins, H. and S. Monro (1951). “A Stochastic Approximation Method.” In The Annals of Mathematical Statistics 22.3, pp. 400-407).

[0094] The training data used for training each facial emotion classification model includes 201 image feature samples. Each 201 image feature sample includes action units as explained previously.

[0095] For training the first facial emotion classification model, each feature sample (training data) comprises action units that discriminate to determine whether the image feature sample (training data) represents a facial emotion associated with a first facial emotion class or whether the image feature sample (training data) represents a facial emotion from a group of facial emotion classes. The output data corresponding to each image feature sample is a binary value indicating whether the image feature sample represents (binary value equal to 1) the facial emotion associated with the first emotion class or whether the image feature sample represents a facial emotion associated with the group of facial emotion classes.Ground reality data includes samples of image features that are associated with binary data whose values ​​are determined experimentally.

[0096] During the inference phase of the first trained facial emotion classification model, the input data to the model can be either the image feature sample 201 (step 22 of [Fig. 2]) or a filtered image feature sample 201 (step 215) of [Fig. 3]. The output data of the first trained facial emotion classification model is then the binary data 221.

[0097] For training the second facial emotion classification model, each sample of image features (training data) includes action units that are discriminating in determining facial emotions from the group of facial emotion classes. The output data corresponding to each sample of image features is a vector of probability values, each probability value quantifying the veracity of the sample of image features in representing a facial emotion associated with a facial emotion class from a subgroup of facial emotion classes determined from the group of facial emotion classes. The real-world data includes samples of features image which are associated with probability value vectors whose values ​​are determined experimentally.

[0098] During the inference phase of the second trained facial emotion classification model, the input data to the model can be either the sample of image features 201 (step 22 of [Fig. 2]) or a sample of filtered image features 201 (step 235) of [Fig. 3]. The output data of the second trained facial emotion classification model is then the vector of probability values ​​231.

[0099] For training each third facial emotion classification model corresponding to a subgroup of facial emotion classes, each sample of image features (training data) includes action units that are discriminatory in determining which facial emotion associated with a facial emotion class from the subgroup of facial emotion classes is best represented by the sample of features. The output data corresponding to each sample of image features is an identification data point for a third facial emotion class, the value of which indicates that the sample of image features represents a facial emotion associated with a facial emotion class, the third facial emotion class.The ground reality data includes samples of image features that are associated with third-class facial emotion identification data whose values ​​are determined experimentally.

[0100] During the inference phase of the third trained facial emotion classification model, the input data to the model can be either the image feature sample 201 or a filtered image feature sample 201 (step 252i) from [Fig. 3]. The output data of the third trained facial emotion classification model is then the identification data for the third facial emotion class 25 hj.

[0101] Figure 6 schematically illustrates a device 3 configured for the control of a vehicle embedded system based on a prediction of facial emotions of a vehicle driver implemented in the form of a neural network, according to examples of embodiments of the invention.

[0102] Device 3 advantageously corresponds to a data processing device embedded in a vehicle, for example a computer.

[0103] Device 3 is, for example, configured to carry out the steps of the processes described opposite Figures 2 to 5. Examples of such a device 3 include, but are not limited to, embedded electronic equipment such as a vehicle's on-board computer or an electronic control unit such as an ECU (Electronic Control Unit). The elements of device 3, individually or in combination, can be integrated into a single integrated circuit, into several integrated circuits, and / or into discrete components. Device 3 can be implemented as electronic circuits or software (or computer) modules, or a combination of electronic circuits and software modules.

[0104] The device 3 comprises one (or more) processor(s) 30 configured to execute instructions for carrying out the steps of the process and / or for executing instructions from the software embedded in the device 3. The processor 30 may include integrated memory, an input / output interface, and various circuits known to those skilled in the art. The device 3 further comprises at least one memory 31, for example, volatile and / or non-volatile memory, and / or includes a memory storage device that may include volatile and / or non-volatile memory, such as EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic disk, or optical disk.

[0105] The computer code of the embedded software(s) including the instructions to be loaded and executed by the processor is for example stored on memory 31.

[0106] According to various particular and non-limiting embodiments, the device 3 is coupled in communication with other similar devices or systems and / or with communication devices, for example a TCU (Telematic Control Unit), for example via a communication bus or through dedicated input / output ports.

[0107] According to a particular and non-limiting embodiment, the device 3 includes a block 32 of interface elements for communicating with external devices. The interface elements of the block 32 include one or more of the following interfaces: - radio frequency RF interface, for example of the Wi-Fi® type (according to IEEE 802.11), for example in the 2.4 or 5 GHz frequency bands, or of the Bluetooth® type (according to IEEE 802.15.1), in the 2.4 GHz frequency band, or of the Sigfox type using UBN (Ultra Narrow Band) radio technology, or LoRa in the 868 MHz frequency band, LTE (Long-Term Evolution), LTE-Advanced; - USB interface (from the English "Universal Serial Bus" or "Universal Serial Bus" in French); - HDMI interface (from the English "High Definition Multimedia Interface", or "High Definition Multimedia Interface" in French); - LIN interface (from the English "Local Interconnect Network", or in French "Réseau interconnecté local").

[0108] According to another particular and non-limiting embodiment, the device 3 includes a communication interface 33 which enables communication with other devices (such as other computers in the embedded system) via a communication channel 34. The communication interface 33 corresponds, for example, to a transmitter configured to transmit and receive information and / or data via the communication channel 34. The communication interface 33 corresponds, for example, to a wired network of the CAN (Controller Area Network), CAN FD (Controller Area Network Flexible Data-Rate), FlexRay (standardized by ISO 17458) or Ethernet (standardized by ISO / IEC 802-3) type.

[0109] Internal parameters (weights) of the pre-trained neural network can for example be received via the communication interface 33.

[0110] According to a particular and non-limiting embodiment, the device 3 can provide output signals to one or more external devices, such as a display screen 35, touch or not, one or more speakers 36 and / or other peripherals 37 via output interfaces 38, 39 and 40 respectively. According to a variant, one or more of the external devices is integrated into the device 3.

[0111] Of course, the present invention is not limited to the embodiments described above but extends to a method for controlling an on-board vehicle system based on a prediction of the facial expressions of a vehicle driver, which would include secondary steps without departing from the scope of the present invention. The same would apply to a device configured for implementing such a method.

[0112] The present invention also relates to a vehicle, for example an automobile or more generally an autonomous land-powered vehicle, comprising the device 3 of [Fig.2].

Claims

1. Demands A method for controlling an on-board system of a vehicle based on a prediction of the facial emotions of a vehicle driver, said method being implemented by at least one on-board computer in said vehicle, said method comprising the following steps: - obtaining (21) a sample of image features from at least one image of the face of the vehicle driver, said sample of image features comprising action units representing contractions or relaxations of facial muscles of the vehicle driver resulting in movements of parts of the vehicle driver's face; - obtaining (22) a binary data from a first facial emotion classification model trained to associate, with the sample of image features obtained, the binary data whose value indicates whether the sample of image features obtained represents a facial emotion associated with a first class of facial emotion or whether the sample of image features obtained represents a facial emotion from a group of facial emotion classes; - if the value of the binary data indicates that the sample of image features obtained represents the facial emotion associated with the first class of facial emotion then a first prediction of facial emotions of the vehicle driver is given by the facial emotion associated with the first class of facial emotion; - if the value of the binary data indicates that the sample of image features obtained represents a facial emotion from the group of facial emotion classes, obtaining (23) a vector of probability values ​​from a second facial emotion classification model trained to associate, with the sample of image features obtained, the vector of probability values, each probability value quantifying the veracity of the sample of image features obtained to represent a facial emotion associated with a facial emotion class from a subgroup of facial emotion classes determined from the group of facial emotion classes, the dimension of the vector of values ​​of

2. probability being defined by a number of subgroups of facial emotion classes; - selection (24), based on the probability value vector obtained, of a subgroup of facial emotion from among the subgroups of facial emotion determined from the group of facial emotion classes; - if the selected facial emotion class subgroup includes only one facial emotion class, called the second facial emotion class, then a second prediction of the vehicle driver's facial emotions is given by a facial emotion associated with the second facial emotion class; - otherwise, obtaining (25i) an identification data of a third class of facial emotion from a third facial emotion classification model trained to associate with the sample of image features obtained the identification data of the third class of facial emotion whose value indicates that the sample of image features obtained represents a facial emotion associated with a class of facial emotion, called the third class of facial emotion, of the selected subgroup of classes of facial emotion classes, a third prediction of facial emotions of the vehicle driver then being the facial emotion associated with the third class of facial emotion; - obtaining (26) a prediction of the facial emotions of the vehicle driver from the first prediction of facial emotions of the vehicle driver if it exists, from the second prediction of facial emotions of the vehicle driver if it exists, and from the third prediction of facial emotions of the vehicle driver if it exists, and - control (27) of said vehicle on-board system based on the prediction of the vehicle driver's facial emotions obtained. Method according to claim 1, wherein the image feature sample obtained is filtered before obtaining the binary data to limit the number of action units of the image feature sample obtained to those representing contractions or relaxations of facial muscles of the vehicle driver which are discriminating to determine whether the image feature sample obtained represents the first class of facial emotion or not.

3. A method according to claim 1, wherein the image feature sample obtained is filtered before obtaining the probability value vector to limit the number of action units of the image feature sample obtained to those representing contractions or relaxations of facial muscles of the vehicle driver which are discriminating to determine whether the image feature sample represents one of the facial emotions of the facial emotion class group.

4. A method according to claim 1, wherein the image feature sample obtained is filtered, before obtaining the identification data of the third facial emotion class, to limit the number of action units to those representing contractions or relaxations of facial muscles of the vehicle driver that are discriminating to distinguish the facial emotions associated with the facial emotion classes of the selected subgroup of facial emotion classes.

5. A method according to claim 1, wherein the first facial emotion classification model is a binary classification model and the third facial emotion classification model of a subgroup of facial emotion classes is a multi-value classification model.

6. A method according to claim 1, wherein the first, second and third classification models each comprise a trained neural network.

7. A computer program comprising instructions for carrying out the method according to any one of the preceding claims, when such instructions are executed by at least one processor.

8. A computer-readable recording medium on which is recorded a computer program comprising instructions for carrying out the steps of the process according to any one of claims 1 to 6

9. 1 d O. Device (3) for controlling an on-board system of a vehicle from a prediction of facial emotions of a driver of the vehicle, said device (3) comprising a memory (31) associated with at least one processor (30) configured for the implementation of at least one step of the method according to any one of claims 1 to 6.

10. Vehicle comprising the device (3) according to claim 9.

Citation Information

Patent Citations

  • Vehicle and control method for the same

    US20200215970A1