Method for generating a temporal sequence of evaluations of a situation and associated device
Patent Information
- Application Number
- EP2024712077
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-04-05
- Filing Date
- 2024-03-20
- Publication Date
- 2026-02-11
AI Technical Summary
Existing methods for multimodal learning models fail to effectively handle missing input data during training and validation, leading to reduced performance and statistical bias, especially in applications like emotion recognition and event detection from multimedia data.
A robust evaluation model using a neural network architecture with a multimodal 'Transformer' encoding model and auto-regressive self-attention mechanism, which processes positional, modality, and temporal data to generate reliable temporal sequences of evaluations even with missing modalities, by focusing on recent information and modality importance.
The solution enables reliable situation evaluation and service triggering in environments with missing data, improving model performance and reducing statistical bias, while maintaining robustness across training, validation, and exploitation phases.
Smart Images

Figure EP2024057494_10102024_PF_FP_ABST
Abstract
Description
Method for generating a time sequence of evaluations of a situation and associated device
[0001] The present invention belongs to the general field of data analysis, and more particularly to the processing of data to assess the extent to which a particular situation occurs. It relates more particularly to a method for generating a temporal sequence of evaluations of a situation by a robust evaluation model in the event of missing input data. It also relates to an electronic device configured to implement such a method.
[0002] The invention finds a particularly advantageous, although in no way limiting, application in the case where the model is used to evaluate the emotional state of a person or the evolution of the autonomy of a fragile person at home, for example with a view to determining a suitable measure to provide to them in the context of a remote assistance service. The invention also applies in the case where the model is used in the context of a video surveillance service, for example with a view to determining a suitable protection measure for an environment equipped with sensors and actuators.
[0003] In order to adapt to the continuous and ever-increasing growth of data transmitted, different technologies are currently being implemented and are the subject of research and improvement with a view to optimal exploitation in the years to come.
[0004] Among these technologies, multimodal learning is enjoying growing success, as it offers significantly better performance than methods that consider only a single modality. The objective of multimodal learning is to build models adapted to process information from different modalities. Multimodal learning has been successfully used for many applications, such as emotion recognition from multimodal data and event detection from multimedia data.
[0005] Early work in the field has primarily considered the case where complete observations are provided as input to the model, whether during training, validation, or exploitation. However, in practice, it is common for some modalities to be missing (e.g., input data to be affected by missing values at certain times), which disrupts model learning and exploitation.The absence, at least temporary, of data for at least one of the modalities classically expected as input to the model can be caused by different reasons, such as a communication problem between a sensor and an electronic device in which the previously mentioned model is embedded, a problem in the capture of data by one of the sensors, a movement of a sensor in the environment resulting in the absence of data capture for a certain period, a movement of a person within their environment, etc.
[0006] Recently, different approaches have been proposed to deal with these missing modalities. A relatively classic solution is simply to not consider the input data (or samples) for which at least one of the modalities is missing. However, this solution has the disadvantage of reducing the time periods during which the model is trained or validated, which leads to a significant drop in model performance.
[0007] Other so-called imputation methods aim to infer missing data based on heuristics, so as to be able to provide input data as input to the model whose modalities correspond to all those expected by this model. However, these imputation methods fail to restore the original distribution of missing data, and therefore have the disadvantage of introducing statistical bias. Furthermore, these imputation methods make data processing and analysis more laborious.
[0008] Finally, other methods aim to process all input data, even if some modalities are missing at certain times. However, the performance of these methods is significantly reduced.
[0009] The present invention aims to remedy all or part of the drawbacks of the prior art, in particular those set out above, by proposing a model which is robust to missing data, both during the model training phase and during the model validation or exploitation phase.
[0010] More particularly, the present invention proposes a model considering the complementarity of the information included in the input data through their different modalities, while also taking into account the previous evaluations carried out by this model. In this way, even if at least one of the modalities expected by the model is not available, the latter will still be able to reliably evaluate a situation.
[0011] To this end, and according to a first aspect, the invention relates to a method for generating a temporal sequence of evaluations (or predictions) of a situation by an evaluation model (based on a neural network and) robust in the event of missing input data, the method being implemented by an electronic device and comprising: a step of obtaining a plurality of temporal sequences of input data from sensors equipping an environment and / or a user, each sequence being associated with a modality, each element of a sequence being representative of a state, at an instant, of the modality associated with the sequence;a step of encoding, according to a "Transformer" multimodal encoding model, data resulting from a combination, at each instant, of positional encoding data, modality encoding data and temporal data determined from the plurality of temporal sequences of input data, so as to obtain a plurality of temporal sequences of encoded multimodal representations; a decoding step comprising a sub-step of applying an auto-regressive self-attention model; and a sub-step of implementing a cross-attention model taking as input the plurality of temporal sequences of encoded multimodal representations, so as to generate a temporal sequence of intermediate evaluations;and, a step of converting the time sequence of intermediate evaluations into the time sequence of evaluations of a situation, said time sequence of evaluations of a situation making it possible to control a triggering of a service adapted to the environment and / or to the user.;
[0012] For the purposes of the invention, the notion of modality is defined as a type of data captured by an electronic data capture device (i.e., a sensor), this type of data being expected as input to the evaluation model according to the invention.
[0013] Thus, a method is proposed for generating a temporal sequence of evaluations of a situation, implemented by (at least one) neural network having an architecture consisting of an encoder and a decoder.
[0014] As is well known, the encoder is composed of a set of layers of neurons, which process the data in order to construct so-called "encoded" representations in the sense that the dimensions of these representations are smaller than those of the input data (or samples). The decoder is also composed of layers of neurons which receive these representations and process them.
[0015] Using a multimodal "Transformer" encoding model offers the advantage of combining information from different modalities. In general, those skilled in the art can refer to the following document for more details regarding the implementation of a "Transformer" encoding model: "Attention Is All You Need", Ashish Vaswani & Al., Advances in Neural Information Processing Systems, volume 30, NIPS, 2017.
[0016] The remainder of the description relates more specifically to an evaluation model having an encoder-decoder type architecture. The invention nevertheless remains applicable regardless of the nature of the neural network considered (convolution, perceptron, auto-encoder, recurrent, etc.).
[0017] Furthermore, it is important to note that no limitation is attached to the type of training technique used to obtain the evaluation model. Any technique implementing a learning algorithm ("machine learning") and providing, as output, a representative evaluation of the probability that a certain situation occurs given observations (corresponding to input data), can be considered in the context of the invention (e.g., support vector machine, logistic regression, etc.). In other words, the evaluation model is independent of the training method considered to train this model.
[0018] Furthermore, any training criterion known to those skilled in the art may be considered during the training phase of the evaluation model, such as the least squares method or cross-entropy minimization.
[0019] Furthermore, no limitation is attached to the type or modality of the data processed by the evaluation model. Similarly, no limitation is attached to the nature of the situation evaluated (i.e., the nature of the situation(s) evaluated is not a limiting factor of the invention).
[0020] Thus, in a particular example implementation, the input samples to the model include images (or features extracted from raw images), sound data (or features extracted from raw sound data) synchronized with the images, and physiological signals (or features extracted from physiological signals) also synchronized with the other input data.
[0021] In another particular example of implementation, the inputs of the evaluation model correspond to Internet resources (e.g., web pages, tweets) or documents (audio, video, text, images), and the outputs of the model then correspond to the probability of an event occurring.
[0022] Using an encoder that includes an autoregressive self-attention model is advantageous in that it forces the evaluation model to focus its attention on evaluations previously made by this same model, for a situation that is assumed not to change abruptly over time.
[0023] Finally, the use of a cross-attention model offers the advantage of constraining the evaluation model to focus on the different representations at a certain time (and consequently on the different modalities processed by the model), without temporal consideration.
[0024] Generally speaking, it is considered that the steps of a process should not be interpreted as being linked to a notion of temporal succession.
[0025] In particular embodiments, the generation method may further comprise one or more of the following characteristics, taken individually or in all technically possible combinations.
[0026] In particular modes of implementation, the self-attention model is autoregressive in that, at a current time , the self-attention model takes as input the intermediate evaluations previously determined by the evaluation model from input data associated with past moments .
[0027] In particular embodiments, the positional encoding data represents, for each temporal sequence of input data, the importance of the position of at least one of said elements in said sequence.
[0028] In particular implementations, the modality encoding data represents the importance of a modality among the modalities of the plurality of temporal sequences of input data.
[0029] In particular embodiments, the generation method further comprises a step of determining the temporal data by applying at least one temporal convolution network to the plurality of temporal sequences of input data.
[0030] In particular implementations, a separate temporal convolutional network is applied to each of the input data time sequences.
[0031] In particular implementation modes, the encoding step comprises a sub-step of filtering the input data of the "Transformer" multimodal encoding model, so as to encode only the data included in a sliding time window relative to a current instant. .
[0032] This filtering sub-step is advantageous in that it forces the evaluation model to focus on recent information, which is more likely to influence the current situation than information associated with the distant past.
[0033] In particular embodiments, the temporal sequence of intermediate evaluations comprises a plurality of multidimensional elements, and the conversion step comprises applying a layer of a "fully-connected" neural network to the temporal sequence of intermediate evaluations, so as to convert each multidimensional element into an evaluation value of a situation.
[0034] In particular embodiments, the method further comprises a step of determining positional encoding data, and of combining the intermediate evaluations previously determined by said evaluation model from input data associated with instants with the determined positional encoding data, so as to obtain the intermediate evaluations.
[0035] In particular implementations, the cross-attention model takes as input encoded representations and associated with a current moment , with the number of time sequences of input data.
[0036] In particular embodiments, the generation method according to the invention further comprises: a step of obtaining a plurality of temporal sequences of labeled training data, each sequence being associated with a modality, each element of a sequence being representative of a state, at an instant, of the modality associated with the sequence; a step of training the evaluation model by minimizing a defined cost function such that with the concordance correlation coefficient.
[0037] In particular embodiments, the generation method according to the invention further comprises: a step of iterative generation of evaluations of the situation, by the evaluation model, by processing input data of the same modality at each iteration, so as to classify the modalities according to their impact on the performance of the evaluation model; a step of retraining the evaluation model, by taking as input learning data whose modality has an impact on the performance of the evaluation model below a predetermined threshold.
[0038] In particular modes of implementation, the situation corresponds to an emotional state of a user.
[0039] In particular embodiments, the sensors comprise a sound capture device and the associated temporal sequence of input data comprises elements representative of speech emitted by the user and / or of the user's prosody (sound modality); and / or the sensors comprise a physiological sensor, and the associated temporal sequence of input data comprises elements ( ) representative of the user's heart rate, respiratory rate and / or brain electrical activity (physiological modality); the sensors comprise an image capture device, and the associated temporal sequence of input data comprises elements representative of the user's facial and / or bodily expression (visual modality).
[0040] According to a second aspect, the invention relates to a computer program comprising instructions for implementing a generation method according to the invention, when said program is executed by a processor.
[0041] According to a third aspect, the invention relates to a computer-readable information or recording medium on which the computer program according to the invention is recorded.
[0042] The information or recording medium may be any entity or device capable of storing the program. For example, the medium may include a storage medium, such as a ROM, for example a CD ROM or a microelectronic circuit ROM, or a magnetic recording medium, for example a floppy disk or a hard disk.
[0043] On the other hand, the information or recording medium may be a transmissible medium such as an electrical or optical signal, which may be conveyed via an electrical or optical cable, by radio or by other means. The program according to the invention may in particular be downloaded from a network such as the Internet.
[0044] Alternatively, the information or recording medium may be an integrated circuit in which the program is incorporated (or embedded), the circuit being adapted to carry out or to be used in carrying out the method in question.
[0045] According to a fourth aspect, the invention relates to an electronic device comprising an evaluation model configured to determine a temporal sequence of evaluations of a situation, the model being robust in the event of missing input data, the device comprising: a module for obtaining a plurality of temporal sequences of input data from sensors equipping an environment and / or a user, each sequence being associated with a modality, each element of a sequence being representative of a state, at an instant, of the modality associated with the sequence; a multimodal encoder "Transformer" configured to encode data from a combination, at each instant, of positional encoding data, modality encoding data and temporal data determined from the plurality of temporal sequences of input data, so as to obtain a plurality of temporal sequences of encoded multimodal representations;a multimodal "Transformer" decoder configured to generate a temporal sequence of intermediate evaluations, the multimodal decoder comprising an auto-regressive multi-head self-attention sub-module; and a multi-head cross-attention sub-module taking as input the plurality of temporal sequences of encoded multimodal representations; a module for converting the temporal sequence of intermediate evaluations into the temporal sequence of evaluations of a situation; and, a module for controlling the triggering of a service adapted to the environment and / or to the user, as a function of the temporal sequence of evaluations of a situation.;
[0046] Other characteristics and advantages of the present invention will emerge from the description given below, with reference to the appended drawings which illustrate an exemplary embodiment thereof without any limiting character. In the figures:
[0047] lais an illustration of input data to an evaluation model, when observations are complete, and lais an illustration of input data to the evaluation model, when data are missing.
[0048] The represents an example of an evaluation system comprising an electronic device as proposed;
[0049] Schematically represents an example of the architecture of an encoder of the evaluation model as proposed;
[0050] Schematically represents an example of the architecture of a decoder of the evaluation model as proposed;
[0051] Schematically represents an example of hardware architecture of an electronic evaluation device as proposed;
[0052] Illustrates, in the form of a flowchart, the main steps of a general method for evaluating a situation, according to an example of implementation of the invention.
[0053] Theis an illustration of input data to the evaluation model, when observations are complete, and theis an illustration of input data to the evaluation model, when data are missing.
[0054] In this example, at each instant , the evaluation model expects as input data samples (or elements) associated with two different modalities And . The modality corresponds to the visual modality, and the input data associated with this modality correspond to characteristics from images of a person's face. The modality corresponds to the so-called physiological modality, and the input data associated with this modality correspond to signals representative of the respiratory, cardiac and / or cerebral activity of this same person.
[0055] Figure 1A illustrates a first example where all observations are complete, that is, at each instant , the model processes as input samples including data associated with all expected modalities.
[0056] Figure 1B illustrates the case of missing data. More precisely, at the time the model receives as input a sample corresponding to a complete observation since it includes data from both modalities And . Then at the moment the model receives as input a sample including data from the modality only. At this moment , the model therefore receives missing data as input. This situation is, for example, the consequence of a problem on the telecommunications network linking the capture device (e.g., the video camera) and the electronic device in which the evaluation model is incorporated or embedded. At the moment the model receives again as input a sample corresponding to a complete observation. Then from the instant the model receives as input a sample including only data from the modality
[0057] The present invention represents an example of an assessment system comprising an electronic assessment device (10) as proposed. In this example, the electronic device (10) comprises an assessment model used to assess the emotional state of a fragile person (e.g., senior, person suffering from a disability, isolated person, etc.) at home, with a view to determining a suitable measure to provide to him / her as part of a teleassistance service.
[0058] Such a system is particularly advantageous for maintaining vulnerable people at home. Indeed, in this context, it appears necessary to have a better understanding of the emotions of these people, in order to provide a measure of autonomy and services adapted to the situation. This assessment of the emotional state is one of the strong indicators for assessing the good physical, social and moral health of vulnerable people by the medical profession through teleassistance services. Furthermore, a person's emotions can also reveal changes in their autonomy.
[0059] Also, the evaluation model is configured to process raw data emitted by multimodal sensors equipping the environment (eg, the home) in which a person is located, to determine an emotion score representative of an emotional state, then to trigger (or not) a certain service in order to influence this emotional state.
[0060] More specifically, in a particular example of implementation, the evaluation model is configured to evaluate "negative" emotional situations of people from raw data from sensors equipping a connected home, and, in the event of evaluation of such a situation, to trigger services encouraging these people to have "more positive" emotions. In general, a person skilled in the art can refer to the following work for more details concerning the classification of emotions: "Emotions in Social Psychology", Key Readings in Social Psychology, W. Gerrod Parrott, Psychology Press, 2000.
[0061] In a particular example of implementation, fifteen emotion classes are defined: neutral (1), disgusted (2), panicked (3), anxious (4), angry (5), cold anger (6), desperate (7), sad (8), enthusiastic (9), happy (10), interested (11), bored (12), ashamed (13), proud (14), and contemptuous (15). The classes disgusted (2), panicked (3), anxious (4), angry (5), cold anger (6), desperate (7), sad (8), bored (12), and ashamed (13) can be considered negative, and the emotion classes enthusiastic (9), happy (10), interested (11) can be considered positive.
[0062] The services that can be provided by the evaluation model in response to an evaluation of a situation (eg, a negative emotion) are for example identified by links (eg, web addresses) and stored in a correspondence table associating said services with predetermined values of respective scores.
[0063] If a threshold value is exceeded, at least one link to a corresponding digital service is accessed and then its activation is offered to this person via a human-machine interface (not shown). Alternatively, this service is triggered automatically.
[0064] As illustrated in, the evaluation system comprises an electronic evaluation device (10) configured to continuously obtain or receive a set of signals. Each signal of the set carries data of at least one modality.
[0065] These signals are generated by connected sensors, then transmitted to the electronic evaluation device (10) via a telecommunications network (not shown). In this example, the connected sensors comprise an image sequence device (e.g., a video camera or a photo camera) (10) configured to capture images of a person detected in a capture zone of this camera.
[0066] A video recording of the person is then obtained (visual modality). This recording is analyzed by an analysis module ("Video Content Analysis" according to English terminology) of the video camera (20) or alternatively of the electronic evaluation device (10), so as to collect different characteristics relating to the face or facial expressions, the different gestures or postures, the movements of the person, the ambient brightness, etc.
[0067] The electronic evaluation device (10) is also configured to simultaneously receive other signals emitted by other sensors fitted to the environment or the vulnerable person.
[0068] Thus, the evaluation system further comprises an audio capture device (40) configured to collect an audio recording. The processing of this audio recording is for example implemented by said audio capture device (40), in order to extract therefrom signals representative of words emitted by the person and / or of the prosody of this person, of an ambient sound level, of an estimation of the position of the person, etc. Thus, in this example, the audio capture device (40) is configured to transmit a plurality of different signals (but associated with the same sound modality) to the electronic evaluation device (10). Alternatively, these processing operations on the raw signals are carried out by an analysis module of the electronic evaluation device (10).
[0069] Finally, the evaluation system comprises a physiological signal capture device (30) taking the form of a connected watch in this example. This physiological signal capture device (30) is configured to capture and process signals representative of the heart rate, respiratory rate and / or cerebral electrical activity of this person. Alternatively, the processing on the different physiological signals (but associated with the same physiological modality) is implemented by an analysis module of the electronic evaluation device (20).
[0070] Schematically represents an example of the architecture of an encoder of the evaluation model as proposed.
[0071] As illustrated in Figure 3, the encoder takes as input temporal sequences of input data from sensors equipping an environment and / or a user. Each sequence is associated with a modality , and includes a series of elements . Each of these elements is representative of a state, at an instant , of the modality associated with the sequence.
[0072] This encoder includes temporal convolutional networks (TCNs) which take as input the temporal sequences of input data. In other words, a temporal convolution network is dedicated to each of the modalities expected as input to the model. In a manner known per se, each of the temporal convolutional networks (TCN) allow to extract temporal data . Alternatively, the evaluation model includes only a single temporal convolutional network (TCN) that takes as input the time sequences of data.
[0073] More formally, either an element representative of a state, at a time , of the modality , of dimension . Either the length sequence of states associated with modalitym, then
[0074]
[0075] with the output sequence of the TCN. For each of the modalities, the outputs have the same size .
[0076] This encoder further includes positional encoding modules taking the form of neural networks (or corresponding to one or more layers of a neural network), each positional encoding module generating weighted temporal data from the elements . These weights (or positional encoding data ( )) represent, for each time sequence of input data, the importance of the position of at least one of said elements in the said sequence .
[0077] More formally, for a given modality, either the positional encoding data sequence with , then the output of one of the positional encoding modules is expressed in the form
[0078]
[0079] The encoder further includes modal encoding modules also taking the form of a neural network (or corresponding to one or more layers of a neural network), each modal encoding module taking as input the outputs of one of the positional encoding modules, and generating weighted temporal data . These weights (or modality encoding data) ( ) represent the importance of a modality among the terms of the time sequences of input data.
[0080] For each modality data, a modality encoding data is determined and then added to each input of the associated modal encoding module. The output of this modal encoding module is then expressed in the form
[0081]
[0082] More precisely, the elements And are combined by a combination module, so as to obtain only a sequence of "simple" elements. More formally, either , andMthe number of modalities expected by the evaluation model, then the combined sequence is expressed in the form:
[0083]
[0084] THE Combined sequences are processed by a Multimodal Transformer Encoder (MMTE). More precisely, this Multimodal Transformer Encoder is configured to encode data from a combination, at each instant, of positional encoding data , modality encoding data and temporal data determined from the time sequences of input data, so as to obtain temporal sequences of encoded multimodal representations .
[0085] More formally, the encoded representations at the output of the Transformer multimodal encoder are expressed in the form:
[0086]
[0087] with a function implemented by the Multimodal Transformer Encoder (MMTE).
[0088] As is well known, a "Transformer" encoder consists of a stack of identical encoders that do not share their weights. Each encoder of the stack is typically divided into two modules: a self-attention module which forces the encoder to consider the other elements of the same input data sequence when encoding a particular element ("mono-modal" mode); and, a module based on a feed forward neural network.
[0089] The self-attention module of an encoder is configured to take an element (eg, a data vector) as input, and generate an intermediate representation as output , which is itself transmitted to the module based on a forward propagation neural network of this encoder , so as to generate a representation .
[0090] This representation is then passed to the encoder's self-attention module of the stack which generates an intermediate representation as output , which is itself transmitted to the module based on a forward propagation neural network of this encoder , so as to generate a representation . Etc.
[0091] Generally speaking, those skilled in the art may refer to the following document for further details regarding the implementation of a multimodal "Transformer" encoder: "Transformer Encoder With MultiModal Multi-Head Attention for Continuous Affect Recognition", H. Chen & Al., IEEE Transactions on Multimedia, vol. 23, pp. 4171-4183, 2021.
[0092] Schematically represents an example of the architecture of a decoder of the evaluation model as proposed.
[0093] As illustrated in, the decoder of the evaluation model as proposed comprises a multimodal "Transformer" TDL decoder, this multimodal "Transformer" TDL decoder comprising an autoregressive MHSA multi-head self-attention module, a multi-head cross-attention MHCA module, and a first FFN conversion module taking the form of a "fully-connected" neural network.
[0094] The MHSA is connected to the MHCA, which is itself connected to the FFN. The TDL is itself connected to a second conversion module (FC) in the form of a layer of a fully-connected neural network.
[0095] The TDL outputs a time sequence of intermediate evaluations .
[0096] The MHSA is said to be autoregressive in the sense that it takes as input the intermediate evaluations previously generated by the TDL. More precisely, in the case where the decoder seeks to evaluate a situation at the current time , the TDL MHSA takes as input a plurality of previously generated intermediate evaluations ( ) and having been weighted.
[0097] It is important at this stage to recall that the use of an auto-regressive MHSA is advantageous in that it forces the evaluation model to focus its attention on evaluations previously carried out by this same model, for a situation which is assumed not to evolve abruptly over time.
[0098] The decoder further comprises a positional encoding module referenced and taking the form of a neural network (or corresponding to one or more layers of a neural network). This positional encoding module is configured to weight these previously generated intermediate evaluations, and operates in a similar manner to those previously discussed in reference to the.
[0099] More formally, the multi-head attention mechanism of MHSA and MHCA projects a query vector at a first position to a key vector to a second position in order to determine the attention (i.e., the weighting) to give to a vector of values associated with the position of the key vector . The final value corresponds to the weighted sum of the value vectors at different positions. The multi-head attention mechanism is then expressed as follows:
[0100]
[0101] with Q, K and V the sequences used as query, key and value.
[0102] In case a situation at the moment is evaluated, the TDL has previously generated the sequence with The decoder input is then expressed as follows:
[0103]
[0104] with a vector initialized using random or predetermined values.
[0105] This input is weighted by the positional encoding module which generates the sequence next
[0106]
[0107] Within the TDL, the sequence is first processed by the MHSA. It is important to remember at this point that the MHSA uses a self-attention mechanism aimed at focusing its attention on evaluations previously carried out by this same evaluation model. Consequently, the query, key and value vectors are all three determined from the input sequence . The sequence of characteristics at the output of the MHSA is then expressed in the form:
[0108]
[0109] These sequences of characteristics are then processed by the MHCA module which offers the advantage of constraining the evaluation model to focus on the different representations at a certain time (and consequently on the different modalities processed by the model), without temporal consideration.
[0110] More precisely, the MHCA takes as input the sequence of encoded representations . In other words, the key and value vectors are determined from this sequence of encoded representations, and the query vector corresponds to the output of the MHSA. More formally, the output of the MHCA is expressed in the form
[0111]
[0112] So, in the event that a situation at the moment is evaluated, the MHCA takes as input the encoded representations associated with this instant only, so as to force the evaluation model to focus on the different representations, without temporal consideration.
[0113] The TDL further includes a first conversion module (FFN) taking the form of a "fully-connected" neural network. , with . Then, the temporal sequence of intermediate evaluations [ ] at the output of this first conversion module is expressed as follows:
[0114]
[0115] And with .
[0116] The invention has so far been described in the case where the decoder comprises only one TDL. These developments can however be generalized without difficulty by those skilled in the art to the case where the decoder comprises a stack of TDLs. In this particular case, the developments previously mentioned apply to the first TDL (to be entered into the stack), and the sequence corresponds to the entry of the second TDL (to be entered into the stack). The last TDL (to be entered into the stack) determines the sequence of intermediate data [ ] such as with .
[0117] The decoder further comprises the second conversion module (FC) which takes this sequence as input , and converts each element of the sequence into a value representative of a situation (or into a vector of values representative of one or more different situations), so as to obtain a temporal sequence where each element corresponds, at a certain instant, to an evaluation of a particular situation. As mentioned previously, this second conversion module takes the form of a layer of a "fully-connected" neural network applied to each of the elements of the sequence .
[0118] In a particular implementation, the same second conversion module is applied to each element.
[0119] Schematically represents an example of hardware architecture of an electronic evaluation device (10) as proposed.
[0120] As illustrated by the, the electronic evaluation device 10 has the hardware architecture of a computer. Thus, the electronic evaluation device 10 comprises, in particular, a processor 1, a random access memory 2, a read-only memory 3 and a non-volatile memory 4. It further comprises a communication module 5.
[0121] The read-only memory 3 of the wireless communication device 10 constitutes a recording medium as proposed, readable by the processor 1 and on which is recorded a computer program PROG according to the invention, comprising instructions for executing steps of the generation method as proposed below. The program PROG defines one or more functional modules of the electronic evaluation device, which rely on or control the hardware elements 1 to 5 cited above, and which comprise in particular: a module for obtaining a plurality ( ) of temporal sequences of input data from sensors equipping an environment and / or a user, each sequence being associated with a modality ( ), each element ( ) of a sequence being representative of a state, at an instant ( ), of the modality ( ) associated with the sequence; a multimodal "Transformer" encoder (MMTE) configured to encode data from a combination, at each instant, of positional encoding data ( ), modality encoding data ( ) and temporal data determined from the plurality ( ) of temporal sequences of input data, so as to obtain a plurality ( ) of temporal sequences of encoded multimodal representations ( ); a multimodal "Transformer" decoder (TDL) configured to generate a temporal sequence of intermediate evaluations ( ), the multimodal decoder comprising an autoregressive multi-head self-attention (MHSA) sub-module; and a multi-head cross-attention (MHCA) sub-module taking as input the plurality ( ) of temporal sequences of encoded multimodal representations ( ); a conversion module (FC) of the time sequence of intermediate evaluations into the time sequence of evaluations of a situation; and, a control module (not shown) for triggering a service adapted to the environment and / or to the user, depending on the time sequence of evaluations of a situation.
[0122] Furthermore, the wireless communication device 10 may also comprise other modules, in particular for implementing particular modes of the full-duplex communication method, as described in more detail later.
[0123] ] Illustrates, in the form of a flowchart, the main steps of a general method for evaluating a situation, according to an example of implementation of the invention;
[0124] As illustrated by the, the general method for evaluating a situation comprises a first phase S1000 of training an evaluation model comprising steps S110 to S140, and a second phase S2000 of validating or exploiting the evaluation model and comprising steps S210 to S360, this second phase corresponding to an example of a generation method as proposed.
[0125] PhaseS1000 training
[0126] The method comprises a first step S110 of obtaining a plurality ( ) of temporal sequences of labeled training data, each sequence being associated with a modality , each element of a sequence being representative of a state, at an instant , of the modality associated with the sequence. This step S110 is implemented by the module for obtaining a plurality of time sequences previously mentioned.
[0127] More concretely, these training data are provided as input to the model in the form of data vectors.
[0128] In a particular implementation example, raw data emitted by different types of sensors – each sensor being associated with a particular modality – are continuously recorded in a database. Samples of this raw data are labeled using a value representative of one or more categories or classes. This representative value is for example in the interval [-1;1].
[0129] In a particular example of implementation, signals associated with a visual modality (and corresponding for example to image sequences captured by an image capture device such as a camera or a photo camera), with a sound modality (and corresponding for example to sound samples captured by a microphone) and with a physiological modality (and corresponding for example to signals representative of a respiratory rate, heart rate and / or cerebral electrical activity captured by a connected watch) are considered.
[0130] In a particular example of implementation, the appraisal model is a model for evaluating an emotional state, and fifteen emotion classes are defined: neutral (1), disgusted (2), panicked (3), anxious (4), angry (5), cold anger (6), desperate (7), sad (8), enthusiastic (9), happy (10), interested (11), bored (12), ashamed (13), proud (14), and contemptuous (15). The classes disgusted (2), panicked (3), anxious (4), angry (5), cold anger (6), desperate (7), sad (8), bored (12), and ashamed (13) can be considered negative, and the emotion classes enthusiastic (9), happy (10), interested (11) can be considered positive.
[0131] This database is accessible by a raw data analysis module, for example embedded in the electronic evaluation device 10, and which is then configured to generate new signals on the basis of the raw data. This analysis module is further configured to segment the raw data into data sections of a predetermined time interval (for example 30 seconds), without overlapping the sections.
[0132] Temporal segmentation allows for chronologically synchronized time series of interest. For example, a temporal segmentation into 30-second time intervals results in all time series of interest having, for each 30-second time interval, an associated value of a quantity of interest. This temporal synchronization allows, for example, the establishment of multimodal correlations between the values of different quantities of interest during the same time interval, or during consecutive time intervals. Thus, it is possible to structure in the model, for example, a correlation between the user's speech rate during a given time interval and a facial expression during a subsequent time interval.
[0133] Then the analysis module identifies features of interest (e.g., a face, words of the individual, etc.) from the raw data, and generates signals of interest. These signals of interest are finally processed by the analysis module so as to generate training data that will be used during the S1000 training phase of the evaluation model.
[0134] This processing generally allows the signals of interest to be shaped with a view to evaluating the situation (e.g., the emotional state of the person).
[0135] In a particular example of implementation, the targeted processing includes the calculation of facial characteristics ("Facial Action Units" according to English terminology), eGeMAPS sound characteristics ("extended Geneva Minimalistic Acoustic Parameter Set" according to English terminology), and / or a concatenation of BPM data ("Beat Per Minutes" according to English terminology), an electrocardiogram, and a respiratory rate).
[0136] In a particular implementation example, the targeted treatments include at least one of: applying low-pass filtering, normalization, or resampling. Low-pass filtering offers the advantage of denoising the information, normalization allows for standardizing the data, and resampling the data allows for synchronizing the sources.
[0137] The training phase further comprises a step S120 during which the evaluation model is trained, by processing the training data obtained during step S110.
[0138] In a particular mode of implementation, the evaluation model is trained by applying a cost function defined such that with the concordance correlation coefficient.
[0139] We recall at this stage that the concordance correlation coefficient is expressed in the form
[0140]
[0141] with the correlation coefficient between the predicted values and the exact values ("ground-truth values"), the standard deviation and the average of either the predicted values or the exact values.
[0142] In a particular embodiment, phase S1000 further comprises steps S130 and S140. During step S130, the modality or modality(s) having a significant impact on the evaluation of a given situation are identified. To do this, the evaluation model is first trained by applying step S120, then this same evaluation model, once trained, is evaluated by processing data from a single modality at a time. In other words, this evaluation model evaluates the same situation iteratively, considering at each iteration input data from the same modality.
[0143] Then the iterations in which the evaluation performance is low (e.g., the lowest or, alternatively, those below a predetermined threshold value) are determined, which correspond to the modalities having a significant impact on the model's evaluation performance. In this way, each of the modalities is ranked according to its impact on the prediction model's performance.
[0144] The training phase S1000 then comprises a step S140 during which the evaluation model is re-trained, but this time by excluding, from the input data, those associated with the modalities determined during step S130 as having a significant impact on the evaluation performance.
[0145] According to a particular example, for each input data sequence, the data associated with the modalities having a significant impact with a probability are excluded from relearning, with a probability that the data is missing or present, and only the modalities having a probability are considered as input to the evaluation model during this step S140.
[0146] Thus, by hiding from the evaluation model the modalities having a significant impact on performance, this forces this model to evaluate the correlations between the modalities having a less significant impact. Thus, these steps S130 and S140 make it possible to obtain a model for evaluating a situation which is more robust in the event of missing data for at least one of the modalities classically expected as input to the evaluation model, and having a significant impact on performance.
[0147] S2000 phase of validation or exploitation of the previously trained model
[0148] This S2000 phase corresponds to a process of generating a temporal sequence of evaluations (or predictions) of a situation by an evaluation model (based on a neural network and) robust in the event of missing input data.
[0149] Phase S2000 includes steps S210 to S250 implemented by the encoder as represented by 1a, and steps S310 to S360 implemented by the decoder as represented by 1a.
[0150] As illustrated in Figure 6, the method of generating a sequence comprises a first step of obtaining input data from a plurality ( ) of temporal sequences of input data from sensors equipping an environment and / or a user, each sequence being associated with a modality ( ), each element ( ) of a sequence being representative of a state, at an instant ( ), of the modality ( ) associated with the sequence.
[0151] The step of obtaining these input data is similar to step S110, and is not re-detailed for the sake of brevity. Since this is a validation or exploitation phase, this step S210 is distinguished from step S110 by the fact that the input data are of course not labeled. This step is implemented by the module for obtaining a plurality of time sequences previously mentioned.
[0152] The method for generating a sequence further comprises a step S220 during which temporal data are determined by applying at least the temporal convolution network TCN of the encoder to the plurality ( ) of time sequences of input data.
[0153] In a step S230, positional encoding data ( ) are determined by one or more of the positional encoding modules of the encoder, these encoding data representing, for each temporal sequence of input data, the importance of the position of at least one of said elements in said sequence. The positional encoding data are then combined with the temporal data, so that the output of this positional encoding module is then expressed for example in the form
[0154]
[0155] Then, during a step S240, modal encoding data ( ) are determined by the modal encoding module of the encoder, these encoding data representing, for each temporal sequence of input data, the importance of a modality among the modalities of the input data. The modal encoding data is then combined with the outputs of the encoder's positional encoding module, so that the output of this modal encoding module is then expressed, for example, in the form
[0156]
[0157] The method further comprises a step S250 of encoding, according to a multimodal “Transformer” encoding model, the output , so as to obtain a plurality ( ) of temporal sequences of encoded multimodal representations ( ). This step is implemented by the previously mentioned MMTE multimodal "Transformer" encoder.
[0158] In a particular embodiment, the encoding step S250 comprises a sub-step of filtering the input data of the multimodal encoding model "Transformer", so as to encode only the data included in a sliding time window relative to a current instant. .
[0159] Then, during a step S310, the decoder obtains as input the intermediate evaluations previously generated by the TDL. Thus, in the case where the decoder seeks to evaluate a situation at the instant , the MHSA of the TDL takes as input a plurality (eg, ) of previously generated intermediate evaluations ([ ]) and having been weighted.
[0160] The generation method further comprises a step S320 during which positional encoding data ( ) are determined by a positional encoding module ( ) of the decoder. The positional encoding data is then combined with the plurality of previously generated intermediate evaluations, so that when evaluating the situation at the instant , the entrance of the multimodal decoder "Transformer" is expressed for example in the form
[0161]
[0162] The generation method further comprises a step S330 during which a self-attention mechanism is applied to the input by the MHSA.
[0163] It is important at this stage to recall that the application of a self-attention mechanism offers the advantage of forcing the evaluation model to focus its attention on evaluations previously carried out by this same evaluation model.
[0164] The generation method further comprises a step S340 during which a cross-attention mechanism is applied by the MHCA. It is important at this stage to recall that the application of a cross-attention mechanism offers the advantage of constraining the evaluation model to focus on the different representations at a certain instant (and consequently on the different modalities processed by the model), without temporal consideration. The output of the MHCA is expressed in the form
[0165]
[0166] The generation method further comprises a step S350 during which the first FFN conversion module generates the time sequence of intermediate evaluations [ ] such as:
[0167]
[0168] And with .
[0169] Then, during a step S360, each element of this sequence is converted into a value representative of a situation (or into a vector of values representative of one or more different situations), so as to obtain a time sequence [ ] where each element corresponds, at a certain moment, to an evaluation of a particular situation.
[0170] The invention has so far been described in the case where the decoder comprises only a single TDL. These developments can however be generalized without difficulty by those skilled in the art to the case where the decoder comprises a stack of TDLs. In this particular case, steps S330, S340 and S350 are repeated.
[0171] Finally, this temporal sequence of evaluations of a situation [ ] is used to determine whether a service, a service appropriate to the environment and / or the user, should be triggered or not.
[0172] For this purpose, according to a particular implementation example, a plurality of digital services that can be suggested to a person and defined by respective links are stored in a correspondence table, in association with respective score thresholds. In case of exceeding a threshold by one or more elements of the time sequence [ ], at least one link of a corresponding digital service is read to suggest said corresponding digital service to the user via a human-machine interface. The assessed situation corresponds for example to the emotional state of the person.
[0173] Thus, a service is only recommended to the person if the intensity of an assessment of a predetermined situation is sufficiently high. This has the effect of reserving the suggestion of services to this person at times when the execution of these services would be most useful to him.
[0174] According to a particular implementation example, an animation routine of a human-machine interface of any communicating equipment of the environment in which a person is located (eg, a connected habitat) is triggered.
[0175] The service recommendation is, for example, carried out on the best broadcast channel for the person, such as the living room TV, a communicating speaker, a smartphone, etc. The most suitable communicating equipment for this purpose can be selected, for example, on the basis of the functionalities offered by the different communicating equipment in the installation.
[0176] Another criterion may be an estimate of the distance, at the current time, between the different communicating devices and the person, making it possible to choose, among the different communicating devices in the installation capable of triggering the services offered, the one closest to the person at the current time.
[0177] According to a non-limiting example, the evaluated situation corresponds to an assessment of the emotional state of a person, and the animation routine is configured to suggest, to the person, at least one predefined digital service associated with the assessment of his emotional state. For example, if the predicted situation is "panicked", the associated services in a preference model may be in an order of preference "an incentive to call a friend", "an incentive to call the telecare service" or "an increase in home comfort by automatically adjusting the brightness".
[0178] Other examples of services can be described for each "negative" emotional situation according to the preferences and habits of the individuals. These preferences can be defined by the individuals themselves or by a trusted third party, such as family, a loved one, or the medical profession. These preferences can be defined, reorganized, or updated automatically by learning from services used in the past, in correlation with the situations evaluated.
[0179] The invention has so far been described in the case where the combination operator corresponds to an addition. The invention nevertheless remains applicable in the case where other combination operators are considered, such as .
Claims
A method for generating a time sequence of evaluations of a situation by a robust evaluation model in the event of missing input data, the method being implemented by an electronic device (10) and comprising: a step of obtaining (S210) a plurality ( ) of temporal sequences of input data from sensors equipping an environment and / or a user, each sequence being associated with a modality ( ), each element ( ) of a sequence being representative of a state, at an instant ( ), of the modality ( ) associated with the sequence; an encoding step (S250), according to a multimodal "Transformer" encoding model (MMTE), of data resulting from a combination, at each instant, of positional encoding data ( ), modality encoding data ( ) and temporal data determined from the plurality ( ) of temporal sequences of input data, so as to obtain a plurality ( ) of temporal sequences of encoded multimodal representations ( ); a decoding step (S300) comprising a sub-step (S330) of applying an auto-regressive self-attention model (MHSA); and a sub-step (S340) of implementing a cross-attention model (MHCA) taking as input the plurality ( ) of temporal sequences of encoded multimodal representations ( ), so as to generate a temporal sequence of intermediate evaluations ( ); and, a step of converting (S360) the time sequence of intermediate evaluations into the time sequence of evaluations of a situation, said time sequence of evaluations of a situation making it possible to control a triggering of a service adapted to the environment and / or to the user. Generation method according to claim 1, the self-attention model (MHSA) being auto-regressive in that, at a current time , the said self-attention model takes as input the intermediate evaluations previously determined by said evaluation model from input data associated with past moments . A generation method according to claim 1 or 2, wherein the positional encoding data ( ) represent, for each time sequence of input data, the importance of the position of at least one of said elements in said sequence. Generation method according to one of claims 1 to 3, in which the modality encoding data ( ) represent the importance of a modality among the ( ) modalities of plurality ( ) of time sequences of input data. Generation method according to one of claims 1 to 4, further comprising a step of determining the temporal data by applying (S220) at least one temporal convolution network (TCN) to the plurality ( ) of time sequences of input data. Generation method according to one of claims 1 to 5, in which the encoding step (S250) comprises a sub-step of filtering the input data of the multimodal encoding model "Transformer" (MMTE), so as to encode only the data included in a sliding time window relative to a current instant. . Generation method according to one of claims 1 to 6, further comprising a step of determining positional encoding data ( ), and combination (S320) of intermediate assessments previously determined by said evaluation model from input data associated with instants with the positional encoding data ( ) determined, so as to obtain the intermediate evaluations. Generation method according to one of claims 1 to 7, in which the cross-attention model takes as input encoded representations and associated with a current moment , with the number of time sequences of input data. Generation method according to one of claims 1 to 8, further comprising: a step of obtaining (S110) a plurality ( ) of temporal sequences of labeled training data, each sequence being associated with a modality ( ), each element ( ) of a sequence being representative of a state, at an instant ( ), of the modality ( ) associated with the sequence; a training step (S120) of the evaluation model by minimizing a cost function ( ) defined as with the concordance correlation coefficient. Generation method according to claim 9, further comprising: a step of iterative generation (S130) of evaluations of the situation, by the evaluation model, by processing input data of the same modality at each iteration, so as to classify the modalities according to their impact on the performance of the evaluation model; a step of retraining (S140) the evaluation model, by taking as input training data whose modality has an impact on the performance of the evaluation model lower than a predetermined threshold. Generation method according to one of claims 1 to 10, in which the situation corresponds to an emotional state of a user. A generation method according to one of claims 1 to 11, wherein: the sensors comprise a sound capture device and the associated input data time sequence comprises elements ( ) representative of speech emitted by the user and / or the user's prosody; and / orthe sensors comprise a physiological sensor, and the associated temporal sequence of input data comprises elements ( ) representative of the user's heart rate, respiratory rate and / or brain electrical activity; the sensors comprise an image capture device, and the associated input data time sequence comprises elements ( ) representative of the user's facial and / or bodily expression. Computer program comprising instructions for implementing a determination method according to any one of claims 1 to 12, when said program is executed by a computer. A computer-readable recording medium on which a computer program according to claim 13 is recorded. Electronic device (10) comprising an evaluation model configured to determine a temporal sequence of evaluations of a situation, the model being robust in the event of missing input data, the device comprising:a module for obtaining a plurality ( ) of temporal sequences of input data from sensors equipping an environment and / or a user, each sequence being associated with a modality ( ), each element ( ) of a sequence being representative of a state, at an instant ( ), of the modality ( ) associated with the sequence; a multimodal "Transformer" encoder (MMTE) configured to encode data from a combination, at each instant, of positional encoding data ( ), modality encoding data ( ) and temporal data determined from the plurality ( ) of temporal sequences of input data, so as to obtain a plurality ( ) of temporal sequences of encoded multimodal representations ( ); a multimodal "Transformer" decoder (TDL) configured to generate a temporal sequence of intermediate evaluations ( ), the multimodal decoder comprising an autoregressive multi-head self-attention (MHSA) sub-module; and a multi-head cross-attention (MHCA) sub-module taking as input the plurality ( ) of temporal sequences of encoded multimodal representations ( ); a conversion module (FC) of the time sequence of intermediate evaluations into the time sequence of evaluations of a situation; and, a module for controlling the triggering of a service adapted to the environment and / or to the user, depending on the time sequence of evaluations of a situation.