A method for machine learning a multimedia data prediction model and a method for detecting anomalies from such a model
A multi-prediction model addresses the inefficiencies in anomaly detection in multimedia data by using multiple predictors to reconstruct normal data, enhancing the model's ability to distinguish anomalies and capturing the diversity of normal behaviors.
Patent Information
- Application Number
- FR2023012204
- Authority / Receiving Office
- FR · FR
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-09
- Publication Date
- 2025-05-16
AI Technical Summary
Existing automatic learning methods for anomaly detection in multimedia data are inefficient due to the imbalance between normal and abnormal data classes, high annotation costs, and limited ability to capture the diversity of normal data behavior.
A multi-prediction model is developed that uses multiple predictors to reconstruct masked normal data, improving the model's ability to distinguish between normal and abnormal data by capturing the diversity of normal behaviors through multiple predictions.
The multi-prediction model enhances anomaly detection by improving the power of discrimination between normal and abnormal data, reducing the non-detection rate of anomalies, and adapting better to normal data patterns.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Title of the invention: Method for automatic learning of a multimedia data prediction model and method for detecting anomalies from such a model
[0001] The invention relates to the field of machine learning methods and relates to a new method for learning a multimedia data prediction model and a method for detecting anomalies using such a multimedia data prediction model. The data considered are, for example, images, image sequences, videos, audio sequences, multi-spectral images or more generally data which can be multidimensional and multimodal.
[0002] Anomaly detection consists of identifying so-called “abnormal” data for a given application context. Data is said to be “abnormal” when it is unusual, unpredictable or unwanted. More generally, “abnormal” data can be defined as data that deviates significantly from “normal” data in a given application context.
[0003] The abnormal nature of data depends on the type of data and the intended application as well as the context. For example, if the data are images of a component or a product from an industrial production line, an anomaly corresponds to a visible defect on the component or product.
[0004] In the case where the data are video sequences, an anomaly corresponds, for example, to unusual behavior of a pedestrian in a given area.
[0005] Anomaly detection is particularly of interest in the field of video surveillance to identify unusual behavior of a person in a public or private place, in the field of autonomous driving, to identify an unforeseen obstacle on the road, in the field of industrial control to identify a defect on a manufactured product.
[0006] In the field of machine learning, the task of detecting a category of data is carried out by a detector optimized following learning or training, supervised by data representative of the categories to be detected.
[0007] Machine learning methods can therefore be used to develop an anomaly detection model in multimedia data.
[0008] By definition, abnormal data are inherently rare and diverse compared to normal data which are abundant. However, to avoid detection bias, it is advisable to train a detector on data sets balanced between the different categories.
[0009] For this reason, supervised machine learning methods are poorly suited to addressing the problem of anomaly detection because the task of annotating abnormal data is extremely costly and requires covering a vast domain of heterogeneous anomalies depending on the context and the intended application. Furthermore, these approaches are not very efficient given the natural imbalance of normal data classes compared to abnormal data classes.
[0010] A general problem that the invention aims to solve therefore consists of developing an unsupervised method of automatic learning for the detection of anomalies in multimedia data.
[0011] One-class classification machine learning methods are better suited to the problem of anomaly detection because they use only "normal" data as input and aim to predict or reconstruct this data. The trained model is thus based on the extraction of relevant characteristics from normal data and is subsequently used to infer the degree of abnormality of new data to be evaluated.
[0012] Unsupervised learning approaches for anomaly detection are based on training an artificial intelligence model (e.g., a deep neural network) to perform a pretext task on normal data only. In other words, the model is not directly trained to detect anomalies but is trained on another task, for example, reconstructing normal data, and then using the model indirectly to solve the anomaly detection task. During the inference phase, an anomaly score can be deduced from the model's inability to perform the task correctly.
[0013] In order for the trained model to be able to effectively characterize the normality of the data and distinguish it from anomalies, the pretext tasks must satisfy two necessary conditions. They must be correctly performed by normal data and they must be poorly performed in the presence of anomalies. In other words, the chosen pretext tasks must induce poor generalization of the model to anomalies.
[0014] Anomaly detection methods using unsupervised learning can be grouped essentially into two categories: reconstruction methods and prediction methods.
[0015] Reconstruction-based approaches aim to train a model to reconstruct, as output, the normal training data received as input. One assumption of these approaches is that the reconstruction model will not be able to correctly generalize the reconstruction to anomalies, i.e., abnormal data will not be well reconstructed.
[0016] Unlike reconstruction-based methods, approaches by Prediction models learn to predict missing information, such as hidden parts of normal data, in order to better learn their characteristics.
[0017] Methods based on a reconstruction approach have the disadvantage of sometimes also correctly reconstructing abnormal data.
[0018] Reference [1] describes a method for detecting anomalies by reconstruction which proposes learning the distribution of data using a multi-hypothesis autoencoder. In addition, the model is evaluated by a discriminator, which prevents the generator from producing implausible predictions. Autoencoders have the disadvantage of being able to reconstruct abnormal characteristics because of their extrapolation capabilities. Thus, they induce a non-negligible non-detection rate since certain anomalies will be detected as corresponding to normal behaviors.
[0019] Methods based on a prediction approach generally predict anomalies poorly because training is carried out only on normal data and the information to be predicted does not exist in the input data. However, these methods are less well adapted to normal data because the missing information to be predicted is not present in the normal data, which can lead to prediction difficulties.
[0020] Most anomaly detection methods by prediction are based on learning a single prediction, which has the disadvantage of not taking into account the diversity of normality. Indeed, a single prediction often does not allow characterizing the diversity of so-called normal behavior. For example, if we consider a simple scenario of a camera observing a vehicle moving on a road and arriving at an intersection with three possibilities: turn right, go straight or turn left, these three possible future states can all be described as "normal". In such a scenario, a model based on a single predictor will not be able to predict the different future states of the vehicle's trajectory with a single prediction. On the contrary, the prediction generated will correspond to an average of the three possible "normal" states.If normal data is not correctly predicted by such a model, then it will also not be able to detect anomalies by comparison.
[0021] One solution to this problem is to design a multiple prediction model to predict all "normal" states of a data item.
[0022] Reference [2] describes a method for detecting anomalies by prediction which proposes training a model to produce several different predictions of the same masked data. Learning several predictions makes it possible to better cover the diversity of so-called normal behaviors in the masked input data.
[0023] The authors propose to stochastically predict normal video data using a conditional variational autoencoder. The method's predictions are made stochastically, which means that the samples are not necessarily representative of the learned distribution, and the anomaly score considered does not precisely quantify membership in the normal data distribution.
[0024] There is therefore a need for a new method of anomaly detection which overcomes the drawbacks of reconstruction or prediction approaches.
[0025] The proposed invention makes it possible to combine the advantages of reconstruction methods and prediction methods. The invention consists of training a multi-prediction model that does not generalize predictions well in the presence of anomalies, which improves the power of discrimination of anomalies. Furthermore, due to the use of several predictors, the proposed model adapts better to normal data than a single-prediction model. The predictions made are deterministic, which ensures the repeatability of the anomaly scores given by the system. The predictions are also diversified, which makes it possible to cover the entire distribution of normal data, each predictor specializing in a particular pattern among the set of patterns corresponding to a normal characteristic.
[0026] The subject of the invention is a method, implemented by computer, for training a multimedia data reconstruction model represented by at least one modality, the model being composed of a set of several different predictors for each modality of the data, the training method comprising the steps of, for each data item in a set of training data not comprising any anomaly, • for each data modality: - Hide at least part of the data modality, - Train each predictor of the set associated with said modality to calculate a different prediction of the same masked data, each predictor being specialized in a possible prediction of the masked data among different credible alternatives, - Select the predictor from the ensemble that provides the closest prediction to a reference data extracted from the training data, - Calculate a distance between said prediction and the reference data, • Calculate a first cost function equal to the sum of said distances for all modalities, • Update the parameters of the selected predictors for each modality so as to minimize the first cost function.
[0027] In a particular embodiment, the method further comprises, • for each data modality: - Select a subset of the set of predictors that were not optimized in a previous iteration of training, - Calculate the sum of the distances between the respective predictions provided by the predictors of said subset and the reference data, • Calculate a second cost function equal to the sum of said distances for all modalities and modify the first cost function by adding the second cost function weighted by a weighting coefficient, • Update the parameters of the predictor that provides the prediction closest to the reference data and of the predictors of said subset for each modality so as to minimize the first modified cost function.
[0028] According to a particular aspect of the invention: • multimedia data are represented, according to a basic modality, by a temporal sequence of successive images, • The step of masking at least part of the basic modality comprises applying a predefined spatial mask to each image of the sequence so as to mask at least one area of the image, • The reference data corresponds to a successive image in the time sequence compared to the current image provided as input to the model, • Each predictor is trained to predict said successive image from the current image masked by means of said mask.
[0029] According to a particular aspect of the invention: • Multimedia data is further represented, according to an additional modality, by an optical flow sequence, • The step of masking at least part of the additional modality includes masking the entire optical flow, • The reference data is the optical flow, • Each predictor is trained to predict the optical flow from the current image masked according to the base modality.
[0030] According to a particular aspect of the invention: • The multimedia data are further represented, according to an additional modality, by a sequence comprising, for each image, a set of classes of objects detected in the image, • The step of hiding at least part of the additional modality includes hiding the entire sequence of object classes, • The reference data is the sequence of object classes, • Each predictor is trained to predict the sequence of object classes from the current image masked according to the base modality.
[0031] According to a particular aspect of the invention: • The multimedia data is further represented, according to an additional modality, by an audio sequence synchronized with the image sequence, • The step of masking at least part of the additional modality comprises removing at least part of the audio sequence, • The reference data is the audio sequence, • Each predictor is trained to predict the audio sequence from the current masked image according to the base modality.
[0032] According to a particular aspect of the invention: • multimedia data is represented, according to a basic modality, by a set of images, • The step of masking at least part of the basic modality comprises applying a predefined spatial mask to each image so as to mask at least one area of the image, • The reference data corresponds to the current image provided as input to the model but not masked, • Each predictor is trained to predict said current image from the current image masked by means of said mask.
[0033] According to a particular aspect of the invention: • The multimedia data are further represented, according to a second modality, by a set of multispectral images, • The step of masking at least part of the second modality comprises removing images at at least one given wavelength, • The reference data is the set of multispectral images, • Each predictor is trained to predict a multispectral image from the current image masked according to the base modality.
[0034] According to a particular aspect of the invention, the model comprises: • a first projector neural network receiving masked data as input and trained to transform the input into a latent representation, • a recurrent neural network trained to produce a sequence of states, in a recurrent manner, the initial state being equal to the latent representation, • several predictive neural networks each corresponding to a modality, each predictor network receiving as input the masked data and a state provided by the recurrent network, the number of states generated by the recurrent network being equal to the number of predictors for a modality.
[0035] The invention also relates to a method, implemented by computer, for detecting anomalies in a set of multimedia data having at least one modality, the method comprising the steps of: • Execute, for said data set, the machine learning model trained using the training method according to any one of the preceding claims, the model receiving as input the data masked using masks identical to those used for training said model and producing as output several predictions, • Calculate the first cost function from the predictions generated by the model, • Calculate an anomaly score from a distance between the first cost function and a representative value of the first cost function for a data set containing no anomalies, • Compare the anomaly score to a predetermined detection threshold and deduce the presence or absence of anomalies in the data.
[0036] The invention also relates to a computer program comprising code instructions for implementing the invention as well as a computer-readable recording medium on which the computer program according to the invention is recorded.
[0037] Other characteristics and advantages of the present invention will appear better on reading the description which follows in relation to the following appended drawings.
[0038] [Fig. 1] represents a diagram of a first example of architecture of a multi-prediction machine learning model according to an embodiment of the invention,
[0039] [Fig.2] represents a flowchart detailing the steps of a method for training an anomaly detection model according to an embodiment of the invention,
[0040] [Fig.3] represents a diagram of a second example of architecture of a multi-prediction model capable of being trained using the method of [Fig.2],
[0041] [Fig.4] illustrates a gradient backpropagation step to update the model parameters of [Fig.3],
[0042] [Fig.5] illustrates the operation of gradient backpropagation for a model taking into account three distinct modalities of the input data,
[0043] [Fig.6] represents an example of a machine learning model to implement the different networks of the example of [Fig.3].
[0044] [Fig.7] represents a flowchart detailing the steps of an anomaly detection method from the model trained according to the method of [Fig.2],
[0045] [Fig.l] schematically represents an example of a multi-prediction machine learning model according to an embodiment of the invention.
[0046] The model receives as input data masked according to a predefined mask M. It comprises a first encoder network E capable of converting the masked input data into a latent representation and several predictor networks D(1), D(2)... D(n) which are trained to each determine a possible prediction of the original data from the masked data.
[0047] More generally, the n predictions can be generated by n distinct predictor networks or by a single network capable of producing n distinct predictions.
[0048] The model of [Fig.l] is trained using so-called “normal” training data, i.e. data which does not contain any anomalies in the targeted application context. For example, if the targeted application consists of detecting defects on parts during manufacture, the training data only includes images of parts having no defects.
[0049] If the intended application consists of the detection of abnormal behaviors in surveillance videos, the training data only includes images of normal behaviors in the context of the monitored area.
[0050] Thus, the model in [Fig.l] is trained to generate different predictions of the masked parts of the input data, all the predictions being assumed to correspond to normal data. Once the model is trained, it thus makes it possible to detect anomalies because if an anomaly is masked, the model will not reconstruct the anomaly but on the contrary predict a reconstruction corresponding to normal behavior. By comparing the predictions with the real data, it is therefore possible to detect an anomaly when the predictions deviate from the data.
[0051] The use of several different predictors allows optimal training in the sense that the model will not simply learn to generate a prediction corresponding to an average of all possible (credible) reconstructions but on the contrary, each predictor will specialize in a type of possible reconstruction.
[0052] For example, in the case of an image of a car at a crossroads comprising several possible roads, each predictor will specialize in predicting a trajectory of the car towards one of the roads. In this context, all these predictions correspond to a possible normal behavior of the car. Conversely, a car traveling between two roads (on a sidewalk or more generally an area not corresponding to a road) corresponds to an abnormal behavior and therefore to an anomaly in this context.
[0053] Generally speaking, the invention applies to any type of multimedia data re- presented by at least one modality. For example, it applies to video sequences, still images, audio sequences, RGB or multispectral images or a combination of these different media.
[0054] Subsequently, the invention is described for data corresponding to video sequences, but it can be generalized to the other types of data mentioned above.
[0055] [Fig.2] shows in a flowchart the steps of implementing a method for training the model of [Fig.1] according to one embodiment of the invention.
[0056] The training is carried out using training data 201 which do not contain any anomalies so as to train the model to reconstruct so-called “normal” data in the sense that they do not contain any anomalies. Thus, the model is trained to reconstruct partially masked “normal” data.
[0057] In step 202, the training data are partially masked by means of a predefined mask M. The masking step 202 can take different forms. It consists of masking or altering at least a part of at least one modality of the data. More precisely, if the input data are represented by several modalities, each of the modalities can be masked totally or partially with the exception that at least one modality must be masked only partially and this in order to provide a minimum of information at the input of the model. Examples of different modalities will be explained later.
[0058] For example, in the case of a sequence of images, the masking step 202 may consist of masking one or more areas of each image according to a predefined spatial mask M. The mask may be solely spatial or spatio-temporal in which case it also depends on the temporal index of the image in the sequence. The masking step may consist of entirely removing an area of an image or applying noise, for example white noise, to certain areas of an image.
[0059] The masked data are then provided as input to the model to be trained. In step 203, the different predictions of the original data are calculated from the parameters of the model and the masked input data. Depending on the chosen embodiment, the predictions aim to predict the current masked image or a future image in the video sequence, for example the image following the masked image in the sequence.
[0060] The model therefore provides n AF predictions then at step 204 we calculate a first cost function or LNN(Y) loss function. The chosen cost function consists of selecting, among the n predictions, the one that is closest to a reference Iref corresponding to the original data and calculating the error between this selected prediction and the reference. For example, if the model aims to predict a image It+i at time t+1 from a masked image It at time t, then the reference is the original image It+i. Alternatively, if the model aims to directly predict the current image It then the reference is the original image It.
[0061]
[0062]
[0063] The first cost function is thus given by the following relation: L w (r)= min || r w -r„HI (1) âæIILh]] Y denotes a representation of the data according to a modality. For example, ^7 is a prediction of an image It+i and Yref is the original image It+i.
[0064] The cost function Y ) can be calculated for one or more modalities of the data as will be illustrated later.
[0065] This first cost function aims to encourage the diversity of predictions via the selection, at each iteration, of the prediction closest to the reference.
[0066] In step 206, a backpropagation algorithm based on a gradient calculation is applied so as to update the parameters of the model to minimize the first cost function. At this stage, only the parameters of the predictor of index k selected to calculate the cost function are optimized with the parameters of the encoder E during the backpropagation. The unselected predictors are not optimized.
[0067] In an alternative embodiment, a second cost function is calculated in step 205 via the following relationship: 100681 <2>
[0069] UT is the set of predictors that have not been selected or have been selected only slightly to calculate the first cost function during a previous iteration of the training. To determine the UT set, a selection threshold is for example set below which it is considered that a predictor has been selected only slightly during an iteration.
[0070] More precisely, the training is carried out in several iterations with at each iteration a set of training data called epoch. At the end of the processing of the previous epoch, the predictors which have never been selected are noted to calculate the first cost function and the second cost function is calculated for these predictors.
[0071] It is possible that the predictors selected to calculate the second cost function include the predictor selected to calculate the first cost function for a current epoch.
[0072] In step 206, the gradient backpropagation algorithm is applied so as to minimize a combination of the two cost functions: L=LNN +XLNP with / . a weighting coefficient. Preferably, the parameter X is a positive number strictly less than 1, for example equal to 0.1 in order to give more weight to the first LNN cost function. In this way, we optimize (via the first function of cost Lnn) the predictor whose prediction is closest to the real data while optimizing (via the additional term X. LNP) the predictors which did not participate enough in the training in the previous epoch. The predictors which participated enough in the previous epoch but which are too far from the real data are not optimized (in fact, for them the calculated gradient is substantially zero with respect to the two cost functions LNN and LNP).
[0073] The backpropagation is performed so as to update the parameters of the predictor selected to calculate the first cost function and of the predictors selected to calculate the second cost function, as well as the parameters of the encoder E common to all the predictors.
[0074] The optimization of the second LNP cost function aims to allow the optimization of all predictors, even those which are never or very rarely selected, so as to promote the diversity of predictions.
[0075] In an alternative embodiment, the same machine learning model may be separately optimized to make different predictions relating to different modalities of the data.
[0076] For example, the input data may be organized as a primary or base modality and one or more additional modalities. In other words, the input data may be represented by multiple modalities, the importance of which may vary depending on the applications.
[0077] Multimodal data refers to a set of information that combines several different modes or data sources. For example, for an audio / video sequence, sound and image can be considered as two modalities of this data.
[0078] For example, for an application of detecting anomalies in visual appearance and movements of objects present in a video, the input data is a sequence of images that is partially masked. The basic modality here corresponds to the sequence of images. In this scenario, an additional modality is, for example, an optical flow sequence calculated on the sequence of images, or a sequence of classes of objects present in each image (detected by means of an object detector applied to the sequence of images). The optical flow is information that characterizes the displacement of each pixel of an object between two successive images. It makes it possible to characterize the movement of objects over time in a video sequence. The objects detected in a video sequence can also be classified into object categories. Thus, each image can be accompanied by information on the classes of the objects present in this image.
[0079] In this application example, the machine learning model can also be trained to predict the optical flow from the masked images received in input, the optical flow being completely masked, i.e. removed in the sense that it is not provided as input to the model. In this case, a second set of n predictors is optimized from the same cost function and the same learning procedure as that described in [Fig.2]. The only difference is that the predictions are predictions of the optical flow and are compared to a reference which is the original optical flow.
[0080] Similarly, by replacing the optical flow with a sequence of object classes, a third set of n predictors can be optimized.
[0081] The machine learning model can thus be globally optimized by means of the minimization of a cost function which is a combination of the different cost functions calculated for each modality, this combination is, for example, a sum.
[0082] The principle of learning the multi-predictor model for multi-modal data can be applied to other modalities.
[0083] For example, when the intended application is the detection of anomalies in a still image, the input of the model is unimodal and corresponds to an image. The predictions provided by the model are predictions of the current image from the spatially masked image.
[0084] In the case where the images are multispectral, the input to the model is multimodal, each modality corresponding to an image at a wavelength or a range of wavelengths. In this scenario, a portion of the wavelengths is removed as input to the model and the predictors are trained to predict the images at the removed wavelengths from the images at the other wavelengths. For example, the removed wavelengths correspond to infrared.
[0085] In another example application, the data is multimedia and multimodal, comprising both a video sequence and an audio sequence. In this scenario, the audio sequence can be removed as input to the model and the predictors are trained to predict the audio sequence from the video sequence.
[0086] Without departing from the scope of the invention, other multimodal data may be considered. A general objective of the invention is to train the model to reconstruct so-called normal data from multiple predictions for one or more modalities of the data.
[0087] When the data are multimodal, a basic modality can be defined which corresponds to the modality according to which the data are provided as input to the model, partially masked, the additional modalities being, for example, totally masked.
[0088] A particular example embodiment of the invention is now described which is based on video data and a particular learning model.
[0089] In this example illustrated in [Fig.3], the machine learning model comprises a first projector network P which receives as input each image Ylt of the video sequence masked by a spatial mask M. The projector network P is trained to convert its input into a latent representation in another projection space making it possible to characterize the information contained in the input. In this example, the input data are represented by a single modality which is the image sequence I.
[0090] The model also includes a recurrent neural network R which is trained to produce a sequence of states h; in a recurrent manner, the initial state h0 being equal to the latent representation provided at the output of the projector network P.
[0091] The succession of states h; for i varying from 1 to n makes it possible to characterize the input data according to n different representations.
[0092] Each output state h; of the recurrent network R is provided, with the masked image, to a predictor network h to provide as output n different predictions of the following image Ylt+1.
[0093]
[0094]
[0095]
[0096]
[0097]
[0098] As will be described later, when multiple data modalities are considered, a different predictor network is used for each of the modalities. For example, if m modalities are considered, m predictor networks {fT} therefore generate m*n predictions, or n different predictions for each of the m modalities. The cost function is calculated as described above by comparing the predictions generated by the predictor network to the actual image Ylt+i of the sequence. Figure 4 illustrates an example of the application of the gradient backpropagation algorithm when the prediction is selected as the closest to the real image. The parameter set of the predictor network fj corresponding to the predictor of rank k* is updated as well as the parameters of the recurrent network and the projector network P. In an alternative embodiment of the example of [Fig.3], the data are multimodal and the video sequence is accompanied by an optical flow designated YF and / or a vector of object classes Yc which indicates, for each image, the classes of the objects present in the image. This variant is represented in [Fig.5] which illustrates, on a diagram, the operation of the gradient backpropagation algorithm in the case where three modalities I (sequence of images), F (sequence of optical flows), C (sequence of object class vectors) of the data are exploited. In this variant, the model is trained for the three modalities I,F,C so that n*3 different predictions are obtained by the three predictors corresponding to the three modalities of the data. Each predictor has its specific set of parameters.
[0099] Thus, for each modality of the data, the model provides a number n of predictions which can be different for each modality. These n predictions can be generated by n distinct predictor networks or by a single network with n branches.
[0100] During the gradient backpropagation phase, the parameters of each predictor network fb fF, fc are optimized from the sum of the respective cost functions calculated for each modality, i.e. the sum
[0101] l =
[0102] The calculated gradients impact the parameters of each predictor network fb fF, fc according to the respective variations of the cost functions calculated for each of the modalities. Thus, the parameters of each predictor network are updated independently.
[0103] In the example of [Fig.5], predictor number 2 is retained to calculate the first cost function for the first modality I, predictor number n is retained to calculate the first cost function for the second modality F and predictor number 2 is retained to calculate the first cost function for the third modality C.
[0104] Predictor number 1 is retained to calculate the second cost function for each of the modalities I,F,C.
[0105] This example is given for purely illustrative purposes, it being understood that the predictors retained may be different or identical for each of the modalities and each of the cost functions. They are determined independently for each modality according to the criteria relating to the calculations of the cost functions described previously.
[0106] The calculated gradients are backpropagated in each of the three predictor networks fb fF, fc up to their input layer then the results are summed via a adder S before being propagated to the set consisting of the projector network P and the recurrent network R which is common to all the modalities.
[0107] In other words, the entire model is affected by the sum of the three cost functions calculated for each modality, but in practice, each respective part of the model (fT predictors, recurrent network R and projector P) is affected differently. The overall cost function is calculated as the sum of the cost functions for each modality, itself being a weighted sum of two cost functions LNN and Lnp. This induces, in the different fT predictors, backpropagation gradients of different amplitude depending on the role played by the parameters of these predictors for the calculation of the predictions. In other words, each fT predictor network is optimized only by the gradient generated by the cost function calculated for the corresponding modality. The P and R networks ultimately receive the sum of the gradients coming from the three predictor networks.
[0108] [Fig.6] illustrates an example of possible architecture for the projector network P, the recurrent network R and the three predictor networks fbfF and fc corresponding to the basic modality (the sequence of images), and the two secondary modalities (optical flow F and object class C).
[0109] In this example, the same predictor network is used for the main modality and the optical flow and a predictor network of different architecture is used for the object class. This example is not limiting, it is possible to use the same predictor network for all modalities with different parameter sets for each modality or different predictor networks for each modality or a combination of the two approaches.
[0110] In the example of [Fig.6], the projector network P is composed of a series of four convolution layers conv, three maxpooling layers max pool, four activation layers implementing a ReLU activation function and two fully connected convolution layers FC arranged in the order shown in [Fig.6].
[0111] The recurrent network R is composed of 7 fully connected convolution layers (of identical dimensions) FC alternated by an activation layer implementing a ReLU activation function. The input is added to the output via a residual connection.
[0112] The f^fp predictor network is composed of two stacked autoencoders. The two autoencoders are identical and each composed of an encoder with four convolution layers conv and four ReLU activation layers and then a decoder implementing the same layers in reverse order. The last activation layer of the decoder implements a sigmoid activation function.
[0113] The input is added to the output of the first autoencoder via a residual connection. The hk state provided by the recurrent network is concatenated to the outputs of each encoder.
[0114] Also, the last layer of each decoder can be adapted for each of the two modalities I and F, in particular the dimension of the convolution layer can vary according to the modality.
[0115] The predictor network fc comprises an encoder consisting of four convolution layers conv, three maxpooling layers maxpool, five activation layers implementing a ReLU activation function and two fully connected convolution layers FC arranged in the order shown in [Fig.6].
[0116] The output of the encoder is concatenated to the hk state provided by the recurrent network and is provided as input to a decoder composed of two fully connected convolution layers, a ReLU activation layer and an activation layer implementing the Softmax activation function.
[0117] The architectural examples given in [Fig.6] are illustrative and not limiting.
[0118]
[0119]
[0120]
[0121]
[0122]
[0123]
[0124]
[0125]
[0126]
[0127]
[0128]
[0129]
[0130]
[0131] Once the machine learning model is trained to reconstruct “normal” data through several credible predictions, it can be used to detect anomalies in a new dataset that may include anomalies. For this, an anomaly detection method involving the implementation of the model described above is proposed in [Fig.7]. The method is applied to 701 data of the same nature and modalities as those used for training, with the difference that they may now contain anomalies. In step 702, the data is masked using the same masking procedure as in step 202 of the training and then the masked data is provided as input to the previously trained model. At step 703, the model calculates several predictions in the same way as at step 203 of training. In step 704, only the first LNN cost function (Y) is calculated using equation (1) by selecting the prediction closest to the reference data. The cost function is calculated by taking into account one or more modalities of the available data. Finally, in step 705 an anomaly score is calculated which is based on the difference between the LNN cost function (Y) and a characteristic value of the average of this cost function calculated during training. Indeed, if the input data 701 does not contain any anomaly, the predictors will correctly reconstruct the masked parts of the data and the calculated cost function will be close to the average of this function calculated during training. Conversely, if the input data contains an anomaly in the masked areas, the predictors will not reconstruct this anomaly and the cost function calculated in step 704 will have a value significantly different from that calculated during training. An example of an anomaly score formula is given in relation (3): 5 ( K ) - L Te r ---- T= { / , F, C} is the set of modalities that the variable T goes through (YT) is the cost function calculated for each modality. and aT are respectively the mean and standard deviation of the same cost function calculated during training for training data that contains no anomalies. is a weighting coefficient which allows different weights to be given to each modality. The anomaly score formula can be substituted for any other metric making it possible to measure the difference between the cost function calculated in step 604 and a value representative of the cost function calculated on training data without anomalies.
[0132] In step 706, the anomaly score is compared to a detection threshold to deduce the presence or absence of anomalies in the input data.
[0133] The invention may be implemented as a computer program comprising instructions for its execution. The computer program may be recorded on a recording medium readable by a processor.
[0134] Reference to a computer program that, when executed, performs any of the functions described above, is not limited to an application program running on a single host computer. Rather, the terms computer program and software are used herein in a general sense to refer to any type of computer code (e.g., application software, firmware, microcode, or any other form of computer instruction) that can be used to program one or more processors to implement aspects of the techniques described herein. The computing means or resources may notably be distributed (“Cloud computing”), possibly using peer-to-peer technologies.The software code may be executed on any suitable processor (e.g., a microprocessor) or processor core or a set of processors, whether provided in a single computing device or distributed among several computing devices (e.g., as may be accessible in the device environment). The executable code of each program enabling the programmable device to implement the processes according to the invention may be stored, for example, in the hard disk or in read-only memory. Generally, the program(s) may be loaded into one of the storage means of the device before being executed. The central unit may control and direct the execution of the instructions or portions of software code of the program(s) according to the invention, instructions which are stored in the hard disk or in read-only memory or in the other aforementioned storage elements.
[0135] The invention can be implemented on a computing device based, for example, on an embedded processor. The processor can be a generic processor, a specific processor, an application-specific integrated circuit (also known as an ASIC for “Application-Specific Integrated Circuit”) or an in situ programmable gate array (also known as an FPGA for “Field-Programmable Gate Array”). The computing device can use one or more dedicated electronic circuits or a general-purpose circuit. The technique of the invention can be implemented on a reprogrammable computing machine (a processor or a microcontroller for example) executing a program comprising a sequence of instructions, or on a dedicated computing machine (e.g. a set of logic gates like an FPGA or ASIC, or any other hardware module).
[0136] The invention makes it possible to improve both reconstruction and prediction learning approaches by using a trained multi-prediction model to reconstruct masked normal data.
[0137] Furthermore, the use of several trained predictors to perform different pretext tasks for different modalities of multimodal data makes it possible to better capture the diversity of the “normal” character of the training data. Indeed, learning predictions according to different modalities: spatial, temporal, optical flow or other makes it possible to better characterize and discriminate the different credible contents of normal data. References
[0138] [1] Duc Tarn Nguyen, Zhongyu Lou, Michael Klar, Thomas Brox. Anomaly Detection with Multiple-Hypotheses Predictions. In ICML, 2018
[0139] [2] Zhian Liu, Yongwei Nie, Chcngjiang Long, Qing Zhang, and Guiqing Li. A hybrid video anomaly détection framework via memory-augmented flow reconstruction and flow-guided frame prédiction. In ICCV, 2021
Claims
Claims
1. A computer-implemented method for training a multimedia data reconstruction model represented by at least one modality (201), the model being composed of a set of several different predictors for each modality of the data, the training method comprising the steps of, for each data item in a training data set comprising no anomalies, - for each modality of the data: • Masking (202) at least part of the modality of the data, • Training each predictor of the set associated with said modality to calculate (203) a different prediction of the same masked data item, each predictor being specialized in a possible prediction of the masked data item from among different credible alternatives, • Selecting the predictor of the set that provides the prediction closest to a reference data item extracted from the training data,• Calculate a distance between said prediction and the reference data, - Calculate (204) a first cost function equal to the sum of said distances for all the modalities, - Update (206) the parameters of the predictors selected for each modality so as to minimize the first cost function.,
2. A method for training an anomaly detection model according to claim 1 further comprising, - for each modality of the data: • Selecting a subset of the set of predictors which have not been optimized during a previous iteration of the training, • Calculating the sum of the distances between the respective predictions provided by the predictors of said subset and the reference data, - Calculate (205) a second cost function equal to the sum of said distances for all the modalities and modify the first cost function by adding to it the second cost function weighted by a weighting coefficient, - Update (206) the parameters of the predictor which provides the prediction closest to the reference data and of the predictors of said subset for each modality so as to minimize the first modified cost function.
3. Method for training an anomaly detection model according to any one of the preceding claims wherein: - the multimedia data (201) are represented, according to a basic modality, by a temporal sequence of successive images, - The step of masking (202) at least part of the basic modality comprises the application of a predefined spatial mask to each image of the sequence so as to mask at least one area of the image, - The reference data corresponds to a successive image in the temporal sequence with respect to the current image provided as input to the model, - Each predictor is trained to predict said successive image from the current image masked by means of said mask.
4. Method for training an anomaly detection model according to claim 3 in which: - The multimedia data (201) are further represented, according to an additional modality, by an optical flow sequence, - The step of masking (202) at least part of the additional modality comprises masking the entire optical flow, - The reference data is the optical flow, - Each predictor is trained to predict the optical flow from the current image masked according to the basic modality.
5. Method for training an anomaly detection model according to any one of claims 3 or 4 wherein: - The multimedia data (201) are further represented, according to an additional modality, by a sequence comprising, for each image, a set of object classes detected in the image, - The step of masking (202) at least part of the additional modality comprises masking the entire sequence of object classes, - The reference data is the sequence of object classes, - Each predictor is trained to predict the sequence of object classes from the current image masked according to the basic modality.
6. Method for training an anomaly detection model according to any one of claims 3 to 5: - The multimedia data (201) are further represented, according to an additional modality, by an audio sequence synchronized with the image sequence, - The step of masking (202) at least part of the additional modality comprises the deletion of at least part of the audio sequence, - The reference data is the audio sequence, - Each predictor is trained to predict the audio sequence from the current image masked according to the basic modality.
7. Method for training an anomaly detection model according to any one of claims 1 or 2 wherein: - the multimedia data (201) are represented, according to a basic modality, by a set of images, - The step of masking (202) at least part of the basic modality comprises the application of a predefined spatial mask to each image so as to mask at least one area of the image, - The reference data corresponds to the current image provided at the input of the model but not masked, - Each predictor is trained to predict said current image from the current image masked by means of said mask.
8. A method of training an anomaly detection model according to claim 7 wherein: - The multimedia data (201) are further represented, according to a second modality, by a set of multi-spectral images, - The step of masking (202) at least a part of the second modality comprises the suppression of images at at least one given wavelength, - The reference data is the set of multi-spectral images, - Each predictor is trained to predict a multi-spectral image from the current image masked according to the basic modality.
9. A method of training an anomaly detection model according to any one of the preceding claims wherein the model comprises: - a first projector neural network (P) receiving masked data as input and trained to transform the input into a latent representation (h0), - a recurrent neural network (R) trained to produce a sequence of states (hi„„ hn), in a recurrent manner, the initial state (h0) being equal to the latent representation, - several predictor neural networks (fT) each corresponding to a modality, each predictor network receiving as input the masked data and a state provided by the recurrent network, the number of states generated by the recurrent network being equal to the number of predictors for a modality.
10. A computer-implemented method of detecting anomalies in a multimedia data set having at least one modality, the method comprising the steps of: - Execute (703), for said data set, the machine learning model trained using the training method according to any one of the preceding claims, the model receiving as input the masked data (702) using masks identical to those used for training said model and producing as output several predictions, - Calculate (704) the first cost function from the predictions generated by the model, - Calculate (705) an anomaly score from a distance between the first cost function and a representative value of the first cost function for a data set including no anomalies, - Compare (706) the anomaly score to a predetermined detection threshold and deduce the presence or absence of anomalies in the data.
11. A computer program comprising code instructions for implementing the methods according to any one of claims 1 to 10, when said program is executed on a computer.
12. A computer-readable recording medium on which the computer program according to claim 11 is recorded.