Method and device for automatic analysis of unlearning at least one class using a data classification model
The automatic analysis process for unlearning in machine learning classification models addresses the challenge of efficiently removing sensitive classes by using control models to verify the unlearning operation, thereby enhancing security and reducing resource usage.
Patent Information
- Application Number
- FR2023012497
- Authority / Receiving Office
- FR · FR
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-15
- Publication Date
- 2025-05-16
AI Technical Summary
Existing systems using machine learning classification models face challenges in effectively unlearning sensitive classes without retraining the entire model, which is resource-intensive and lacks verification mechanisms for ensuring the removal of sensitive classes.
An automatic analysis process and device that determine whether candidate classes are forgotten classes by calculating homogeneity values in a latent space and using control models to refine classification, allowing for the verification of unlearning operations without requiring the initial classification model.
Enables efficient verification of unlearning operations, ensuring that sensitive classes are properly removed from machine learning models, thereby reducing resource usage and enhancing security in applications like industrial, medical, and military fields.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Title of the invention: Method and device for automatic analysis of unlearning of at least one class by a data classification model
[0001] The present invention relates to a method for automatic analysis of unlearning of at least one class by a data classification model.
[0002] The invention also relates to an associated device and an associated computer program.
[0003] The invention lies in the field of security for applications using artificial intelligence.
[0004] Many applied systems, for example in the industrial, medical or military fields, use automatic data classification models, for example models based on artificial neural networks, which are trained by machine learning.
[0005] This type of classification model comprises a very large number of parameters, the values of which are calculated and updated dynamically. It is necessary to learn the values of the parameters defining a classification model, the learning (also called training) of the parameters being carried out in a learning phase on very large quantities of data. Conventionally, such classification models take as input data, for example in vector or matrix form, of known type, and provide as output a classification into a plurality of output classes, for example represented in the form of a vector whose components are output values, also called "logits", each component of the vector corresponding to a class. The higher the output value, the more the corresponding class is predicted by the classification model.
[0006] For example, the input data are images, of predetermined dimensions, and the classification labels represent predetermined classes, for example representative of types of object present in the scene represented by the image, according to the application envisaged.
[0007] The input data used for training, called training data, comes from either public databases or private databases. To perform training, it is necessary to provide example data, also called instances, of each of the output classes.
[0008] In certain applications, requiring a certain level of security, it may be observed a posteriori that certain classes whose classification is initially learned by a classification model, called the initial model, are so-called sensitive classes. It is then planned to carry out an operation called unlearning, which allows, starting from an initial model, to "unlearn" certain classes, which will be called hereinafter forgotten classes, among the initial classes, while continuing to learn new classes. Thus, unlearning makes it possible to obtain a classification model in a new set of output classes, while taking advantage of the learning of the initial classification model. This method is more efficient, and therefore less expensive in terms of resource use, than completely re-learning a classification model.
[0009] Machine unlearning methods for training data, classes or variables of a machine learning classification model have been developed recently.
[0010] It is necessary for a legitimate owner of a machine learning classification model, obtained by applying an unlearning method to an initial classification model, to be able to verify that the unlearning is effective. Similarly, for the legitimate owner of data of one or more sensitive classes, it is critical to ensure that the sensitive class instances have not been used for machine learning of a classification model provided by a third party.
[0011] The aim of the invention is then to propose a method for automatic analysis of unlearning of at least one class by a data classification model, making it possible to verify the unlearning of one or more candidate classes.
[0012] For this purpose, the subject of the invention is a method for automatically analyzing the unlearning of at least one class by a data classification model, called the target model, the target model having been previously trained by machine learning for a classification of data into P output classes, the method receiving as input the target model in the form of a target module for extracting characteristics represented in a latent space, and a target module for classifying the characteristics into P output classes, and a database comprising instances of each of the P output classes and instances of Q candidate classes distinct from the P output classes, the target model having been obtained from an initial model for classifying the data into initial classes, by unlearning at least one class called the forgotten class, among said initial classes, the initial classification model being not provided.
[0013] The method is implemented by a calculation processor and comprises, to determine whether at least one of said candidate classes is one of said forgotten classes, steps of:
[0014] - calculation by machine learning of a plurality of witness models, each witness model comprising a witness module for extracting characteristics represented in a latent space, and a witness classification module for classifying input data into a chosen number of witness classes, from instances of at least a first subset of the output classes and / or at least a second subset of the candidate classes,
[0015] for the or each candidate class:
[0016] - first determination, based on homogeneity values calculated in the latent space for the target model and for each control model, providing as output a first binary prediction and a first associated prediction score, and / or
[0017] - second determination, based on a re-learning characteristic of refining additional classes among said candidate classes by the target model and the control models, providing as output a second binary prediction and associated prediction score;
[0018] - calculation of a final score based on the first prediction score and / or the second prediction score.
[0019] Advantageously, the proposed method makes it possible to determine whether one or more of the candidate classes are part of the forgotten classes, in the absence of any knowledge of the initial classification model.
[0020] The automatic unlearning analysis method according to the invention may also have one or more of the characteristics below, taken independently or in any technically conceivable combination.
[0021] The method implements said first determination and said second determination.
[0022] The first determination comprises, for each candidate class, for a plurality of instances of the candidate classes: - implementation of the target model feature extraction module to obtain a feature vector in the latent space for each instance of each candidate class, - for each candidate class, calculation of a first homogeneity value based on the calculated characteristic vectors.
[0023] The first determination further comprises, for each candidate class and each witness model, an implementation of the witness module for extracting the characteristics of said witness model to obtain a vector of characteristics in the associated latent space, and a calculation of a second homogeneity value per candidate class and per witness model.
[0024] The method further comprises training a normality algorithm on the second homogeneity values calculated for the candidate classes for all the control models.
[0025] It also includes an implementation of the normality algorithm on the first homogeneity values calculated with the target model per candidate class to obtain the first binary prediction and the first candidate class prediction score.
[0026] The second determination implements, for said target model and each of the control models, a re-learning of refinement of at least one candidate class.
[0027] The method comprises an evaluation by candidate class, of a first refinement re-learning speed by the target model, and of second refinement re-learning speeds by the control models.
[0028] According to another aspect, the invention relates to a device for automatic analysis of unlearning of at least one class by a data classification model, called the target model, the target model having been previously trained by machine learning for a classification of data into P output classes, the device receiving as input the target model in the form of a target module for extracting characteristics represented in a latent space, and a target module for classifying the characteristics into P output classes, and a database comprising instances of each of the P output classes and instances of Q candidate classes distinct from the P output classes, the target model having been obtained from an initial model for classifying the data into initial classes, by unlearning at least one class called the forgotten class, among said initial classes, the initial classification model being not provided.
[0029] The device comprising a calculation processor configured to implement, to determine whether at least one of said candidate classes is one of said forgotten classes: - a module for calculating by machine learning a plurality of witness models, each witness model comprising a witness module for extracting characteristics represented in a latent space, and a witness classification module for classifying input data into a chosen number of witness classes, from instances of at least a first subset of the output classes and / or at least a second subset of the candidate classes,
[0030] for the or each candidate class:
[0031] - a first determination module, as a function of homogeneity values computed in the latent space for the target model and for each control model, providing as output a first binary prediction and a first associated prediction score, and / or
[0032] - a second determination module, as a function of a characteristic of re learning to refine additional classes among said classes candidates by the target model and the control models, providing as output a second binary prediction and associated prediction score;
[0033] - a module for calculating a final score based on the first prediction score and / or the second prediction score.
[0034] The device is advantageously configured to implement an automatic unlearning analysis method as briefly described above.
[0035] According to another aspect, the invention relates to an information recording medium, on which are stored software instructions for the execution of an automatic unlearning analysis method as briefly described above, when these instructions are executed by a programmable electronic device.
[0036] According to another aspect, the invention relates to a computer program comprising software instructions which, when implemented by a programmable electronic device, implement an automatic unlearning analysis method as briefly described above.
[0037] The invention will appear more clearly on reading the description which follows, given solely by way of non-limiting example, and made with reference to the drawings in which:
[0038] [Fig-1] [Fig.l] is a schematic representation of the main modules of a automatic unlearning analysis device according to one embodiment;
[0039] [Fig.2] [Fig.2] is a synopsis of the main stages of an analysis process at unlearning automation according to one embodiment;
[0040] [Fig.3] [Fig.3] is a synopsis of the main stages of a first determination reduction of unlearning of one or more of the candidate classes;
[0041] [Fig.4] [Fig.4] is a synopsis of the main stages of a second determination reduction of unlearning of one or more of the candidate classes.
[0042] [Fig.l] schematically represents the main modules of a device 2 for automatic analysis of unlearning of at least one candidate class by a target classification model.
[0043] The device 2 is a programmable electronic device, e.g. a computer. In a variant not shown, the device 2 is formed from a plurality of programmable electronic devices connected to each other.
[0044] The device 2 comprises at least one calculation processor 4, at least one electronic memory unit 6, adapted to communicate via a communication bus 5. For the sake of simplification, only one calculation processor 4 and only one electronic memory unit 6 are shown.
[0045] In addition, the device 2 also comprises a communication interface with remote devices, by a chosen communication protocol, for example a wired protocol and / or a radio communication protocol, as well as a human-machine interface; these interfaces are implemented in a conventional manner and are not shown in [Fig.l].
[0046] The target model MC, referenced 8 in [Fig.l], is provided as input to the device, and stored in the electronic memory unit 6.
[0047] The target model MC is an input data classification model previously trained by machine learning on a training data set, the training data set used not being provided.
[0048] The target model has been previously trained to classify input data into P output classes CSi.. .CSP, P being a positive integer.
[0049] The target model MC is provided in the form of a target feature extraction module 10, the features being represented in a latent space 12, and a target classification module 14, configured for a classification of the features, represented in the latent space 12, into P output classes.
[0050] Each of the modules 10; 14 is for example provided in the form of executable code.
[0051] Thus, it is possible, by executing the target model MC on an input data Ei, presented in vector or matrix form, to obtain by executing the target feature extraction module 10, features in the form of a feature vector of predetermined dimension T (or number of components), for example T=1024 or T=2048 or more, associated with the input data in the latent space 12.
[0052] The dimension T of the latent space is dependent on the target model applied.
[0053] In one embodiment, the target model is a multi-layer neural network.
[0054] In a known manner, a neural network comprises an ordered succession of layers of neurons, each of which takes its inputs from the outputs of the previous layer.
[0055] More precisely, each layer comprises neurons taking their inputs from the outputs of the neurons of the previous layer, or from the input variables for the first layer.
[0056] Alternatively, more complex neural network structures can be envisaged with a layer that can be connected to a layer further away than the immediately preceding layer.
[0057] Each neuron is also associated with an operation, i.e. a type of processing, to be carried out by said neuron within the corresponding processing layer.
[0058] Each layer is connected to the other layers by a plurality of synapses. A synaptic weight is associated with each synapse, and each synapse forms a link between two neurons. It is often a real number, which takes both positive and negative values. In some cases, the synaptic weight is a complex number.
[0059] Each neuron is capable of performing a weighted sum of the value(s) received from the neurons of the previous layer, each value then being multiplied by the respective synaptic weight of each synapse, or link, between said neuron and the neurons of the previous layer, then applying an activation function, typically a non-linear function, to said weighted sum, and delivering at the output of said neuron, in particular to the neurons of the following layer which are connected to it, the value resulting from the application of the activation function. The activation function makes it possible to introduce a non-linearity into the processing carried out by each neuron. The sigmoid function, the hyperbolic tangent function, the Heaviside function are examples of activation functions.
[0060] As an optional addition, each neuron is also capable of applying, in addition, a multiplicative factor, and an additive bias, to the output of the activation function, and the value delivered at the output of said neuron is then the product of the multiplicative factor value and the value resulting from the activation function, added to the bias.
[0061] A convolutional neural network is also sometimes called a convolutional neural network or by the acronym CNN which refers to the English term “Convo-lutional Neural Networks”.
[0062] In a convolutional neural network, each neuron in the same layer has exactly the same connection pattern as its neighboring neurons, but at different input positions. The connection pattern is called a convolution kernel or, more often, a "kernel" in reference to the corresponding English term.
[0063] A fully connected layer of neurons is a layer in which the neurons of said layer are each connected to all the neurons of the previous layer.
[0064] Such a type of layer is more often referred to as “fully connected” and sometimes referred to as “dense layer”.
[0065] The values of the weights, multiplier factors and biases where applicable are learned during a machine learning phase to carry out the classification task.
[0066] The invention applies to all types of neural networks.
[0067] In one embodiment, the target model MC is a multi-layer neural network, and the target feature extraction module 10 comprises the layers of the neural network up to the penultimate layer of the multi-layer neural network, and the target classification module 14, also called "classification head" comprises the last layer.
[0068] The structure and parameters forming the target MC model, and in particular the structure and parameters of the target feature extraction module, are also provided as input to the unlearning analysis device and method.
[0069] In addition to the target model, the device 2 receives as input a database 16 comprising data of the type of input data of the target model. The database 16 comprises class example data, also called instances, which are, on the one hand, instances 18 of each of the P output classes, denoted CSi .. .CSP, and instances 20 of Q candidate classes, denoted CCi.. .CCQ.
[0070] The instances 18 of the database 16 are a priori different from the instances of the training database used during the training of the target model, the training database being a priori not provided.
[0071] The numbers P and Q are positive integers.
[0072] The number Q of candidate classes is greater than or equal to 1.
[0073] The candidate classes are all distinct from the P output classes.
[0074] The candidate classes are the classes on which the automatic unlearning analysis method is based.
[0075] It is considered that the target model is likely to have been obtained by unlearning at least one class, called the forgotten class, from an initial data classification model of the same type as the input data of the target model. The initial classification model itself having been trained by machine learning to classify the data into initial classes, the initial classes being at least partly distinct from the P output classes of the target model.
[0076] The automatic unlearning analysis method described does not take as input either the initial classification model or the initial classes.
[0077] In other words, the method is executed in the absence of knowledge about the initial classification model, and also in the absence of knowledge of the unlearning method and, where appropriate, relearning, implemented.
[0078] The processor 4 is configured to implement a module 30 for generating MTi... MTN control models by machine learning, referenced by references 22i • • -22n in the figure, the number N being a variable of the method implemented.
[0079] Each witness model MT; is a classification model composed of a witness module for extracting characteristics 24; represented in the form of a vector of characteristics in a latent space 12i, of the same dimension as the latent space 12. and of a witness classification module 26;, configured for a classification of the characteristics, represented in the latent space 12, into a chosen number R of witness classes.
[0080] The witness classes are chosen from the output classes CSi.. .CSP and the candidate classes CSi...CSQ.
[0081] For example, the number R of control classes is between P and P+Q.
[0082] The parameters of each 22j witness model are learned by machine learning on a separate partition from the database data 16.
[0083] Preferably, each witness model 22j is learned from a partition comprising instances of at least a first subset of the output classes CSi.. .CSP and at least a second subset of the candidate classes CSi.. .CSQ. In other words, the R witness classes comprise output classes and / or candidate classes.
[0084] Preferably, at least one of the control models is trained to learn the P output classes.
[0085] The training of the control model(s) for which the control classes are the same as the output classes is carried out with the objective of obtaining a classification that is substantially identical, in particular with a similar statistical distribution, to that of the target model on the output classes forming part of the control classes. In other words, each control model for which the control classes are the same as the output classes “imitates” as best as possible the target model for the common classification task.
[0086] Preferably, the other control models are trained for the classification of the R chosen control classes by a learning process analogous to that implemented for the control models for which the control classes are the same as the output classes.
[0087] The witness models thus generated are used for learning determination methods to determine whether a candidate class is one of the classes forgotten by an unlearning method applied to the initial classification model from which the target model is derived.
[0088] Such a determination method applies for each candidate class and provides as output a binary prediction (forgotten class or not) and an associated prediction score, the prediction score indicating the reliability of the prediction and making it possible to determine a probability that the candidate class is a forgotten class.
[0089] The processor 4 further comprises:
[0090] - a first determination module 32 implemented for one or each candidate class, as a function of homogeneity values calculated in the latent space for the target model and in the latent space of each control model, providing as output a first prediction and a first associated prediction score; and
[0091] - a second determination module 34, implemented for the or each candidate class, as a function of a re-learning characteristic of additional classes among the candidate classes by the target model and by each of the control models, providing as output a second prediction and a second associated prediction score.
[0092] The processor 4 is configured to implement the first determination module 32. mination and / or module 34 of second determination.
[0093] The processor 4 further comprises a module 36 for calculating a final score as a function of the first prediction score and / or the second prediction score.
[0094] Preferably, the first determination module 32 and the second determination module 34 are implemented, because their results are complementary, and the final score obtained is more reliable.
[0095] In one embodiment, the modules 30, 32, 34, 36 are produced in the form of software instructions forming a computer program, which, when executed by a programmable electronic device, implements a method of automatic analysis of unlearning of at least one class by a data classification model according to the invention.
[0096] In a variant not shown, the modules 30, 32, 34, 36 are each produced in the form of programmable logic components, such as FPGAs (Field Programmable Gate Arrays) of microprocessors, GPGPU components (General-purpose processing on graphics processing), or even dedicated integrated circuits, such as ASICs (Application Specific Integrated Circuits)•
[0097] The computer program comprising software instructions is further capable of being recorded on a non-transitory, computer-readable information recording medium. This computer-readable medium is, for example, a medium capable of storing electronic instructions and of being coupled to a bus of a computer system. For example, this medium is an optical disk, a magneto-optical disk, a ROM memory, a RAM memory, any type of non-volatile memory (for example EPROM, EEPROM, FLASH, NVRAM), a magnetic card or an optical card.
[0098] [Fig.2] is a synopsis of the main steps of a method for automatic analysis of unlearning of at least one class according to one embodiment.
[0099] The method comprises a step 40 of recovering, on the one hand, the target model provided in the form of a target module for extracting target characteristics and a target module for classifying into P output classes, and instances of a database, comprising instances of the P output classes of the target model and instances of Q candidate classes, distinct from the output classes.
[0100] It should be noted that optionally, the method also comprises a step of increasing the number of instances (not shown) for example by image processing when the instance data are digital images (e.g. rotation, translation, vertical or horizontal flipping). This makes it possible in particular to ensure that a sufficient number of instances is available for the calculation of control models, described below.
[0101] The method then comprises a step 42 of partitioning the instances into a plurality of partitions, so as to generate distinct partitions for the training of the witness models.
[0102] For example, in step 42 N distinct partitions are calculated, for the training of N witness models, each partition containing instances of R classes chosen from the output classes CSi...CSP and the candidate classes CCi...CCQ.
[0103] The sizes of the distinct partitions may be different, depending on the number Ri of classes for each control model MTi.
[0104] Each partition comprises a first subset of instances intended for training a witness model and a second subset of test or validation instances of a training model.
[0105] The method then comprises a learning 44 of N witness models.
[0106] Each witness model MT; is trained by machine learning to classify instances into R; witness classes.
[0107] Preferably the Ri witness classes comprise classes among the output classes CSi...CSP and classes among the candidate classes CCi...CCQ.
[0108] The method implements, using the target model, the calculated witness models and the instances received as input, respectively a step 46 of first determination and / or a step 48 of second determination, which are described in detail below.
[0109] Preferably, the method implements the first determination 46 and the second determination 48, these steps each providing a respective result RES1 and RES2, for each of the candidate classes.
[0110] Preferably, the respective results are binary predictions and prediction scores, for each of the candidate classes, the binary prediction indicating the prediction for the candidate class to be part of the classes said to be forgotten by the target model.
[0111] Preferably, the first determination 46 is carried out as a function of homogeneity values calculated for each candidate class, in the respective latent space of each model, for the target model and for each control model.
[0112] Preferably, the second determination 48 is carried out as a function of a characteristic of re-learning additional classes among said candidate classes by the target model and the control models.
[0113] The method then comprises a step 50 of calculating a final score per candidate class. In the case where the first determination 46 and the second determination 48 are implemented, step 50 implements an aggregation of the respective results.
[0114] For example, the aggregation is carried out by calculating an average of the first and second scores.
[0115] In the case where one or other of the determination steps 46, 48 is implemented, step 50 implements a simple resumption of the probability scores calculated by the determination step implemented.
[0116] The final score calculated then makes it possible to determine, for each candidate class, whether it is one of the classes forgotten during re-learning of the target model from an initial model not provided.
[0117] An embodiment of the first determination 46 is described with reference to [Fig.3].
[0118] In this embodiment, the first determination is carried out by homogeneity values calculated for each candidate class, on the one hand in the latent space of the target model, on the other hand in the latent space for each control model.
[0119] Starting from the instances 20 of the candidate classes, the method comprises a step 54 of extracting the instances of one of the candidate classes.
[0120] For the processed candidate class CQ, each of its instances Ejk is provided as input to the target feature extraction module, which is executed (step 56) to obtain a T-dimensional feature vector Vcjk in the latent space.
[0121] Step 56 is repeated for each of the candidate classes, and the obtained feature vectors are stored.
[0122] The method then comprises a calculation 58 of a first homogeneity value, for each candidate class, from the set of characteristic vectors obtained with the target model, for all the instances of the candidate classes CCk CCQ
[0123] The first homogeneity value, for each candidate class, is representative of the spatial homogeneity of the set of feature vectors (or cluster) of the candidate class in the latent space.
[0124] For example, in one embodiment, the calculation of a calculated homogeneity value is a Silhouette coefficient. In known manner, the Silhouette coefficient, for a given instance Ejk, depends on:
[0125] -of the average of the distances between the characteristic vector Vcjk and the characteristic vectors Vcjm, w & of the same candidate class CCjprocessed, and
[0126] -of the minimum distance between the characteristic vector Vcjk and the characteristic vectors Vchn, hj, n being an index varying from 1 to the number of instances of the candidate class CCh.
[0127] Substantially in parallel or sequentially, the method comprises the implementation for each witness model, selected in step 60, the implementation of the witness module for extracting the characteristics on each instance of the candidate class (step 62) to obtain respective characteristic vectors of dimension T, and the iteration of step 62 for each candidate class, then the calculation 64 of a second homogeneity value per candidate class and per control model, calculation 64 being analogous to the calculation carried out in step 58.
[0128] Thus, first and second homogeneity values are obtained in the latent space for each candidate class CCi ... CCQ, by applying the target model and each of the control models.
[0129] The first determination then comprises a step 66 of learning a normality algorithm on the second homogeneity values previously calculated on the control models.
[0130] For example, an "isolation forest" algorithm is implemented, which calculates an anomaly score for each observation in a data set.
[0131] In its application in the method described here, knowing the classes learned by each witness model, step 66 performs a learning of anomaly scores for each candidate class, as a function of the second homogeneity values calculated, depending on whether the candidate class has been learned or not by the witness models.
[0132] It is then possible to determine, from a homogeneity value calculated for a candidate class in a latent space of a classification model, a score representative of the fact that the candidate class has been learned by the classification model or not learned by the classification model. The method finally comprises a step 68 of applying the algorithm with the first homogeneity values obtained for each candidate class CQ to calculate an associated anomaly score. The anomaly score per class is then transformed into a first prediction score of the fact that the candidate class CQ is a forgotten class.
[0133] According to variants, step 66 of learning a normality algorithm implements the DBSCAN algorithm (for “Density-Based Spatial Clustering of Applications with Noise” or the LOF algorithm (for “Local Outlier Factor”), known in the field of artificial intelligence for unsupervised anomaly detection.
[0134] An embodiment of the second determination 48 is described with reference to [Fig.4],
[0135] In this embodiment, the second determination is performed based on a characteristic of re-learning additional classes among said candidate classes by the target model and the control models.
[0136] Starting from the instances 16 of the output classes of the target model and candidate classes, the method comprises, for subsets of the instances selected in step 70 forming a training data set, a refinement re-learning (in English “fine-tuning”).
[0137] In one embodiment, the method implements raf relearning finely of all candidate classes for a plurality of complete passes of the training datasets, each pass also being called an epoch.
[0138] The performance of the refinement relearning is then analyzed for each candidate class.
[0139] For example, a learning performance is represented by a curve representing the classification accuracy of instances of the candidate class as a function of the epochs.
[0140] The refinement retraining is applied, from the target model in step 72, using the previously trained target feature extraction module.
[0141] The refinement re-learning then modifies the target classification module to move from P output classes to P' output classes, P' being equal to P increased by the number of classes included in the subset of instances selected in step 70.
[0142] In one embodiment, P'=P+Q, or in other words, the subset of instances selected in step 70 comprises instances for all Q candidate classes.
[0143] The method further comprises an evaluation 74 of the re-learning speed of each candidate class by the target model, called first re-learning speed.
[0144] For example, the re-learning speed is evaluated by the average learning accuracy obtained by the model considered at each epoch.
[0145] Substantially in parallel or sequentially, the method comprises the implementation for each witness model, selected in step 76, of a refinement re-learning (in English “fine-tuning”) in step 78, using the previously trained witness feature extraction module.
[0146] The refinement re-learning then modifies the classification witness module of the witness model considered to move from R witness classes to R' witness classes, R' being equal to R increased by the number of classes included in the subset of instances selected in step 70.
[0147] In one embodiment, R'=R+Q, or in other words, the subset of instances selected in step 70 includes instances for all Q candidate classes.
[0148] For each control model, an evaluation of a second relearning speed, similar to that carried out in step 74 for the target model, is carried out in step 80.
[0149] The method then comprises a comparison 82 of the second relearning speeds, for each candidate class, between the control models, in knowledge of the set of witness classes of each witness model.
[0150] In particular, it is found that the re-learning speed is higher in the refinement re-learning phase when the candidate class considered is part of the control classes, i.e. was learned during the learning of the control model considered compared to the learning speed in the refinement re-learning phase when the candidate class considered is not part of the control classes of the control model considered.
[0151] For example, we consider the following values representative of the relearning speed:
[0152] -the epoch index from which the classification accuracy reaches a plateau, i.e. the slope of the classification accuracy curve is horizontal with a margin of 2 to 3%, and / or
[0153] -a maximum precision value reached after a given number of epochs, and / or
[0154] -a minimum value of a loss function minimized during learning, for example the cross-entropy function.
[0155] Of course, other values representative of the re-learning speed within the reach of those skilled in the art are applicable.
[0156] The method then comprises a step 84 of calculating, for each candidate class, one or more values representative of the first re-learning speed of the target model, for example among the values listed above.
[0157] The method then comprises a step 86 of comparing, for each candidate class, the value(s) representative of the first relearning speed of the model with the values representative of the second relearning speed depending on whether or not the candidate class belonged to the classes learned by the control models.
[0158] Depending on a result of the comparison, the method comprises, for each candidate class, a determination of the second binary prediction (forgotten class or not) and of a second associated prediction score. For example, if, for a candidate class, the or each representative value of the first re-learning speed is closer to a representative value of the second re-learning speed corresponding to the control models having initially learned the candidate class, the second binary prediction predicts that the candidate class is a forgotten class.
Claims
1. Claims Method for automatic analysis of unlearning of at least one class by a data classification model, called target model, the target model having been previously trained by machine learning for a classification of data into P output classes, the method receiving as input the target model in the form of a target module for extracting characteristics represented in a latent space, and a target module for classifying the characteristics into P output classes, and a database comprising instances of each of the P output classes and instances of Q candidate classes distinct from the P output classes, the target model having been obtained from an initial model for classifying the data into initial classes, by unlearning at least one class called forgotten class, among said initial classes, the initial classification model being not provided, the method being implemented by a calculation processor and comprising,to determine whether at least one of said candidate classes is one of said forgotten classes, steps of:, - calculation (44) by machine learning of a plurality of witness models, each witness model comprising a witness module for extracting characteristics represented in a latent space, and a witness classification module for classifying input data into a chosen number of witness classes, from instances of at least a first subset of the output classes and / or at least a second subset of the candidate classes, for the or each candidate class: - first determination (46), based on homogeneity values calculated in the latent space for the target model and for each control model, providing as output a first binary prediction and a first associated prediction score, and / or - second determination (48), based on a relearning characteristic of refining additional classes among said candidate classes by the target model and the control models, providing as output a second binary prediction and associated prediction score; - calculation (50) of a final score based on the first prediction score and / or the second prediction score.
2. A method according to claim 1, implementing said first determination and said second determination.
3. Method according to one of claims 1 or 2, in which the first determination comprises, for each candidate class, for a plurality of instances of the candidate classes: - implementation of the target module for extracting the characteristics of the target model to obtain a vector of characteristics in the latent space for each instance of each candidate class, - For each candidate class, calculation of a first homogeneity value as a function of the calculated vectors of characteristics.
4. Method according to claim 3, in which the first determination further comprises, for each candidate class and each witness model, an implementation of the witness module for extracting the characteristics of said witness model to obtain a vector of characteristics in the associated latent space, and a calculation of a second homogeneity value per candidate class and per witness model.
5. The method of claim 4, further comprising training a normality algorithm on the second homogeneity values calculated for the candidate classes for all the control models.
6. Method according to claim 5, comprising an implementation of the normality algorithm on the first homogeneity values calculated with the target model per candidate class to obtain the first binary prediction and the first candidate class prediction score.
7. Method according to any one of claims 1 to 5, in which the second determination implements, for said target model and each of the control models, a refinement re-learning of at least one candidate class.
8. The method of claim 7, comprising an evaluation per candidate class of a first refinement retraining rate by the target model, and second refinement retraining rates by the control models.
9. A computer program comprising software instructions which, when executed by a programmable electronic device, implement a method for automatic analysis of unlearning of at least one class in accordance with claims 1 to 8.
10. Device for automatic analysis of unlearning of at least one class by a data classification model, called target model, the target model having been previously trained by machine learning for a classification of data into P output classes, the device receiving as input the target model in the form of a target module for extracting characteristics represented in a latent space, and a target module for classifying the characteristics into P output classes, and a database comprising instances of each of the P output classes and instances of Q candidate classes distinct from the P output classes, the target model having been obtained from an initial model for classifying the data into initial classes, by unlearning at least one class called forgotten class, among said initial classes, the initial classification model not being provided,the device comprising a calculation processor configured to implement, to determine whether at least one of said candidate classes is one of said forgotten classes:, - a calculation module (30) by machine learning of a plurality of witness models (22i,...,22n), each witness model comprising a witness module for extracting characteristics (24i,...,24n) represented in a latent space (12i,... 12n), and a witness classification module (26x,... .,26N) for classifying input data into a chosen number of witness classes, from instances of at least a first subset of the output classes and / or at least a second subset of the candidate classes, for the or each candidate class: - a first determination module (32), as a function of homogeneity values calculated in the latent space for the target model and for each control model, providing as output a first binary prediction and a first associated prediction score, and / or - a second determination module (34), as a function of a re-learning characteristic of refinement of additional classes among said candidate classes by the target model and the models witnesses, providing as output a second binary prediction and associated prediction score; - a calculation module (36) of a final score as a function of the first prediction score and / or the second prediction score.