Apparatus, method and system for providing signature-based machine forgetting
By introducing auxiliary tasks and Fisher information matrices into the machine learning model, we can efficiently remove data contributions from data providers without retraining the model. This addresses the challenges of data privacy protection and forgetting in machine learning models, reduces costs, and improves model flexibility.
Patent Information
- Application Number
- CN202511056847.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-07-31
- Filing Date
- 2025-07-30
- Publication Date
- 2026-02-03
AI Technical Summary
In existing technologies, the process of deleting training samples and retraining the model in machine learning models is expensive and resource-intensive, especially for complex models and large datasets, making it difficult to effectively meet user privacy protection and data forgetting requirements.
By introducing auxiliary tasks into the machine learning model, the signatures of data providers are mapped to identifiers, and the sensitivity of model parameters to data is calculated using the Fisher information matrix or equivalent statistical measures. The model parameters are then updated to enable machine forgetting, which is applicable to both continuous learning and batch learning models.
It enables the efficient removal of data contributions from specific data providers from machine learning models without retraining the model, protecting data privacy, reducing storage and computing costs, and improving model flexibility and adaptability.
Smart Images

Figure CN121457656A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The disclosed subject matter generally relates to machine forgetting, data privacy, and continual learning. BACKGROUND
[0002] With the increasing popularity of machine learning (ML), consumers and data providers express more concerns about the privacy and misuse of their data sets. As a result, recent data protection regulations (e.g., GDPR: General Data Protection Regulation, CCPA: California Consumer Privacy Act) introduce new laws to protect the privacy of users by granting them the “right to be forgotten.” These laws mandate the deletion of data upon request: specified training samples must be discarded from the training set (if stored) and the training model(s). However, simply deleting samples from the training data and retraining the ML model from scratch using the updated data is an expensive and resource-intensive process, especially for complex ML models and large data sets.
[0003] Some example embodiments
[0004] Accordingly, there is a need for machine forgetting that can remove the impact of a requested training sample on an individual consumer or data provider from a machine learning (ML) model without having to retrain the model from scratch.
[0005] According to one example embodiment, an apparatus includes means for configuring a ML model to learn at least one primary task and one auxiliary task. The auxiliary task maps at least one signature associated with at least one data provider to at least one identifier associated with the at least one data provider, and trains the ML model using training data labeled with the at least one signature. The apparatus also includes means for computing at least one data structure (e.g., a data structure representing a measure of information, such as a Fisher Information Matrix (FIM) or equivalent statistical measure of information sensitivity) that represents the sensitivity of at least one parameter of the ML model to training data associated with the at least one data provider. The apparatus further includes means for updating one or more model parameters of the ML model based on the at least one data structure (e.g., representing a measure of information) to perform machine forgetting on training data associated with the at least one data provider indicated in a forgetting request.
[0006] According to another embodiment, an apparatus includes at least one processor, and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to configure an ML model to learn at least one primary task and an auxiliary task. The auxiliary task maps at least one signature associated with at least one data provider to at least one identifier associated with the at least one data provider, and the ML model is trained using training data labeled with the at least one signature. The apparatus is also caused to compute at least one data structure (e.g., a data structure representing a measure of information, such as a Fisher Information Matrix (FIM) or an equivalent statistical measure of information sensitivity) that represents a sensitivity of at least one parameter of the ML model to training data associated with the at least one data provider. The apparatus is further caused to update one or more model parameters of the ML model based on the at least one data structure (e.g., representing a measure of information) to perform machine forgetting of training data associated with the at least one data provider indicated in a forgetting request.
[0007] According to another embodiment, a method includes configuring an ML model to learn at least one primary task and an auxiliary task. The auxiliary task maps at least one signature associated with at least one data provider to at least one identifier associated with the at least one data provider, and the ML model is trained using training data labeled with the at least one signature. The method also includes computing at least one data structure (e.g., a data structure representing a measure of information, such as a Fisher Information Matrix (FIM) or an equivalent statistical measure of information sensitivity) that represents a sensitivity of at least one parameter of the ML model to training data associated with the at least one data provider. The method further includes updating one or more model parameters of the ML model based on the at least one data structure (e.g., representing a measure of information) to perform machine forgetting of training data associated with the at least one data provider indicated in a forgetting request.
[0008] According to another embodiment, a computer program comprising instructions which, when executed by an apparatus, cause the apparatus to configure an ML model to learn at least one primary task and an auxiliary task. The auxiliary task maps at least one signature associated with at least one data provider to at least one identifier associated with the at least one data provider, and trains the ML model using training data labeled with the at least one signature. The apparatus is further caused to compute at least one data structure (e.g., a data structure representing a measure of information, such as a Fisher Information Matrix (FIM) or an equivalent statistical measure of information sensitivity) representing a sensitivity of at least one parameter of the ML model to training data associated with the at least one data provider. The apparatus is further caused to update one or more model parameters of the ML model based on the at least one data structure (e.g., representing a measure of information) to perform machine forgetting learning on training data associated with the at least one data provider indicated in a forgetting request.
[0009] According to another embodiment, a computer program comprising instructions which, when executed by an apparatus, cause the apparatus to configure an ML model to learn at least one primary task and an auxiliary task. The auxiliary task maps at least one signature associated with at least one data provider to at least one identifier associated with the at least one data provider, and trains the ML model using training data labeled with the at least one signature. The apparatus is further caused to compute at least one data structure (e.g., a data structure representing a measure of information, such as a Fisher Information Matrix (FIM) or an equivalent statistical measure of information sensitivity) representing a sensitivity of at least one parameter of the ML model to training data associated with the at least one data provider. The apparatus is further caused to update one or more model parameters of the ML model based on the at least one data structure (e.g., representing a measure of information) to perform machine forgetting learning on training data associated with the at least one data provider indicated in a forgetting request.
[0010] According to another embodiment, a non-transitory computer-readable storage medium comprising program instructions that, when executed by an apparatus, cause the apparatus to configure an ML model to learn at least one primary task and an auxiliary task. The auxiliary task maps at least one signature associated with at least one data provider to at least one identifier associated with the at least one data provider, and trains the ML model using training data labeled with the at least one signature. The apparatus is further caused to compute at least one information matrix (e.g., a data structure representing a measure of information, such as a Fisher Information Matrix (FIM) or an equivalent statistical measure of information sensitivity) representing a sensitivity of at least one parameter of the ML model to training data associated with the at least one data provider. The apparatus is further caused to update one or more model parameters of the ML model based on the at least one data structure (e.g., representing a measure of information) to perform machine forgetting of training data associated with the at least one data provider indicated in a forgetting request.
[0011] According to one example embodiment, an apparatus comprises ML circuitry configured to cause an ML model to learn at least one primary task and an auxiliary task. The auxiliary task maps at least one signature associated with at least one data provider to at least one identifier associated with the at least one data provider, and trains the ML model using training data labeled with the at least one signature. The ML circuitry is further caused to compute at least one information matrix (e.g., a data structure representing a measure of information, such as a Fisher Information Matrix (FIM) or an equivalent statistical measure of information sensitivity) representing a sensitivity of at least one parameter of the ML model to training data associated with the at least one data provider. The ML circuitry is further caused to update one or more model parameters of the ML model based on the at least one data structure (e.g., representing a measure of information) to perform machine forgetting of training data associated with the at least one data provider indicated in a forgetting request.
[0012] According to further embodiments, an apparatus comprises at least one processor; and at least one memory including computer program code for one or more programs, the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus to configure an ML model to learn at least one primary task and an auxiliary task. The auxiliary task maps at least one signature associated with at least one data provider to at least one identifier associated with the at least one data provider, and trains the ML model using training data labeled with the at least one signature. The apparatus is further caused to compute at least one information matrix (e.g., a data structure representing a measure of information, such as a Fisher Information Matrix (FIM) or equivalent statistical measure of information sensitivity) representing a sensitivity of at least one parameter of the ML model to training data associated with the at least one data provider. The apparatus is further caused to update one or more model parameters of the ML model based on the at least one data structure (e.g., representing a measure of information) to perform machine forgetting of training data associated with the at least one data provider indicated in a forgetting request.
[0013] Further, for various example embodiments of the invention, the following applies: a method comprising facilitating (1) a process and / or (2) a computation, and / or (3) a process and / or process steps and / or a computation and / or a process step, solely and / or resulting from (1) a piece of data and / or (2) information and / or (3) at least one signal, and / or any combination thereof.
[0014] For various example embodiments of the invention, the following also applies: a method comprising facilitating creating and / or facilitating modifying (1) at least one device user interface element and / or (2) at least one device user interface functionality, the (1) at least one device user interface element and / or (2) at least one device user interface functionality based at least in part on data and / or information resulting from (and / or computed and / or generated and / or derived from) one or any combination of methods disclosed in this application as relevant to any embodiment of the invention.
[0015] For various example embodiments of the invention, the following also applies: a method comprising facilitating creating and / or facilitating modifying (1) at least one device user interface element and / or (2) at least one device user interface functionality, the (1) at least one device user interface element and / or (2) at least one device user interface functionality based at least in part on data and / or information resulting from (and / or computed and / or generated and / or derived from) one or any combination of methods disclosed in this application as relevant to any embodiment of the invention.
[0016] For various example embodiments of the application, the following applies: A method comprising creating and / or modifying (1) at least one device user interface element and / or (2) at least one device user interface functionality, the (1) at least one device user interface element and / or (2) at least one device user interface functionality based at least in part on data and / or information resulting from and / or at least one signal resulting from one or any combination of the methods (or processes) disclosed in this application related to any embodiment of the application.
[0017] In various example embodiments, the methods (or processes) can be implemented at the service provider side or at the mobile device side or in any shared manner between the service provider and the mobile device, and the actions performed on both sides.
[0018] For various example embodiments, the following applies: An apparatus comprising means for performing the methods of the claims.
[0019] According to some aspects, the subject matter of the independent claims is provided. Some additional aspects are defined in the dependent claims.
[0020] Other aspects, features, and advantages of the present application will become apparent from the following detailed description, taken in conjunction with the accompanying drawings, which illustrate, by way of example, the principles of the application. The application can be implemented in various embodiments and can take on many different forms. There are a number of non-limiting examples of the application as described herein. BRIEF DESCRIPTION OF DRAWINGS
[0021] Example embodiments of the application are illustrated by way of example and not limitation in the figures of the accompanying drawings in which:
[0022] Figure 1 is a diagram of a system that can provide signature-based machine forgetting according to one example embodiment;
[0023] Figure 2 is a diagram of a continual learning model according to one example embodiment;
[0024] Figure 3A and Figure 3B is a diagram of an example of task interference in multi-task learning according to one example.
[0025] Figure 4 is a diagram of components of a model manager according to one example embodiment;
[0026] Figure 5is a flowchart of a process for signature-based machine forgetting according to one example embodiment;
[0027] Figure 6 is a diagram of example signatures for labeling data according to one example embodiment;
[0028] Figure 7 is a diagram of a model undergoing machine forgetting according to one example embodiment;
[0029] Figure 8 is a diagram of a model with affected parameters after machine forgetting according to one example embodiment;
[0030] Figures 9A-9D is a diagram of a timing diagram for signature-based forgetting for a continuous model according to one example embodiment;
[0031] Figure 10 is a diagram of hardware that can be used to implement example embodiments; and
[0032] Figure 11 is a diagram of a chip set that can be used to implement example embodiments.
[0033] Description of some embodiments
[0034] Examples of methods, apparatuses, and computer programs for providing signature-based machine forgetting according to one example embodiment are disclosed hereinafter. In the following description, for the purposes of explanation, numerous specific details and examples are set forth in order to provide a thorough understanding of the embodiments of the application. It is apparent, however, to one skilled in the art that the embodiments of the present application can be practiced without many of these specific details or with an equivalent arrangement. In other instances, well-known structures and devices are shown in block diagram form in order to avoid unnecessarily obscuring the embodiments of the application.
[0035] References in the specification to “one embodiment,” “an embodiment,” “example embodiment,” “embodiment,” or “the embodiment” mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the disclosure. The appearances of the phrase “in one embodiment” or “in an example embodiment” in various places in the specification are not necessarily all referring to the same example embodiment, nor are they necessarily mutually exclusive, or alternative to one another, to other embodiments. Furthermore, the described embodiments are provided by way of example and so “one embodiment” can equally mean “one example embodiment.” Additionally, the terms “a” and “an” herein do not denote a limitation of quantity, but rather denote the presence of at least one of the referenced items. Furthermore, various features of the described embodiments can be exhibited by some embodiments and not by others. Similarly, various requirements can be imposed by some embodiments and not others.
[0036] As used herein, “at least one of ,” “one or more of ,” “at least one of or combinations thereof,” and similar phrases, where a list of two or more elements is present, means that at least one, or at least one of any two or more elements, or at least all of the elements are present.
[0037] Figure 1 is a diagram of a system capable of providing signature-based machine forgetting according to one example embodiment. As described above, with the increasing popularity of machine learning (ML), consumers and data providers (e.g., user equipment (UE) devices 101a-101n - also referred to collectively as UEs 101) express more concerns about privacy and misuse of their data sets used in ML applications. As a result, recent data protection regulations (e.g., GDPR: General Data Protection Regulation, CCPA: California Consumer Privacy Act) introduce new laws that protect the privacy of users by giving users a “right to be forgotten.” These laws mandate the deletion of data upon request, including requiring the owner of the associated ML model to discard any specified training samples from the training set (if stored) and the training model(s) (e.g., ML model 103). However, simply deleting samples from the training data and retraining the ML model from scratch with the updated data is an expensive process, especially for complex ML models and large data sets. Machine forgetting addresses this problem by removing the effect of the requested samples from the ML model without retraining the model from scratch.
[0038] More specifically, machine forgetting aims to modify the trained model so that it behaves as if the data that is to be forgotten (i.e., data that is to be deleted upon request) was not used for training by ensuring (1) indistinguishability between model distributions, or (2) indistinguishability between model outputs. The level of forgetting request also varies in different settings. For example, the deletion request can be (1) sample (item) removal, (2) class removal, (3) feature removal, (4) sequence removal, (5) graph removal (e.g., especially for graph neural networks), (6) task removal, or (7) data provider (e.g., especially a user or client such as a UE 101 belonging to a user) removal. In one embodiment, the various embodiments described herein contemplate a scenario in which machine forgetting aims for indistinguishability between the output of the forgotten and retrained model (e.g., based on a model accuracy test), and the level of forgetting is “data provider,” where each endpoint (e.g., UE 101 or its user) can request to eliminate the influence of its data from the ML model(s), or where forgetting is necessary due to privacy regulations (e.g., the data provider / UE 101 leaves the system 100).
[0039] In one embodiment, the various described methods contemplate a scenario in which the data provider makes its (private) data available to the organization and enables the construction of ML models in a continual learning manner only. Thus, the data provider does not have any local model at the client and its data evolves constantly. As used herein, continual learning refers to enabling ML models to integrate new data without explicit retraining. For example, in batch learning-based ML, a system has access to a dataset used to train (fit) the ML model on. The system then deploys the model and assumes that data the model will see in the future is drawn from the same underlying distribution as the training data, so the model can perform down-stream predictions. However, unlike such batch learning in which a model is trained on a fixed dataset and then deployed, continual learning enables the model to learn from new tasks while retaining knowledge from previous tasks or experiences. It addresses the challenge of acquiring and retaining knowledge (e.g., stability-plasticity dilemma) in dynamic and evolving environments with limited access to past data. Continual learning algorithms are designed to incrementally update model parameters, adapt to new information, and avoid catastrophically forgetting previously learned knowledge. This allows the model to stay updated, handle concept drift (e.g., changes in patterns in different data segments), and effectively incorporate new data without needing to retrain from scratch.
[0040] First, the concept of machine forgetting seems to be in conflict with continual learning, as even if the old data has been discarded, the ML model in continual learning must maintain its performance on both new and old data. However, if the input data belongs to different resources, or users / UEs 101, immediate machine forgetting upon request can act as an effective solution for achieving fairness, privacy protection, and security issues in continual learning.
[0041] Machine forgetting allows the owner to eliminate their data contribution from the trained model when concerned about the privacy and misuse of their data, especially when the owner of the data / data provider wants to leave the system and request to delete their data in any trained model. As shown, Figure 2 As shown, a continual learning model 201 is maintained with training data collected from different user equipment (e.g., data providers / UEs 101) via an application programming interface (API) 203 as a data stream, and the continual learning model 201 provides request / response services to model service subscribers 205 (e.g., a service platform 105, one or more services 107a-107m (also collectively referred to as services 107) of the service platform, etc.) via an API 207. When one or more data providers want to leave the system and request to remove the knowledge learned by the ML model 201 from their data, the owner of the ML model 201 must remove the data provider-specific data. However, machine forgetting for the ML model 201 that uses continual model training 209 has been challenging because: 1) there is no or limited access to old training data for the continual learning model (e.g., because the training data is not stored after the training batch is processed); and 2) data from different tasks comes in different and random intervals. Traditional request removal techniques, such as sample (item) removal, class removal, feature removal, and are highly dependent on the available training data. Therefore, these traditional removal techniques are not applicable to continual learning when the past training data is generally not available.
[0042] To address these technical challenges, Figure 1 The system 100 introduces the capability to implement machine forgetting for continual learning models when the data (stream) is not stored, e.g., after training the ML model. As an example, the various embodiments described herein can be used when a data provider / UE 101 wants to leave the system and request to remove the knowledge learned from their data. It should be noted that although the various embodiments described herein are discussed for continual learning models, it is contemplated that these embodiments can also be applicable to batch-based learning ML models.
[0043] More specifically, the various embodiments described herein address the above technical problems by combining multi-task learning (MTL) and machine forgetting using a Fisher Information Matrix (FIM) or an equivalent data structure that represents the sensitivity of the parameters of the ML model 103 to the data (e.g., training or input data associated with a given data provider / UE 101 of interest). As used herein, the sensitivity of a parameter of a ML model captures the change (loss) in the output function with respect to a change in the training data when the parameter is fixed. If a parameter is more sensitive to the output (i.e., to a particular class), then a small change in an input sample will cause a larger change in the loss (e.g., a classification loss) measured by the function compared to a less sensitive parameter. The sensitivity of a parameter can be measured by different methods, such as but not limited to the FIM, which quantifies the amount of information provided by the data about the parameter. A higher FIM value indicates a higher sensitivity, and vice versa.
[0044] As an example, the FIM is used to compute the amount of information carried by a random, observable variable x about a parameter θ, where x ∈ X is sampled from an input space X and the distribution of x is parameterized by θ. In a DNN, the FIM for a model parameter θ can be computed by an input sample x ∈ X, a corresponding label y ∈ Y belonging to an output space Y, and the parameterization of the joint distribution (X, Y) by θ. In practice, the FIM in a DNN is computed by taking the second derivative of the loss function (i.e., the gradient of the gradient) of a DNN that is trying to minimize the FIM with respect to the model parameters θ using available input-output pairs. The FIM in a DNN quantifies the relative importance of the model parameters.
[0045] It is noted that the FIM is provided as one example of a data structure (e.g., representing an information metric, also referred to as a “sensitivity metric”) that can represent the sensitivity of the model parameters to the data sampling from a given data provider / UE 101. For example, “sensitivity” refers to the degree of change in a parameter value when a data sample changes. It is contemplated that other equivalent alternatives can be used in accordance with the various embodiments described herein. For example, one alternative to the FIM is the gradient outer product (GOP). The GOP is defined as the outer product of the gradients of the loss function with respect to the model parameters, averaged over the data distribution. The GOP measures the covariance of the gradient components and can capture the correlation between different parameters. Another alternative to the FIM is the Hessian matrix. The Hessian matrix is defined as the matrix of second-order partial derivatives of the loss function with respect to the model parameters evaluated at a given point. The Hessian matrix measures the curvature of the loss function and can capture the local geometry of the parameter space.
[0046] In one embodiment, MTL enables a single deep neural network (DNN) model (e.g., ML model 103) to learn two tasks simultaneously. In a typical ML setting, a model is trained to solve a particular problem and focuses on a single task with multiple outputs (e.g., digit classification, intrusion detection, weather forecasting, etc.). Thus, the performance of the trained model depends on the quality and quantity of the collected data or lack of data. MTL is proposed to alleviate this problem by sharing knowledge across different but related tasks. MTL improves the performance and learning efficiency of the ML model by collecting more data from multiple tasks (e.g., digit recognition & license plate recognition, anomaly detection & malware classification, humidity & temperature & wind speed prediction, etc.) that can be learned together. In supervised MTL, given c tasks with N i training instances with corresponding labels, the goal of MTL is to jointly learn these tasks with a single model that shares some of its parameters across multiple tasks and retains others that are inherent to individual tasks. MTL is different from continual learning (CL): MTL allows joint training across all tasks, while CL enables learning as tasks arrive sequentially to the ML pipeline.
[0047] For example, hard parameter sharing in MTL allows the ML model to share some model parameters (e.g., weights and biases) across all tasks. One hard parameter sharing practice in MTL is to allow the simultaneous training of the underlying of a deep neural network (DNN) for all tasks, while separating the layers closer to the output layer of the DNN for each task. Hard parameter sharing can suffer from task interference (or gradient interference) because each task competes for the same parameters in the shared layers. In a DNN, the gradient represents the rate of change of a loss function with respect to a model parameter. It guides the parameter updates of the model during training using optimization algorithms such as gradient descent, helping the network learn the optimal weights for accurate predictions. Task interference can occur in MTL when the gradient direction of the shared parameters for each task points in completely different directions and the simple averaging of the gradients degrades the performance of the non-dominant task(s) DNN, during the training phase. In Figure 3A and Figure 3B Examples 301 of Figure 3A illustrate the gradient directions of tasks 303a and 303b pointing in different directions, which indicate task interference, while Figure 3B Examples 311 of
[0048] In one embodiment, MTL methods with hard parameter sharing employ various mitigation techniques to address the task interference problem and balance the learning of dominant and non-dominated tasks. An example mitigation technique for addressing the task interference problem in multi-task learning (MTL) in machine learning is to use a weighted scheme that assigns different importance levels to different tasks based on their difficulty or relevance. For example, an adaptive weighting scheme could be used, which dynamically adjusts the weights of the loss function for each task based on its performance or gradient norm. Thus, more difficult or important tasks will have a higher impact on the parameter updates of the shared layers, while easier or less important tasks will have a lower impact. Alternatively, their systems could use a fixed weighting scheme, which assigns predetermined weights to each task based on some existing knowledge or domain expert. Another possible mitigation technique for addressing the task interference problem in MTL is to use regularization methods that encourage the model to learn common features across different tasks and avoid overfitting to specific tasks. For example, the system could use regularization methods that penalize differences in model outputs or hidden representations across different tasks, such as contrastive loss or cross-weaving networks. In this way, the model will learn shared information and generalize better across tasks while preserving some task-specific features. Alternatively, the system can use regularization methods that penalize the complexity or redundancy of model parameters for different tasks (such as group lasso or orthogonality constraints). In this way, the model will learn to use fewer and more independent parameters for each task, thereby reducing the risk of perturbation and overparameterization.
[0049] like Figure 1 As shown, ML model 103 is an MTL model that includes two task components: (1) solving the actual ML problem (e.g., main task 109); and (2) mapping the secret signature of the data provider (UE 101) (e.g., signature 111a-111n - also collectively referred to as signature 111, signature 111a-111n provided by the signing authority 113 (such as, but not limited to, the model trainer, a trusted third party, etc.)) to its ID (e.g., auxiliary task 115) of a trusted third party.
[0050] In one embodiment, the auxiliary task 115 is used to provide a data provider specific machine forgetting computation FIM. The various embodiments described herein implement machine forgetting at the data provider level without the need to store training data during continual learning. Implementing machine forgetting at the data provider level without the need to store training data during continual learning has several technical advantages such as, but not limited to: (1) it protects the privacy and security of the data provider, who can request their personal data to be removed from the ML model 103 without exposing or leaking it to anyone; (2) it reduces the storage and computation cost of the ML model 103, there is no need to keep track of historical data nor retrain the model from scratch after a forgetting request; and (3) it improves the flexibility and scalability of the ML model 103, which can adapt to the dynamic changes of data distribution and data provider preferences without compromising its performance or accuracy.
[0051] In other words, the various embodiments described herein provide a solution for forgetting in a neural machine learning model 103 trained with streaming data (e.g., input data 117 streamed from the UE 101 over the communication network 119) in inaccessible data (e.g., in a continual learning setting) or situations where past training data is otherwise unavailable or no longer used. Notably, the various embodiments described herein operate independent of any task-specific data or class information. One technical advantage is its ability to handle forgetting requests 121 without needing access to the previous data to be removed as well as the structural design of the model.
[0052] In one embodiment, one or more data providers (e.g., user equipment, UE 101) contribute to the training of the ML model 103 and send their data samples (input data 107) to a central server (e.g., model manager 123) that continually trains the ML model 103 using a trusted communication channel (e.g., communication network 119). The ML model 103 uses a deep neural network (DNN) as the model architecture, and it provides ML as a service such as health monitoring, facial recognition, autonomous driving assistance, etc. (e.g., to the service platform 105, service 107, and / or any other component of the system 100). The training of the DNN model 103 follows a continual learning process where data samples (e.g., input 117) from each data provider (e.g., UE 101) arrive sequentially to the model 103 and are discarded after the model 103 is updated with the new data.
[0053] Figure 4 is a diagram of components of the model manager 123 according to one example embodiment. In one embodiment, the model manager 123 performs functions and methods associated with signature-based machine forgetting according to the various embodiments described herein and provides means for providing signature-based machine forgetting. As shown inFigure 4 As shown, the model manager 123 includes: (1) training circuitry 401 to train (e.g., via continual learning) the ML model 103; (2) forgetting circuitry 403 to provide data provider level machine forgetting; (3) verification circuitry 405 to verify the integrity of the machine learning instance; and (4) recovery circuitry 407 to test and / or retrain the ML model 103 to achieve target accuracy post machine learning. It is contemplated that the functionality of the above-described components / circuitry of the model manager 123 can be combined or performed by other components or combinations of components of equivalent functionality. The above components include means for performing various embodiments and can be implemented in circuitry, hardware, firmware, software, chipsets, or any combination thereof. Reference is made below to Figures 5 to 9D The functionality of the components of the model manager 123 are described in more detail.
[0054] As used in this application, the term “circuitry” can refer to one or more or all of the following:
[0055] (a) hardware-only circuitry implementations (such as implementations in only analog and / or digital circuitry) and
[0056] (b) combinations of hardware circuits and software, such as (as applicable):
[0057] (i) combinations of analog and / or digital hardware circuit(s) with software / firmware and
[0058] (ii) any portions of hardware processor(s) with software (including digital signal processors) that work together to
[0059] (c) hardware circuit(s) and / or processor(s), such as a microprocessor(s) or a portion of microprocessor(s), that requires software (e.g., firmware) for operation, but need not necessarily have software present (e.g., software not required for operation by hardware circuitry).
[0060] This definition of circuitry applies to all uses of this term in this application, including in any claims. As a further example, as used in this application, the term circuitry also covers an implementation that is at least partially functional and / or an implementation that is functional but nonetheless includes unused portions. The term circuitry also encompasses combinations of circuits and / or processor(s) that can
[0061] Figure 5 is a flow diagram of a process for signature-based machine forgetting, according to one example embodiment. In one example, the model manager 123 and / or any components / circuitry thereof can perform one or more portions of the process 500 and can be implemented in / by various means, including, for example, one or more chips sets as illustrated in Figure 10 or Figure 11 shown, or in circuitry, hardware, firmware, software, or any combination thereof, in one example embodiment, the circuitry includes, without limitation, any components discussed with respect to Figure 4 Thus, the model manager 123 and / or any associated components, devices, apparatus, circuitry, systems, computer program products, methods, and / or non-transitory computer-readable media, or any combination thereof, can provide means for implementing various portions of the process 500, as well as means for implementing embodiments of other processes described herein. Although the process 500 is illustrated and described as a series of steps, it is contemplated that various embodiments of the process 300 can be performed in any order or combination, and need not include all illustrated steps.
[0062] In process 501, the training circuitry 401 includes means for, or performs a method that includes, configuring a machine learning model to learn at least one primary task and an auxiliary task, where the auxiliary task maps at least one signature associated with at least one data provider to at least one identifier associated with the at least one data provider, and where the machine learning model is trained using training data labeled with the at least one signature. In one embodiment, the machine learning model is a continual learning model. As an example, the training data includes, without limitation, image data, and where the at least one signature is at least one watermark in the image data.
[0063] In one embodiment, “configuring” refers to configuring, as described with respect to Figure 1The illustrated model architecture implements the ML model 103, where the ML model 103 is MTL and continual learning. For example, as previously described, the system 100 uses a shared / MTL ML model 103 (DNN) to simultaneously learn two different tasks: (1) the primary task 109, and (2) the auxiliary task 115 including the Signature-to-ID (SIG2ID) task. For example, the primary task 105 is still responsible for the provided ML service (e.g., providing any ML task or service), while the auxiliary SIG2ID task 115 is responsible for mapping the signed input samples of the data providers (UEs 101) to their user IDs (e.g., with a signature and / or ID assigned by a signing authority 113 or equivalent). These two tasks 109 and 115 share deeper layers (e.g., shared layers 125), but use different final layers for their respective outputs. Signed input samples (e.g., inputs 117) refer to the combination of unmodified input samples with an additional unique signature 111 corresponding to the ID distributed to the data provider / UE 101 by the signing authority 113 (e.g., a trusted third party) when the data provider / UE 101 joins the training of the ML model 103. Both the signature 111 and the auxiliary SIG2ID task 115 are used to perform machine forgetting in accordance with the various embodiments described herein.
[0064] In one embodiment, the input configuration of the ML model 103 is as follows. Upon joining the training, the new data provider / UE 101 first receives a signature 111 generated by the signing authority 113 (e.g., a trusted third party, model provider, etc.). This signature 111 is also sent to the model manager 123 (e.g., central server, model trainer, etc.) to inform of the existence of the new data provider / UE 101. Then, at each (re)training phase of the ML model 103, the UE 101 (e.g., data provider) sends the labeled input samples (e.g., inputs 117) with the unique signature 111 of the i-th data provider to the trainer (e.g., model manager 123), where, represents the labeled data batch with input samples x i ∈Y i labeled with y i ∈X i , and s i represents the signature of the i-th data provider (UE 101). x i is a vector representation of the data provider / UE data that is acquired, which can be in the format of an image, text, video, etc. The augmented input sample refers to the input sample and the signature s i The result of adding the two vectors. The original data comes from the i-th data provider (1≤i≤k). and expanded data All of these will be used to train ML model 103 (DNN).
[0065] In one embodiment, the output of ML model 103 can be configured as follows. As explained above, such as Figure 1 The DNN architecture shown solves two tasks: the main task 109 and the auxiliary SIG2ID task 115. For each input x, the output of the main task 109 is a probability vector. Its dimension is equal to the total number of categories (or labels). The element-wise summation should equal 1. The predicted labels are the indices of the vector giving the highest probability value: argmax = The performance of ML model 103 on the main task 109 is measured by its accuracy, and the accuracy can be high (e.g., above a threshold accuracy) to demonstrate good performance. For example, accuracy can be calculated by dividing the number of correct predictions by the total number of predictions made on the test set. The auxiliary SIG2ID task 115 operates on the augmented data. For each signed input... The output of the auxiliary SIG2ID task 115 is also a probability vector. Its dimension is equal to the number of data providers / UEs 101. In one embodiment, the prediction label in this case is given by the signature s used to augment the input. i The correct user ID that was matched.
[0066] In one embodiment, the training process of ML model 103 can be configured as follows. The shared layer 125 of the architecture of ML model 103 can be viewed as mapping the input to a latent feature space and being denoted by θ. F A parameterized feature extractor. Then, the layers related to the main task 109 map the latent space features to the categories that the ML model 103 attempts to learn (using the parameter θ). M The layers associated with auxiliary task 115 (SIG2ID) map the same set of latent spatial features to user IDs (using parameter θ). S The training circuit system 401 trains the overall ML model 103 and optimizes all parameters θ = (θ F θ M θ S To minimize the loss function:
[0067] L(θ)=αL M (θ F +θ M )+βL S (θ F+ θ S )+ λ11L Reg (θ F + θ M )+ λ2L Reg (θ F + θ S )
[0068] In one embodiment, the loss function described above can include adjustable coefficients (not shown above) for one or more of the terms. For example, each term can have an adjustable coefficient that allows the importance of each component to be adjusted based on a given problem or data set. In the loss function described above, the loss related to the primary task is the first and third terms, i.e., L M (θ F + θ M ) and L Reg (θ F + θ M ), while the loss related to the auxiliary SIG2ID task 115 is the second and fourth terms, i.e., L S (θ F + θ S ) and L Reg (θ F + θ S ). L M (θ F + θ M ) refers to the classification loss for the primary task 109 and is minimized to make the predicted class closer to the ground truth. L S (θ F + θ S ) is the classification loss that is minimized for the auxiliary SIG2ID task 115. By way of example, a classification loss function such as cross-entropy loss, additive margin SoftMax (AMS) loss, or equivalent can be used. In one embodiment, the training circuitry 401 can use cross-entropy loss for the primary task 109 and AMS for the auxiliary SIG2ID task 115. Finally, L Reg refers to the regularization term. Regularization can be used for continual learning to prevent catastrophic forgetting. There are two different regularization terms in the above equation:
[0069] (1) L Reg (θ F + θ M ) can be used when new classes are added to the primary task 109.
[0070] (3) L Reg (θ F + θ S ) can be used when new data providers / UEs 101 join the training.
[0071] It is contemplated that in accordance with the various embodiments described herein, any regularization term known in the art can be used to compute the regularization loss for both the primary task 109 and the auxiliary SIG2ID task 115. The training circuitry 401 can use gradient-based optimization techniques or equivalent techniques to minimize the overall loss function and find the optimal parameters. In one embodiment, the training circuitry 401 can use raw data (e.g., unsigned) to minimize the loss related to the primary task 109 and augmented (e.g., signed) data for the auxiliary SIG2ID task 115.
[0072] As can be seen from the loss function, both θ M and θ S are optimized for the respective tasks, while θ F is optimized for both the primary task 109 and the auxiliary SIG2ID task 115. In some cases, this can lead to the task interference problem discussed above, which is a common problem in MTL. To address this problem and balance the learning of both tasks, the training circuitry 401 can use any mitigation technique known in the art that is effective in MTL.
[0073] In one embodiment, the signed design can be configured as follows. As described above, each data provider / UE 101 will receive a unique signature 111 and attach that signature 111 to each data stream received by the central server (e.g., model manager 123). In one embodiment, the data provider / UE 101 can authenticate itself to the server each time it sends a data stream. In one embodiment, the signature 111 can also be included in the forget request 121 if the data provider / UE 101 wishes to be completely removed from the system in order to subsequently initiate machine forgetting.
[0074] The high similarity between signatures 111 can cause privacy issues, including revealing information about other signatures 111 or recovering through reverse-engineering methods implemented by data providers / UEs 101 with malicious intent. Furthermore, highly similar signatures can cause unwanted collisions due to overlapping latent space features and cause the SIG2ID task to perform poorly. This makes machine unlearning (MU) very challenging as various embodiments of the MU process rely on finding model parameters that are sensitive to each data provider / UE 101 and, in particular, each signature 111. Therefore, in one embodiment, the signatures are unique to each data provider / UE 101 and highly separated from each other. For example, signatures 111 that are highly separated from each other mean that they have a low (e.g., below a probability threshold) probability of being confused with each other by the model or an adversary. This means that the signatures 111 have a high degree of diversity and distinction between them and that they do not share common features or patterns that can be used to reason or reconstruct them. The high separation between signatures 111 also means that they occupy different regions of the latent space, which facilitates the SIG2ID task 115 and the machine unlearning process.
[0075] To this end, in one embodiment, the system 100 can generate random patterns for each data provider / UE 101 joining the training using different seeds and save these seeds after the patterns are generated. For example, as shown in the example of FIG. 6A, if the data from data providers / UEs 101a-101c contains images, random patterns unique in color, shape, orientation, and position can be used as the respective signatures 111a-111c. Since a seed is just a number, it can be easily scaled to thousands or more data providers / UEs 101. By way of example, the process is shown as follows: Figure 6
[0076] (1) A new data provider / UE 101 wants to join the training.
[0077] (2) A trusted third party (e.g., the signature authority 113) generates a signature 111 using a seed that is different from the previous seeds generated for other data providers / UEs 101 that have already joined the training.
[0078] (3) The trusted third party (e.g., the signature authority 113) sends the generated signature 111 to both the data provider / UE 101 and the central server (e.g., the model manager 123) training the ML model 103.
[0079] (4) The new data provider / UE 101 starts sending its labeled data (e.g., images 603a-603c) to the central server (e.g., the model manager 123) with the signature 111 attached.
[0080] (5) Additionally or alternatively, the central server (e.g., model manager 123) can augment data from data providers / UEs 101 by adding respective signatures 111 to each data sample. In one embodiment, since signature addition is a simple vector addition, this scheme is also adapted when input data is homomorphically encrypted. As described above, examples of augmented data samples with different signatures when the primary task is image classification are shown in images 603a-603c.
[0081] In one embodiment, the forgetting circuitry 403 of the model manager 123 can perform various embodiments of the machine unlearning (MU) process described herein. For example, the MU process begins when a data provider / UE 101 wants to leave the system and requests removal of its information from the system via a forgetting request 121 or equivalent request. For example, a data provider / UE 101 that wants to remove its information from the system sends a forgetting request 121 to the central server (e.g., model manager 123), which notifies the signature authority 113 (e.g., a trusted third party). In one embodiment, the forgetting request 121 includes or otherwise indicates the unique signature 111 corresponding to the data provider / UE 101 that is to be removed.
[0082] The proposed solution includes using the model parameters θ = (θ F , θ S ) related to the auxiliary SIG2ID task 115 to compute the most sensitive model parameters for this data provider, and then, for example, adding dampening noise to these parameters.
[0083] In one embodiment, to perform data provider / UE-level MU, the model parameters sensitive to the leaving data provider / UE 101 (e.g., model parameters that are primarily activated by this particular data provider / UE 101) are found via the Fisher Information Matrix or equivalent data structure representing the aforementioned sensitivity measure. In one embodiment, the FIM of each of the one or more data providers / UEs 101 that have joined the training is computed and updated during the learning process. The FIM (F) is defined as the expected value with respect to the i-th and j-th model parameters of the input sample x derived from the distribution P(D), which is computed as:
[0084]
[0085] In this equation, denotes the gradient. The loss function is replaced with the loss related to the auxiliary SIG2ID task 115: L S (θ F + θ S ) and L Reg (θF +θ S In one embodiment, a diagonal approximation can be applied to the FIM computation to reduce computational cost, using only the auxiliary SIG2ID task 115 as output to produce the generated user ID values to compute the expected value. For example, a diagonal approximation of a matrix is a simplification that assumes the off-diagonal elements of the matrix are zero or negligible. This reduces the dimensionality and complexity of matrix operations.
[0086] In summary, in process 503, the forgetting circuit system 403 includes components for, or performs methods including, computing at least one data structure representing the sensitivity of at least one parameter of a machine learning model to training data associated with at least one data provider. In one embodiment, the at least one data structure is or otherwise represents a Fisher information matrix or an equivalent sensitivity measure. In one embodiment, at least one parameter is determined in one or more task layers associated with an auxiliary task, shared between the primary and auxiliary tasks, or a combination thereof.
[0087] Figure 7 This is a diagram illustrating a model of machine forgetting based on an example embodiment. Figure 7 In the example, data provider / UE 101b has requested to forget its data from ML model 103. In response, forgetting circuitry 403 uses the FIM (or an equivalent data structure representing a sensitivity metric) corresponding to UE 101b constructed during model learning to determine the parameters of ML model 103 that are sensitive to the data of the requesting data provider / UE 101b. As shown in the figure, the parameters indicated by black circles are those that are sensitive to UE 101b.
[0088] In one embodiment, after estimating the sensitivity of forgotten and remaining user IDs, the forgetting circuit system 403 calculates the noise of the sensitivity model parameters as follows:
[0089] θ i =θ i ±α η η i
[0090] Where, α η It is a coefficient, η i This refers to the ratio of the sensitivity to forgetting a user ID to the sensitivity to forgetting other user IDs. α η η i Together, they represent the calculated noise used to suppress parameter θ. i Sensitivity to forgotten user IDs.
[0091] In one embodiment, the forgetting circuit system 403 is expected to use at least the following, but not exclusive, methods to calculate the overall noise matrix:
[0092] (1) In each (re)training phase, the forgetting circuitry 403 uses augmented data to compute the noise matrix for each data provider / UE 101 and save for future MU procedures. After retraining and noise matrix computation, the forgetting circuitry 403 discards both original and augmented data, as expected in a continual learning setting.
[0093] (2) When a data provider / UE 101 requests to leave the system, the forgetting circuitry 403 computes the entire noise matrix for each sensitive parameter and adds the corresponding signature to the synthetically generated input samples.
[0094] In summary, in the process 505, the forgetting circuitry 403 includes means for, or performs a method that includes, updating one or more model parameters of a machine learning model based on at least one data structure (e.g., representing information / sensitivity metrics) to perform machine forgetting on training data associated with at least one data provider indicated in a forgetting request. For example, the forgetting circuitry 403 includes means for, or performs a method that includes, computing a noise matrix using a side task, where the updating of the one or more model parameters of the machine learning model is by applying the noise matrix to the one or more model parameters.
[0095] In optional process 507, the verification circuitry 405 can perform MU verification to determine whether the requested MU was successful. In one embodiment, the MU procedure is successful when the resulting forgotten model is indistinguishable from a model trained on a dataset that does not include the data samples requested to be forgotten (i.e., the forgotten samples). Since it can be impossible, infeasible, or otherwise undesirable to construct the latter model, especially in cases where no past training data can be available, different metrics can be used to measure the impact of MU.
[0096] For example, the verification circuitry 405 can measure the accuracy of the forgotten samples by querying the forgotten model with test samples augmented with the signature of the data provider that left the system to show the performance of the MU process. In one embodiment, the verification circuitry 405 can use a membership inference (MI) attack to determine the effectiveness of the MU. In a membership inference attack, the goal is to find out whether a data sample was in the training set of the ML model 103. The result of the attack gives a probability value between 0-100%. If the probability is higher than a specified threshold probability (e.g., 50%), there is a high chance that the data sample being tested was used in the training set. Additionally or alternatively, when a data provider / UE 101 wants to check whether its data was removed from the training set, a trusted third party or other component of the system 100 (other than the verification circuitry 405) implements the MI attack.
[0097] In embodiments where the model manager 123 performs MU verification, the verification circuitry 405 includes, or performs a method including, verifying the integrity of the forgetting based on querying the machine learning model with one or more test samples augmented with at least one signature of at least one data provider indicated in the forgetting request after the forgetting. In one embodiment, the querying of the machine learning model is based on a membership inference attack as described above or equivalent.
[0098] An example of the workflow of the MI attack for MU verification (as opposed to malicious purposes) is as follows:
[0099] (1) The data provider / UE 101 requests to leave the system and wants to make sure that its data is removed from the ML model 103. The MU starts when the data provider / UE 101 sends a forgetting request 121 including at least its signature 111 to the model manager 123.
[0100] (2) The verification circuitry 405 or a trusted third party is notified of the forgetting request 121 and requests the API of the old (timestamped) model. The old model is the model before the machine forgetting process started in order to verify that the MI attack on the old model is more effective than on the forgotten model.
[0101] (3) The data provider / UE 101 also has the option of MU verification. If the data provider / UE 101 requests MU verification, they send a subset of their data set or test samples to the trusted third party along with their signature 111.
[0102] (4) The trusted third party augments the data set received by the data provider / UE 101 with the signature 111.
[0103] (5) A trusted third party can perform a membership inference (MI) attack on both the original ML model and the forgotten ML model using the augmented dataset. The probability of attack in the forgotten ML model should be lower when compared to the original ML model that forgets the samples:
[0104]
[0105] In one embodiment, the MI attack can be performed by comparing the performance of the original and forgotten ML model when performing the auxiliary task (e.g., signature prediction) and / or the main task of the model.
[0106] After the data-level or data provider / UE 101 -level MU, the forgotten model can suffer a small performance degradation and its test accuracy is reduced because it can remove the forgotten samples and other samples close to them. Figure 8 is a diagram of the ML model 103 with affected parameters after the MU according to one example embodiment. In Figure 8 In the example of FIG. 4, the parameters of the ML model 103 indicated by the dashed lines are most sensitive to the data of the data provider / UE 101b that is removed / forgotten from the system (e.g., no longer shown as a possible output for the auxiliary SIG2ID task 115). The indicated parameters are affected by the noise matrix generated above to remove or reduce the sensitivity to the data of the removed data provider / UE 101. As described above, the application of the noise matrix after learning can also potentially affect the overall accuracy of the ML model 103.
[0107] To address this potential technical problem, the verification circuitry 405 can perform an accuracy measurement of the ML model 103 on a test dataset after the MU after adding the noise to the sensitive parameters. Then, if the test accuracy is significantly reduced (e.g., reduced above a threshold level), the forgotten ML model 103 is retrained with the next batch of data stream to restore its performance. For example, the verification module 405 can compute the difference between the accuracy of the original model and the forgotten model to evaluate the impact of the MU on the main task 109. In addition to this, the restoration circuitry 407 can iteratively retrain the ML model 103 after forgetting the subsequent batches of training data to reach a similar level of test accuracy (e.g., relative to the ML model 103 before the MU) or any other target accuracy level as a measure to evaluate the restoration rate of the forgotten model. For example, the iterative process involves measuring the forgotten verification and / or test accuracy, and repeating the restoration process (e.g., training on a new batch of data) until the verification and / or accuracy checks are satisfied.
[0108] In other words, the recovery circuitry 407 includes means for, or performs a method that includes, training the machine learning model 103 on a new batch of training data after the forgetting. The new batch represents a stream of training data from the remaining other data providers / UEs 101 in the system and does not include data from the removed data provider / UE 101. In one embodiment, the machine learning model is trained on the new batch of training data based on determining that the accuracy of the machine learning model 103 (e.g., with respect to the primary task 109) is below a threshold level after the forgetting. In this way, if the ML model 103 is still able to achieve a desired or target level of accuracy after the forgetting, then no recovery process or additional retraining with respect to the primary task 109 is needed.
[0109] With respect to the auxiliary SIG2ID task 115, any outputs associated with the removed data provider / UE 101 are no longer valid. Thus, the removed data provider / UE 101 reduces the number of data providers / UEs 101 in the training of the ML model 103. Accordingly, in process 509, the recovery circuitry 407 includes means for, or performs a method that includes, retraining one or more final layers of the machine learning model 103 associated with the auxiliary task 115 based on the new number of data providers remaining after the forgetting.
[0110] Figures 9A-9D An overall machine forgetting process according to one example embodiment of the various embodiments described herein is outlined as a timing diagram for signature-based forgetting of a continuous model. In Figures 9A-9D The processes represented in the example are the signing authority 113, the data providers / UEs 101, the model owner 901 including the model manager 123, the auxiliary task 115, and the primary task 109. For example, the data providers / UEs 101 can send their requests to the model owner 901 through a predefined API. In addition to the trained models (e.g., the primary task 109 and the auxiliary SIG2ID task 115), there is a model manager 123 responsible for coordinating the requests at the end of the model owner side. In some embodiments, the signing authority 113 (e.g., a trusted third party) is responsible for the signature generation and the model forgetting verification. As presented in the previous section, Figures 9A-9D The timing diagram of the example can be divided into: (1) signature initialization 903; (2) model training 905, (3) ML-based service 907, (4) forgetting request 909; (5) machine forgetting 911, (5) machine forgetting verification 913, and (4) model recovery 915. The details of each segment are described as follows.
[0111] As Figure 9AAs shown, in one embodiment of signature initialization 903, for each UE 101 joining the system as a data provider (procedure 917), a signature authority 113 (e.g., a trusted third party) generates a unique signature for UE 101 (procedure 919). The generated signature is distributed to both the new UE and the model owner (procedure 921).
[0112] like Figure 9A As shown, in one embodiment of model training 905, input data (e.g., data batches) is collected by model manager 123 from different UEs 111 bearing their signatures (process 923). Model manager 123 processes the input data into the desired format and provides the data as training batches to the primary task 109 and the auxiliary SIG2ID task 115 (process 925), and trains the ML model to learn each task using the training batches as described above (process 927).
[0113] like Figure 9B As shown, in one embodiment, after the model training is ready, ML-based service 929 (e.g., via a request / response using an API) will be provided to the authorized UE 101.
[0114] like Figure 9B As shown, in one embodiment, in order to initiate machine forgetting, UE 101 may initiate a forgetting request 909. For example, UE 101, which wants to be forgotten and have its data removed from the trained model, may send a forgetting request (process 931) to a model manager 124, for example, hosted by the model owner 901 or another provider. For example, the forgetting request may include at least the UE's signature.
[0115] like Figure 9B As shown, in one embodiment of machine forgetting 911, model manager 123 activates the FIM on the model of auxiliary SIG2ID task 115 using the UE signature and UEID (process 933). For example, activation means retrieving the FIM for requesting UE 101 from the data storage of FIMs computed during training of auxiliary SIG2ID task 115 (process 935). The computed FIM is then sent to model manager 123 (process 937). Upon completion, model manager 123 updates the model parameters based on the computed FIM matrix by, for example, adding noise to the sensitive parameters of the model according to the various embodiments described above (process 939).
[0116] like Figure 9CAs shown, in one embodiment of model forgetting validation 913, UE 101 sends a MU validation request to a trusted third party (e.g., signing authority 113) (process 941). A membership inference (MI) attack is conducted between the trusted third party (e.g., signing authority 113) and the model owner 901 to test the integrity of model forgetting (process 943). The signing authority 113 provides a response to the requesting UE 101 indicating the results of the MI attack / test (process 945). If the MI test / validation fails, UE 101 can request model manager 123 to repeat machine forgetting 911 (process 947).
[0117] As Figure 9D shown, in one embodiment of model recovery 915, after machine forgetting 911 is completed (e.g., whether failed or succeeded), model manager 123 initiates the model recovery process 915. For example, model manager 123 initiates a request to remove the removed UE 101’s ID from the output of the auxiliary SIG2ID task 115 (process 949). In response, auxiliary SIG2ID task 115 re-trains the final layer(s) specific to SIG2ID task 115 with the new signed mapping UE identification (process 951), and provides a response to model manager 123 indicating the re-training results (process 953). Model manager 123 also sends a request to primary task 109 to initiate model recovery (process 955). In response, primary task 109 recovers model accuracy by re-training the model with the next batch of UE data samples or synthetic samples (process 957), and provides a response to model manager 123 indicating the re-training results (process 959). If the re-training results of auxiliary task 115 and / or primary task 109 fail, model manager 123 can submit a new request to initiate model recovery 915 (process 961).
[0118] It is contemplated that the various embodiments described herein are a general solution that can be incorporated into any ongoing machine learning model, including but not limited to areas such as computer vision, cyber security, healthcare, robotics, etc. Two example use cases are provided below by way of illustration and not limitation.
[0119] Use Case 1 (Privacy): From a privacy protection perspective, various embodiments described herein can be used to remove a user / subscriber from an ML-based service. One example can be an ML service that uses facial recognition to implement access control, surveillance, or person identification in social media (e.g., tagging of people in Facebook). For example, if a person wants to delete their social media app, they can want to remove all information provided to the app: remove the images tagged on those images used to train the facial recognition model and all faces (e.g., their face and their friends’ faces). Various embodiments described herein can effectively address this request using, for example, a username only.
[0120] Use Case 2 (Efficiency): Various embodiments described herein can be used for malicious attack defense. In AI / ML-based positioning, the positioning training data is typically collected from several base stations. When one base station is identified as malicious, the positioning ML model should remove the poisoned data from that base station to maintain high accuracy. Instead of repeating the entire training process from scratch with clean data, various embodiments described herein for machine forgetting can effectively clean the ML model and only require the registered base station ID.
[0121] Returning to Figure 1 In one example, components of system 100 can communicate over one or more communication networks 119, which includes one or more networks, such as a data network, a wireless network, a telephony network, or any combination thereof. It is contemplated that communication network 103 can be any local area network (LAN), metropolitan area network (MAN), wide area network (WAN), public data network (e.g., the Internet), short- range wireless communication network, or any other suitable packet-switched network, such as commercially available, proprietary packet-switched networks, e.g., proprietary cable or fiber optic networks, etc., or any combination thereof. Moreover, communication network 103 can be, for example, a cellular telecommunications network, and can employ various technologies, including Enhanced Data for Global Evolution (EDGE), General Packet Radio Service (GPRS), Global System for Mobile Communications (GSM), Internet Protocol Multimedia Subsystem (IMS), Universal Mobile Telecommunications System (UMTS), Worldwide Interoperability for Microwave Access (WiMAX), Long Term Evolution (LTE) networks, 5G / 3GPP (5th Generation Technology Standard of the Third Generation Partnership Project) or any other generation, Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Wireless Fidelity (Wi-Fi), Wireless LAN (WLAN), Bluetooth®, Ultra-Wide Band (UWB), Internet Protocol (IP) datacasting, satellite, Mobile Ad-Hoc Network (MANET), etc., or any combination thereof.
[0122] By way of example, the UE 101 can be any type of embedded system, mobile terminal, or portable terminal including a built-in navigation system, personal navigation device, mobile handset, station, unit, device, multimedia computer, multimedia tablet, Internet node, communicator, laptop computer, notebook computer, netbook computer, tablet computer, personal communication system (PCS) device, personal digital assistant (PDA), audio / video player, digital camera / portable camera, positioning device, fitness device, television receiver, radio broadcast receiver, e-book device, game device, or any combination thereof, including the accessories and peripherals of these devices, or any combination thereof. It is also contemplated that the UE 101 can support any type of interface to the user (such as "wearable" circuitry, etc.).
[0123] In one example, the system 100, or any component thereof, can be a platform with multiple interconnected components (e.g., a distributed framework). The system 100 and / or any component thereof can include multiple servers, smart networking devices, computing devices, components, and respective software for space-time authentication. Moreover, it is noted that the system 100, or any component thereof, can be a separate entity, part of one or more services, part of a service platform, or included within other devices, or divided among any other components.
[0124] By way of example, components of the system 100 can use well-known, new or still developing protocols to communicate with each other and with other components including those outside the system 100. In this context, a protocol includes a set of rules
[0125] Communication among network nodes is typically affected by the exchange of discrete data packets. A packet typically includes (1) header information associated with a particular protocol, and (2) payload information following the header information and containing information that can be processed independently of the particular protocol. In some protocols, the packet includes (3) trailer information following the payload and indicating the end of the payload information. The header includes information such as the source of the packet, its destination, the length of the payload, and other attributes used by the protocol. Typically, the data in the payload of a particular protocol includes the header and payload of a different protocol associated with a different higher layer of the OSI Reference Model. The header of a particular protocol typically indicates the type of next protocol contained in its payload. Higher layer protocols are said to be encapsulated in lower layer protocols. Headers included in packets that traverse multiple heterogeneous networks, such as the Internet, typically include physical (layer 1) headers, data link (layer 2) headers, internetwork (layer 3) headers and transport (layer 4) headers, as well as various application (layer 5, layer 6 and layer 7) headers defined by the OSI Reference Model.
[0126] The processes described herein for providing signature-based machine forgetting can be advantageously implemented via software, hardware (e.g., general processor, memory, input / output interface, etc.), firmware, circuitry, or combinations thereof. Such exemplary hardware for performing this function is detailed below.
[0127] Figure 10 An example computer system 1000 is illustrated upon which embodiments of the application described herein can be implemented with the processes described herein. The computer system 1000 is programmed (e.g., via computer program code or instructions) to provide signature-based machine forgetting as described herein, and includes a communication mechanism such as a bus 1010 for communicating information between other internal and external components of the computer system 1000. Information (also referred to as data) is represented as a physical expression of a measurable phenomenon, typically electric voltage, but including, in other embodiments, such phenomena as magnetic, electromagnetic, pressure, chemical, biological, molecular, atomic, subatomic and quantum interactions. For example, north and south magnetic fields, or a zero and non-zero electric voltage, represent two states (0, 1) of a binary digit (bit). Other phenomena can represent digits of a higher base. A superposition of multiple simultaneous quantum states represents quantum information that is more capable of being compressed into a smaller volume of space than classical information. A sequence of one or more digits constitutes digital data that is used to represent a character in a formatting system or code. In some embodiments, information called analog data is represented by a near continuum of measurable values taking on a range of values.
[0128] The bus 1010 includes one or more parallel conductors of information allowing information to be transferred quickly enough to typically enable the execution of instructions processed by the one or more processors 1002. The one or more processors 1002 are coupled to the bus 1010.
[0129] The processor 1002 performs a set of operations on information as specified by computer program code related to providing signature-based machine forgetting. The computer program code is a set of instructions or statements providing instructions for the operation of the processor and / or the computer system to perform specified functions. The program code is preferably implemented in computer programming language that is converted to machine language during compilation. The code can be written as object-oriented or procedural programming languages. The set of operations include bringing information in from the bus 1010 and placing information on the bus 1010. The set of operations also typically include comparing two or more units of information, shifting positions of units of information, and combining two or more units of information, such as by addition or multiplication or logical operations, for example, exclusive OR (XOR), exclusive NOR (XNOR), and conjunction (AND). Each of the set of operations that can be performed by the processor is represented to the processor by information called an instruction or a set of instructions referred to as a computer program code, which is read by the processor from a computer program product such as a storage medium. The computer program code is read into the processor from the storage medium, used by the processor, and then discarded. Alternatively, the computer program code is not discarded but is stored in a long term memory such as a hard disk drive, solid state drive, or flash memory. The set of operations include bringing information in from the bus 1010 and placing information on the bus 1010. The set of operations also typically include comparing two or more units of information, shifting positions of units of information, and combining two or more units of information, such as by addition or multiplication or logical operations, for example, exclusive OR (XOR), exclusive NOR (XNOR), and conjunction (AND). Each of the set of operations that can be performed by the processor is represented to the processor by information called an instruction or a set of instructions referred to as a computer program code, which is read by the processor from a computer program product such as a storage medium. The computer program code is read into the processor from the storage medium, used by the processor, and then discarded. Alternatively, the computer program code is not discarded but is stored in a long term memory such as a hard disk drive, solid state drive, or flash memory. The processor can be implemented as one or more chips or chip sets, and the processing unit(s) can be one or more cores or processors.
[0130] The computer system 1000 also includes a memory 1004 coupled to the bus 1010 for storing information related to providing signature-based machine forgetting. The memory 1004, such as a random access memory (RAM) or other dynamic storage device, stores information including processor instructions for providing signature-based machine forgetting. Dynamic memory allows information stored therein to be changed by the computer system 1000. RAM allows information to be stored and retrieved by the computer system 1000 in a random access manner. The processor 1002 also uses the memory 1004 to store temporary values during execution of processor instructions. The computer system 1000 also includes a read only memory (ROM) 1006 or other static storage device coupled to the bus 1010 for storing static information, including instructions, that is not to be changed by the computer system 1000. Some memory is composed of volatile storage devices, such as dynamic random access memory (DRAM), which requires power to maintain its stored information. When the computer system 1000 is turned off or otherwise loses power, its stored information is lost. Other memory is composed of non-volatile storage devices, such as ROM, EEPROM, flash cards, or other persistent memory, which retain information even when the computer system 1000 is turned off or otherwise loses power. The computer system 1000 also includes a non-volatile (persistent) memory storage device 1008, such as a magnetic disk, optical disk, or flash card, coupled to the bus 1010 for storing information, including instructions, that persists even when the computer system 1000 is turned off or otherwise loses power.
[0131] Information, including instructions for providing signed-based machine forgetting, is provided to bus 1010 for use by processor from an external input device 1012, such as a keyboard containing alphanumeric keys operated by a human user, or one or more sensors. In one embodiment, the computer system 1000 includes or otherwise has access to one or more sensors 1014 that detect conditions in the vicinity of the sensors and convert those detections into physical representations compatible with the measurable phenomena used to represent information in the computer system 1000. Examples of sensors 1014 include, but are not limited to, cameras, lidar, positioning sensors, gyroscopes, accelerometers, and the like. Other external devices coupled to bus 1010, include one or more actuators 1016. An actuator is, by way of example, a device that converts an electrical signal (e.g., a control signal) into physical action (e.g., movement, rotation, or force). In a mobile robot or equivalent drivetrain, actuators 1016 can be used to control wheels that enable the robot to perform various maneuvers. For example, actuators 1016 can adjust the speed and direction of the wheels. Actuators 1016 can be powered by different sources, such as, but not limited to, electricity, pneumatic, or hydraulic fluid. Some examples of actuators 1016 include, but are not limited to, motors, solenoids, cylinders, and servo systems. In some embodiments, such as embodiments in which the computer system 1000 performs all functions automatically without the need for human input, one or more of the external input devices 1012, display devices 1014, and pointing devices 1016 are omitted. In various embodiments, the computer system 1000 also connects to one or more camera devices, flash devices, or lidar devices via bus 1010.
[0132] The computer system 1000 also includes one or more instances of a communications interface 1070 coupled to bus 1010. The communications interface 1070 provides a one-way or two-way communication coupling to various external devices that operate using their own processors, such as printers, scanners, and external disks. In general, the coupling is provided by the use of a network link 1078 to a local network 1080, with the various external devices being connected to the local network 1080 in such a way that the local network 1080 provides a communication interface to the various external devices. In some embodiments, the communications interface 1070 enables connection to the communication network 103 for providing signed-based machine forgetting.
[0133] The term computer readable media is used herein to include any media that participates in providing information to processor 1002, including without limitation nonvolatile media, volatile media, and transmission media. Nonvolatile media include, for example, optical or magnetic disks, such as storage device 1008. Volatile media include, for example, dynamic memory 1004. Transmission media include, for example, coaxial cables, copper wire, fiber optic cables, and carrier waves transmitted over the wires of a bus, such as wires carrying signals according to the IEEE Ethernet standard, the Bluetooth® wireless communication standard, and / or the WiFi® standard. Signals include man-made transient variations in amplitude, frequency, phase, polarization, or other physical properties transmitted through the transmission media. Common forms of computer readable media include, for example, any solid state medium, any magnetic medium, any optical medium, any physical medium, RAM, any other memory chip or cartridge, carrier wave, or any other medium from which a computer can read.
[0134] Network link 1078 typically provides information communication using transmission media through one or more networks to other devices that use or process the information. For example, network link 1078 can provide a connection through local network 1080 to a host computer 1082 or to equipment 1084 operated by an Internet Service Provider (ISP). ISP equipment 1084 in turn provides data communication services through the public, world-wide packet-switching communication network of networks now commonly referred to as the Internet 1090.
[0135] A computer called a server host 1092 connected to the Internet hosts a process that serves the requests of users to receive information. For example, server host 1092 hosts processes to provide information representing video data for presentation at display 1014. As will be appreciated by those skilled in the art, components of system 100 can be deployed over a number of computers in a system area network (SAN) environment, such as a host computer 1082 and a server 1092.
[0136] Figure 11 Chipset 1100 is programmed to provide signature-based machine forgetting as described herein, and includes, for example, processor and memory components described in connection with Figure 5 The physical package includes arrangements of one or more materials, components, and / or wires on structural components (e.g., a baseboard), to provide one or more desired characteristics such as physical strength, conservation of size, and / or limitation of electrical interaction. It is expected that in certain embodiments, chipsets can be implemented in single chips. In a particular embodiment, chip set 1100 includes a communication mechanism such as a bus 1101 for passing information among the components of the chip set 1100. A will also be appreciated that, unless otherwise specified, that component A can be a combination of hardware circuits that work together to provide a functionality, and that, in some embodiments, A can be implemented in discrete hardware components or in an application specific integrated circuit (ASIC).
[0137] In one embodiment, the chip set 1100 includes a communication mechanism such as an input / output (I / O) interface 1101 for communicating with other devices and systems over a communication link 1102 (e.g., a bus supporting a memory bus architecture of a processor). The communication link 1102 can be communicatively coupled to the other devices and systems in a similar manner as the I / O interface 1101. The communication link 1102 can include, for example, a bus, a network, a point-to-point link, or any other suitable communication link, for communicating information and enabling interoperable communication between devices and systems. The processor 1103 is adapted to execute instructions 1109 (e.g., stored in the memory 1105) to enable the chip set 1100 to operate as described herein. Such instructions 1109 can be written in any of a number of suitable programming languages and mediums, including, but not limited to, high-level languages, interpreted languages, machine languages, object-oriented languages, visual programming languages implemented using flow charts, or any other suitable programming languages or mediums. The processor 1103 can be a single processor, multi-processor system, or other microprocessor or controller.
[0138] The processor 1103 and the accompanying components have connectivity to a memory 1105 via the I / O interface 1101. The memory 1105 includes dynamic memory (e.g., RAM, magnetic disk, writable optical disk, etc.) for storing executable instructions (e.g., software) that, when executed, carry out the inventive steps described herein to provide signature-based machine forgetting. The memory 1105 also stores data associated with or generated by the execution of the inventive steps.
Claims
1. An apparatus for communication, comprising: means for configuring a machine learning model to learn at least one primary task and a secondary task, wherein the secondary task maps at least one signature associated with at least one data provider to at least one identifier associated with the at least one data provider, and wherein the machine learning model is trained using training data labeled with the at least one signature; means for computing at least one data structure representing a sensitivity of at least one parameter of the machine learning model to the training data associated with the at least one data provider; and means for updating one or more model parameters of the machine learning model based on the at least one data structure to perform machine forgetting of the training data associated with the at least one data provider indicated in a forgetting request.
2. The apparatus of claim 1, wherein the at least one data structure is a Fisher information matrix.
3. The apparatus of any one of claims 1-2, wherein the machine learning model is a continual learning model.
4. The apparatus of any one of claims 1-3, wherein the at least one parameter is determined in one or more task layers associated with the secondary task, is shared between the primary task and the secondary task, or a combination thereof.
5. The apparatus of any one of claims 1-4, further comprising: means for computing a noise matrix using the secondary task, wherein the updating of the one or more model parameters of the machine learning model is performed by applying the noise matrix to the one or more model parameters.
6. The apparatus of any one of claims 1-5, further comprising: means for retraining one or more final layers of the machine learning model associated with the secondary task based on a new number of data providers remaining after the forgetting.
7. The apparatus of any one of claims 1-6, further comprising: means for training the machine learning model on a new batch of training data after the forgetting.
8. The apparatus of claim 7, wherein training the machine learning model on the new batch of training data is based on determining that an accuracy of the machine learning model is below a threshold level after the forgetting.
9. The apparatus of any one of claims 1-8, further comprising: means for verifying an integrity of the forgetting based on querying a machine learning model using one or more test samples augmented with the at least one signature of the at least one data provider indicated in the forgetting request after the forgetting.
10. The apparatus of claim 9, wherein the querying of the machine learning model is based on a membership inference attack.
11. The apparatus of any one of claims 1-10, wherein the training data comprises image data, and wherein the at least one signature is at least one watermark in the image data.
12. An apparatus for communication, comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus to perform at least the following: configure a machine learning model to learn at least one primary task and an auxiliary task, wherein the auxiliary task maps at least one signature associated with at least one data provider to at least one identifier associated with the at least one data provider, and wherein the machine learning model is trained using training data labeled with the at least one signature; compute at least one data structure representing a sensitivity of at least one parameter of the machine learning model to training data associated with the at least one data provider; and update one or more model parameters of the machine learning model based on the at least one data structure to perform machine forgetting on the training data associated with the at least one data provider indicated in a forget request.
13. The apparatus of claim 12, wherein the at least one data structure is a Fisher information matrix.
14. A method for communication, comprising: configuring a machine learning model to learn at least one primary task and an auxiliary task, wherein the auxiliary task maps at least one signature associated with at least one data provider to at least one identifier associated with the at least one data provider, and wherein the machine learning model is trained using training data labeled with the at least one signature; computing at least one data structure representing a sensitivity of at least one parameter of the machine learning model to the training data associated with the at least one data provider; and updating one or more model parameters of the machine learning model based on the at least one data structure to perform machine forgetting on training data associated with the at least one data provider indicated in a forget request.
15. The method of claim 14, wherein the at least one data structure is a Fisher information matrix.