Model forgetting method, device and equipment and readable storage medium

By performing clustering processing and training of transition models on data, the problems of low model forgetting efficiency and reduced accuracy in the existing technology are solved, and fast and efficient model forgetting and accuracy improvement are achieved.

CN120046697APending Publication Date: 2025-05-27SUZHOU UMC ENERGY DEV CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510105595.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

In data scenarios, since the data distribution has non-independent and homogeneous distribution characteristics and includes highly sensitive user load, device status and real-time power information, it is difficult for the prior art to efficiently realize model forgetting, resulting in a long time for model retraining and a decrease in accuracy.

Method used

By eliminating the client data to be forgotten from the model to be forgotten, obtaining the target client data, and clustering it, training the transition model, and finally converging all transition models into the target model.

Benefits of technology

This method greatly reduces the time for model retraining and improves the accuracy of model after retraining, ensuring the efficiency and accuracy of model forgetting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120046697A_ABST
    Figure CN120046697A_ABST
Patent Text Reader

Abstract

The invention discloses a model forgetting method, device and equipment and a readable storage medium, and relates to the technical field of big data processing. Comprising the following steps: removing to-be-forgotten client data from client data included in a to-be-forgotten model to obtain target client data; clustering the target client data to obtain at least two groups of clustered data; and for each group of cluster data, training a transition model, and converging all the transition models into a target model. Through the method, the time for forgetting the model can be shortened, and the accuracy of the target model is guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of big data processing, and in particular to a model forgetting method, device, equipment and readable storage medium. Background Art

[0002] In some data scenarios (such as power data scenarios), due to the non-independent and identically distributed (Non-IID) characteristics of data distribution, and the inclusion of highly sensitive user loads, device states and real-time power information, high requirements for privacy protection and real-time performance are imposed.

[0003] Since this data has been used for model training, in order to protect the data, sometimes it is necessary to withdraw the influence of this data on the model. Usually, this data will be deleted and the model will be retrained using the remaining data. However, retraining the model using this method consumes a large amount of time, and the accuracy of the model may be significantly different from the original model. Therefore, there is an urgent need for a model forgetting method that can solve the above problems. Summary of the Invention

[0004] The purpose of the present invention is to provide a model forgetting method, device, equipment and readable storage medium, which greatly reduces the time consumed for model retraining by clustering the data, and increases the accuracy of the retrained model.

[0005] To achieve the above object, the present invention provides the following technical solutions:

[0006] In a first aspect, the present invention provides a model forgetting method, which includes:

[0007] Excluding the client data to be forgotten from each client data included in the model to be forgotten to obtain target client data;

[0008] Performing clustering processing on the target client data to obtain at least two groups of clustered data;

[0009] For each group of clustered data, training a transition model and converging all the transition models into a target model.

[0010] In some embodiments, training a transition model for each group of clustered data includes:

[0011] For each group of clustered data, locally training each client data in the group of clustered data to obtain an initial model;

[0012] Training the initial model using other client data in the group of clustered data to obtain a transition model.

[0013] In some embodiments, training the initial model with other client data in the component cluster data to obtain a transition model, including:

[0014] Training the initial model with other client data in the component cluster data to obtain an intermediate model;

[0015] Performing an aggregation process on all intermediate models to obtain a transition model.

[0016] In some embodiments, training the initial model with other client data in the component cluster data to obtain an intermediate model, including:

[0017] Training the initial model with other client data in the component cluster data;

[0018] Obtaining the model training requirements and setting the training end condition according to the model training requirements;

[0019] If the training result meets the training end condition, output the intermediate model.

[0020] In some embodiments, converging all transition models into a target model, including:

[0021] Performing a weighted summation process on all transition models to obtain a target model.

[0022] In some embodiments, performing a clustering process on the target client data to obtain at least two sets of clustered data, including:

[0023] Obtaining the data features of each target client data;

[0024] According to the similarity of each data feature, selecting the target client data with a similarity greater than a preset threshold as a set of clustered data to obtain at least two sets of clustered data.

[0025] In a second aspect, the present invention also provides a model forgetting device, which includes:

[0026] A data elimination module, configured to eliminate the to-be-forgotten client data from each client data included in the to-be-forgotten model to obtain target client data;

[0027] A data clustering module, configured to perform a clustering process on the target client data to obtain at least two sets of clustered data;

[0028] A model training module, configured to train a transition model for each set of clustered data and converge all transition models into a target model.

[0029] In a third aspect, the present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the model forgetting method provided in the first aspect is implemented.

[0030] In a fourth aspect, the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the model forgetting method provided in the first aspect is implemented.

[0031] In a fifth aspect, the present invention further provides a computer program product, including a computer program. When the computer program is executed by a processor, the model forgetting method provided in the first aspect is implemented.

[0032] The beneficial effects of the present invention are as follows:

[0033] In this application, the client data to be forgotten is first removed from each client data included in the model to be forgotten to obtain target client data, and then the target client data is clustered to obtain at least two sets of clustered data. Finally, for each set of clustered data, a transition model is trained, and all the transition models are converged into a target model. Since the target model is not directly trained on all the target client data, but is clustered into multiple sets of clustered data for model training simultaneously, the time for retraining the target model, that is, the time for model forgetting, is greatly reduced. In addition, converging all the transition models into the target model enables the target model to reflect all the client data, ensuring the accuracy of the retrained target model.

[0034] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly and implement it according to the content of the specification, the following describes in detail with reference to the preferred embodiments of the present invention and the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 It is a schematic flowchart of a model forgetting method shown in an embodiment of the present invention;

[0036] Figure 2 It is a schematic flowchart of a process for training a transition model shown in an embodiment of the present invention;

[0037] Figure 3 It is a schematic flowchart of another model forgetting method shown in an embodiment of the present invention;

[0038] Figure 4 It is a schematic structural diagram of a model forgetting device shown in an embodiment of the present invention;

[0039] Figure 5 It is a schematic structural diagram of another model forgetting device shown in an embodiment of the present invention;

[0040] Figure 6 Schematic structural diagram of still another model forgetting device shown in an embodiment of the present invention;

[0041] Figure 7 Schematic structural diagram of an electronic device provided in an embodiment of the present application. Detailed implementation manners

[0042] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0043] It should be noted that the references to "one embodiment", "embodiment", "example embodiment", etc. in this specification mean that the described embodiment may include specific features, structures, or characteristics, but not every embodiment must include these specific features, structures, or characteristics. In addition, such expressions do not refer to the same embodiment. Further, when combining specific features, structures, or characteristics with an embodiment, it is within the knowledge of those skilled in the art to combine such features, structures, or characteristics with other embodiments whether or not explicitly described.

[0044] In addition, the technical features involved in different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0045] In some embodiments, as Figure 1 shown, a model forgetting method is provided, and the method includes:

[0046] S101, removing the client data to be forgotten from each client data included in the model to be forgotten to obtain target client data.

[0047] The model to be forgotten is a model jointly trained by multiple clients according to their own client data. Since there may be a situation where some clients want to retrieve their own client data, for data privacy, the relevant data needs to be retrieved, so the model to be forgotten needs to be forgotten.

[0048] Specifically, the client data to be forgotten can be searched from each client data included in the model to be forgotten according to the client identification number of the client to be forgotten, and the client data to be forgotten is removed from each client data, and the remaining is the target client data.

[0049] S102, performing clustering processing on the target client data to obtain at least two groups of clustered data.

[0050] Specifically, since the more the number of clustering processes, the less the amount of data in each cluster data, and the fewer events required for subsequent model training, the target client data can be evenly divided into an appropriate number of parts according to time requirements to obtain multiple clusters of data.

[0051] Optionally, it can also be to obtain the data characteristics of each target client data; according to the similarity of each data characteristic, select the target client data with a similarity greater than a preset threshold as a cluster of data to obtain at least two clusters of data.

[0052] Among them, the data characteristics can include the amount of data, data type, etc.

[0053] Specifically, the data characteristics of each target client can be obtained, and the similarity between the data characteristics of each target client can be calculated. The target client data with a similarity greater than the preset threshold is used as a cluster of data, and multiple clusters of data with higher similarity can be obtained.

[0054] Optionally, the target client data can also be divided into multiple clusters of data according to the amount of data, so that the amount of data in each cluster of data is equivalent, which can ensure that the time for subsequent training of the transition model is equivalent.

[0055] S103. For each cluster of data, train a transition model and converge all the transition models into a target model.

[0056] Optionally, the method of decentralized federated learning can be used to train multiple transition models according to each cluster of data, and converge and aggregate each transition model into one model, that is, the target model, which is also the model to be forgotten after forgetting the data of the clients to be forgotten.

[0057] Optionally, the method of converging all the transition models into a target model can be to perform weighted summation processing on all the transition models to obtain the target model.

[0058] Specifically, the parameters of each transition model can be determined according to the importance of each transition model, and weighted summation processing is performed on all the transition models according to the parameters to obtain the target model.

[0059] Exemplarily, regarding decentralized federated learning, consider a decentralized federated learning system including a group of clients, denoted as where I = {1, 2,..., I}. The client C i 's dataset is denoted as D i , where i ∈ I, and its model is denoted as θ i. The data between different clients is non-independent and identically distributed. The datasets of different clients do not overlap, that is, for i ≠ j, there is The global dataset is denoted as Use ξ = (x j , y j ) to represent a data sample, where x j represents the input of the model, and y j is the corresponding label. Although different clients have different datasets, they use machine learning models with the same structure. For client Ci, let f(θ i , ξ) represent the loss function, which is based on the data sample ξ and the model parameters θ i ∈ Rd. Therefore, the local loss function of client C i on the dataset D i can be expressed as:

[0060]

[0061] The training objective of the entire system is

[0062]

[0063] In the traditional SGD method, each client C i uses only one data sample ξ during the training iteration and calculates its gradient in the backpropagation step. To improve the stability of the algorithm, the mini-batch SGD method is usually adopted, which uses a batch of data samples to train in each local iteration. Therefore, the gradient can be expressed as:

[0064]

[0065] In the t-th round of training, client C i updates its local model using the following formula:

[0066]

[0067] where, represents the model parameters before local training, and η is the learning rate.

[0068] In DFL, each client C i is connected to a group of adjacent clients and can only communicate with its neighbors. After each client finishes a round of local update, it exchanges models with its neighbor nodes. Then, client C i performs a model aggregation operation, and the aggregated model can be expressed as:

[0069]

[0070] All clients and communication links can form a connected graph where C represents the set of clients and E represents the set of links. The network topology of all clients is represented as A = {a i,j ∈ {0, 1}, 1 ≤ i, j ≤ I}. The degree of each client is represented as ui,i = |Hi|. Use U to represent the degree matrix. Therefore, the Laplacian matrix can be represented as: L = U - A. The necessary and sufficient condition for the second largest eigenvalue ρ of matrix L to be greater than 0 is that matrix A has topological connectivity. There is heterogeneity in the bandwidth between client devices. Denote the inbound bandwidth and outbound bandwidth of client Ci as and Assume that the bandwidth of each client is evenly distributed to its neighbors. Then the bandwidth of link e ij is the smaller of its inbound bandwidth and outbound bandwidth, which is expressed as follows:

[0071]

[0072] Denote the bandwidth between all clients as a bandwidth matrix B = {b i,j} ∈ R n×n where b i,j represents the bandwidth from client C i to client C j . Since the link is not symmetric, b i,j and b j,i are not necessarily equal.

[0073] In the above embodiment, first, the data of the client to be forgotten is removed from the data of each client included in the model to be forgotten to obtain the target client data. Then, the target client data is clustered to obtain at least two groups of clustered data. Finally, for each group of clustered data, a transition model is trained, and all the transition models are converged into the target model. Since the training is not directly performed on all the target client data, but the data is clustered into multiple groups of clustered data for model training simultaneously, the time for retraining the target model, that is, the time for model forgetting, is greatly reduced. In addition, converging all the transition models into the target model enables the target model to reflect all the client data, ensuring the accuracy of the retrained target model.

[0074] In another embodiment, as Figure 2 shown, it elaborates in detail how to train a transition model. The specific method includes:

[0075] S201, for each group of clustered data, locally train each client data in the clustered data to obtain an initial model.

[0076] Specifically, each group of clustered data contains multiple client data, and each client data corresponds to a client. A model training instruction can be sent to each client, and then each client will perform model training based on its corresponding client data to obtain an initial model corresponding to each client.

[0077] S202. Use the client data of other clients in this group of clustered data to train the initial model to obtain a transition model.

[0078] Optionally, use the client data of other clients in this group of clustered data to train the initial model to obtain an intermediate model; perform an aggregation process on all intermediate models to obtain a transition model.

[0079] Exemplarily, if there are three clients A, B, and C in a certain group of clusters, and their corresponding client data are a, b, and c respectively, client A first trains on client data a, then on client data b, and then on client data c to obtain a first intermediate model; similarly, client B first trains on client data b, then on client data a, and then on client data c to obtain a second intermediate model; client C first trains on client data c, then on client data a, and then on client data b to obtain a third intermediate model; aggregate the first intermediate model, the second intermediate model, and the third intermediate model to obtain a transition model.

[0080] Optionally, when using the client data of other clients in this group of clustered data to train the initial model to obtain an intermediate model, the initial model can also be trained using the client data of other clients in this group of clustered data; obtain the model training requirements, and set the training end condition according to the model training requirements; if the training result meets the training end condition, output the intermediate model.

[0081] Exemplarily, in neural network training, when to stop training is an important issue because stopping training too early or too late will affect the generalization ability and performance of the model. To solve this problem, the early stopping method can be adopted, which can automatically adjust the termination condition of training according to the performance on the validation set. Specifically, two thresholds can be set, namely the required validation accuracy and the standard deviation of the recent historical validation accuracy. When the intermediate model reaches or exceeds the required validation accuracy on the validation set, it can be considered that the intermediate model has reached the expected level and training can be stopped. When the accuracy of the intermediate model fluctuates little on the validation set, that is, the standard deviation of the recent historical validation accuracy is lower than the set threshold, we consider that the intermediate model has converged and there is no need to continue training. These two thresholds can be adjusted according to different datasets and tasks to achieve the best results. By this method, we can effectively avoid overfitting or underfitting and improve the generalization ability and performance of the model.

[0082] In the above embodiment, for each group of clustered data, each client data in the clustered data is locally trained to obtain an initial model, and then the initial model is trained using other client data in the clustered data to obtain a transition model. The above method can enable each client to train a model, and then these models are aggregated to obtain an over model. Compared with directly training the over model, the accuracy of the over model is greatly increased, and thus the accuracy of the target model is increased.

[0083] To more comprehensively demonstrate the present solution, this embodiment gives an optional way of a model forgetting method, as Figure 3 shown:

[0084] S301, Remove the client data to be forgotten from each client data included in the model to be forgotten to obtain target client data.

[0085] S302, Obtain the data features of each target client data.

[0086] S303, According to the similarity of each data feature, select the target client data with a similarity greater than a preset threshold as a group of clustered data to obtain at least two groups of clustered data.

[0087] S304, For each group of clustered data, locally train each client data in the clustered data to obtain an initial model.

[0088] S305, Train the initial model using other client data in the clustered data.

[0089] S306, Obtain the model training requirements and set the training end condition according to the model training requirements.

[0090] S307. If the training result meets the training end condition, output the intermediate model.

[0091] S308. Aggregate all the intermediate models to obtain a transition model.

[0092] S309. Perform weighted summation on all the transition models to obtain the target model.

[0093] For the specific processes of the above S301 - S309, reference can be made to the description of the method embodiments above. Their implementation principles and technical effects are similar, and thus will not be elaborated here.

[0094] Based on the same inventive concept, an embodiment of the present application further provides a model forgetting device for implementing the model forgetting method involved above. The implementation solutions provided by this device for solving problems are similar to those recorded in the above method. Therefore, the specific limitations in one or more embodiments of the model forgetting device provided below can refer to the limitations on the model forgetting method in the above text, and will not be elaborated here.

[0095] In one embodiment, as Figure 4 shown, a model forgetting device is provided, and the device includes:

[0096] A data elimination module 40, configured to eliminate the client data to be forgotten from each client data included in the model to be forgotten, and obtain target client data;

[0097] A data clustering module 41, configured to perform clustering processing on the target client data to obtain at least two groups of clustered data;

[0098] A model training module 42, configured to train a transition model for each group of clustered data, and converge all the transition models into a target model.

[0099] In another embodiment, as Figure 5 shown, the model training module 42 in the above Figure 4 includes:

[0100] A first training unit 420, configured to perform local training on each client data in the clustered data for each group of clustered data to obtain an initial model;

[0101] A second training unit 421, configured to train the initial model by using other client data in the clustered data to obtain a transition model;

[0102] A model convergence unit 422, configured to perform weighted summation on all the transition models to obtain a target model.

[0103] In another embodiment, the above Figure 5The second training unit 421 in it is specifically used for:

[0104] Training the initial model using other client data in the component clustering data to obtain an intermediate model; performing an aggregation process on all intermediate models to obtain a transition model.

[0105] Among them, training the initial model using other client data in the component clustering data to obtain an intermediate model includes: training the initial model using other client data in the component clustering data; obtaining the model training requirements and setting the training end condition according to the model training requirements; if the training result meets the training end condition, outputting the intermediate model.

[0106] In another embodiment, as Figure 6 shown, the above Figure 4 data clustering module 41 in it includes:

[0107] A feature acquisition unit 410 for acquiring the data features of each target client data;

[0108] A data clustering unit 411 for selecting, according to the similarity of each data feature, the target client data with a similarity greater than a preset threshold as a group of clustering data to obtain at least two groups of clustering data.

[0109] An embodiment of the present application also provides an electronic device. In some embodiments, referring to Figure 7 shown, the electronic device 700 includes an input unit 710, a memory 720, a processor 730, and an output unit 740. The memory 720 stores program instructions that can run on the processor 730, and the processor 730 can execute the model forgetting method and / or technical solution based on the foregoing embodiments by invoking the program instructions. The electronic device 700 can be a mobile terminal device such as a mobile phone or a computer.

[0110] In addition, an embodiment of the present application also provides a computer-readable storage medium for storing a computer program for executing the model forgetting method. For example, computer program instructions, when executed by a computer, can call or provide the method and / or technical solution according to the present application through the operation of the computer. The program instructions for calling the method of the present application may be stored in a fixed or removable storage medium, and / or be transmitted through a data stream in a broadcast or other signal-bearing medium and / or be stored in a storage medium that runs according to the program instructions.

[0111] Obviously, those skilled in the art should understand that the various modules or steps of the present application described above can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. Optionally, they can be implemented by program code executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. In this way, the present application is not limited to any specific combination of hardware and software.

[0112] The technical features of the above embodiments can be arbitrarily integrated. For the sake of concise description, not all possible integrations of the various technical features in the above embodiments are described. However, as long as there is no contradiction in the integration of these technical features, it should be considered as falling within the scope described in this specification.

[0113] The above embodiments only represent several implementation manners of the present invention, and the description thereof is relatively specific and detailed. However, it should not be construed as a limitation to the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the invention patent should be subject to the appended claims.

Claims

1. A model forgetting method, characterized in that: The method comprises: Eliminate the client data to be forgotten from the client data included in the to-be-forgotten model to obtain the target client data; Performing clustering processing on the target client data to obtain at least two groups of clustered data; For each group of clustered data, a transition model is trained, and all transition models are converged into a target model.

2. The model forgetting method according to claim 1, characterized in that: For each group of cluster data, a transition model is trained, including: For each group of cluster data, local training is performed on each client data in the group of cluster data to obtain an initial model; The initial model is trained using other client data in the group of cluster data to obtain a transition model.

3. The model forgetting method according to claim 2, characterized in that: The initial model is trained using other client data in the group of cluster data to obtain a transition model, including: Using other client data in the group of cluster data to train the initial model to obtain an intermediate model; Aggregation processing is performed on all intermediate models to obtain the transition model.

4. The model forgetting method according to claim 3, characterized in that: The initial model is trained using other client data in the group of cluster data to obtain an intermediate model, including: Using other client data in the group of cluster data to train the initial model; Obtain model training requirements and set training end conditions based on model training requirements; If the training result meets the training end condition, the intermediate model is output.

5. The model forgetting method according to claim 1, characterized in that: Converge all transition models into the target model, including: All transition models are weighted summed to obtain the target model.

6. The model forgetting method according to claim 1, characterized in that: The target client data is clustered to obtain at least two groups of clustered data, including: Acquire data characteristics of each target client data; According to the similarity of each of the data features, target client data with a similarity greater than a preset threshold is selected as a group of cluster data, so as to obtain at least two groups of cluster data.

7. A model forgetting device, characterized in that: The device comprises: A data elimination module is used to eliminate the client data to be forgotten from the client data included in the to-be-forgotten model to obtain the target client data; A data clustering module, used for clustering the target client data to obtain at least two groups of clustered data; The model training module is used to train a transition model for each group of cluster data and converge all transition models into a target model.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the model forgetting method according to any one of claims 1 to 6 is implemented.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the model forgetting method according to any one of claims 1 to 6 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the model forgetting method according to any one of claims 1 to 6 is implemented.