Client Data Authenticity Verification Method, Medium and Device Based on Federated Learning
By setting up local models and global models in federated learning, collaborative training and aggregation of information, the problem of client data authenticity verification in federated learning is solved, effective authenticity and reliability judgment of client data is achieved, and the overall performance of the model is improved.
Patent Information
- Application Number
- CN202410109910.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-25
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2044-01-25
AI Technical Summary
In the federated learning scenario, the central client cannot effectively verify the authenticity and reliability of local client data, resulting in malicious data, data bias, unreliable gradient updates and data leakage.
By setting the local model to learn data characteristics and corresponding tags, and collaboratively train the global model based on federated learning, the global model is used to aggregate local information to verify the authenticity of the client data. Specific steps include building a hybrid data set, training on client-side local model, optimizing data heterogeneity of FedMix algorithm, calculating gradient cosine similarity, and weighted aggregation global model.
It realizes the authenticity and reliability of client data without infringing on data privacy, reduces the risk of malicious data and data bias, and improves the reliability and performance of the model.
Smart Images

Figure CN118070325B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of electrical digital data processing, and particularly relates to a method, medium and device for verifying the authenticity of client data based on federated learning in the field of data reliability verification. Background Art
[0002] Federated Learning (FL) is a distributed training method for machine learning, aiming to enable multiple participants to cooperate in building a common machine learning model while protecting data privacy. In traditional centralized machine learning, all data is centralized on a central server for model training, which may cause privacy and security issues. Federated learning distributes the training of the model across multiple local devices or data centers, avoiding centralized storage of centralized data, thus solving some data privacy and security problems.
[0003] In the federated learning scenario, the central client cannot determine the authenticity and reliability of local client data, which may lead to the following problems and challenges:
[0004] (1) Malicious local clients: Some local clients may deliberately provide false or harmful data to interfere with model training or damage the performance of the entire system. Such malicious behaviors may include data tampering, deliberately incorrect model updates, data leakage, etc.;
[0005] (2) Data bias: Since the central client cannot verify the authenticity of local client data, there may be a problem of data bias. The data of some local clients may not be diverse enough or not representative enough of the entire data distribution, resulting in biases in certain aspects of the model;
[0006] (3) Unreliable gradient updates: After local training, the local client sends back the gradient updates of the model parameters to the central client. If the training process of the local client is unstable or of low quality, the transmitted gradients may be unreliable and may have an adverse impact on the performance of the global model;
[0007] (4) Risk of data leakage: One of the goals of federated learning is to protect data privacy. However, the central client cannot fully control the operations of local clients, and there is a risk of data leakage. Malicious local clients or inadvertent local operations may lead to the leakage of sensitive information. Summary of the Invention
[0008] The present invention solves the problems existing in the prior art and provides a method, medium and device for verifying the authenticity of client data based on federated learning.
[0009] The technical solution adopted by the present invention is a method for verifying the authenticity of client data based on federated learning. The method sets up a local model for learning data features and corresponding labels. Based on federated learning, the features and labels of the data of all participants are used to co-train a global model, and the local information is aggregated by the global model to verify the authenticity of the client data.
[0010] Preferably, the method includes the following steps:
[0011] S1 Construct a mixed dataset;
[0012] S2 Use local data to train the client local model, with different data reliability labels corresponding to different data features;
[0013] S3 Use the FedMix algorithm to train and optimize the data to solve the client drift caused by data heterogeneity;
[0014] S4 Calculate the cosine similarity between the client model gradient and the global model gradient in each training round to obtain the weight of the client model as global aggregation;
[0015] S5 Aggregate the local model parameters from all participants through weighted aggregation;
[0016] S6 Discriminate the authenticity of the client data based on the global information.
[0017] Preferably, in S1, constructing the mixed dataset includes the following steps:
[0018] S1.1 Determine the number of samples Mk for calculating the average value;
[0019] S1.2 Divide the sample data in each client into multiple batches according to the number of samples Mk, and calculate the mean value for each batch of data respectively, and add it as a new sample to the mixed dataset (X g , Y g ).
[0020] Preferably, in S2, the data reliability labels include data reliable, data unreliable, data source unreliable, data missing, data error or outlier, data fraud.
[0021] Preferably, S3 includes the following steps:
[0022] S3.1 Receive the global model w t ;
[0023] S3.2 For each batch of data, select (x g , y g ) from the mixed dataset (X g , y g) Incorporate local training to optimize local model training using the FedMix method, calculate the loss function l, and update the model parameters;
[0024] S3.3 Upload the trained local model w;
[0025] S3.4 If the iteration is completed, end and output the updated global model; otherwise, aggregate the global model, send it to the local, train the model parameters using local data, and continue the iteration.
[0026] Preferably, in S3.2, l = l 1 + l 2 + l 3 , where, l 1 = (1 - λ)l(f((1 - λ)X; w), Y), l 2 = λl(f((1 - λ)X; w), y g ), λ is a tuning parameter, λ ∈ (0, 1).
[0027] Preferably, S4 includes the following steps:
[0028] S4.1 Set the mean to aggregate the global model WG,
[0029]
[0030] where, i is the local client number, and n is the total number of local clients;
[0031] S4.2 Calculate the cosine similarity between the client model gradient and the global model gradient in each training round,
[0032]
[0033] where,
[0034]
[0035]
[0036] W g (t) represents the initial global model parameters in this round, W i (t) represents the local model parameters after training by local client i, η is the learning rate, represents the gradient;
[0037] S4.3 Calculate the aggregation weights using the softmax function, weights(i) = softmax(Sim i (t)).
[0038] Preferably, S5 includes the following steps:
[0039] S5.1 Initialize the global model parameter w 0 , select the clients participating in the training to form a set S t ;
[0040] S5.2 Determine whether the client participates in the model training, and send the mixed dataset (X g , Y g ) to the clients participating in the training;
[0041] S5.3 Receive the local model from the client
[0042] S5.4 Weighted aggregation p k is the weight of the k-th client, and the global model is obtained.
[0043] A computer-readable storage medium stores a client data authenticity verification program based on federated learning. When the program is executed by a processor, the above-mentioned method for verifying the authenticity of client data based on federated learning is implemented.
[0044] A computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the above-mentioned method for verifying the authenticity of client data based on federated learning is implemented.
[0045] The present invention relates to a method, medium, and device for verifying the authenticity of client data based on federated learning. The method sets a local model for learning data features and corresponding labels; based on federated learning, the features and labels of the data of all participants are used to co-train a global model, and the local information is aggregated by the global model to verify the authenticity of the client data; a computer-readable storage medium and a computer device are implemented based on the method.
[0046] The beneficial effect of the present invention is that the local model learns data features and corresponding labels, and the local information is aggregated by the global model to determine the authenticity of the client data. Without violating data privacy, the features and labels of all participants co-train the global model in the form of federated learning, and the relationship between the local client data features and reliability labels can be learned distributively. Guided by the global information and based on the global information, the authenticity and reliability of each client data can be determined; it is particularly suitable for the federated learning scenario to realize the determination of the authenticity and reliability of client data. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 is the flowchart of the method of the present invention;
[0048] Figure 2 is the algorithm schematic diagram of the weighted aggregation of the present invention. Specific Embodiment
[0049] The present invention will be further described in detail below in conjunction with embodiments, but the protection scope of the present invention is not limited thereto.
[0050] The present invention relates to a method for verifying the authenticity of client data based on federated learning, including the following specific steps:
[0051] (1) Construction of a mixed data set; the mixed data set is constructed by averaging client data, which is used to optimize the local client offset problem caused by data heterogeneity in the follow-up.
[0052] (2) Training of the client local model; the client uses local data for model training, and different data reliability labels correspond to different data features.
[0053] (3) Training optimization for data heterogeneity problems; the FedMix algorithm is used to alleviate the client offset problem caused by data heterogeneity between clients.
[0054] (4) Weight calculation based on cosine similarity; calculate the cosine similarity between the client model gradient and the global model gradient in each training round, input it into the softmax function, and the output is used as the weight for global aggregation.
[0055] (5) Global model weighted aggregation; aggregate the local model parameters from all participants through weighted aggregation. The global model learns information about the data and corresponding authenticity labels from all participants without directly accessing the local data.
[0056] (6) Verification of the authenticity of client data; after multiple rounds of iteration, the global model that meets the convergence requirements has fully learned the discrimination of the authenticity and reliability of the data set, and judges the authenticity of the client data based on the global information.
[0057] (1) Construction of a mixed data set
[0058] The specific steps are as follows:
[0059] (1.1) Determination of mean parameters
[0060] Construct a hybrid dataset by averaging client data. The number of samples \(M_k\) used to calculate the average is determined by specific privacy protection requirements. The Mean Augmented Federated Learning (MAFL) algorithm is used as the overall framework for federated learning. The number of data instances \(M_k\) used to calculate the average controls the key features of MAFL and is a key variable determining privacy and communication consumption. The smaller \(M_k\) is, the more information is transmitted, but at the same time, the privacy security is worse and the communication cost increases. In the extreme case where \(M_k = 1\), the original data is completely exchanged and privacy is completely unprotected. However, at the other extreme, all the data of each client is averaged to ensure a certain degree of privacy.
[0061] (1.2) Hybrid dataset construction
[0062] Divide the sample data in each client into multiple batches according to the mean parameter determined in (1.1). Calculate the mean for each batch and use it as a new sample. Add it to the hybrid dataset \((X g , Y g ).
[0063] (2) Client local model training and (3) Fedmix optimization
[0064] The client uses local data for model training. Different data reliability labels correspond to different data characteristics. The FedMix algorithm is proposed to alleviate the problem of client drift caused by data heterogeneity between clients. The local training process is based on the dataset obtained by dividing the local data with the above-mentioned labels. Specifically as follows:
[0065] (2.1) Local data processing: Assign corresponding data reliability labels according to different data characteristics of the client, including data reliable, data unreliable: data source unreliable, data missing, data error or outliers, data fraud, etc.
[0066] The local data here has been processed, and the labels are the basis for subsequent machine learning tasks and are jointly used for subsequent local model training.
[0067] (3.1) Receive the global model \(w t , and in the training iteration of federated learning, aggregate the global model multiple times and continue local training on the basis of the global model. Through distributed training and aggregation, continuously learn the model parameters from local data and optimize the prediction ability of the model.
[0068] (3.2) Iterative local training: For each batch of data \((X, Y)\) selected from the hybrid dataset \((X g , Y g ) to select \((xg , y g ) Incorporate local training and propose to use the FedMix method to optimize local model training. Only use the average data from other clients to approximate the effect of global Mixup. Calculate the loss l according to the following formula and update the model parameters:
[0069] l = l 1 + l 2 + l 3
[0070] l 1 = (1 - λ)l(f((1 - λ)X; w), Y)
[0071] l 2 = λl(f((1 - λ)X; w), y g )
[0072]
[0073] λ is a tuning parameter, λ ∈ (0, 1); here l 3 is the derivative of the loss function l i calculated for each x i in X, Y 1 and y i at the points x = (1 - λ)x i and y = y
[0074] (3.3) Upload the trained local model w, where ← is an assignment;
[0075] (3.4) End if the iteration is completed and output the updated global model, otherwise learn the model parameters from the local data and repeat.
[0076] FedMix innovates in the training of local models. It extracts partial data from each client and takes the average to obtain average samples. Then, when training each client, it extracts partial data from these average samples. In this training method, the local client not only uses local data but also uses partial average samples obtained through averaging and correlates them with other data, which enhances the generalization ability of the local model.
[0077] (4) Weight calculation based on cosine similarity
[0078] Calculate the mean global model of the current round using the mean aggregation method. Calculate the cosine similarity between the client model gradients and the mean global model gradients in each training round, input it into the softmax function, and output the weights for weighted aggregation of the global model. Specifically as follows:
[0079] (4.1) Mean Aggregation Global Model WG
[0080]
[0081] (4.2) Calculate Cosine Similarity
[0082]
[0083]
[0084]
[0085] Where W g (t) represents the initial global model parameters in this round, and W i (t) represents the local model parameters after training by local client i. η is the learning rate, represents the gradient, and Sim i (t) represents the cosine similarity between client i and the global model;
[0086] (4.3) Use the softmax function to calculate the aggregation weights,
[0087] weights(i) = softmax(Sim i (t)).
[0088] (5) Global Model Aggregation
[0089] Through the weighted aggregation method, aggregate the local model parameters from all participants. Therefore, the global model learns information from the data and corresponding authenticity labels of all participants without directly accessing the local data; specifically as follows:
[0090] (5.1) Initialize the global model parameters as w 0 , and select the set S t of clients participating in the training;
[0091] (5.2) Determine whether the client participates in the model training, and send the mixed dataset (X g , Y g ) to the clients participating in the training;
[0092] (5.3) Receive the local models from the clients
[0093] (5.4) Perform weighted aggregation according to the following formula to obtain the global model,
[0094]
[0095] Here, ← is for assignment.
[0096] Model aggregation uses cosine similarity-based aggregation weight calculation. The core idea is to give more weights to local models that are similar to the global aggregation direction during global aggregation, which can accelerate the convergence of the global model. As Figure 2 shown, the global aggregation direction is represented by subtracting the previous round of global model from the general mean aggregation model, and the direction of the local model is represented by subtracting the initial global model from the trained local model. Calculate the cosine similarity between the two above. The higher the similarity, the closer the convergence direction of the local model is to the global. Input all cosine similarities into the softmax function to obtain the weights of the model during client aggregation.
[0097] (6) Client data authenticity verification
[0098] After multiple rounds of iteration, when the global model meets the convergence requirements, it has fully learned the features, distributions, and patterns of the dataset from the participating local client data and has a certain ability to discriminate the authenticity and reliability of the data. At this stage, the global model can be used to verify the authenticity of local client data. If the data of a certain local client is inconsistent with the global information during the global model verification process, this may indicate that there is a problem with the data of this client.
[0099] The present invention also relates to a computer-readable storage medium in application, on which a client data authenticity verification program based on federated learning is stored. When the program is executed by a processor, the above-mentioned client data authenticity verification method based on federated learning is implemented.
[0100] The present invention also relates to a computer device in application, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the above-mentioned client data authenticity verification method based on federated learning is implemented.
[0101] Those skilled in the art should understand that the embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0102] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general purpose computers, special purpose computers, embedded processors, or other programmable data processing devices to produce a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices create means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in one block or multiple blocks.
[0103] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufacture including instruction means that implement the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in one block or multiple blocks.
[0104] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in one block or multiple blocks.
[0105] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the present invention.
[0106] Obviously, those skilled in the art can make various changes and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.
Claims
1. A method for verifying the authenticity of client data based on federated learning, characterized in that: The method sets up a local model for learning data features and corresponding labels; Based on federated learning, the global model is trained collaboratively with the features and labels of the data of all participants, and the global model is used to aggregate local information to verify the authenticity of the client data. The method comprises the following steps: S1 builds a mixed dataset, including the following steps: S1.1 Determine the number of samples Mk used to calculate the mean; S1.2 Divide the sample data in each client into multiple batches according to the sample number Mk, calculate the mean of each batch of data, and add it to the mixed data set (X) as a new sample. g , Y g ); S2 uses local data to train the client's local model, with different data reliability labels corresponding to different data features; data reliability labels include reliable data, unreliable data, unreliable data source, missing data, data errors or outliers, and fake data; S3 uses the FedMix algorithm to train and optimize data to solve client deviations caused by data heterogeneity; S4 calculates the cosine similarity between the client model gradient and the global model gradient in each training round, and obtains the client model as the weight of the global aggregation; S4 includes the following steps: S4.1 Set the mean aggregation global model WG, Where i is the local client serial number, and n is the total number of local clients; S4.2 The cosine similarity between the client model gradient and the global model gradient in each training round, in, W g (t) represents the initial global model parameters of this round, W i (t) represents the local model parameters after training of local client i, η is the learning rate, represents the gradient; S4.3 uses the softmax function to calculate the aggregation weights, weights(i) = softmax(Sim i (t)) S5 aggregates the local model parameters from all participants through weighted aggregation; S5 includes the following steps: S5.1 Initialize the global model parameter w0 and select the clients participating in the training to form a set S t ; S5.2 determines whether the client participates in model training and sends the mixed data set (X g , Y g ) is sent to the clients participating in the training; S5.3 Receive local model from client S5.4 Weighted aggregation p k is the weight of the kth client, and the global model is obtained; S6 determines the authenticity of client data based on global information.
2. According to a method for verifying the authenticity of client data based on federated learning according to claim 1, it is characterized in that: S3 includes the following steps: S3.1 Receiving the global model w t ; S3.2 For each batch of data, from the mixed data set (X g , Y g )Select (x g ,y g ) Join local training, optimize local model training using the FedMix method, calculate the loss function l, and update the model parameters; S3.3 upload the trained local model w; S3.4 ends when the iteration is completed and outputs the updated global model. Otherwise, the global model is aggregated and sent to the local, and the model parameters are trained using local data to continue the iteration.
3. According to a method for verifying the authenticity of client data based on federated learning according to claim 2, it is characterized in that: In S3.2, l=l1+l2+l3, where l1=(1-λ)l(f((1-λ)X;w),Y), l2=λl(f((1-λ)X;w),y g ), λ is the adjustment parameter, λ∈(0,1).
4. A computer-readable storage medium, characterized in that: A client data authenticity verification program based on federated learning is stored thereon, and when the program is executed by the processor, the client data authenticity verification method based on federated learning described in one of claims 1 to 3 above is implemented.
5. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the client data authenticity verification method based on federated learning described in any one of claims 1 to 3 is implemented.
Citation Information
Patent Citations
Federal learning security aggregation method and system for distributed load prediction
CN117010483A