Multi-party collaborative data availability security detection method and system

Through the multi-party collaborative data availability security detection method, the gradient data is disturbed and verified by using encryption matrix and shared mask, which solves the problems of gradient privacy and detection efficiency in the prior art, and realizes efficient and secure gradient data aggregation and detection in federated learning.

CN120342738AActive Publication Date: 2025-07-18XIDIAN UNIV
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202510606053.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-07-18
Estimated Expiration
2045-05-12

AI Technical Summary

Technical Problem

The existing data availability security detection scheme is difficult to achieve efficient and accurate detection and aggregation under the premise of protecting client gradient privacy, and it has huge computing and communication overhead, and is easily confused by low-level adversarial samples, resulting in reduced defense effects.

Method used

A multi-party collaborative data availability security detection method is adopted to generate encryption matrix and shared masks through the initialization stage, and disturb the gradient data is imposed in the gradient upload stage, and two verifications are conducted in the data availability detection stage to ensure the consistency of detection and aggregated gradient data. The scrambled gradient data is used for cluster analysis, and the toxic gradient data of malicious clients is eliminated, and the aggregation results of available gradient data are restored.

Benefits of technology

While protecting gradient privacy, it reduces computing and communication overhead, improves detection accuracy, ensures the security and efficiency of model training, can efficiently complete data detection and model aggregation when there is a malicious client, and restores the secure aggregation results of availability data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120342738A_ABST
    Figure CN120342738A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-party collaborative data availability security detection method and system. The method comprises the following steps: an initialization stage: preprocessing local data by a client; the verification server constructs a projection matrix, distributes a first encryption matrix generated according to the projection matrix to each client, and uploads a generated second encryption matrix and a private key corresponding to the first encryption matrix to the central server; a gradient uploading stage: adjacent clients negotiate a shared mask, train local data to obtain gradient data, apply disturbance to the gradient data by using the shared mask, and upload the scrambled gradient data to the central server; in the data availability detection stage, a multi-party cooperation mode is adopted, two rounds of verification are conducted, it is guaranteed that shared masks uploaded by different adjacent clients are consistent, gradient data used for aggregation and uploaded by the same client are consistent with gradient data used for availability detection, and a result obtained through clustering analysis detection is meaningful. According to the method, the gradient privacy is protected, and meanwhile, the detection is more accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of information security, and particularly relates to a multi-party collaborative data availability security detection method and system. Background Art

[0002] In the open and interconnected environment, the integration and utilization of data resources have become an important foundation for effectively releasing the value of data elements and supporting the construction of the digital economy. As an effective way of data security integration, federated learning has received extensive attention from all walks of life since its proposal. The defense against data poisoning attacks in federated learning has always been a hot research field. The server may be damaged by poisoned data uploaded by some malicious clients, resulting in the failure of model training to converge or a significant decline in performance. To effectively resist data poisoning attacks and protect the security of clients, a data availability detection and aggregation scheme against data poisoning attacks has emerged. This scheme mainly includes two detection and aggregation methods: robust aggregation and anomaly detection. However, these schemes still have certain limitations in terms of training accuracy, communication efficiency, and security, and need to be further optimized and improved.

[0003] Specifically, in the detection scheme using robust aggregation, the inherent attributes of model updates are often further utilized to identify and weaken the effects of malicious model updates to complete the detection and resistance to poisoned data. For example, the paper "Robust aggregation for federated learning" published in 2022 proposed a new type of robust federated learning method based on the traditional Krum aggregation algorithm. By introducing the geometric median to replace the traditional arithmetic mean aggregation mechanism, it effectively resists the damage of model updates caused by malicious use of poisoned data or uploading of poisoned gradients by clients. This method designs an iterative optimization framework based on the smoothed Weiszfeld algorithm, which reduces the communication overhead using a secure multi-party computation protocol while protecting user privacy. This method regards the model as a vector and extracts information using its statistical features, and the calculation is relatively simple, which is suitable for detecting attacks that have a greater impact on model updates. If the change amplitude generated by the attack is too small, or the evaluation criteria of statistical features and similarity cannot well distinguish malicious gradients, the defense effect will be greatly reduced.

[0004] In addition, the patent applied for by Beijing Institute of Technology is "A robust aggregation method for federated learning based on backdoor attack defense" (application number: CN202410776571.8, application publication number: CN118965415A). The present invention analyzes the similarity of key parameters of the global model of federated learning, reduces the dimension of the model parameters, performs unsupervised clustering and calculates the local proxy model. After the dimension reduction, the cosine distance between the model parameters is used to divide the malicious model from the benign model to ensure the performance of outlier detection and unsupervised clustering. By trimming the local proxy model through the Euclidean distance, it can effectively resist malicious backdoor attacks with high amplitude values and improve the robustness of the aggregation.

[0005] In the aggregation scheme using anomaly detection, statistical and analytical methods are often used to identify the model's training mode, data set or related events by defining one or more loss functions. If an unexpected pattern, abnormal behavior or abnormal data is detected, the system will issue an early warning and take responsive measures. The defense method of anomaly detection is similar to robust aggregation, and there is a certain overlap. The difference is that the latter only detects model updates and ultimately obtains a global aggregate gradient, while the former detects maliciously injected data or false models without considering aggregation. For example, "Anomaly Detection in Time Series with Robust Variational Quasi-Recurrent Autoencoders" published in 2022 proposed a new deep learning method called Variational Quasi-Recurrent Autoencoder (VQRAE) and its bidirectional extended version (BiVQRAE), and designed a robust objective function based on α, β and γ divergence. By suppressing the weight of anomalies in the loss calculation, the model is more focused on learning normal patterns in unlabeled training, which makes it show strong robustness and better overall performance in detection accuracy and training efficiency.

[0006] Although the above existing aggregation schemes mitigate to some extent the harm caused by malicious clients uploading poisoned gradient data or using poisoned gradient data for server model aggregation, researchers still hope to impose requirements for the detection of poisoned gradient data while protecting the privacy of the gradients uploaded by clients, and remain skeptical about whether the model can be efficiently aggregated when clients use or upload poisoned gradient data. First, most existing robust aggregation schemes use the uploaded parameters as vectors and extract their statistical features as the evaluation method, which causes the gradient data uploaded by clients to be saved on the server in plaintext, thus posing a great threat to the privacy of client data. Second, when performing certain differentiation operations on the gradient data uploaded by each client in a non-disclosed manner, the huge computational and communication overheads greatly reduce the practicality of the entire data availability security detection scheme. Finally, since all operations are performed without disclosing the gradient data, ensuring the consistency between the detected data and the actual aggregated data is also a major problem. How to accurately and efficiently detect the availability of the gradient data uploaded by clients while ensuring the privacy of the uploaded gradients has become the main problem to be solved urgently. Summary of the Invention

[0007] To solve the above problems existing in the prior art, the present invention provides a multi-party collaborative data availability security detection method and system. The technical problems to be solved by the present invention are realized through the following technical solutions:

[0008] In a first aspect, an embodiment of the present invention provides a multi-party collaborative data availability security detection method, the method comprising:

[0009] The initialization stage includes: each client preprocesses the local data and loads the global model parameters issued by the central server; the verification server constructs a projection matrix, generates a first encryption matrix and a second encryption matrix according to the projection matrix, distributes the first encryption matrix to each client, and uploads the second encryption matrix and the private key corresponding to the first encryption matrix to the central server;

[0010] The gradient upload stage includes: each client negotiates with adjacent clients to share a mask, trains to obtain gradient data according to the global model parameters and the preprocessed local data, applies a perturbation to the gradient data using the shared mask to obtain scrambled gradient data, and uploads the scrambled gradient data to the central server;

[0011] The data availability detection phase includes: each client uses the first encryption matrix to perform feature mapping on the preprocessed local data to generate a first verification parameter, uploads the first verification parameter to the central server, and uploads the shared mask to the verification server; the central server calculates a new first verification parameter based on the first verification parameter and the corresponding private key, calculates a second verification parameter according to the second encryption matrix, and uploads the new first verification parameter and the second verification parameter to the verification server; the verification server verifies the shared mask uploaded by each client. If the verification passes, it verifies the consistency between the gradient data for detection and the gradient data for aggregation according to the second verification parameter. If the verification passes, the central server performs clustering analysis on the first verification parameter to obtain secure and available scrambled gradient data, aggregates all the secure and available scrambled gradient data, and issues new global model parameters to each client, or ends the model training to complete data security aggregation.

[0012] In an embodiment of the present invention, when there is a situation where the verification fails during the data availability detection phase, to determine malicious clients, the corresponding method further includes:

[0013] The data aggregation phase includes: the central server performs clustering analysis on the first verification parameter to obtain the scrambled gradient data of the malicious client as poisoned gradient data; the clients adjacent to the malicious client upload the shared masks negotiated with the malicious client to the central server; the central server performs aggregation recovery processing on the scrambled gradient data uploaded by the clients other than the malicious client according to the poisoned gradient data and all the shared masks negotiated with the malicious client.

[0014] In a second aspect, an embodiment of the present invention provides a multi-party collaborative data availability security detection system, the system includes clients, a verification server, and a central server; wherein,

[0015] The client includes: a model parameter receiving module for receiving the global model parameters issued by the central server; a preprocessing module for preprocessing local data; a model parameter loading module for loading the global model parameters issued by the central server; a shared mask negotiation module for each client to negotiate a shared mask with adjacent clients; a shared mask uploading module for uploading the shared mask to the verification server; a gradient training module for training based on the global model parameters and the preprocessed local data to obtain gradient data; a gradient scrambling processing module for applying perturbations to the gradient data using the shared mask to obtain scrambled gradient data; a scrambled gradient uploading module for uploading the scrambled gradient data to the central server; an encryption matrix receiving module for receiving the first encryption matrix issued by the verification server; a first verification parameter generation module for generating a first verification parameter by performing feature mapping on the preprocessed local data using the first encryption matrix; a first verification parameter uploading module for uploading the first verification parameter to the central server;

[0016] The verification server includes: an encryption matrix construction module for constructing a projection matrix and generating a first encryption matrix and a second encryption matrix based on the projection matrix; an encryption matrix distribution and uploading module for distributing the first encryption matrix to each client and uploading the second encryption matrix and the private key corresponding to the first encryption matrix to the central server; a first shared mask receiving module for receiving the shared mask uploaded by each client; a second verification parameter receiving module for receiving the second verification parameter uploaded by the central server; a shared mask verification module for verifying the shared mask uploaded by each client; a consistency confirmation module for verifying the consistency between the gradient data for detection and the gradient data for aggregation based on the second verification parameter;

[0017] The central server includes: an encryption matrix and key receiving module for receiving the second encryption matrix and the private key corresponding to the first encryption matrix issued by the verification server; a gradient receiving module for receiving the scrambled gradient data uploaded by each client; a first verification parameter receiving module for receiving the first verification parameter uploaded by each client; a first detection parameter update module for calculating a new first verification parameter based on the first verification parameter and the corresponding private key; a second verification parameter generation module for calculating a second verification parameter based on the second encryption matrix; a second verification parameter uploading module for uploading the second verification parameter to the verification server; an availability security detection module for performing clustering analysis on the first verification parameter to obtain secure and available scrambled gradient data; a secure aggregation module for aggregating all the secure and available scrambled gradient data; a model parameter distribution module for distributing global model parameters to each client; a model parameter update module for updating the global model parameters distributed to each client.

[0018] In one embodiment of the present invention, the central server further includes:

[0019] An availability security detection module, which is further configured to perform clustering analysis on the first test parameter to obtain scrambled gradient data of malicious clients as poisoning gradient data when there are malicious clients;

[0020] A second shared mask receiving module, which is configured to receive the shared masks negotiated with the malicious clients respectively uploaded by the clients adjacent to the malicious clients;

[0021] An aggregation recovery module, which is configured to perform aggregation recovery processing on the scrambled gradient data uploaded by the remaining clients except the malicious clients by using all the shared masks negotiated with the malicious clients.

[0022] Advantages of the present invention:

[0023] The multi-party collaborative data availability security detection method proposed by the present invention efficiently realizes the secure aggregation of client gradient data on the premise of availability security detection and protecting user data. Specifically:

[0024] In terms of gradient privacy protection, in traditional data availability security detection schemes, most of them use the plaintext of the uploaded gradient data or directly sample the client local data, and it is difficult to guarantee privacy. In order to protect gradient privacy, security detection technologies such as homomorphic encryption upload are further proposed. Although gradient privacy protection is achieved to a certain extent, there is a huge computational overhead and it is difficult to be practically applied. However, the data availability security detection scheme proposed by the present invention scrambles the gradients uploaded by the clients, makes a security check on the data while ensuring that only the scrambled gradient data is uploaded, guarantees gradient privacy while detecting the uploaded scrambled gradient data, effectively prevents unauthorized access and leakage, and thus guarantees the confidentiality of the client's private data. Even when gradient aggregation is performed on the central server, the original gradients remain scrambled to ensure the secure progress of the entire training process under the premise of privacy protection.

[0025] In terms of availability security detection and overhead, in traditional data availability security detection schemes, either feature extraction operations need to be performed on gradient data, or multiple rounds of communication need to be added, resulting in a large amount of computational overhead or communication overhead. Moreover, traditional data availability security detection schemes are easily confused by low-level adversarial samples and cannot maintain good detection accuracy under slight perturbations. This causes the defense effect to be greatly reduced when the change amplitude caused by the attack during model training is too small, or when the evaluation criteria for statistical features and similarity cannot well distinguish malicious gradients. Although some schemes have achieved data detection operations in ciphertext environments such as homomorphic encryption, they are often accompanied by huge computational overheads. The data availability security detection scheme proposed by the present invention designs verification parameters and adopts a multi-party collaboration method. After two verifications, it respectively ensures the consistency of the shared masks uploaded by adjacent clients and at the same time ensures the consistency of the gradient data used for detection and the gradient data used for aggregation, so that the results obtained by clustering analysis detection are meaningful. Through such careful design, the availability of the gradient data used by the central server for aggregation can be guaranteed, and a strong binding relationship between the aggregated gradient and the detection gradient is achieved, making the detection results more accurate. At the same time, through the secure upload method implemented using a non-homomorphic scrambling scheme, the communication overhead and computational overhead are reduced, enabling efficient data detection and model aggregation under the premise of small-scale computational communication overhead.

[0026] The present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 is a schematic flowchart of a multi-party collaborative data availability security detection method provided by an embodiment of the present invention;

[0028] Figure 2 is a schematic flowchart of another multi-party collaborative data availability security detection method provided by an embodiment of the present invention;

[0029] Figure 3 is a schematic structural diagram of a central server in a multi-party collaborative data availability security detection system provided by an embodiment of the present invention;

[0030] Figure 4 is a schematic structural diagram of each client in a multi-party collaborative data availability security detection system provided by an embodiment of the present invention;

[0031] Figure 5 is a schematic structural diagram of a verification server in a multi-party collaborative data availability security detection system provided by an embodiment of the present invention;

[0032] Figure 6It is a schematic structural diagram of a central server in another multi - party collaborative data availability security detection system provided by an embodiment of the present invention. Detailed implementation manners

[0033] The present invention will be further described in detail below with reference to specific embodiments, but the implementation manners of the present invention are not limited thereto.

[0034] In order to make an accurate and efficient availability detection of the gradient data uploaded by clients while ensuring the privacy of the uploaded gradients, in a first aspect, please refer to Figure 1 , an embodiment of the present invention provides a multi - party collaborative data availability security detection method, which specifically includes the following steps:

[0035] S10. The initialization stage includes: each client pre - processes the local data and loads the global model parameters sent by the central server; the verification server constructs a projection matrix, generates a first encryption matrix and a second encryption matrix according to the projection matrix, distributes the first encryption matrix to each client, and uploads the second encryption matrix and the private key corresponding to the first encryption matrix to the central server.

[0036] The initialization stage in the embodiment of the present invention specifically includes: the global model parameters sent by the central server; each client pre - processes the local data and loads the global model parameters sent by the central server; the verification server randomly selects two global parameters, generates multiple groups of hyperplane groups according to the two selected global parameters, each group of hyperplane groups includes multiple hyperplanes, and orthogonalizes each group of hyperplane groups to generate a projection matrix; the verification server generates corresponding first encrypted vectors according to each column vector in the projection matrix, integrates all the first encrypted vectors to obtain a first encryption matrix, distributes the first encryption matrix to each client, and at the same time integrates and uploads the private key corresponding to the first encryption matrix to the central server; the verification server extends the projection matrix to obtain an extended projection matrix, generates corresponding second encrypted vectors according to each column vector in the extended projection matrix, integrates all the second encrypted vectors to obtain a second encryption matrix, and uploads the second encryption matrix to the central server. Among them, the verification server extends the projection matrix to obtain an extended projection matrix, including: the verification server adds a row vector of all 1s after the projection matrix to obtain the extended projection matrix. More specifically for the initialization stage:

[0037] Each client pre - processes the local data, such as normalization, data augmentation, and missing value processing, to improve the data quality and ensure that it meets the requirements of model training, that is, the local data mentioned later for use is the pre - processed local data. At the same time, each client receives and stores the global model parameters sent by the central server to ensure that the model corresponding to the latest global model parameters is loaded for training.

[0038] The verification server randomly selects two global parameters num and num_v, which are used to determine the projection matrix W. That is, num groups of hyperplane groups are determined by the global parameter num, and each group of hyperplane groups has num_v hyperplanes, so that the projection matrix can be denoted as:

[0039] W = [V1, V2, ……, V I , ……, V num ;

[0040] Each group of hyperplane groups randomly generates num_v column vectors to form num_v hyperplanes, denoted as:

[0041] V I = [v I,1 , v I,2 , ……, v I,l , ……, v I,num_v ;

[0042] v I,l = rand(z, 1);

[0043] Among them, V I represents the I-th group of hyperplane groups, where the value range of I is 1 to num, and v I,l represents the l-th column vector in V I , where the value range of l is 1 to num_v, z represents the dimension of each hyperplane, which is determined by the dimension of the gradient data during the training process, and rand() represents the random number generation function.

[0044] Next, the verification server orthogonalizes each group of hyperplane groups to make it more comprehensively reflect the features:

[0045]

[0046] Among them, the projection is defined as:

[0047]

[0048] Among them, represents the transpose operation of v I,l , and the projection matrix W can be represented as the space spanned by multiple vectors u j .

[0049] Furthermore, after generating the projection matrix W in the embodiment of the present invention, the verification server further processes each column vector in each group of hyperplane groups of the projection matrix W: randomly select two large prime numbers α and P, select a large random number s x ∈Z p as the key, and select z + 2 random numbers, denoted as b = [b1, b2, ……, b k, ……, b z+2 , b k is the k-th element in b, where k ranges from 1 to z + 2, and denote a k as the k-th element in a certain column vector, and expand this column vector a z+1 = a z+2 = 0, and calculate as follows to obtain the first encrypted vector E x = (α, P, e x,1 , e x,2 , ……, e x,k , ……, e x,z+2 ), where x ranges from 1 to num * num_v, and the calculation method of e x,k is as follows:

[0050] e x,k = s x (a k ·α + b k ) mod P;

[0051] The verification server integrates all the first encrypted vectors into the first encrypted matrix G = (E1, E2, ……, E x , ……, E num*num_v ), distributes the first encrypted matrix G to each client. Each client prepares for data feature extraction by receiving the first encrypted matrix G distributed by the verification server; and integrates the private key S = (s1, s2, ……, s x , ……, s num*num_v ) corresponding to the first encrypted matrix G and uploads two large prime numbers α and P to the central server.

[0052] Similarly, the verification server expands the projection matrix W to obtain the expanded projection matrix that is, add a row vector of all 1s after the projection matrix W, and then perform the above processing similar to the projection matrix W again to first obtain the projection matrix the second encrypted vectors corresponding to all column vectors, integrate all the second encrypted vectors into the second encrypted matrix G s , upload the second encrypted matrix G s to the central server, and the verification server integrates and retains the private key S s corresponding to the second encrypted matrix G s .

[0053] S20. The gradient upload stage includes: each client negotiates with adjacent clients to share masks, trains to obtain gradient data based on the global model parameters and preprocessed local data, applies perturbations to the gradient data using the shared masks to obtain scrambled gradient data, and uploads the scrambled gradient data to the central server.

[0054] Suppose there are n clients, and the total amount of data of the n clients is N. Each client negotiates a common random number with its adjacent clients as a shared mask. For example, the i-th client negotiates the shared mask with its adjacent clients respectively, and the shared mask between the i-th client and the (i - 1) mod n client is [r i-1modn,i , and the shared mask between the i-th client and the (i + 1) mod n client is [r i+1modn,i .

[0055] Each client uses the existing optimization algorithm to train the gradient data based on the global model parameters and the preprocessed local data, and applies the perturbation to the gradient data using the shared mask to obtain the scrambled gradient data. For example, the gradient data obtained after the training of the i-th client is in plaintext, and this gradient data is denoted as m i , with the amount of data being n i , and the gradient data for uploading detection is denoted as Furthermore, according to the set security protocol, calculate the scrambled gradient data to ensure the privacy and security of the gradient data, and upload it to the central server for the central server to perform subsequent detection and aggregation. Each client performs this operation.

[0056] S30. The data availability detection stage includes: Each client uses the first encryption matrix to perform feature mapping on the local data to generate the first verification parameter, uploads the first verification parameter to the central server, and uploads the shared mask to the verification server; The central server calculates the new first verification parameter according to the first verification parameter and the corresponding private key, and calculates the second verification parameter according to the second encryption matrix, and uploads the new first verification parameter and the second verification parameter to the verification server; The verification server verifies the shared mask uploaded by each client. If the verification passes, it verifies the consistency between the gradient data for detection and the gradient data for aggregation according to the second verification parameter. If the verification passes, the central server performs clustering analysis on the first verification parameter to obtain the secure and available scrambled gradient data, aggregates all the secure and available scrambled gradient data, and issues the new global model parameters to each client, or ends the model training to complete the data security aggregation.

[0057] In the data availability detection phase of the embodiments of the present invention, it specifically includes: the central server broadcasts a notice to each client to perform data availability detection; each client negotiates with adjacent clients to share masks, and uses the first encryption matrix to perform feature mapping on local data to generate the first verification parameter, uploads the first verification parameter to the central server, and uploads the shared mask to the verification server; the central server receives the first verification parameter uploaded by each client, decrypts the corresponding first verification parameter using the private key uploaded by the verification server, randomly generates a random number, calculates a new first verification parameter according to the random number and the decryption result, and issues it to the verification server; the central server constructs a random number vector according to the random number, expands the scrambled gradient data uploaded by each client according to the random number vector to obtain expanded scrambled gradient data, calculates a second verification parameter according to the second encryption matrix and the expanded scrambled gradient data, and issues the second verification parameter to the verification server; the verification server receives the shared mask uploaded by each client and conducts a verification. If the verification fails, it determines that the corresponding client is a malicious client. If the verification passes, based on the linear additivity property of the dot product, it verifies the consistency between the gradient data for detection and the gradient data for aggregation according to the new first verification parameter and the second verification parameter. If the verification fails, it determines that the corresponding client is a malicious client. If the verification passes, the central server performs clustering analysis on the first verification parameter to obtain secure and available scrambled gradient data, aggregates all the secure and available scrambled gradient data, and issues new global model parameters to each client, or ends the model training to complete data security aggregation. More specifically:

[0058] The central server broadcasts a notice to each client to perform data availability detection.

[0059] Each client, such as client i, for its gradient data used for detection and the shared masks [r i-1modn,i and [r i+1modn,i negotiated with adjacent clients, perform an F1 operation using the first encryption matrix G to obtain the first verification parameter D i , upload the first verification parameter D i to the central server for the central server to detect the scrambled gradient data and other related parameters uploaded by the client, and upload the shared masks [r i-1modn,i and [r i+1modn,i to the verification server. The formula for the F1 operation is as follows:

[0060]

[0061] The following specifically explains the operation F1(T, G), where T is the data Assume that the k-th element in the vector T is t k , first expand the vector T by tz+1 = t z+2 = 0, for each E x perform the following calculations. First, calculate to obtain j x,k :

[0062]

[0063] where r k is a randomly selected random number.

[0064] Further calculate:

[0065]

[0066] Integrate to obtain the first verification parameter D i = (J1, J2, ……, J x , ……, J num*num_v ), and upload this first verification parameter D i to the central server. Each client performs the above first verification parameter generation process.

[0067] The central server receives the first verification parameter from each client. For example, after receiving the first verification parameter D i , it decrypts it using the corresponding private key S uploaded by the verification server and performs the F2 operation to obtain the verification parameter and randomly generates a random number r, calculates the new first verification parameter d i + r and sends it to the verification server. The F2 operation formula is expressed as:

[0068] d i = F2(D i , S);

[0069] The following is a specific explanation of the operation F2(D, S), where D is the first verification parameter D i : Assume J x is the x-th element in D i , calculate:

[0070] q x = s x -1 · J x mod P;

[0071] Further calculate:

[0072]

[0073] Integrate to obtain the verification parameter d i = (Q1, Q2, ……, Q x , ……, Q num*num_v ).

[0074] Further, the central server constructs a random number vector [r] using the previously generated random number r, i.e., each element in the vector is r, and extends the scrambled gradient data C of each client i to obtain the extended scrambled gradient data S di :

[0075]

[0076] The central server uses the second encryption matrix G s and the extended scrambled gradient data S di to perform the F1 operation to obtain the second verification parameter D si , and sends it to the verification server:

[0077] D si = F1(S di , G s );

[0078] The verification server receives the shared masks of each client, such as the shared masks [r i-1modn,i and [r i+1modn,i of client i, and verifies its shared masks:

[0079] r (i-1)+1mod n,i-1 = r i-1 mod n,i ;

[0080] r i+1 mod n,i = r (i+1)-1 mod n,i+1 ;

[0081] In a multi-party collaborative manner, ensure the authenticity of the features of [r ii1 mod n,i and [r i+1 mod n,i uploaded by each client. If the verification fails, the client is determined to be a malicious client.

[0082] The verification server uses the reserved private key S s to perform the F2 operation on the second verification parameter D si to obtain the verification parameter where the F2 operation formula is expressed as:

[0083] d si = F2(D s i , S s ).

[0084] Furthermore, based on the linear additivity property of the dot product, the verification server performs the following calculations to verify the consistency between the gradient data uploaded by the client for detection and the gradient data for aggregation:

[0085]

[0086] Among them, d i is the gradient data for detection, and d si is the gradient data for aggregation. The verification server verifies whether the calculation result is equal to d si . If they are equal, it indicates that the i used to calculate d is consistent with the si contained in the one used to calculate d , to ensure the consistency of the gradient data for detection and the gradient data for aggregation. If they are not equal, it is determined as a malicious client.

[0087] Finally, the central server performs clustering analysis on the first verification parameters uploaded by each client, decrypts the first verification parameters. For example, the first verification parameter D i uploaded by the normal client i after decryption and then analyzes to detect the secure and available scrambled gradient data The first verification parameter D j uploaded by the malicious client j after decryption and then analyzes to detect the poisoned gradient data deviating from the secure and available For the detected secure and available scrambled gradient data perform data aggregation on all the secure and available scrambled gradient data; for the detected poisoned gradient data deviating from the secure and available continue to execute the subsequent data aggregation stage.

[0088] The central server updates the weights of the global model, adjusts the global model parameters using an optimization algorithm to improve the overall performance of the global model, and sends the new global model parameters to each client, or ends the model training to complete the secure data aggregation. Among them, the updated model is distributed to each client for local training to ensure that all clients always use the latest model parameters, thereby enhancing the collaborative learning ability of the system and improving the overall performance of the model. After completing the model update, conduct a comprehensive evaluation of the new model to detect its performance on the validation set or test set. By calculating the accuracy rate, loss value, and other key indicators, analyze the improvement of the model to ensure that the optimized model has better generalization ability and stability in practical applications.

[0089] Furthermore, if there is a situation where the verification fails during the data availability detection stage, determine the malicious client. Please refer to Figure 2 , and the corresponding method also includes:

[0090] S40. The data aggregation stage includes: the central server performing clustering analysis on the first test parameter to obtain the scrambled gradient data of malicious clients as poisoned gradient data; the clients adjacent to the malicious clients uploading their respective shared masks negotiated with the malicious clients to the central server; the central server performing aggregation and recovery processing on the scrambled gradient data uploaded by the remaining clients except the malicious clients according to the poisoned gradient data and all the shared masks negotiated with the malicious clients.

[0091] In the embodiment of the present invention, the data aggregation stage specifically includes: the central server performing cumulative summation on the scrambled gradient data uploaded by each client to obtain the first cumulative perturbation gradient data, and calculating the second cumulative perturbation gradient data according to the first cumulative perturbation gradient data and all the shared masks negotiated with the malicious clients; the central server performing clustering analysis on the first test parameter to obtain the scrambled gradient data of malicious clients as poisoned gradient data, calculating the new total data volume except the malicious clients according to the poisoned gradient data, and performing aggregation and recovery processing on the scrambled gradient data uploaded by the remaining clients except the malicious clients according to the new total data volume and the second cumulative perturbation gradient data. More specifically:

[0092] The central server receives the scrambled gradient data C uploaded by each client i , and performs cumulative summation calculation on all the scrambled gradient data C i to obtain the first cumulative perturbation gradient data:

[0093]

[0094] where M temp1 represents the first cumulative perturbation gradient data.

[0095] When there is a malicious client j, ask the two clients adjacent to the malicious client j for the shared masks [r j-1modn,j and [r j+1modn,j negotiated with the malicious client j.

[0096] The central server uses the shared masks [r j-1modn,j and [r j+1modn,j to process the first cumulative perturbation gradient data M temp1 to obtain the second cumulative perturbation gradient data:

[0097]

[0098] The central server performs clustering analysis on the first test parameter to obtain the scrambled gradient data C of the malicious client j j as the poisoned gradient data, and counts the data volume of the poisoned gradient data, denoted as n j , then calculates the new total data volume N ’ = N - n j, and further process the second accumulated perturbation gradient data M temp2 to obtain the aggregated recovery result of the scrambled gradient data of the remaining available clients except the malicious clients:

[0099]

[0100] In the data aggregation stage of the embodiments of the present invention, a multi-party collaborative availability security data aggregation result recovery method is adopted to detect poisoned gradient data and then recover the aggregation result of the remaining available gradient data. Assume that there is a malicious client j that negotiates a shared mask [r j-1modn,j and [r j+1modn,j with adjacent clients, and the remaining clients have negotiated shared masks [r i-1modn,i and [r i+1modn,i with adjacent clients. The central server has the scrambled gradient data C i of each client. After performing the multi-party collaborative availability security data aggregation result recovery process, the central server will obtain the true value after data aggregation processing of the scrambled gradient data uploaded by honest clients

[0101] It should be noted here that the embodiments of the present invention can detect and judge a dropped client as a malicious client, and can realize the recovery of the security aggregation result of the availability data of the remaining clients when there is a dropped client.

[0102] As can be seen, in a multi-party collaborative data availability security detection method provided by the embodiments of the present invention:

[0103] Central server: The central server hopes to train an efficient global model by fusing multi-party data. However, limited by the lack of data faced by a single data set and the computing and storage pressure brought by centralized storage, it is difficult for traditional training methods to break through the performance bottleneck. To solve this problem, the central server adopts the federated learning method to obtain the gradients or model parameters generated by local training from multiple clients (such as terminal devices or institutions), and performs secure aggregation on the central server. This can not only fully exploit the heterogeneous data and computing resources in distributed devices, but also improve the generalization ability of the model, enabling it to still maintain good adaptability in more complex and variable environments.

[0104] Client: The client participates in the joint training of the global model by providing model gradients or parameter updates calculated during the local training process, hoping to obtain a model with superior performance and stronger generalization capabilities in this collaborative process to meet its own actual application needs. However, for privacy reasons, clients are usually reluctant to share local data directly to prevent potential leakage risks caused by data transmission. At the same time, in the open collaborative training process, there may be malicious clients that affect the training of the global model by uploading gradients or parameter updates with malicious interference (i.e., "poisoning" data). For example, an attacker may inject adversarial samples to make the model perform abnormally on a specific task, or even deliberately reduce the accuracy of the model in key scenarios, thereby achieving the purpose of misleading decision-making.

[0105] Verification server: As a neutral third party in federated learning, the verification server is responsible for consistency verification of the model gradients or parameter updates uploaded by the client without obtaining any privacy information, ensuring that they are consistent with the versions participating in the aggregation during the detection phase, thereby preventing malicious clients from performing poisoning attacks by uploading forged or tampered model updates, and improving the credibility and security of global model training.

[0106] The application scenario of the present invention has two important features: data availability security detection and aggregation result recovery after excluding poisoned data. Specifically, during the federated learning process, it cannot be guaranteed that every client is honest, and there may be malicious clients uploading poisoned data to interfere with the convergence of the global model parameters of the central server. Therefore, data availability security detection is required. After detecting the poisoned data, in order to obtain the true aggregation result, the aggregation result needs to be recovered. Based on these two features, in the present invention, a multi-party collaborative data availability security detection process (data availability detection phase) and a multi-party collaborative availability security data aggregation result recovery process (data aggregation phase) are designed. By using the client to perform feature mapping on the uploaded scrambled gradient data and the shared mask, the consistency between the encrypted gradient for detection and the gradient for aggregation is ensured, and then clustering analysis is performed on the detection data to achieve data availability security detection. Compared with the traditional solution that can only collect and aggregate client data but cannot perform data availability security detection to resist client poisoning attacks, the present invention is more secure; compared with using homomorphic encryption to perform consistency detection in the ciphertext state, the present invention is more efficient. At the same time, the embodiments of the present invention adopt a multi-party collaborative method to enhance the binding between the encrypted gradient data for detection and the gradient data for aggregation. Compared with the traditional solution, the problem that it is difficult to bind the encrypted gradient data for detection and the gradient data for aggregation caused by the misbehavior of a single malicious client is reduced. In addition, when poisoned gradient data is found, a multi-party collaborative method is adopted to recover the shared mask, and then the aggregation result is recovered. Compared with re-training in a traditional solution, the utilization rate of client gradient data in a single training is improved. These elaborate designs can ensure the availability of the gradient data for aggregation and enhance the consistency between the encrypted gradient data for detection and the gradient data for aggregation as much as possible.

[0107] In summary, the multi-party collaborative data availability security detection method proposed in the embodiments of the present invention efficiently realizes the secure aggregation of client gradient data on the premise of availability security detection and protecting user data. Specifically:

[0108] In terms of differential privacy protection, in traditional data availability security detection schemes, most of them utilize the plaintext of the uploaded gradient data or directly sample the client's local data, making it difficult to guarantee privacy. To protect differential privacy, technologies such as secure detection with homomorphic encryption upload have been further proposed. Although differential privacy protection has been achieved to a certain extent, it has a huge computational overhead and is difficult to apply in practice. The data availability security detection scheme proposed by the present invention scrambles the gradients uploaded by the client. While ensuring that only the scrambled gradient data is uploaded, it conducts a security check on the data, ensuring differential privacy while detecting the uploaded scrambled gradient data, effectively preventing unauthorized access and leakage, thereby ensuring the confidentiality of the client's private data. Even when gradient aggregation is performed on the central server, the original gradients remain scrambled to ensure the security of the entire training process under the premise of privacy protection.

[0109] In terms of availability security detection and overhead, in traditional data availability security detection schemes, either feature extraction operations need to be performed on the gradient data or multiple rounds of communication need to be added, resulting in a large amount of computational overhead or communication overhead. Moreover, traditional data availability security detection schemes are easily confused by low-level adversarial samples and cannot maintain good detection accuracy under minor perturbations. This means that when the change amplitude caused by the attack during model training is too small, or the evaluation criteria for statistical features and similarity cannot well distinguish malicious gradients, the defense effect is greatly reduced. Although some schemes have achieved data detection operations in ciphertext environments such as homomorphic encryption, they are often accompanied by a huge computational overhead. The data availability security detection scheme proposed by the present invention designs verification parameters and adopts a multi-party collaboration method. After two verifications, it ensures the consistency of the shared masks uploaded by adjacent clients and the consistency of the gradient data used for detection and the gradient data used for aggregation, respectively, making the results of the clustering analysis detection meaningful. Through such careful design, the availability of the gradient data used for aggregation by the central server can be guaranteed, and a strong binding relationship between the aggregated gradients and the detected gradients is achieved, making the detection results more accurate. At the same time, by using a non-homomorphic scrambling scheme to achieve a secure upload method, the communication overhead and computational overhead are reduced, enabling efficient data detection and model aggregation under the premise of a small-scale computational communication overhead.

[0110] In terms of the robustness of availability detection, considering the existence of malicious clients, the embodiments of the present invention further perform clustering analysis on the scrambled gradient data to screen out the poisoned gradient data. Further, in order to enable the aggregation result of the gradient data of the remaining clients to be restored after removing the poisoned gradient data of the malicious clients, the present invention also designs a multi-party collaborative availability security data aggregation result restoration scheme. Through such careful design, the recoverability of the aggregation result can be guaranteed, so that it can not only remove the gradient data of the malicious clients and restore the availability data security aggregation result when there are malicious clients, but also realize the restoration of the availability data security aggregation result of the remaining clients when there are clients dropping out. Since the overhead of re-performing a new round of training is reduced, the utilization efficiency of data is improved. The present invention is applicable to the scenario of malicious clients in federated learning, designs a multi-party collaborative availability security data aggregation result restoration scheme, detects poisoned data in the scenario where malicious clients deliberately use malicious data to interfere with model training, and can obtain the aggregation restoration result of the remaining available data, so as to safely and effectively execute federated learning.

[0111] In a second aspect, please refer to Figures 3 to 5 , the embodiments of the present invention provide a multi-party collaborative data availability security detection system, which includes clients, a verification server, and a central server; wherein,

[0112] The client includes: a model parameter receiving module, configured to receive the global model parameters issued by the central server; a preprocessing module, configured to preprocess local data; a model parameter loading module, configured to load the global model parameters issued by the central server; a shared mask negotiation module, configured to negotiate a shared mask with adjacent clients for each client; a shared mask uploading module, configured to upload the shared mask to the verification server; a gradient training module, configured to perform training based on the global model parameters and local data to obtain gradient data; a gradient scrambling processing module, configured to apply perturbations to the gradient data by using the shared mask to obtain scrambled gradient data; a scrambled gradient uploading module, configured to upload the scrambled gradient data to the central server; an encrypted matrix receiving module, configured to receive the first encrypted matrix issued by the verification server; a first verification parameter generating module, configured to generate a first verification parameter by performing feature mapping on local data by using the first encrypted matrix; a first verification parameter uploading module, configured to upload the first verification parameter to the central server;

[0113] The verification server includes: an encryption matrix construction module, configured to construct a projection matrix and generate a first encryption matrix and a second encryption matrix according to the projection matrix; an encryption matrix distribution and upload module, configured to distribute the first encryption matrix to each client and upload the second encryption matrix and the private key corresponding to the first encryption matrix to the central server; a first shared mask receiving module, configured to receive the shared masks uploaded by each client; a second verification parameter receiving module, configured to receive the second verification parameter uploaded by the central server; a shared mask verification module, configured to verify the shared masks uploaded by each client; a consistency confirmation module, configured to verify the consistency between the gradient data for detection and the gradient data for aggregation according to the second verification parameter.

[0114] The central server includes: an encryption matrix and key receiving module, configured to receive the second encryption matrix and the private key corresponding to the first encryption matrix sent by the verification server; a gradient receiving module, configured to receive the scrambled gradient data uploaded by each client; a first verification parameter receiving module, configured to receive the first verification parameter uploaded by each client; a first detection parameter update module, configured to calculate a new first verification parameter according to the first verification parameter and the corresponding private key; a second verification parameter generation module, configured to calculate the second verification parameter according to the second encryption matrix; a second verification parameter upload module, configured to upload the second verification parameter to the verification server; an availability security detection module, configured to perform clustering analysis on the first verification parameter to obtain secure and available scrambled gradient data; a secure aggregation module, configured to aggregate all the secure and available scrambled gradient data; a model parameter distribution module, configured to send the global model parameters to each client; a model parameter update module, configured to update the global model parameters sent to each client.

[0115] Please refer to Figure 6 , the central server in the embodiment of the present invention further includes:

[0116] The availability security detection module is further configured to, when there is a malicious client, perform clustering analysis on the first verification parameter to obtain the scrambled gradient data of the malicious client as poisoned gradient data;

[0117] A second shared mask receiving module, configured to receive the respective shared masks negotiated with the malicious client uploaded by the clients adjacent to the malicious client;

[0118] An aggregation recovery module, configured to perform an aggregation recovery process on the scrambled gradient data uploaded by the clients other than the malicious client using all the shared masks negotiated with the malicious client.

[0119] Please refer to again Figure 6 , the central server in the embodiment of the present invention further includes:

[0120] A model performance evaluation module is used to evaluate the model corresponding to the new global model parameters to ensure the generalization ability and stability of the model corresponding to the new global model parameters.

[0121] For the system embodiments of the second aspect, since they are basically similar to the method embodiments of the first aspect, the description is relatively simple. For the relevant parts, refer to the partial description of the method embodiments of the first aspect.

[0122] In the description of the present invention, it should be understood that the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, "a plurality" means two or more unless otherwise specifically defined.

[0123] Although the present invention has been described in conjunction with various embodiments herein, however, in the process of implementing the claimed invention, those skilled in the art can understand and achieve other variations of the disclosed embodiments by referring to the description of the specification and its accompanying drawings. In the specification, the word "comprising" does not exclude other components or steps, and "a" or "one" does not exclude a plurality. Certain measures are recited in different embodiments, but this does not mean that these measures cannot be combined to produce good results.

[0124] The above content is a further detailed description of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can be made, and all should be regarded as belonging to the protection scope of the present invention.

Claims

1. A multi-party collaborative data availability security detection method, characterized in that The method includes: The initialization phase includes: each client preprocesses local data and loads the global model parameters sent by the central server; the verification server constructs a projection matrix, generates a first encryption matrix and a second encryption matrix according to the projection matrix, distributes the first encryption matrix to each client, and uploads the second encryption matrix and the private key corresponding to the first encryption matrix to the central server; The gradient upload phase includes: each client negotiates and shares masks with adjacent clients, trains to obtain gradient data according to the global model parameters and the preprocessed local data, applies perturbations to the gradient data using the shared masks to obtain scrambled gradient data, and uploads the scrambled gradient data to the central server; The data availability detection phase includes: each client uses the first encryption matrix to perform feature mapping on the preprocessed local data to generate a first verification parameter, uploads the first verification parameter to the central server, and uploads the shared mask to the verification server; the central server calculates a new first verification parameter according to the first verification parameter and the corresponding private key, calculates a second verification parameter according to the second encryption matrix, and uploads the new first verification parameter and the second verification parameter to the verification server; the verification server verifies the shared mask uploaded by each client, if the verification passes, then verifies the consistency between the gradient data for detection and the gradient data for aggregation according to the second verification parameter, if the verification passes, then the central server performs clustering analysis on the first verification parameter to obtain secure and available scrambled gradient data, aggregates all the secure and available scrambled gradient data, and distributes new global model parameters to each client, or ends the model training to complete the secure data aggregation.

2. The multi-party collaborative data availability security detection method according to claim 1, characterized in that The initialization phase specifically includes: The global model parameters sent by the central server; Each client preprocesses local data and loads the global model parameters sent by the central server; The verification server randomly selects two global parameters, generates multiple groups of hyperplane groups according to the two selected global parameters, each group of hyperplane groups includes multiple hyperplanes, and performs orthogonalization processing on each group of hyperplane groups to generate a projection matrix; The verification server generates corresponding first encryption vectors according to each column vector in the projection matrix, integrates all the first encryption vectors to obtain a first encryption matrix, distributes the first encryption matrix to each client, and at the same time integrates and uploads the private key corresponding to the first encryption matrix to the central server; The verification server expands the projection matrix to obtain an expanded projection matrix, generates corresponding second encryption vectors according to each column vector in the expanded projection matrix, integrates all the second encryption vectors to obtain a second encryption matrix, and uploads the second encryption matrix to the central server.

3. The multi-party collaborative data availability security detection method according to claim 2, wherein The verification server expands the projection matrix to obtain an expanded projection matrix, including: The verification server adds a row vector of all 1s after the projection matrix to obtain an expanded projection matrix.

4. The multi-party collaborative data availability security detection method according to claim 1, characterized in that The data availability detection phase specifically includes: The central server broadcasts a notice to each client to perform data availability detection; Each client negotiates a shared mask with adjacent clients, uses the first encryption matrix to perform feature mapping on the preprocessed local data to generate a first verification parameter, uploads the first verification parameter to the central server, and uploads the shared mask to the verification server; The central server receives the first verification parameters uploaded by each client, decrypts the corresponding first verification parameter using the private key uploaded by the verification server, randomly generates a random number, calculates a new first verification parameter according to the random number and the decryption result, and issues it to the verification server; The central server constructs a random number vector according to the random number, expands the scrambled gradient data uploaded by each client according to the random number vector to obtain expanded scrambled gradient data, calculates a second verification parameter according to the second encryption matrix and the expanded scrambled gradient data, and issues the second verification parameter to the verification server; The verification server receives the shared masks uploaded by each client and conducts a verification. If the verification fails, the corresponding client is determined to be a malicious client. If the verification passes, based on the linear additivity property of the dot product, the verification server verifies the consistency between the gradient data for detection and the gradient data for aggregation according to the new first verification parameter and the second verification parameter. If the verification fails, the corresponding client is determined to be a malicious client. If the verification passes, the central server performs a clustering analysis on the first verification parameter to obtain secure and available scrambled gradient data, aggregates all the secure and available scrambled gradient data, and issues new global model parameters to each client, or ends the model training to complete data security aggregation.

5. The multi-party collaborative data availability security detection method according to claim 1, wherein When there is a situation where the verification fails during the data availability detection phase, the malicious client is determined. The corresponding method further includes: The data aggregation phase includes: The central server performs a clustering analysis on the first verification parameter to obtain the scrambled gradient data of the malicious client as poisoned gradient data; The clients adjacent to the malicious client upload the shared masks negotiated with the malicious client to the central server; The central server performs an aggregation recovery process on the scrambled gradient data uploaded by the remaining clients except the malicious client according to the poisoned gradient data and all the shared masks negotiated with the malicious client.

6. The multi-party collaborative data availability security detection method according to claim 5, wherein The data aggregation phase specifically includes: The central server accumulates and sums the scrambled gradient data uploaded by each client to obtain a first accumulated perturbation gradient data, and calculates a second accumulated perturbation gradient data according to the first accumulated perturbation gradient data and all the shared masks negotiated with the malicious client; The central server performs a clustering analysis on the first verification parameter to obtain the scrambled gradient data of the malicious client as poisoned gradient data, calculates a new total data volume except for the malicious client according to the poisoned gradient data, and performs an aggregation recovery process on the scrambled gradient data uploaded by the remaining clients except the malicious client according to the new total data volume and the second accumulated perturbation gradient data.

7. A multi-party collaborative data availability security detection system, characterized in that, The system includes clients, a verification server, and a central server; wherein, The client includes: a model parameter receiving module for receiving the global model parameters sent by the central server; a preprocessing module for preprocessing local data; a model parameter loading module for loading the global model parameters sent by the central server; a shared mask negotiation module for each client to negotiate a shared mask with adjacent clients; a shared mask uploading module for uploading the shared mask to the verification server; a gradient training module for training using the global model parameters and the preprocessed local data to obtain gradient data; a gradient scrambling processing module for applying a perturbation to the gradient data using the shared mask to obtain scrambled gradient data; a scrambled gradient uploading module for uploading the scrambled gradient data to the central server; an encrypted matrix receiving module for receiving the first encrypted matrix sent by the verification server; a first verification parameter generation module for generating a first verification parameter by performing feature mapping on the preprocessed local data using the first encrypted matrix; a first verification parameter uploading module for uploading the first verification parameter to the central server; The verification server includes: an encrypted matrix construction module for constructing a projection matrix and generating a first encrypted matrix and a second encrypted matrix according to the projection matrix; an encrypted matrix distribution and uploading module for distributing the first encrypted matrix to each client and uploading the second encrypted matrix and the private key corresponding to the first encrypted matrix to the central server; a first shared mask receiving module for receiving the shared mask uploaded by each client; a second verification parameter receiving module for receiving the second verification parameter uploaded by the central server; a shared mask verification module for verifying the shared mask uploaded by each client; a consistency confirmation module for verifying the consistency between the gradient data for detection and the gradient data for aggregation according to the second verification parameter; The central server includes: an encrypted matrix and key receiving module for receiving the second encrypted matrix and the private key corresponding to the first encrypted matrix sent by the verification server; a gradient receiving module for receiving the scrambled gradient data uploaded by each client; a first verification parameter receiving module for receiving the first verification parameter uploaded by each client; a first detection parameter update module for calculating a new first verification parameter according to the first verification parameter and the corresponding private key; a second verification parameter generation module for calculating a second verification parameter according to the second encrypted matrix; a second verification parameter uploading module for uploading the second verification parameter to the verification server; an availability security detection module for performing clustering analysis on the first verification parameter to obtain secure and available scrambled gradient data; a secure aggregation module for aggregating all secure and available scrambled gradient data; a model parameter distribution module for sending global model parameters to each client; a model parameter update module for updating the global model parameters sent to each client.

8. The multi-party collaborative data availability security detection system according to claim 7, characterized in that, The central server further includes: The availability security detection module is further configured to, when there is a malicious client, perform clustering analysis on the first verification parameter to obtain the scrambling gradient data of the malicious client as the poisoning gradient data; The second shared mask receiving module is configured to receive the shared masks negotiated with the malicious client by the clients adjacent to the malicious client respectively; The aggregation recovery module is configured to perform aggregation recovery processing on the scrambled gradient data uploaded by the clients other than the malicious client by using all the shared masks negotiated with the malicious client.

9. The multi-party collaborative data availability security detection system according to claim 7, characterized in that, The central server further includes: The model performance evaluation module is configured to evaluate the model corresponding to the new global model parameters to ensure the generalization ability and stability of the model corresponding to the new global model parameters.

Citation Information

Patent Citations

  • Federal learning robust aggregation method based on backdoor attack defense

    CN118965415A

  • Verifiable gradient security aggregation method and system based on multi-party security computing

    CN115189950A

  • Federal learning security aggregation method based on cosine similarity and homomorphic encryption

    CN117216779A

  • Privacy protection federated learning method with verifiable aggregation result and verifiable gradient quality

    CN117521853A

  • Information protection method and system based on queue data desensitization and differential privacy protection

    CN117708868A