Cloud-edge Collaborative Federated Learning Method, System, and Storage Medium Against Covert Adversaries
By introducing random secret sharing and probabilistic verification mechanisms in cloud-edge collaborative commissioned learning, the problems of data privacy and calculation results security under hidden enemy attacks in the existing technology are solved, and an efficient and secure commissioned learning solution is achieved.
Patent Information
- Application Number
- CN202410626553.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-20
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2044-05-20
AI Technical Summary
Existing commissioned learning programs are difficult to ensure data privacy and correctness and security of computational results when faced with semi-honest or malicious concealed opponents, and are difficult to weigh between security and efficiency.
A cloud-edge collaborative delegation learning method that resists hidden enemies is proposed, and the client shares data secretly and sends it to two groups of unconspiring edge servers for calculation. The edge server performs model calculation tasks, and the cloud server sums and verifys the intermediate data. If the results are consistent or within the error range, the data will be stored; otherwise, the guaranteed output delivery protocol is executed, and the malicious computing party will be detected together with the edge cluster.
It realizes that in the presence of hidden enemies, ensures that the client obtains the correct output results, reduces the risks caused by server fraud, and improves the system's security and privacy protection capabilities.
Smart Images

Figure CN118586515B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to related technical fields such as privacy-preserving machine learning, cloud-edge collaboration, covert security, probabilistic verification, etc. in the transmission of digital information, and particularly relates to a cloud-edge collaborative federated learning method, system and storage medium against covert adversaries. Background Art
[0002] Privacy-Preserving Machine Learning (PPML) refers to the process of training and predicting machine learning models on data without revealing personal or sensitive data. This method aims to protect data privacy while allowing valuable information and knowledge to be extracted from the data.
[0003] Most of the current implementations of privacy-preserving machine learning are based on the federated learning method. Specifically, it refers to a method in which the data owner entrusts the data to a server with computing power for data training and inference, and the server charges a computing fee according to the service method. This model is sometimes referred to as Machine Learning as a Service (MLaaS). In this method, the privacy and security of data become important issues. Therefore, researchers have proposed various technologies to protect data privacy, including but not limited to homomorphic encryption, Secure Multi-Party Computation (SMC), differential privacy, etc.
[0004] Homomorphic Encryption allows direct computation on encrypted data, and the computation result is still encrypted. This means that the training and inference of machine learning models can be performed on the data without decrypting it, thus protecting data privacy. Secure Multi-Party Computation (SMC) is a cryptographic technique that allows multiple distrusted parties to jointly compute the result of a certain function without revealing their respective private inputs. In machine learning, SMC can be used to perform joint training of data while protecting privacy. Differential Privacy protects personal privacy by adding noise to the data, ensuring that the published or shared data does not reveal information about any single record. This method can be applied to the training process of machine learning models to reduce the risk of privacy leakage ("A Review of Privacy Protection Research in Machine Learning", Liu Junxu et al., Journal of Computer Research and Development, Vol. 57, No. 2, 2020).
[0005] In addition, for example, Chinese Patent Document CN114780999A discloses a deep learning data privacy protection method, system, device and medium. This method constructs the objective function and parameter training method of a noise generator, achieving the maximization of the noise intensity added to the training data while minimizing the model performance difference, and automatically balancing the model feasibility and privacy protection intensity.
[0006] In the field of privacy-preserving machine learning, a covert adversary refers to an individual or entity that attempts to extract sensitive information from a machine learning model or system. These adversaries may have different capabilities, ranging from passive observers to attackers capable of actively intervening in the system operation. One of the goals of privacy protection is to ensure that even in the presence of covert adversaries, sensitive data can be protected. The key challenge in federated learning is to ensure the privacy of data during the training process while guaranteeing the integrity and accuracy of the training results. In this mode, the data owner (the client) does not want to directly expose their data to the service provider (the server), and the service provider needs to train the model without directly accessing this data.
[0007] However, existing federated learning schemes are difficult to balance between security and efficiency and cannot guarantee the correctness of the results. This is because adversaries may attempt to obtain sensitive information of the training data through various means, even if the data is encrypted during transmission and storage. For example, an adversary may try to determine whether a specific data record is used to train the model, which may lead to the leakage of personal privacy, or introduce malicious data during the training phase to affect the learning process of the model, resulting in a decrease in the accuracy of the model or the introduction of bias. Moreover, even if they cannot directly access the training data, adversaries may infer information about the training data by analyzing the learned model parameters, which is called a model stealing attack, or may try to determine whether a specific data record is used to train the model, which may lead to the leakage of personal privacy. If encryption technology is used, an adversary may try to discover the weaknesses of the encryption algorithm or implement a cryptographic attack to break the encryption, and may even exploit the security vulnerabilities of the system to bypass the privacy protection measures and directly access or tamper with the data.
[0008] Therefore, when considering the behavior of adversaries, the federated learning schemes designed for the current semi-honest adversaries cannot meet the higher-level security requirements, and the malicious adversary schemes often require higher overhead. Therefore, existing federated learning schemes are difficult to balance between security and efficiency and cannot guarantee the correctness of the results. Summary of the Invention
[0009] The object of the present invention is to provide a cloud-edge collaborative delegated learning method against covert adversaries, so as to overcome the problems that the existing delegated learning schemes designed for semi-honest adversaries cannot meet higher-level security requirements, and malicious adversary schemes often require higher overheads.
[0010] Based on the first main aspect of the present invention, there is provided a cloud-edge collaborative delegated learning method against covert adversaries, which proposes a verifiable cloud-edge collaborative delegated learning framework against covert adversaries that can guarantee the output results. The framework includes at least a delegator, an edge server, and a cloud server; where:
[0011] Delegator: It is the data owner. Due to limited computing power, the data is delegated to the server for calculation. To protect the privacy of the data, the delegator randomly and secretly shares the data and then sends it to two non-colluding covert edge servers.
[0012] Edge server: As a server providing computing functions, it calculates addition and multiplication through private data interaction, completes a linear regression model, and sends the results to the cloud server. Among them, the edge server acts as a covert adversary and hopes to successfully cheat without being detected.
[0013] Cloud server: The cloud server sums and verifies the intermediate data sent by each group of edge servers. If the results are equal or within the error range, the cloud server stores the current data. If the sum results are not equal and exceed the error range, the guaranteed output delivery protocol is executed. The cloud server sends replacement information to the edge cluster, sums and verifies the recalculated results, and jointly detects malicious computing parties with the edge cluster. After being tested by the cloud server, the model finally obtains the correct learning result.
[0014] Based on the above framework, the method includes at least the following steps:
[0015] The delegator randomly and secretly shares the data and sends the shared data to two covert edge servers;
[0016] The edge server executes the model calculation task, including the calculation of the linear regression model, and sends the calculation results to the cloud server;
[0017] The cloud server sums and verifies the intermediate data sent by the edge server. If the results are consistent or within the preset error range, the data is stored; otherwise, the guaranteed output delivery protocol is executed, and malicious computing parties are jointly detected and located with the edge cluster.
[0018] Furthermore, a verifiable secure computing module based on the re-sharing mechanism is adopted for data sharing and interactive computing between the edge server and the cloud server. Through the secret sharing mechanism, data is split into multiple shares and stored dispersedly on different servers. In this way, even if a part of the data is leaked or accessed by unauthorized parties, the attacker cannot restore the original data, thus effectively protecting the privacy of the data.
[0019] The verifiable secure computing module allows the data owner or other participating parties to verify the correctness of the computing process without knowing the specific computing content, which increases the transparency and credibility of the computing process. The malicious adversary detection module can identify and deal with possible malicious behaviors of the edge server, such as disconnection or intentional incorrect computing, thus ensuring the integrity and correctness of the computing results. Through probabilistic verification, even in the face of potential covert adversary attacks, it can ensure that the principal obtains the correct output results, thus reducing the risks brought by server fraud behaviors.
[0020] In terms of design, this module aims to be as efficient as existing semi-honest attacker schemes while providing higher security, which means that stronger security guarantees can be obtained without significantly sacrificing performance. Even when some edge servers are dishonest, through the cloud-edge collaborative working mechanism and the guaranteed output delivery protocol, it can ensure that the principal gets the correct computing results.
[0021] Moreover, this module design takes into account the assumption of covert adversaries, which makes the scheme more in line with the security requirements in the real world and can adapt to a wider range of security threat models. By providing verifiable computing processes and results, it enhances the trust of data owners in the server, which is particularly important for scenarios such as delegated learning that rely on third-party computing resources. This module supports complex computing tasks, such as the training and prediction of linear regression models, while ensuring the security and privacy of the computing.
[0022] As a further preferred solution, this method includes a covert security model, where the covert attacker has limited deception ability and will face significant costs once the deception behavior is discovered. By introducing the concept of covert adversaries, the model can cope with more complex security threats, where the adversary has both the ability to deceive and is worried about the costs after being discovered, which increases the risk of the adversary's malicious behaviors. The covert security model is more in line with the security scenarios in the real world, where the attacker may have both motives and concerns, and this model can better simulate and defend against attacks in the real world. Since the covert attacker will face significant costs once the deception behavior is discovered, this potential deterrence can prevent the adversary from carrying out malicious behaviors, thus protecting the security of the system.
[0023] In the framework of federated learning, the covert security model can increase the data owner's trust in the server because they know that any fraudulent behavior will be identified and punished. This model provides a flexibility that allows for a trade-off between different security requirements and costs to adapt to different application scenarios. Compared with the fully malicious adversary model, the covert adversary model may reduce false positives because it takes into account the possible reasonable behaviors of the adversary.
[0024] In the cloud-edge collaborative environment, the covert security model can promote cooperation between the edge server and the cloud server because they jointly bear the responsibility of detecting and defending against malicious behaviors. Compared with the malicious adversary model that requires higher overhead, the covert security model provides a more efficient solution because it does not need to strictly verify each possible malicious behavior. Even when facing attacks from covert adversaries, this model can ensure that the federated learning model produces correct output results, thus enhancing the robustness of the entire system.
[0025] On the other hand, by introducing the covert security model, it can stimulate the innovation of new security technologies and algorithms to better address security challenges in the real world.
[0026] As a further preferred solution, this method adopts a probabilistic verification cloud-edge collaboration scheme, combining the traditional federated learning framework with the cloud-edge collaborative computing mode. In this mode, the edge server performs model calculations, and the cloud server assists in verification and storage. The probabilistic verification cloud-edge collaboration scheme can, through probabilistic verification, increase the supervision of the edge server's computing behavior without significantly increasing the computational and communication overhead, thereby improving the security of the overall system. And the edge server performs model calculations without knowing the original data, which can protect the privacy of the data owner because the data will not be leaked during the calculation process.
[0027] On the other hand, the cloud-edge collaborative computing mode allows the edge server to share the computing tasks, which can reduce the burden on the cloud server and improve the computing efficiency. Since the edge server is closer to the data source, it can reduce the distance for data to be transmitted to the cloud server, thereby reducing the communication cost and latency. Moreover, the edge server can be scaled as needed to adapt to the growing computing requirements, while the cloud server can focus on data verification and storage. The cloud server's assistance in verification can serve as a secondary guarantee for the computing results of the edge server to ensure the accuracy and integrity of the computing results.
[0028] Therefore, this scheme can flexibly adjust the workload between the edge server and the cloud server according to the actual computing requirements and resource conditions. Through probabilistic verification, even if a covert adversary tries to deceive, there is a high probability of being detected, thus protecting the system from attacks by covert adversaries.
[0029] Moreover, by enhancing security and privacy protection, the trust of users in using this technical solution for data processing and analysis can be increased. The edge server can utilize its computing power to execute more localized tasks, while the cloud server can focus on providing higher-level services such as data storage and long-term analysis.
[0030] As a further preferred solution, the interaction between the principal, the edge server, and the cloud server is carried out through a secure communication channel; and the security is defined by comparing the real and ideal interactions to ensure the security of the protocol. The secure communication channel ensures that the data transmitted between the principal, the edge server, and the cloud server is not intercepted and interpreted by unauthorized third parties. Through the secure communication channel, it can be verified that the data has not been tampered with during transmission, ensuring the integrity of the data. Moreover, the secure communication channel usually includes an authentication mechanism to ensure that all parties involved in the communication are the expected participants. The secure communication protocol can provide proof of the source and integrity of the message, preventing any party from denying the operations that have been performed. By ensuring the security of the communication, the trust of the principal in the edge server and the cloud server is enhanced.
[0031] By comparing the real and ideal interactions, the security of the protocol can be formally verified to ensure that there are no security vulnerabilities. In the case where the security has been formally verified, the risks brought by security vulnerabilities can be significantly reduced. Even in the case where the edge server may not be completely trustworthy, the privacy data of the principal can be protected from being leaked.
[0032] This technical solution allows the intensity of the secure communication channel and the verification mechanism to be adjusted according to different security requirements and application scenarios. As the system scale expands, this solution can be extended to maintain the security of the communication. Moreover, since the security is achieved through a standardized communication channel and formal verification, the cost of maintenance and update is low. Within a known secure communication framework, developers can develop and deploy new applications and services with more confidence. Even in the face of potential attacks, the system can still operate stably because the security has been strictly verified. More complex security policies, such as multi-factor authentication and encrypted communication, can be implemented on this basis.
[0033] In some embodiments, as a further preferred solution, the cloud-edge collaborative delegated learning method against covert adversaries designs a probabilistically verifiable secure computing building block and a malicious adversary detection protocol to increase the probability of correct computation by the covert edge server; wherein, the probabilistically verifiable secure computing building block includes a secret sharing protocol and a verifiable secure computing building module. Through the probabilistic verification mechanism, the motivation of the edge server to correctly execute the computing task is increased, thereby improving the overall computing correctness. The probabilistically verifiable secure computing building block and the malicious adversary detection protocol provide an additional security layer for the system, making it more difficult for covert adversaries to carry out undetected malicious behaviors. Even in the presence of covert adversaries, it can ensure that the delegator obtains the correct computation result, reducing the risk brought by fraudulent behaviors.
[0034] Moreover, probabilistic verification, as an incentive mechanism, encourages the edge server to perform computations honestly because dishonest behaviors have a high probability of being detected. The secret sharing protocol ensures the privacy of data during the computation process and does not leak sensitive information even among edge servers. Even if the behaviors of some edge servers are untrustworthy, the system can still continue to operate and provide correct outputs through probabilistic verification and the malicious adversary detection protocol.
[0035] Through probabilistic verification, computing resources and security measures can be more effectively allocated because security resources can be adjusted according to the probability of adversary behaviors. Compared with traditional malicious adversary models, probabilistic verification generally requires lower computational and communication overheads, thereby improving the overall computing efficiency. This enables the scheme to adapt to different security requirements and threat models and provide appropriate security levels for different types of adversaries.
[0036] On the other hand, users can trust the cloud-edge collaborative delegated learning framework more because the framework can provide a high level of security guarantee. The scheme supports complex computing tasks such as the training and prediction of machine learning models while ensuring the security and privacy of the computation process. Through theoretical analysis and experimental data, the effectiveness of the scheme can be verified, providing a solid foundation for practical applications. In addition, the scheme considers the possible behavior patterns of adversaries in the real world and provides a security solution that better meets the actual application requirements.
[0037] Based on the second main aspect of the present invention, a system for the aforementioned cloud-edge collaborative delegated learning method against covert adversaries, the system at least includes the following modules:
[0038] A delegator module for performing random secret sharing of data; through random secret sharing of data, the security of the data is increased, so that even if the data is intercepted during transmission, the original data content cannot be restored.
[0039] Edge server module for performing model calculation tasks; the edge server module is responsible for performing model calculation tasks, which can utilize the advantages of edge computing to reduce latency and speed up data processing.
[0040] Cloud server module for data summation and verification; the cloud server module performs data summation and verification, ensuring the correctness of the calculation results and enhancing the trustworthiness of the system.
[0041] Communication channel module for implementing secure communication. The communication channel module implements secure communication, protecting the privacy and integrity of data during transmission and preventing potential eavesdropping and tampering.
[0042] The above system design takes into account the existence of covert adversaries, provides a more realistic and powerful security protection measure, and adapts to a wider range of security threat models. The modular design of the system allows for flexible addition or upgrade of individual modules as needed to adapt to different application scenarios and computing requirements. Through cloud-edge collaboration, more computing tasks can be completed on the edge server, reducing the dependence on cloud server resources and thus lowering the overall operating cost. Even if some edge servers fail or are attacked by adversaries, the system can still ensure the completion of computing tasks through the collaborative work of other modules.
[0043] Based on the third main aspect of the present invention, a computer-readable storage medium, when the program is executed, implements the aforementioned cloud-edge collaborative delegated learning method against covert adversaries.
[0044] Advantages and beneficial effects of the present invention:
[0045] The present invention proposes a cloud-edge collaborative delegated learning scheme integrating probabilistic verification, which combines the traditional delegated learning framework with the cloud-edge collaborative computing mode. In this scheme, the edge server is responsible for performing the model calculation tasks, while the cloud server provides auxiliary probabilistic verification and data storage services. Through this cloud-edge collaborative working mechanism, the present invention realizes an efficient and secure delegated learning solution.
[0046] The present invention constructs a series of verifiable secure computing modules and malicious adversary detection modules based on the re-sharing mechanism. In view of the situation where the adversary's disconnection or malicious calculation behavior during the delegated learning process leads to the abortion of model training, a cloud-edge collaborative verifiable four-party server learning protocol ensuring output reachability is designed. It not only ensures the privacy protection of the data of the delegator, effectively guards against the fraud behavior of the server, but also ensures that the delegated learning model can still produce correct output results even when facing attacks from malicious adversaries.
[0047] When considering the behavior of adversaries, the present invention finds that the existing federated learning schemes designed for current semi - honest adversaries cannot meet the higher - level security requirements, while malicious adversary schemes often incur higher overheads. Therefore, the present invention first applies the concept of covert adversaries to privacy - preserving machine learning, which improves security compared to the semi - honest adversary model and is more efficient than the verification method for malicious adversaries.
[0048] The present invention conducts comparative experiments on the linear regression model for verification, and through theoretical analysis and experimental data, proves the effectiveness of the proposed scheme of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following - described drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, obtaining other drawings based on these drawings still falls within the scope of the present invention.
[0050] Figure 1 Shows the composition of the cloud - edge collaborative federated learning framework in an example of the present method. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0051] The following will describe in detail the preferred embodiments of the present invention with reference to the accompanying drawings to more clearly understand the purpose, features, and advantages of the present invention. It should be understood that the embodiments shown in the drawings are not a limitation on the scope of the present invention, but only to illustrate the essential spirit of the technical solution of the present invention.
[0052] In the following description, for the purpose of explaining various disclosed embodiments, certain specific details are set forth to provide a thorough understanding of the various disclosed embodiments. However, those skilled in the relevant art will recognize that the embodiments can be practiced without one or more of these specific details. In other instances, well - known devices, structures, and technologies associated with the present application may not be shown or described in detail to avoid unnecessarily obscuring the description of the embodiments.
[0053] References to "an embodiment" or "one embodiment" throughout the specification mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment. Thus, appearances of "in an embodiment" or "in one embodiment" throughout the specification are not necessarily all referring to the same embodiment. Additionally, the particular features, structures, or characteristics may be combined in any manner in one or more embodiments.
[0054] The cloud-edge collaborative federated learning method against covert adversaries of the present invention has an execution subject including but not limited to at least one of electronic devices such as a server, a terminal, etc. that can be configured to execute the method provided in the embodiments of the present application. In other words, the cloud-edge collaborative federated learning method against covert adversaries can be executed by software or hardware installed on a terminal device or a server device, and the software can be a blockchain platform.
[0055] The server includes but is not limited to: a single server, a server cluster, a cloud server, or a cloud server cluster, etc. The server can be an independent server or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.
[0056] 1. Composition of the federated learning framework
[0057] As Figure 1 shown, the present invention proposes a federated learning framework composed of 3 main entities, including a delegator D, an edge server ES, and a cloud server C. The specific responsibilities of each unit are as follows:
[0058] Delegator: It is the data owner. Due to limited computing power, the data is entrusted to the server for calculation. To protect the privacy of the data, the delegator randomly and secretly shares the data and then sends it to two groups of non-colluding covert edge servers.
[0059] Edge server: As a server providing computing functions, it calculates addition and multiplication through private data interaction, completes a linear regression model, and sends the result to the cloud server. Among them, the edge server acts as a covert adversary and hopes to successfully cheat without being discovered.
[0060] Cloud server: The cloud server sums and verifies the intermediate data sent by each group of edge servers. If the results are equal or within the error range, the cloud server stores the current data. If the sum results are not equal and exceed the error range, the guaranteed output delivery protocol is executed. The cloud server sends replacement information to the edge cluster, sums and verifies the recalculated results, and jointly detects malicious computing parties with the edge cluster. After being tested by the cloud server, the model finally obtains the correct learning result.
[0061] 2. Security model
[0062] In the solution, the present invention uses a covert security model where a covert attacker has the ability to deceive, but will pay a huge price once discovered. In the covert security where a covert attacker may deceive, the present invention uses an output-reachable protocol so that the delegator can obtain the correct result regardless of whether the edge server is honest.
[0063] In addition, it is assumed that the edge servers used for sharing data and interactive computing are untrusted but do not collude. This means that the edge servers will not disclose more information to each other than the protocol. This assumption is more realistic because edge servers may be deployed on competitive service providers. Since the delegator is the initiator of the protocol, it hopes to obtain the correct calculation result. Therefore, it can be reasonably assumed that the delegator is always honest and acts as a trusted third party to generate the auxiliary random number in the Beaver triple. In addition, the present invention assumes that there is a secure communication channel between entities.
[0064] In the work of the present invention, security is defined by comparing real and ideal interactions. The ideal target function of the protocol of the present invention is as follows: real function:
[0065] Parameters: Delegators D1, D2,..., D n and edge servers E1, E2,..., E4, cloud server C.
[0066] Upload data: The delegators upload private data to the edge servers, the edge servers upload the intermediate calculation results to the cloud server, and the cloud server verifies and stores the intermediate calculation results.
[0067] Calculation: The edge servers and the cloud server jointly calculate and verify the function, and finally send the obtained output result to the delegator.
[0068] Definition 1: The present invention states that a protocol is secure if for any real adversary A, there always exists a probabilistic polynomial-time simulator S (acting as an ideal adversary) to simulate the interaction in the real environment, such that the view of the real environment and the view of the ideal environment are computationally indistinguishable.
[0069] Lemma 1: For the security requirements under the universal composability framework, if each sub-protocol in the solution is fully simulatable and the ideal function is indistinguishable from the real function, then the solution composed of multiple sub-protocols is securely simulatable and satisfies the ideal function and the real function.
[0070] 3. Design Goals
[0071] The goal of the delegated learning protocol is to ensure the correctness of the server's calculation result, the privacy of the data of the delegated party, and the efficiency of the learning process.
[0072] Correctness: The main objective of the present invention is to ensure the correctness of linear regression training and prediction. In the framework of the present invention, the learner can correctly train the linear regression results and make correct predictions from the randomly scrambled data of the entrusted party. At the same time, the four edge servers of the present invention can compare the results obtained from pairwise training and prediction, thereby further ensuring the reliability of the results.
[0073] Privacy protection: The present invention establishes a privacy protection scheme based on secret sharing, and its main purpose is to protect the privacy of the data of the entrusting party. The data is randomly divided into two groups of different random numbers and sent to two groups of edge servers for model training respectively. Without collusion, the covert edge servers cannot obtain the real data of the entrusting party, and the covert edge servers will not collude in the protocol setting.
[0074] Efficiency: The covert edge servers perform calculations according to the protocol requirements. To prevent malicious calculations by the covert edge servers, the present invention designs a probabilistically verifiable secure computing building block and a malicious adversary detection protocol. The covert edge servers will increase the probability of correct calculation under deterrence. The anti-covert adversary scheme proposed by the present invention is as efficient as the existing semi-honest attackers, and its training efficiency is significantly higher than that of malicious attackers.
[0075] 4. Model of the Present Invention
[0076] In the model of the present invention, there are an entrusting party, an edge cluster (including multiple edge servers and playing a management role), and a cloud server. The edge cluster will allocate edge servers and cloud servers for collaborative computing to complete the entrusted learning process.
[0077] 4.1 Probabilistically Verifiable Secure Computing Building Block
[0078] The building block of the entrusted learning of the present invention includes a secret sharing protocol and a verifiable secure computing building module.
[0079] Protocol 1: Secret Sharing Protocol
[0080] Input: The entrusting party has data X, Y;
[0081] Output: The edge servers respectively receive the shared data <x> i , <y> i , and a multiplication triple i , <v> i , <f> i , <V′> i <F′> i ;
[0082] S1: The client randomly and uniformly generates \(x_0, y_0, x′_0, y′_0\in Z\), and calculates \(x_1 = X - x_0, y_1 = Y - y_0, x′_1 = X - x′_0, y′_1 = Y - y′_0\) such that <x> 0=x0, <x> 1=x′0, <x> 2=x1, <x> 3=x′1, <y> 0=y0, <y> 1=y′0,
[0083] <y> 2=y1, <y>3 = y1', satisfying X = <x> 0+ <x> 2= <x> 1+ <x> 3∪ <x> 0≠ <x> 1≠ <x> 2≠ <x>3 and Y = <y> 0+ <y> 2= <y> 1+ <y> 3∪ <y> 0≠ <y> 1≠ <y> 2≠ <y>3;
[0084] S2: The consignor will generate <x> i , <y> i Send them to two different groups of edge servers (E0, E2) and (E1, E3) respectively;
[0085] S3: The principal first randomly generates data U and V, calculates U×V = F, U T ×V′ = F′, and obtains multiplication triples through secret sharing and re-secret sharing i , <v> i , <f> i , <V′> i <F′> i , where \(i\in\{0, 1, 2, 3\}\) and satisfies ( 0+ 2)×( <v> 0+ <v> 2)= <f> 0+ <f> 2,( 1+ 3)×( <v> 1+ <v> 3)= <f> 1+ <f>3, meanwhile
[0086] 0+ 2= 1+ 3, <v> 0+ <v> 2= <v> 1+ <v> 3, <f> 0+ <f> 2= <f> 1+ <f>3;
[0087] S4: The consignor will generate the i , <v> i , <f> i , <V′> i <F′> i i ∈ {0, 1, 2, 3} are sent to the edge server E i and the cloud server respectively.
[0088] Protocol II: Verifiable addition and subtraction computation of edge servers
[0089] S1: The edge server E i computes and sends it to the cloud server for verification with probability ρ, or directly sends it to the principal without verification with probability 1 - ρ;
[0090] S2: When the cloud server receives the computation result from the edge server with probability ρ, it compares the sum of the results of each group of edge servers;
[0091] S3: If: then
[0092] the cloud server computes and sends to the principal.
[0093] Otherwise: The cloud server determines that there is at least one malicious edge server and invokes the malicious adversary detection protocol.
[0094] S4: When the principal receives the result from the edge server with probability 1 - ρ, the principal computes to obtain the computation result;
[0095] Protocol III: Verifiable multiplication computation of edge servers
[0096] Input: Edge server E i receives the private data sent by the principal <x> i , <y> i i , <v> i , <f> i ;
[0097] Output: The principal obtains the result of the privacy calculation f = x × y;
[0098] S1: The edge server first invokes the addition and subtraction verifiable protocol to calculate <m> i = <x> i - i ,
[0099] <n> i = <y> i - <v> i , obtain
[0100] S2: Edge server E i , i∈{0,1} for local computing For edge server E i , i∈{2,3} for local computing
[0101] And send with probability ρ To the cloud server for verification, or with probability 1 - ρ, without verification, directly send to the consignor;
[0102] S3: When the cloud server receives the computing results from the edge server with probability ρ, it compares the sum of the results of each group of edge servers;
[0103] S4: If: Then
[0104] The cloud server computes Send To the edge server for continued learning or send To the consignor;
[0105] Otherwise: The cloud server determines that there is at least one malicious edge server and invokes the malicious adversary detection protocol;
[0106] S5: When the consignor receives the results from the edge server with probability 1 - ρ, it computes To obtain the final result.
[0107] 4.2. Cloud-edge collaborative malicious attacker detection model
[0108] In the delegated learning model, the goal of the delegator is to obtain the training results of the model. However, many existing works often terminate the protocol and stop training after detecting malicious attackers. The consignor cannot detect and punish malicious computers, and the consignor cannot obtain the final training results.
[0109] The present invention proposes a malicious attacker detection protocol for resisting covert attackers, which detects inconsistent training results while discovering deceivers. The malicious adversary detection protocol will provide deterrence to covert edge servers with the ability to deceive but afraid of being discovered, making them more willing to perform honest computing. However, most covert edge servers will perform honest computing, and there are also a few edge servers that cheat. This protocol can ensure that even if a few edge servers cheat, the privacy-preserving delegated learning model can still output correct training results and ensure output delivery. The cloud server compares and verifies the intermediate training results of the edge servers with probability ρ.
[0110] For example, ρ = 0.1 means that the parameter is updated 10 times, and the edge server will compare and verify once. If the result is within the error range, the cloud server will store the correct intermediate training result. Otherwise, the cloud server will call the malicious adversary detection protocol. The specific malicious adversary detection protocol is as follows:
[0111] Protocol Four: Malicious Adversary Detection Protocol
[0112] Input: The cloud server sends inconsistent summation results Or the edge server sends a report message;
[0113] Output: The cloud server detects the malicious edge server E i ;
[0114] S1: The cloud server uses the latest known values and the dataset to train for j now -j before rounds. The cloud server will compare the training values with the sum of the training values of the two groups of edge servers for comparison;
[0115] S2: If: Then, this means that there is a malicious edge server among a group of edge servers that interact with each other; the cloud server only needs to use private data to detect cheaters in a group of edge servers with calculation errors;
[0116] S3: The cloud server compares The cloud server will detect the malicious edge server E i ; In this case, there is a group of honest edge servers that can continue training to ensure the output;
[0117] S4: If: This means that there are malicious edge servers in both groups of edge servers; the cloud server needs to use privacy data to detect cheaters in each group of edge servers.
[0118] S5, the cloud server compares The cloud server will detect the malicious edge server E i ; In this case, the honest edge servers in each group will be merged, and the cloud server will re - distribute the private data to two honest edge servers to continue training, thus ensuring the output.
[0119] In the malicious adversary detection protocol, when the cloud server verifies the calculation results of the edge server with probability ρ and the results are inconsistent, the cloud server will call the malicious adversary detection protocol.
[0120] First, the cloud server uses the latest known values and the dataset to train for j now -j before Round, the cloud server will send the training value and compare it with the sum of the training values of the two sets of edge servers. If this means that there is a malicious edge server in a group of edge servers that interact with each other. The cloud server only needs to use private data to detect cheaters in a group of edge servers with computational errors.
[0121] In this case, there is a group of honest edge servers that can continue training to ensure the output. If this means that there are malicious edge servers in both sets of edge servers. The cloud server needs to use privacy data to detect cheaters in each set of edge servers. In this case, the honest edge servers in each set will be combined, and the cloud server will re - distribute the private data to two honest edge servers to continue training, thus ensuring the output.
[0122] 4.3 Privacy - Preserving Linear Regression Scheme for Cloud - Edge Collaboration
[0123] The present invention applies the above - mentioned verifiable secure computing protocol to the linear regression model of machine learning. In linear regression, there are r training data samples x, where each x contains d features, corresponding to the output label y. It allows the present invention to obtain a function g such that g(x)=y. Assume that the function g is a linear function, g(x)=xw. The goal of linear regression is to find the best w to minimize the difference between g(x) and y.
[0124] The present invention trains a privacy - preserving four - party verifiable linear regression model, where the delegator first performs additive secret sharing on the data x, y and multiplicative triples i , <v> i , <f> i , <V′> i <F′> i , and then send the shared data to each edge server. The edge server receives the private data of the delegator <x> i , <y> i Initialization Then perform collaborative computing. To prevent edge servers from cheating, an intermediate verification method of the cloud server is added during the training process. The specific protocol is as follows:
[0125] Protocol Five: SGD Linear Regression Protocol with Concealed Privacy Protection
[0126] Input: Edge server E i Have <x> i , <y> i , initialization Meanwhile
[0127] Output: Edge server E i Output separately And the cloud server outputs w;
[0128] S1: The edge server invokes the verifiable addition calculation protocol Obtain the calculation result This step will be verified by the cloud server;
[0129] S2: Execute for j = 0,..., l;
[0130] S3: E i Select a small batch of data
[0131] S4: Edge server E i Invoke verifiable multiplication calculation Obtain the calculation result
[0132] S5: Edge server E i Calculate the error value
[0133] S6: Edge server E i Invoke verifiable multiplication calculation Obtain the calculation result
[0134] S7: Edge server E i Update for i ∈ {0, 1, 2, 3}
[0135] S8: End the loop;
[0136] S9: The cloud server finally obtains w l .
[0137] In the verifiable privacy - preserving linear regression protocol, the delegator first generates a large number of multiplicative triples (U, V, F, V′, F′), and then randomly divides its own data x, y, initial parameters and the generated multiplicative triples (U, V, F, V′, F′) into private data i , <v> i , <f> i , <V′> i <F′> i , For E i i ∈ {0, 1, 2, 3} to use, and finally sent to the edge server E i . These private data are used by the edge server E i for interactive training.
[0138] To improve the training efficiency, the present invention uses mini-batch stochastic gradient descent.
[0139] Based on the SGD linear regression protocol with covert privacy protection, first in step S1, the edge server E i invokes a verifiable secure addition and subtraction protocol to hide private data <x> i and i , this step is verified by the cloud server. In step S2, the edge server determines the mini-batch data used for each training Steps S4 - S5 of the protocol are the forward propagation phase of training. In step S4, the edge server first invokes verifiable secure multiplication calculation and uses it in the multiplication triples <v> i Hide private data <w> i , and then the result <z> i Send them to the cloud server respectively. The cloud server verifies whether the sum of the two sets of edge servers is equal with probability ρ. If it is then send the result to the edge server to continue training. If the cloud server does not verify the result with probability 1 - ρ, it will directly start the next step. In step S5, the edge servers calculate the difference between the predicted value and the true value respectively to obtain the loss value.
[0140] Steps S6 - S7 are the backpropagation stage of the linear regression model. In step S6, the edge server calls the verifiable privacy - preserving multiplication again to calculate the gradient backward to obtain the value of. The cloud server also performs probability verification during this process.
[0141] In step S7, the edge server updates the parameter In a cloud - edge collaborative manner, through iterative repetition of mini - batch stochastic gradient descent in steps S2 - S7, finally the cloud server aggregates the training result w l .
[0142] 5. Security Analysis
[0143] In the covert security model of the present invention, the adversary is allowed to corrupt at most one or two of the four edge servers. To prove that the protocol of the present invention is secure, it is sufficient to prove that given its input and output, the view of the corrupted party is simulatable. In particular, the present invention uses the following definitions.
[0144] Definition 1: If there exists a probabilistic polynomial - time simulator S that can generate the view generated by the adversary a in the real world, and the view is computationally indistinguishable from its true view, then the present invention says the protocol is secure.
[0145] The present invention proves the security of the protocol in the universal composability framework according to Security Definition 1 and Lemma 1 in the above - mentioned Section 2 security model. In the covert security model, at most one or two of the four edge servers are allowed to be corrupted. And due to the existence of verifiable threats, the edge servers are more willing to choose honest computing. To prove the security of the protocol of the present invention, the present invention uses the following lemmas.
[0146] Lemma 1: If all sub - protocols of a protocol are perfectly simulatable, then it is perfectly simulatable.
[0147] Lemma 2: If a random element r is uniformly distributed over Zn and independent of any variable x ∈ Zn
[0148] The present invention refers to the proofs of Lemma 1 and Lemma 2 in this section. Since there is a local computation process in the protocol of the present invention, the present invention mainly proves the security during the interaction process below.
[0149] Theorem 1: The secret sharing building block in the verifiable four-party secure delegation computing protocol is secure in the covert adversary model.
[0150] Proof: For the output of the delegator D in the protocol is
[0151] output D = ( <x> i , <y> i , i , <v> i , <f> i ), i ∈ {0, 1, 2, 3}, the input of edge server si during the protocol execution is Meanwhile, the ideal function and the real function being simulated are indistinguishable. According to Theorem 1 and Lemma 1 in this section, it can be obtained that this building block is secure in the covert adversary model.
[0152] Theorem 2: The verifiable addition calculation building block in the protocol is secure in the covert adversary model.
[0153] Proof: The verifiable addition calculation building block can be divided into local calculation and interactive verification, where the local calculation part is executed by each edge server itself.
[0154] Since the input of edge server Si is uniformly random, according to Lemma 2, it can be known that the output result obtained after addition and subtraction calculations is also uniformly random and simulatable. Meanwhile, the ideal function and the real function being simulated are indistinguishable. According to Theorem 1 and Lemma 1 in this section, it can be obtained that this building block is secure in the covert adversary model.
[0155] Theorem 3: The verifiable multiplication calculation building block in the protocol is secure in the covert adversary model.
[0156] Proof: The present invention analyzes the view of a group of mutually interacting edge servers (E0, E2). The view of edge server E0 is where
[0157] <m> 2= <x> 2- 2, <n> 2= <y> 2- <v>2 is the data sent from edge server E2 to edge server E0. Similarly, the view of edge server E2 is where <m> 0= <x> 0- 0,
[0158] <n> 0= <y> 0- <v>0 is the data that interacts with the edge server E0. Additionally, the edge servers E0 and E2 calculate respectively according to the building block requirements Execute the output The edge server E2 executes the output
[0159]
[0160] Since the above data is uniformly randomly generated by the principal, its calculation results are also uniformly random. Therefore, it can be simulated by the simulator, and the ideal function after simulation is indistinguishable from the real function. The security proof process for the edge server combination (E1, E3) is the same as above.
[0161] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.< / v> < / y> < / n> < / x> < / m> < / v> < / y> < / n> < / x> < / m> < / f> < / v> < / y> < / x> < / z> < / w> < / v> < / x> < / f> < / v> < / y> < / x> < / y> < / x> < / f> < / v> < / v> < / y> < / n> < / x> < / m> < / f> < / v> < / y> < / x> < / f> < / v> < / f> < / f> < / f> < / f> < / v> < / v> < / v> < / v> < / f> < / f> < / v> < / v> < / f> < / f> < / v> < / v> < / f> < / v> < / y> < / x> < / y> < / y> < / y> < / y> < / y> < / y> < / y> < / y> < / x> < / x> < / x> < / x> < / x> < / x> < / x> < / x> < / y> < / y> < / y> < / y> < / x> < / x> < / x> < / x> < / f> < / v> < / y> < / x>
Claims
1. A cloud-edge collaborative delegated learning method against hidden adversaries, characterized in that: The method proposes a cloud-edge collaborative delegated learning framework that guarantees output results and is verifiable against hidden adversaries, the framework at least includes a delegate, an edge server, and a cloud server; and the method at least includes the following steps: The client shares the data randomly and secretly, and sends the shared data to two groups of hidden edge servers; The edge server performs model calculation tasks, including the calculation of the linear regression model, and sends the calculation results to the cloud server; The cloud server sums and verifies the intermediate data sent by the edge server, and stores the data if the results are consistent or within the preset error range; Otherwise, it executes the guaranteed output delivery protocol and works with the edge cluster to detect and locate the malicious computing party; The method designs a probabilistically verifiable secure computing building block and a malicious adversary detection protocol to improve the probability of correct calculation of the hidden edge server; wherein the probabilistically verifiable secure computing building block includes a secret sharing protocol and a verifiable secure computing building module; The edge servers used for sharing data and interactive computing are untrusted but not collusive.
2. The cloud-edge collaborative delegated learning method against hidden adversaries according to claim 1 is characterized in that: The secret sharing protocol includes: Input: The client has data X, Y; Output: The edge servers receive the shared data respectively <x> i , <y> i , and multiplication triples< / y> < / x> i , <v> i , <f> i ,<V′> i <F′> i ;< / f> < / v> S1: The client generates x0, y0, x′0, y′0∈Z uniformly and randomly, and calculates x1=X-x0, y1=Y-y0. x1′=Xx′0,y1′=Yy′0 so that <x> 0=x0, <x> 1=x′0, <x> 2=x1, <x> 3=x1′, <y> 0=y0, <y> 1=y′0,< / y> < / y> < / x> < / x> < / x> < / x> <y> 2=y1, <y>3=y1′, satisfying X= <x> 0+ <x> 2= <x> 1+ <x> 3∪ <x> 0≠ <x> 1≠ <x> 2≠ <x>3 and Y = <y> 0+ <y> 2= <y> 1+ <y> 3∪ <y> 0≠ <y> 1≠ <y> 2≠ <y> 3;< / y> < / y> < / y> < / y> < / y> < / y> < / y> < / y> < / x> < / x> < / x> < / x> < / x> < / x> < / x> < / x> < / y> < / y> S2: The client will generate <x> i , <y> i Send to two different sets of edge servers (E0, E2) and (E1, E3) respectively;< / y> < / x> S3: The client first randomly generates data U, V, and calculates U×V=F, U T ×V′=F′, the multiplication triplet is obtained through secret sharing and re-secret sharing i , <v> i , <f> i ,<V′> i <F′> i , i∈{0,1,2,3}, satisfying ( 0+ 2)×( <v> 0+ <v> 2)= <f> 0+ <f> 2,( 1+ 3)×( <v> 1+ <v> 3)= <f> 1+ <f> 3. At the same time< / f> < / f> < / v> < / v> < / f> < / f> < / v> < / v> < / f> < / v> 0+ 2= 1+ 3, <v> 0+ <v> 2= <v> 1+ <v> 3, <f> 0+ <f> 2= <f> 1+ <f> 3;< / f> < / f> < / f> < / f> < / v> < / v> < / v> < / v> S4: The client will generate i , <v> i , <f> i ,<V′> i <F′> i i∈{0,1,2,3} are sent to the edge server E respectively. i and cloud servers.< / f> < / v> 3. The cloud-edge collaborative delegated learning method against hidden adversaries according to claim 1 is characterized in that: The verifiable secure computation building block includes verifiable addition and subtraction computations of the edge server and verifiable multiplication computations of the edge server; wherein, The edge server's verifiable addition and subtraction calculations include: Input: Edge Server E i Receive the private data sent by the client <x> i , <y> i ;< / y> < / x> Output: The client obtains the privacy calculation result f = x ± y; S1: Edge Server E i calculate And send with probability ρ The data is verified by the cloud server, or sent directly to the client with a probability of 1-ρ without verification. S2: When the cloud server receives the computation result from the edge server with probability ρ, it compares the sum of the results of each group of edge servers; S3: If: but Cloud Server Computing And send To the client; Otherwise: the cloud server determines that there is at least one malicious edge server and invokes the malicious adversary detection protocol; S4: When the client receives the result from the edge server with a probability of 1-ρ, the client calculates Get the calculation results; The edge server's multiplication verifiable computation includes: Input: Edge Server E i Receive the private data sent by the client <x> i , <y> i i , <v> i , <f> i ;< / f> < / v> < / y> < / x> Output: The client obtains the result of the privacy calculation f = x × y; S1: The edge server first calls the addition and subtraction verifiable protocol calculation <m> i = <x> i - i , < / x> < / m> <n> i = <y> i - <v> i ,get < / v> < / y> < / n> S2: Edge Server E i ,i∈{0,1} local calculation For edge server E i ,i∈{2,3} local calculation And send with probability ρ The data is verified by the cloud server, or sent directly to the client with a probability of 1-ρ without verification. S3: When the cloud server receives the computation results from the edge server with probability ρ, it compares the sum of the results of each group of edge servers; S4: If: Cloud Server Computing send Continue learning or sending to the edge server To the client; Otherwise: the cloud server determines that there is at least one malicious edge server and invokes the malicious adversary detection protocol; S5: When the client receives the result from the edge server with probability 1-ρ, it calculates Get the final result.
4. A system for implementing the cloud-edge collaborative delegated learning method against hidden adversaries as described in any one of claims 1-3, characterized in that: The system includes at least the following modules: A client module for performing random secret sharing of data; An edge server module for performing model computing tasks; Cloud server module for data summing and verification; Communication channel module for achieving secure communication.
5. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed, it implements the cloud-edge collaborative delegated learning method against hidden adversaries as described in any one of claims 1-3.
Citation Information
Patent Citations
Deep learning data privacy protection method, system and device and medium
CN114780999A
Device for secret sharing-based multi-party computation
CN113268744A
Privacy protection neural network reasoning method based on three-party server
CN117744798A