Node permission determination method and device for distributed artificial intelligence training, computer device, storage medium and computer program product
Patent Information
- Application Number
- CN202610885596.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-18
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2046-06-18
AI Technical Summary
该类方案虽然避免了单一固定节点的风险,但其权限调整具有滞后性,无法在训练过程中实时调整节点的参与权限,恶意节点在其行为被认定之前仍可能持续影响模型性能
[0061]The aforementioned method, apparatus, computer equipment, storage medium, and computer program product for determining node permissions in distributed artificial intelligence training, wherein the method can elect an audit node from among the nodes based on the current reputation value of each node; confirm the election result, wherein the current reputation value is used to characterize the trustworthiness of the node; the confirmed audit node verifies the trustworthiness of the model update data submitted by the worker nodes, generating an audit result; update the reputation value of each worker node based on the audit result, and control the permissions of each worker node to participate in model training and verification tasks based on the updated reputation value. This scheme, by dynamically electing and confirming audit nodes, verifying model update data in real time, and dynamically updating node reputation values and permissions, effectively solves the problems of vulnerability of permission management nodes to attacks and system security relying on a single node in fixed centralized schemes by adjusting node participation permissions in real time during training. It also overcomes the shortcomings of decentralized post-event traceability schemes, such as delayed permission adjustments and the continuous impact of malicious nodes on model performance. This improves the dynamism, real-time nature, and security of node permission management in distributed artificial intelligence training, and is adaptable to training scenarios in untrusted environments.
Smart Images

Figure CN122413432B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of distributed artificial intelligence technology, and in particular to a method, apparatus, computer device, storage medium, and computer program product for determining node permissions for distributed artificial intelligence training. Background Technology
[0002] In distributed artificial intelligence training, numerous participating nodes are allowed to collaboratively train a model without sharing the original data. In such distributed training architectures, effectively managing the permissions of participating nodes is crucial to ensuring the performance and security of the global model.
[0003] Typically, a fixed centralized server manages the participation permissions of each node, determining which nodes can participate in training and whose updates can be aggregated. However, in this approach, the permission management node remains fixed for a long period. If this fixed node is attacked or malfunctions, the entire system's permission management mechanism will fail. Furthermore, this fixed permission management node is vulnerable to pre-targeted attacks by malicious nodes, making system security heavily reliant on the reliability of a few fixed nodes. Other solutions introduce decentralized post-event traceability mechanisms, recording node behavior through a distributed ledger and assigning responsibility after training is complete or a problem occurs. While this approach avoids the risks associated with a single fixed node, its permission adjustments are delayed, failing to adjust node participation permissions in real-time during training. Malicious nodes may continue to impact model performance until their actions are identified. Therefore, the node permission management methods in these technologies are insufficient in terms of dynamism, real-time performance, and security, making them unsuitable for distributed AI training scenarios in untrusted environments. Summary of the Invention
[0004] Therefore, it is necessary to address the aforementioned technical problems by providing a method, apparatus, computer device, computer-readable storage medium, and computer program product for determining node permissions in distributed artificial intelligence training scenarios applicable to untrusted environments.
[0005] Firstly, this application provides a method for determining node permissions for distributed artificial intelligence training, including:
[0006] Based on a verifiable random function and the current reputation value of each node, an audit node is elected from among the nodes as the election result; the election result is confirmed, wherein the current reputation value is used to characterize the trustworthiness of the node;
[0007] The verified audit node verifies the credibility of the model update data submitted by the working node and generates an audit result.
[0008] The reputation value of each working node is updated based on the audit results, and the permissions of each working node to participate in model training and validation tasks are controlled based on the updated reputation value.
[0009] In one embodiment, the step of electing an audit node from among the nodes based on a verifiable random function and the current reputation value of each node includes:
[0010] A random seed is generated based on a random function, and the election weight of each node is determined based on the current reputation value of each node.
[0011] The nodes are weighted and randomly sampled using the random seed, and the selected nodes are determined as the audit nodes, which is the election result.
[0012] In one embodiment, confirming the election result includes:
[0013] The validity of the proof information of the random function in the election results is verified by the existing audit nodes respectively;
[0014] Existing audit nodes that have passed verification broadcast confirmation messages of the election results;
[0015] When the cumulative number of broadcast election result confirmation messages exceeds a preset threshold as a percentage of the total number of existing audit nodes, the election result is determined to have been confirmed.
[0016] In one embodiment, the step of verifying the credibility of the model update data submitted by the working node by the confirmed audit node and generating an audit result includes:
[0017] Perform at least one of gradient anomaly detection, proof-of-work verification, and collaborative behavior analysis on the model update data to obtain the audit results;
[0018] The audit results include the anomaly types of the work nodes, which include at least one of the following: gradient anomaly, proof-of-work anomaly, and collaboration anomaly.
[0019] In one embodiment, gradient anomaly detection is performed on the model update data to generate audit results, including:
[0020] The gradient data in the model update data of each working node is grouped according to the neural network layer type, and the vector norm of the gradient of each group is calculated to obtain the low-dimensional representation vector of each working node.
[0021] Determine the median baseline vector of the low-dimensional representation vectors of all working nodes;
[0022] The anomaly score of each working node is determined based on the normalized distance between the low-dimensional representation vector of each working node and the median baseline vector.
[0023] Based on the anomaly scores of each working node, an audit result is generated, which includes whether each working node exhibits gradient anomalies.
[0024] In one embodiment, generating audit results based on the anomaly scores of each working node includes:
[0025] Determine the median of the anomaly scores for all working nodes;
[0026] Based on the anomaly score of each working node and the median of the anomaly scores, determine the median of the absolute deviation of all working nodes.
[0027] A dynamic threshold is determined based on the median of the abnormal scores and the median of the absolute deviations.
[0028] Work nodes whose abnormal scores exceed the dynamic threshold are identified as gradient anomalies to obtain the audit results.
[0029] In one embodiment, performing proof-of-work verification on the model update data to obtain the audit result includes:
[0030] Randomly select a subset of validation data from the public validation dataset;
[0031] The model performance metrics are recalculated on the validation subset to obtain the local performance metrics.
[0032] The performance metrics declared by each worker node are obtained from the model update data. Based on the local performance metrics and the performance metrics declared by each worker node, the confidence score of each worker node is determined.
[0033] Work nodes with confidence scores below a preset threshold are identified as proof-of-work anomalies to obtain the audit results.
[0034] In one embodiment, the performance metrics include loss value and accuracy; the confidence score is obtained by combining the relative error of loss and the absolute error of accuracy.
[0035] In one embodiment, verification of the collaborative behavior analysis on the model update data to obtain the audit results includes:
[0036] The gradient data in the model update data of each working node is normalized to obtain the gradient direction vector of each working node.
[0037] Determine the first similarity between every two gradient direction vectors;
[0038] For each worker node, calculate the average similarity between the gradient direction vector of that worker node and the gradient direction vectors of all other worker nodes, and use this as the average similarity of that worker node.
[0039] Work node pairs with a first similarity greater than a first threshold and an average similarity between the two work nodes less than a second threshold are identified as suspected collaborative pairs.
[0040] The suspected collaborative pair and its associated working nodes are identified as having collaborative anomalies to obtain the audit results.
[0041] In one embodiment, determining the associated working node of the suspected collaboration pair as having collaboration anomalies to obtain the audit result includes:
[0042] The transitive closure algorithm is used to merge all suspected cooperative pairs into related cooperative groups.
[0043] Work nodes belonging to the same collaboration group are uniformly identified as collaboration anomalies to obtain the audit results.
[0044] In one embodiment, updating the reputation value of each working node based on the audit results includes:
[0045] Based on the anomaly type of each work node in the audit results, determine the deduction score for each work node;
[0046] The updated reputation value is determined based on the historical reputation value of each working node, the deduction score, and the time decay coefficient.
[0047] In one embodiment, based on the updated reputation value, the permissions of each working node to participate in model training and validation tasks are controlled, including:
[0048] If the reputation value of the target worker node is greater than or equal to the first reputation threshold, the target worker node is determined as a trusted node, and the trusted node is configured to be allowed to participate normally in model training and audit node election.
[0049] If the reputation value of the target working node is less than the first reputation threshold and greater than or equal to the second reputation threshold, the target working node is identified as a suspicious node, and the first weight of the model update data submitted by the target working node in the global model aggregation is reduced to the second weight.
[0050] If the reputation value of the target worker node is less than the second reputation threshold, the target worker node is identified as a malicious node, and the malicious node is prohibited from participating in model training and audit node election.
[0051] In one embodiment, the method further includes:
[0052] Based on the updated reputation value and the status of each working node, a list of safe nodes for the current training round is generated, which includes node information that has not been identified as a malicious node.
[0053] Once the list of secure nodes has been confirmed by multiple signatures from the auditing nodes, it is broadcast to all nodes.
[0054] Secondly, this application also provides a node permission determination device for distributed artificial intelligence training, comprising:
[0055] The dynamic election module is used to elect an audit node from among the nodes based on a verifiable random function and the current reputation value of each node, and the election result is used to confirm the election result through a Byzantine fault tolerance protocol, wherein the current reputation value is used to characterize the trustworthiness of the node;
[0056] The multi-dimensional audit module is used by the confirmed audit nodes to verify the credibility of the model update data submitted by the working nodes and generate audit results.
[0057] The reputation management and policy execution module is used to update the reputation value of each working node based on the audit results, and to control the permissions of each working node to participate in model training and verification tasks based on the updated reputation value.
[0058] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the method described in the first aspect.
[0059] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in the first aspect.
[0060] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the method described in the first aspect.
[0061] The aforementioned method, apparatus, computer equipment, storage medium, and computer program product for determining node permissions in distributed artificial intelligence training, wherein the method can elect an audit node from among the nodes based on the current reputation value of each node; confirm the election result, wherein the current reputation value is used to characterize the trustworthiness of the node; the confirmed audit node verifies the trustworthiness of the model update data submitted by the worker nodes, generating an audit result; update the reputation value of each worker node based on the audit result, and control the permissions of each worker node to participate in model training and verification tasks based on the updated reputation value. This scheme, by dynamically electing and confirming audit nodes, verifying model update data in real time, and dynamically updating node reputation values and permissions, effectively solves the problems of vulnerability of permission management nodes to attacks and system security relying on a single node in fixed centralized schemes by adjusting node participation permissions in real time during training. It also overcomes the shortcomings of decentralized post-event traceability schemes, such as delayed permission adjustments and the continuous impact of malicious nodes on model performance. This improves the dynamism, real-time nature, and security of node permission management in distributed artificial intelligence training, and is adaptable to training scenarios in untrusted environments. Attached Figure Description
[0062] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0063] Figure 1 This is a schematic diagram of a distributed training architecture in one embodiment;
[0064] Figure 2 This is a flowchart illustrating a node permission determination method for distributed artificial intelligence training in one embodiment.
[0065] Figure 3 This is a flowchart illustrating the dynamic election process in one embodiment;
[0066] Figure 4 This is a flowchart illustrating gradient anomaly detection during the auditing process in one embodiment;
[0067] Figure 5 This is a flowchart illustrating the verification process of proof of work during an audit, as shown in one embodiment.
[0068] Figure 6 This is a flowchart illustrating the collaborative behavior analysis process during an audit, as shown in one embodiment.
[0069] Figure 7This is a flowchart illustrating the dynamic reputation management process in one embodiment;
[0070] Figure 8 This is a flowchart illustrating the security policy execution module in one embodiment;
[0071] Figure 9 This is a structural block diagram of a node permission determination device for distributed artificial intelligence training in one embodiment;
[0072] Figure 10 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0073] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0074] In distributed artificial intelligence training, numerous participating nodes are allowed to collaboratively train a model without sharing the original data. In such distributed training architectures, effectively managing the permissions of participating nodes is crucial to ensuring the performance and security of the global model.
[0075] Typically, a fixed centralized server manages the participation permissions of each node, determining which nodes can participate in training and whose updates can be aggregated. However, in this approach, the permission management node remains fixed for a long period. If this fixed node is attacked or malfunctions, the entire system's permission management mechanism will fail. Furthermore, this fixed permission management node is vulnerable to pre-targeted attacks by malicious nodes, making system security heavily reliant on the reliability of a few fixed nodes. Other solutions introduce decentralized post-event traceability mechanisms, recording node behavior through a distributed ledger and assigning responsibility after training is complete or a problem occurs. While this approach avoids the risks associated with a single fixed node, its permission adjustments are delayed, failing to adjust node participation permissions in real-time during training. Malicious nodes may continue to impact model performance until their actions are identified. Therefore, the node permission management methods in these technologies are insufficient in terms of dynamism, real-time performance, and security, making them unsuitable for distributed AI training scenarios in untrusted environments.
[0076] To address the aforementioned issues, embodiments of this application provide a method, apparatus, computer device, storage medium, and computer program product for determining node permissions in distributed artificial intelligence training.
[0077] The node permission determination method for distributed artificial intelligence training provided in this application embodiment can be applied to a node permission determination device or computer device for distributed artificial intelligence training. The node permission determination device for distributed artificial intelligence training can be a functional module or functional entity in the computer device used to implement the above-mentioned node permission determination method for distributed artificial intelligence training.
[0078] For example, the aforementioned computer equipment may be participating node devices (worker nodes, audit nodes, etc.) in distributed artificial intelligence training, cluster (distributed artificial intelligence training cluster) management servers, cloud servers, edge computing nodes, or federated learning participating terminal devices. Such computer equipment has the capabilities of distributed communication, consensus verification, model gradient calculation and anomaly detection, node reputation scoring and permission policy execution, and can independently or collaboratively complete the entire process of audit node election, consensus confirmation of election results, model update data credibility verification, node reputation value update and training and verification task permission control.
[0079] For example, Figure 1 This is a schematic diagram of a distributed artificial intelligence training cluster architecture in one embodiment of this application. The cluster adopts a physically peer-to-peer, logically separated architecture. The physical layer consists of multiple nodes of equal status that are interconnected. Figure 1 The diagram shows N nodes; there is no fixed centralized permission management node. The logic layer dynamically divides the nodes into worker nodes and audit nodes. Figure 1 The worker nodes and audit nodes shown are exemplary divisions, and in practice, this division can change dynamically. Worker nodes are used to execute local model training and submit model update data. Audit nodes are elected by the cluster based on a verifiable random function and node reputation value, and confirmed by network consensus. Audit nodes do not participate in model training in the current training round; they are only used to verify the credibility of model update data, update node reputation values, and control the permissions of nodes participating in training and verification tasks, thus achieving decentralized and dynamic node permission management.
[0080] Physical peering means that all nodes in the cluster are equal in terms of hardware status and network access permissions, and there is no pre-defined central management node. Logical separation means that the role of a node can change dynamically in each round of training; the same node can be a worker node in one round and can be elected as an audit node in the next round.
[0081] In one exemplary embodiment, such as Figure 2 As shown, a method for determining node permissions for distributed artificial intelligence training is provided, the method including the following steps 201 to 204:
[0082] 201. Based on the current reputation value of each node, elect an audit node from among the nodes as the election result.
[0083] The current reputation score mentioned above is used to characterize the trustworthiness of a node. The reputation score can be a rating used to characterize the trustworthiness of a node's historical behavior. For example, the higher the reputation score, the greater the probability of being selected as an audit node.
[0084] In some embodiments, a random seed can be generated first based on a random function, and the election weight of each node can be determined according to the current reputation value of each node. Then, the random seed can be used to perform weighted random sampling on each node, and the selected node is determined as the audit node, which is used as the election result.
[0085] The random seed can be generated by a verifiable random function to ensure that the election process is fair, unpredictable, and reproducible.
[0086] In some implementations, the aforementioned random function can be a verifiable random function (VRF). A verifiable random function is a function that generates random numbers and comes with verifiable proofs, ensuring that the random results are immutable and publicly verifiable.
[0087] In one alternative implementation, the election weight can be calculated from the node's reputation value; the higher the reputation, the greater the weight and the higher the probability of being selected.
[0088] In the above weighted random sampling process, while ensuring randomness, nodes with higher reputation values are more likely to be selected, thus balancing fairness and security.
[0089] 202. Confirm the election results.
[0090] In this process, the audited nodes, once confirmed, will not participate in model training in the current training round, but will participate in the verification task. This verification task includes subsequent trustworthiness verification processes.
[0091] In some implementations, election results can be confirmed using Byzantine fault tolerance protocols.
[0092] The aforementioned Byzantine fault tolerance protocol is a distributed consensus mechanism that enables the network to reach a consistent result even when some nodes are abnormal or malicious. After the election results are confirmed through the Byzantine fault tolerance protocol, it can be guaranteed that the majority of existing audit nodes in the entire network recognize the list of audit nodes for this round, which helps to reduce the possibility of malicious nodes forging election results.
[0093] In this way, audit nodes do not participate in training, which can prevent audit nodes from using their verification permissions to tamper with the model and ensure verification neutrality.
[0094] In some embodiments, the process of confirming the election results through the Byzantine fault tolerance protocol may include, but is not limited to: having existing audit nodes verify the legality of the proof information of the verifiable random function in the election results; having existing audit nodes that have passed the verification broadcast the election result confirmation message; and determining that the election results have been confirmed when the cumulative number of broadcast election result confirmation messages accounts for more than a preset threshold.
[0095] In the process of verifying the legality of the proof information, the audit node can verify whether the random seed is legally generated by a verifiable random function and has not been tampered with.
[0096] The aforementioned broadcast confirmation message is used by the auditing node to declare to the network its acceptance of the election results.
[0097] The aforementioned preset threshold can be set to exceed two-thirds of the total number of audit nodes to meet the security requirements of Byzantine fault tolerance. For example, the preset threshold could be 80%, 85%, etc.
[0098] 203. The verified audit nodes perform credibility verification on the model update data submitted by the working nodes and generate audit results.
[0099] Among them, model update data can refer to data such as model gradients and parameter changes obtained after local training on the worker node.
[0100] The aforementioned credibility verification may include checking whether the model update is genuine, reasonable, and free from malicious tampering.
[0101] For example, the audit results above can record which work nodes are normal, which work nodes are abnormal, and the types of abnormalities.
[0102] In some embodiments, the audit node may perform at least one of the following verifications on the model update data: gradient anomaly detection, proof-of-work verification, and collaborative behavior analysis, to obtain the audit results.
[0103] The audit results include the anomaly types of work nodes, which include at least one of the following: gradient anomaly, proof of work anomaly, and collaboration anomaly.
[0104] For example, gradient anomalies can refer to a significant deviation of the numerical distribution of model updates from the normal range, which may be malicious poisoning. Proof-of-work anomalies can refer to a discrepancy between the model performance declared by worker nodes and the actual computation results; collaboration anomalies can refer to highly similar gradients of multiple nodes that differ significantly from the overall network, which is judged as collusion.
[0105] 204. Update the reputation value of each working node based on the audit results, and control the permissions of each working node to participate in model training and validation tasks based on the updated reputation value.
[0106] In the process of updating the reputation value of each working node based on the audit results, the reputation of normal nodes can be maintained or improved, and the reputation of abnormal nodes can be deducted according to the type of abnormality. Then, based on the latest reputation, nodes can be allowed, restricted or prohibited from participating in subsequent training and election.
[0107] In the above embodiments, distributed permission management that does not rely on a fixed center is achieved by dynamically electing audit nodes, reaching consensus through Byzantine fault tolerance, having audit nodes perform multi-dimensional trusted verification, and updating reputation values and permissions in real time based on verification results. It can identify malicious nodes in real time during training, which helps reduce the risk of single point of failure and single point of attack, and improves the security, real-time performance and robustness of distributed artificial intelligence training.
[0108] Steps 201 and 202 above involve the dynamic election process of the audit node, which will be explained in detail below.
[0109] In one exemplary embodiment, such as Figure 3 The diagram shown is a flowchart of a dynamic election process in one embodiment, which includes the following steps 301 to 305:
[0110] 301. Generate a random seed based on a random function, and determine the election weight of each node based on the current reputation value of each node.
[0111] Among them, the random function can generate unpredictable, tamper-proof and publicly verifiable random seeds, ensuring a fair and just starting point for the election.
[0112] The aforementioned election weights are positively correlated with the node's reputation value. The higher the reputation value, the greater the corresponding election weight, making nodes with trustworthy historical behavior more likely to be elected as audit nodes, thus improving the security and rationality of the election process.
[0113] For example, the calculation of election weight takes into account both the node's historical reputation and online activity. The specific calculation formula (1) is as follows:
[0114] (1)
[0115] In the above calculation formula (1), For nodes The electoral weight; For nodes The current reputation value is used to characterize the credibility of a node's historical behavior; For nodes Current online time; The maximum online time among all nodes; The online duration weighting coefficient is used to balance the impact of reputation and online duration on election results. The arrows in the formula indicate that the values on the right are assigned to the variables on the left. .
[0116] For example, It can be set to 0.2.
[0117] The election weights calculated using the above formula are positively correlated with a node's reputation score and online duration. In other words, nodes with higher reputation scores and longer online durations have greater election weights. This makes nodes with reliable historical behavior and long-term stable online presence more likely to be elected as audit nodes, thus improving the security and rationality of the election process.
[0118] 302. Use a random seed to perform weighted random sampling on each node, and determine the selected node as the audit node, which is the election result.
[0119] Weighted random sampling uses a random seed for fairness and election weights for security. This ensures the election process is unmanipulated while making it easier for high-reputation nodes to be selected, achieving a balance between randomness and trustworthiness. The random seed also guarantees that election results obtained under the same sampling rules are reproducible, facilitating subsequent network-wide verification.
[0120] For example, it can be verified that the random function is implemented based on the Ed25519 elliptic curve, using the concatenated value of the block height and timestamp as input to generate a random seed, i.e. .
[0121] The Ed25519 elliptic curve cryptography algorithm is a high-performance, high-security signature algorithm based on elliptic curve cryptography.
[0122] The random seed ensures the election process is fair, unpredictable, and reproducible, reducing the risk of malicious nodes manipulating the election results. Weighted random sampling uses this random seed as the basis for fairness and the election weights calculated in step 301 as the basis for security. K nodes are selected from all nodes according to their weights as audit nodes (for example, k can be set to 5). This ensures the election process is unmanipulated while making nodes with high reputation and long online duration more likely to be selected, achieving a balance between randomness and trustworthiness.
[0123] 303. The existing audit nodes shall verify the legality of the proof information of the random function in the election results.
[0124] The proof information of the random function is used to verify whether the random seed is legally generated. By verifying this proof information, the audit node can confirm that the election process has not been tampered with and the election list has not been maliciously constructed, thus ensuring the compliance of the election process.
[0125] The aforementioned existing audit nodes refer to the audit nodes that have been elected and confirmed in the current network.
[0126] 304. Existing nodes that have passed verification broadcast the election result confirmation message.
[0127] The broadcast confirmation message is used to indicate that the current audit node has completed the legitimacy verification and recognizes the election results, providing a basis for reaching a consensus across the entire network.
[0128] 305. When the cumulative number of broadcast election result confirmation messages exceeds the preset threshold, the election result is confirmed.
[0129] The preset threshold is usually set to exceed two-thirds of the total number of audit nodes to meet the security requirements of Byzantine fault tolerance and ensure that the election results are still credible, tamper-proof, and recognized by the entire network when there are a small number of abnormal audit nodes.
[0130] In the above embodiments, the election fairness can be guaranteed by using a random function, and the current election results (such as the elected audit nodes) can be confirmed by existing audit nodes, which can dynamically complete the election of audit nodes and eliminate the risk of single point of attack and single point of failure caused by fixed verification nodes.
[0131] In one exemplary embodiment, such as Figure 4 The diagram shown illustrates the gradient anomaly detection process during auditing in one embodiment. The process includes steps 401 to 404:
[0132] 401. Group the gradient data in the model update data of each working node according to the neural network layer type, calculate the L2 norm of the gradient of each group, and obtain the low-dimensional representation vector of each working node.
[0133] The neural network layers can include different types such as convolutional layers and fully connected layers. Grouping them allows for a more accurate reflection of the gradient distribution characteristics of different structural layers. By calculating the L2 norm of each group of gradients, high-dimensional gradient information can be transformed into low-dimensional values, forming a low-dimensional representation vector that can represent the overall characteristics of the gradient of the current node.
[0134] In some possible implementations, gradient data are grouped according to the type of neural network layer, such as convolutional layer or fully connected layer. The vector norm of each group of gradient direction vectors is calculated, and the norms of each group are combined in the order of the layers to form the low-dimensional representation vector of the working node.
[0135] For example, the vector norm mentioned above can be the L2 norm, and the low-dimensional representation vector can be as follows: norm_vec[i] = [||g_conv||, ||g_fc||, ...], where norm_vec[i] represents the low-dimensional representation vector, and ||g_conv|| and ||g_fc|| represent the L2 norm of each group of gradients, respectively.
[0136] 402. Determine the median baseline vector of the low-dimensional representation vectors of all working nodes.
[0137] The median baseline vector is obtained by taking the median of the low-dimensional representation vectors of all working nodes in each dimension. Compared with the mean, it is more resistant to outlier interference and can stably represent the gradient update trend of normal nodes in this round of training.
[0138] In some possible implementations, low-dimensional representation vectors of all working nodes are collected, the median is calculated for each dimension of the vector, and the medians of each dimension are combined to obtain the median baseline vector.
[0139] 403. Determine the anomaly score of each working node based on the normalized distance between the low-dimensional representation vector of each working node and the median baseline vector.
[0140] First, the Euclidean distance between the low-dimensional representation vector and the median baseline vector is calculated. Then, the norm of the baseline vector is used to normalize the distance, and the normalized distance is used as the anomaly score. The higher the anomaly score, the more obvious the gradient of the working node deviates from the overall normal trend.
[0141] In some possible implementations, the Euclidean distance between the low-dimensional representation vector of the working node and the median baseline vector is used, and then the distance is normalized using the L2 norm of the median baseline vector. The normalized result is used as the anomaly score.
[0142] For example, for each node, the Euclidean distance between its low-dimensional representation vector and the median baseline vector can be calculated, and this distance can be normalized using the L2 norm of the median baseline vector to obtain the final normalized distance as the anomaly score: distance[i] = ||norm_vec[i] - baseline||2. Here, distance[i] represents the anomaly score, ||norm_vec[i] represents the low-dimensional representation vector, and baseline||2 represents the L2 norm of the median baseline vector.
[0143] 404. Based on the anomaly scores of each working node, generate audit results, which include whether each working node has gradient anomalies.
[0144] In some embodiments, the process of generating audit results based on the anomaly scores of each working node may include, but is not limited to: determining the median of the anomaly scores of all working nodes; determining the median of the absolute deviation of all working nodes based on the anomaly scores of each working node and the median of the anomaly scores; determining a dynamic threshold based on the median of the anomaly scores and the median of the absolute deviation; and identifying working nodes whose anomaly scores exceed the dynamic threshold as gradient anomalies to obtain audit results.
[0145] For example, the above dynamic threshold is calculated according to the following formula (2).
[0146] threshold=Median(distance)+2.5×MAD(distance) (2)
[0147] In the above formula (2), threshold represents the dynamic threshold, Median(distance) represents the median of the outlier score, and MAD(distance) represents the median of the absolute deviation.
[0148] In some possible implementations, the median of the abnormal scores and the median of the absolute deviation are calculated first, and then the dynamic threshold is calculated according to the above formula (2). Nodes whose scores exceed the dynamic threshold are marked as gradient anomalies.
[0149] In the above embodiments, by using hierarchical gradient grouping, L2 norm dimensionality reduction, median baseline and dynamic threshold judgment, malicious behaviors such as gradient forgery and gradient poisoning can be accurately identified. Without disclosing the privacy of training data, the credibility verification of model update data can be achieved, and the accuracy and robustness of anomaly detection can be improved.
[0150] In one exemplary embodiment, such as Figure 5 The diagram shown illustrates the workflow for verifying proof of work during an audit process in one embodiment. This workflow includes the following steps 501 to 504:
[0151] 501. Randomly select a subset of validation data from the public validation dataset.
[0152] The public validation dataset is a standardized dataset shared by all worker nodes and audit nodes, ensuring data consistency and immutability. It is used to uniformly evaluate model performance, serving as a public validation benchmark in distributed training. Randomly selecting a validation subset reduces the likelihood of worker nodes constructing false results in advance for fixed validation samples, thus improving the anti-cheating capabilities of the validation process.
[0153] In some possible implementations, 10% to 20% of the data samples are randomly selected from the public verification dataset shared by all nodes as a verification subset for independently verifying the performance results claimed by the nodes.
[0154] Since the audit nodes only use a small number of extracted samples for verification, the overall computational cost is significantly reduced while ensuring the accuracy of the verification.
[0155] 502. Recalculate the model performance metrics on the validation subset to obtain the local performance metrics.
[0156] The audit node reconstructs the local model based on the global model parameters and the model update data submitted by the worker node. The reconstruction method is to synthesize the global model and the gradient update submitted by the node to restore the complete model state after the node submits the update. Forward inference calculation is performed on the validation subset to obtain the true and objective loss value and accuracy as local performance indicators. This process only performs inference calculation and does not perform backpropagation training, so the computational cost is much lower than that of the complete training process.
[0157] The aforementioned audit nodes achieve independent and reliable verification of the updated content submitted by the nodes through model reconstruction and forward inference, without relying on intermediate results provided by the nodes themselves.
[0158] For example, the data submitted by the working node to the audit node includes: gradient data Δθ, claimed loss value loss_claimed, claimed accuracy acc_claimed, and Ed25519 digital signatures for the above content, to ensure that the data has not been tampered with and that its source is trustworthy, and to prevent the node from denying its claims.
[0159] In some possible implementations, the local model is reconstructed based on the global model parameters and the model updates submitted by the worker nodes. Forward propagation is then performed on the validation subset to obtain the true loss value and accuracy as local performance metrics.
[0160] 503. Obtain the performance metrics declared by each worker node from the model update data, and determine the confidence score of each worker node based on the local performance metrics and the performance metrics declared by each worker node.
[0161] The performance metrics include loss value and accuracy. The relative error of loss reflects the deviation between the local loss and the claimed loss, while the absolute error of accuracy reflects the deviation between the local accuracy and the claimed accuracy. The confidence score is calculated based on a combination of the relative error of loss and the absolute error of accuracy; a higher score indicates a more realistic node claim.
[0162] For example, the confidence score can be calculated according to the following formula (3):
[0163] confidence=1−(loss_error+acc_error) / 2 (3)
[0164] Here, loss_error represents the relative error of loss, and acc_error represents the absolute error of accuracy.
[0165] For example, the relative error of the loss is loss_error = |loss_local - loss_claimed| / loss_local. Here, loss_local represents the local loss value, and loss_claimed represents the loss value claimed by the worker node.
[0166] For example, the absolute error of accuracy is acc_error = |acc_local - acc_claimed|. Here, acc_local represents the local accuracy, and acc_claimed represents the accuracy claimed by the worker node.
[0167] By comparing local performance metrics with the performance metrics declared by each worker node, abnormal behaviors such as false reporting, concealment, and failure to conduct real training can be accurately identified.
[0168] 504. Work nodes with confidence scores below a preset threshold are identified as work proof anomalies in order to obtain audit results.
[0169] The preset threshold is used to determine whether a node has actually completed training. If the confidence level is too low, it means that the worker node has not honestly performed local computation and has misrepresented performance results, and the model update data it submits is unreliable.
[0170] For example, the preset threshold can be set to 0.95, 0.9, 0.8, etc.
[0171] In some possible implementations, assuming a preset threshold of 0.95, the confidence score is compared with 0.95, and if it is lower than this threshold, the work certificate is judged to be abnormal.
[0172] For work nodes whose work proofs are deemed abnormal, the audit node can directly determine that the submitted model update is invalid and exclude it from the global model update process.
[0173] In the above embodiments, by combining random sampling verification, local recalculation and confidence assessment, it is possible to effectively detect whether the working nodes have false reporting, forged calculation results, or failure to actually perform training, thereby ensuring that the model update data is true and valid and guaranteeing the computational correctness of the distributed training process.
[0174] In one exemplary embodiment, such as Figure 6 The diagram shown illustrates a flowchart of collaborative behavior analysis during an audit process in one embodiment. The flowchart includes the following steps 601 to 605:
[0175] 601. Normalize the gradient data in the model update data of each working node to obtain the gradient direction vector of each working node.
[0176] Normalization is used to eliminate the influence of gradient magnitude and retain only the direction information of model updates, so that anomaly detection focuses more on the consistency of update trends between nodes and improves the effect of collaborative cheating detection.
[0177] In some possible implementations, the model update data for each worker node includes its gradient data. Based on this gradient data, the gradient direction vector of the worker node can be obtained. Dividing the gradient data of the worker node by its own L2 norm completes L2 normalization, resulting in a unit gradient direction vector that only represents the update direction. For example, dir[i] = Δθ[i] / ||Δθ[i]||. Here, dir[i] represents the normalized gradient direction vector of the worker node, and Δθ[i] represents the gradient direction vector of the worker node.
[0178] The purpose of the above normalization is to eliminate amplitude differences and compare only directions. Furthermore, after normalization, the calculation of cosine similarity degenerates into a simple inner product, making the calculation more efficient.
[0179] 602. Determine the first similarity between every two gradient direction vectors.
[0180] The first similarity is used to measure the proximity of two gradient directions. The higher the similarity, the more consistent the model update trends of the two nodes are.
[0181] For example, the first similarity can be cosine similarity.
[0182] In some possible implementations, the cosine similarity is calculated for any two gradient direction vectors, and this similarity is used as the first similarity.
[0183] 603. For each working node, calculate the average similarity between the gradient direction vector of that working node and the gradient direction vectors of all other working nodes, and use this average similarity as the average similarity of that working node.
[0184] The average similarity is used to represent the consistency of the current node's update trend with the overall network. The lower the average similarity, the more the node deviates from the overall normal update pattern.
[0185] In some possible implementations, for each node, the similarity of its gradient directions with all other nodes is calculated and the average value is taken as the average similarity.
[0186] For example, the aforementioned similarity can be cosine similarity.
[0187] 604. Work node pairs whose first similarity is greater than the first threshold and whose average similarity between the two work nodes is less than the second threshold are identified as suspected cooperative pairs.
[0188] Among them, node pairs that meet this condition are highly similar to each other but significantly deviate from the overall network, which is a typical characteristic of coordinated malicious behavior.
[0189] The first similarity mentioned above can also refer to cosine similarity.
[0190] For example, the first threshold can be set to 0.95 and the second threshold can be set to 0.7.
[0191] In some possible implementations, if the first similarity between two nodes is greater than 0.95 and their average similarity is less than 0.7, then they are marked as a suspected cooperative pair.
[0192] 605. Identify the associated work nodes of suspected collaborative pairs as collaborative anomalies to obtain audit results.
[0193] In some implementations, the process of identifying the associated work nodes of the suspected collaborative pair as collaborative anomalies to obtain the audit result may include, but is not limited to: merging all suspected collaborative pairs into a group using a transitive closure algorithm to form an associated collaborative group; and uniformly identifying work nodes belonging to the same collaborative group as collaborative anomalies to obtain the audit result.
[0194] The transitive closure algorithm can merge suspected cooperative pairs into complete cooperative groups, and all nodes in cooperative groups with a size of 2 or greater are marked as cooperative anomalies.
[0195] In some possible implementations, the transitive closure algorithm is used to merge all suspected cooperative pairs, and all nodes in groups of size 2 or larger are identified as cooperative anomalies.
[0196] In one exemplary implementation, collaborative identification and grouping are performed on the model update data submitted by each working node. First, suspected collaborative pairs are constructed between nodes. A suspected collaborative pair is a set of binary relations used to describe which two nodes have a high degree of similarity. Specifically, the cosine similarity of the model update data corresponding to any two nodes is calculated sequentially, and the average similarity between the two nodes is obtained. If the cosine similarity between two nodes is higher than a similarity threshold, and the average similarity between the two nodes is lower than the average threshold, then the two nodes are recorded as a suspected collaborative pair.
[0197] For example, the following initial suspected cooperative pairs are obtained in this embodiment:
[0198] (1) Node A and Node B form a suspected cooperative pair with a cosine similarity of 0.97 and an average similarity of less than 0.7.
[0199] (2) Node B and node C form a suspected cooperative pair with a cosine similarity of 0.96 and an average similarity of less than 0.7.
[0200] (3) Node A and Node C form a suspected cooperative pair with a cosine similarity of 0.95 and an average similarity of less than 0.7.
[0201] Furthermore, based on the suspected cooperative pairs, cooperative groups are constructed. A cooperative group is a set that merges all directly or indirectly connected nodes. Based on the above suspected cooperative pairs, connected component merging is performed. Nodes A, B, and C are directly connected to each other in pairs. Therefore, nodes A, B, and C are merged into the same cooperative group, which can be represented as: Cooperative group = {A, B, C}, with a group size of 3.
[0202] The above methods can accurately identify groups exhibiting collaborative behavior from a large number of work nodes, providing a basis for subsequent identification of abnormal groups and access control.
[0203] In the above embodiments, by using gradient direction similarity, group average similarity and dual threshold judgment, multi-node collaborative cheating behavior can be accurately identified, making up for the deficiency that single anomaly detection cannot identify collusion attacks, and significantly improving the security of distributed training in complex adversarial environments.
[0204] In one exemplary embodiment, such as Figure 7 The diagram shown illustrates a dynamic reputation management process (i.e., updating reputation values) in one embodiment, which includes the following steps 701 to 702:
[0205] 701. Determine the deduction score for each work node based on the anomaly type of each work node in the audit results.
[0206] Different anomaly types correspond to different levels of deduction scores. Gradient anomalies, work certificate anomalies, and collaboration anomalies can be set with progressively increasing deduction scores. If a collaboration anomaly is detected, additional penalty scores will be applied to strengthen the constraint effect on collusive and malicious behavior.
[0207] In some possible implementations, based on the type of anomaly in the audit results, a corresponding deduction score is determined according to preset rules, imposing higher penalties on collaborative abnormal behavior.
[0208] 702. Determine the updated reputation value based on the historical reputation value of each working node, the deduction score, and the time decay coefficient.
[0209] The updated reputation value will be synchronized to the distributed ledger to achieve consistency across the entire network, and a historical behavior decay mechanism will be introduced to prevent nodes from having their reputation permanently reduced due to accidental anomalies, while ensuring that malicious nodes cannot quickly recover their reputation.
[0210] In some possible implementations, the updated reputation value is obtained by subtracting the current round's deduction score from the historical reputation value, and then synchronized to all nodes across the network.
[0211] For example, the formula for updating the reputation value is shown in formula (4) below:
[0212] Rnew=αRold+(100−∑penalty)(1−α)−βIcollusion (4)
[0213] In formula (4), Rnew is the updated reputation value of the working node; Rold is the historical reputation value of the working node; α=0.95 is the time decay coefficient, used to balance the influence of historical reputation and current behavior, and to ensure the continuity and stability of the reputation value; ∑penalty is the sum of the penalty scores corresponding to each anomaly in the current round. For example, each penalty score can be set to penalty=[20, 50, 30], which corresponds to the basic deduction scores of gradient anomaly, proof of work anomaly, and collaboration anomaly, respectively; β=1.5 is the collaboration penalty factor, used to impose additional penalties on collaborative malicious behavior and increase the cost of malicious attacks; Icollusion∈{0,1} is the collaborative malicious flag, which takes the value of 1 when the node is judged to be a collaborative anomaly, and takes the value of 0 otherwise; the updated reputation value is constrained to Rnew∈[0,100], to ensure that the reputation value is always within the effective range of 0 to 100.
[0214] In some possible implementations, the historical reputation value, current penalty score, and collaborative tag are substituted into the above formula to calculate the updated reputation value, which is then written into the distributed ledger and synchronized to all nodes in the network.
[0215] In the above embodiments, a differentiated reputation value update mechanism based on anomaly type is used to incentivize trusted nodes and punish abnormal nodes, enabling the reputation system to be real-time, fair, and secure, and providing data support for dynamic control of node permissions.
[0216] In one exemplary embodiment, such as Figure 8 The diagram shown is a flowchart of a security policy execution module (such as controlling the permissions of each working node to participate in model training and verification tasks) in one embodiment. The flowchart includes the following steps 801 and 802:
[0217] 801. Based on the updated reputation value, determine the node type of each working node.
[0218] Among them, the node types are divided into trusted nodes, suspicious nodes and malicious nodes according to the reputation score range, and different types correspond to different security levels and access control policies.
[0219] In some possible implementations, the updated reputation value is compared with a preset first reputation threshold and a second reputation threshold to classify nodes into three types: trusted nodes, suspicious nodes, and malicious nodes.
[0220] 802. Based on the node type of each working node, control the permissions of each working node to participate in model training and validation tasks.
[0221] The access control system implements a tiered strategy based on reputation level. Trusted nodes participate normally, suspicious nodes have their weight reduced, and malicious nodes are isolated, thereby ensuring security while maintaining the stable operation of the training process.
[0222] In some possible implementations, if the reputation value of the target worker node is greater than or equal to a first reputation threshold, the target worker node is determined as a trusted node, and the trusted node is configured to be allowed to participate normally in model training and audit node election.
[0223] For example, the first reputation threshold mentioned above can be 90 points. Trusted nodes are configured to participate normally in model training and audit node election, in which case the gradient aggregation weights are normal weights.
[0224] In some possible implementations, if the reputation value of the target worker node is less than the first reputation threshold but greater than or equal to the second reputation threshold, the target worker node is identified as a suspicious node, and the first weight of the model update data submitted by the target worker node in the global model aggregation is reduced to the second weight.
[0225] For example, the second reputation threshold mentioned above can be 60 points.
[0226] For example, the first weight can be reduced to the second weight, such as from 0.8 to 0.5. For suspicious nodes, their audit node election probability can be reduced, thus placing them into an observation period.
[0227] In some possible implementations, if the reputation value of the target worker node is less than the second reputation threshold, the target worker node is identified as a malicious node, and the malicious node is prohibited from participating in model training and audit node election.
[0228] Among them, malicious nodes are prohibited from participating in model training and audit node election, their submitted gradients are not included in aggregation, and they are isolated for a preset time, during which they cannot participate in system activities.
[0229] Optionally, after the quarantine period, malicious nodes can be restored to a suspicious state, and their reputation score and full permissions can only be gradually restored after several consecutive rounds of no abnormal behavior.
[0230] In some possible implementations, full permissions are granted to trusted nodes; gradient aggregation weights are reduced and election probabilities are limited for suspicious nodes; training and election are prohibited for malicious nodes and isolation is implemented; reputation and permissions are gradually restored according to rules after the isolation period expires.
[0231] In the above embodiments, a three-level node type classification and dynamic permission control are implemented based on reputation value, which can restrict suspicious nodes and isolate malicious nodes in real time and accurately, while ensuring that trusted nodes can participate in training normally, thus achieving a balance between system security, training efficiency and node participation fairness.
[0232] In some embodiments of this application, a list of safe nodes for the current training round can be generated based on the updated reputation value and the status of each working node. The list of safe nodes includes information on nodes that have not been identified as malicious nodes. The list of safe nodes is then broadcast to all nodes after being confirmed by multiple signatures of the auditing nodes.
[0233] The aforementioned secure node list refers to a list of trusted nodes that have not been identified as malicious, selected based on the updated reputation values and operational status of each working node. Its purpose is to clarify the scope of nodes with legitimate participation permissions in this round of model training and validation. Multi-party signature confirmation refers to multiple audit nodes in the audit node set digitally signing the secure node list to verify its authenticity and legitimacy, preventing tampering or forgery.
[0234] In some possible implementations, after the node auditing process for this round of training is completed, the system filters out nodes with normal reputation and no malicious behavior based on the updated reputation value of each worker node and their online and normal response status, and generates a list of safe nodes. The safe node is then sent to the corresponding audit node for multi-party signature. After the signature verification is successful, the audit node broadcasts the list to all nodes in the distributed training network. All nodes determine the range of nodes that can participate in the model training and verification tasks in this round based on the broadcast list of safe nodes.
[0235] In the above embodiments, generating a safe node list based on reputation value and node status can accurately filter malicious and abnormal nodes, improving the overall security of the training task. Through the existing multi-party signature confirmation mechanism of audit nodes, the authority and immutability of the safe node list can be guaranteed, preventing illegal nodes from entering the training process. Broadcasting the confirmed safe node list to all nodes can realize the synchronization of access control, allowing each node to quickly identify the legitimate participants, improving training efficiency while further ensuring the stability and reliability of the model training process.
[0236] It should be understood that although the steps in the flowcharts of the above embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0237] Based on the same inventive concept, this application also provides a node permission determination device for distributed artificial intelligence training, which implements the node permission determination method for distributed artificial intelligence training described above. The solution provided by this device is similar to the implementation described in the above method. Therefore, the specific limitations in one or more embodiments of the node permission determination device for distributed artificial intelligence training provided below can be found in the limitations of the node permission determination method for distributed artificial intelligence training described above, and will not be repeated here.
[0238] In one exemplary embodiment, such as Figure 9 As shown, a node permission determination device for distributed artificial intelligence training is provided, comprising:
[0239] The dynamic election module 901 is used to elect an audit node from among the nodes based on the current reputation value of each node, and to confirm the election result, wherein the current reputation value is used to characterize the trustworthiness of the node;
[0240] The multi-dimensional audit module 902 is used by the confirmed audit node to verify the credibility of the model update data submitted by the working node and generate audit results.
[0241] The reputation management and policy execution module 903 is used to update the reputation value of each working node based on the audit results, and to control the permissions of each working node to participate in model training and verification tasks based on the updated reputation value.
[0242] In some embodiments, the dynamic election module 901 is specifically used to: generate a random seed based on a random function, and determine the election weight of each node based on the current reputation value of each node;
[0243] The nodes are weighted and randomly sampled using the random seed, and the selected nodes are determined as the audit nodes, which is the election result.
[0244] In some embodiments, the dynamic election module 901 is specifically used for:
[0245] The existing audit nodes verify the legality of the proof information of the verifiable random function in the election results.
[0246] Existing audit nodes that have passed verification broadcast confirmation messages of the election results;
[0247] When the cumulative number of broadcast election result confirmation messages exceeds a preset threshold as a percentage of the total number of existing audit nodes, the election result is determined to have been confirmed.
[0248] In some embodiments, the multi-dimensional audit module 902 is specifically used to: perform at least one of gradient anomaly detection, proof-of-work verification, and collaborative behavior analysis on the model update data to obtain the audit result;
[0249] The audit results include the anomaly types of the work nodes, which include at least one of the following: gradient anomaly, proof-of-work anomaly, and collaboration anomaly.
[0250] In some embodiments, the multi-dimensional audit module 902 is specifically used for:
[0251] The gradient data in the model update data of each working node is grouped according to the neural network layer type, and the vector norm of the gradient of each group is calculated to obtain the low-dimensional representation vector of each working node.
[0252] Determine the median baseline vector of the low-dimensional representation vectors of all working nodes;
[0253] The anomaly score of each working node is determined based on the normalized distance between the low-dimensional representation vector of each working node and the median baseline vector.
[0254] Based on the anomaly scores of each working node, an audit result is generated, which includes whether each working node exhibits gradient anomalies.
[0255] In some embodiments, the multi-dimensional audit module 902 is specifically used for:
[0256] Determine the median of the anomaly scores for all working nodes;
[0257] Based on the anomaly score of each working node and the median of the anomaly scores, determine the median of the absolute deviation of all working nodes.
[0258] A dynamic threshold is determined based on the median of the abnormal scores and the median of the absolute deviations.
[0259] Work nodes whose abnormal scores exceed the dynamic threshold are identified as gradient anomalies to obtain the audit results.
[0260] In some embodiments, the multi-dimensional audit module 902 is specifically used for:
[0261] Randomly select a subset of validation data from the public validation dataset;
[0262] The model performance metrics are recalculated on the validation subset to obtain the local performance metrics.
[0263] The performance metrics declared by each worker node are obtained from the model update data. Based on the local performance metrics and the performance metrics declared by each worker node, the confidence score of each worker node is determined.
[0264] Work nodes with confidence scores below a preset threshold are identified as proof-of-work anomalies to obtain the audit results.
[0265] In some embodiments, the performance metrics include loss value and accuracy; the confidence score is obtained by combining the relative error of loss and the absolute error of accuracy.
[0266] In some embodiments, the multi-dimensional audit module 902 is specifically used for:
[0267] The model update data of each working node is normalized to obtain the gradient direction vector of each working node;
[0268] Determine the first similarity between every two gradient direction vectors;
[0269] For each worker node, calculate the average similarity between the gradient direction vector of that worker node and the gradient direction vectors of all other worker nodes, and use this as the average similarity of that worker node.
[0270] Work node pairs with a first similarity greater than a first threshold and an average similarity between the two work nodes less than a second threshold are identified as suspected collaborative pairs.
[0271] The suspected collaborative pair and its associated working nodes are identified as having collaborative anomalies to obtain the audit results.
[0272] In some embodiments, the reputation management and policy enforcement module 903 is specifically used for:
[0273] The deduction score is determined based on the type of anomaly at each work node in the audit results;
[0274] The updated reputation value is determined based on the historical reputation value of each working node and the deduction score of the current training round.
[0275] In some embodiments, the reputation management and policy enforcement module 903 is specifically used to: merge all suspected collaborative pairs into a group by using a transitive closure algorithm;
[0276] Work nodes belonging to the same collaboration group are uniformly identified as collaboration anomalies to obtain the audit results.
[0277] In some embodiments, the reputation management and policy enforcement module 903 is specifically used for:
[0278] If the reputation value of the target worker node is greater than or equal to the first reputation threshold, the target worker node is determined as a trusted node, and the trusted node is configured to be allowed to participate normally in model training and audit node election.
[0279] If the reputation value of the target working node is less than the first reputation threshold and greater than or equal to the second reputation threshold, the target working node is identified as a suspicious node, and the first weight of the model update data submitted by the target working node in the global model aggregation is reduced to the second weight.
[0280] If the reputation value of the target worker node is less than the second reputation threshold, the target worker node is identified as a malicious node, and the malicious node is prohibited from participating in model training and audit node election.
[0281] In some embodiments, the apparatus further includes:
[0282] The generation module is used to generate a list of safe nodes for the current training round based on the updated reputation value and the status of each working node. The list of safe nodes includes node information that has not been identified as a malicious node.
[0283] Once the list of secure nodes has been confirmed by multiple signatures from the auditing nodes, it is broadcast to all nodes.
[0284] The modules in the aforementioned node permission determination device for distributed artificial intelligence training can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can invoke and execute the operations corresponding to each module.
[0285] In one exemplary embodiment, a computer device is provided, the internal structure of which can be as shown in the figure. Figure 10 As shown, this computer device includes a processor, memory, input / output interfaces (I / O), and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operating system and computer programs stored in the non-volatile storage media. The database stores data. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network. When executed by the processor, the computer program implements a node permission determination method for distributed artificial intelligence training.
[0286] Those skilled in the art will understand that Figure 10 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0287] In one exemplary embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the various processes shown in the above method embodiments.
[0288] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the various processes shown in the above method embodiments.
[0289] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the various processes shown in the above method embodiments.
[0290] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.
[0291] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0292] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A method for determining node permissions for distributed artificial intelligence training, characterized in that, The method includes: Based on the current reputation value of each node, an audit node is elected from among the nodes as the election result; the election result is confirmed, wherein the current reputation value is used to characterize the trustworthiness of the node; The verified audit node performs credibility verification on the model update data submitted by the working node, generating an audit result. This verification includes: performing at least one of gradient anomaly detection, proof-of-work verification, and collaborative behavior analysis on the model update data to obtain the audit result; performing collaborative behavior analysis on the model update data to obtain the audit result includes: normalizing the gradient data in the model update data of each working node to obtain the gradient direction vector of each working node; determining a first similarity between every two gradient direction vectors; for each working node, calculating the average similarity between the gradient direction vector of that working node and the gradient direction vectors of all other working nodes as the average similarity of that working node; identifying working node pairs where the first similarity is greater than a first threshold and the average similarity of both working nodes is less than a second threshold as suspected collaborative pairs; and identifying the working nodes associated with the suspected collaborative pairs as collaborative anomalies to obtain the audit result. The reputation value of each working node is updated based on the audit results, and the permissions of each working node to participate in model training and validation tasks are controlled based on the updated reputation value.
2. The method according to claim 1, characterized in that, The process of electing an audit node from among the nodes based on their current reputation values includes: A random seed is generated based on a random function, and the election weight of each node is determined based on the current reputation value of each node. The nodes are weighted and randomly sampled using the random seed, and the selected nodes are determined as the audit nodes, which is the election result.
3. The method according to claim 2, characterized in that, The confirmation of the election results includes: The validity of the proof information of the random function in the election results is verified by the existing audit nodes respectively; Existing audit nodes that have passed verification broadcast confirmation messages of the election results; When the cumulative number of broadcast election result confirmation messages exceeds a preset threshold as a percentage of the total number of existing audit nodes, the election result is determined to have been confirmed.
4. The method according to claim 1, characterized in that, The audit results include the anomaly types of the work nodes, which include at least one of the following: gradient anomaly, proof-of-work anomaly, and collaboration anomaly.
5. The method according to claim 4, characterized in that, Gradient anomaly detection is performed on the model update data, and audit results are generated, including: The gradient data in the model update data of each working node is grouped according to the neural network layer type, and the vector norm of the gradient of each group is calculated to obtain the low-dimensional representation vector of each working node. Determine the median baseline vector of the low-dimensional representation vectors of all working nodes; The anomaly score of each working node is determined based on the normalized distance between the low-dimensional representation vector of each working node and the median baseline vector. Based on the anomaly scores of each working node, an audit result is generated, which includes whether each working node exhibits gradient anomalies.
6. The method according to claim 5, characterized in that, The audit results generated based on the anomaly scores of each working node include: Determine the median of the anomaly scores for all working nodes; Based on the anomaly score of each working node and the median of the anomaly scores, determine the median of the absolute deviation of all working nodes. A dynamic threshold is determined based on the median of the abnormal scores and the median of the absolute deviations. Work nodes whose abnormal scores exceed the dynamic threshold are identified as gradient anomalies to obtain the audit results.
7. The method according to claim 4, characterized in that, Perform proof-of-work verification on the model update data to obtain the audit results, including: Randomly select a subset of validation data from the public validation dataset; The model performance metrics are recalculated on the validation subset to obtain the local performance metrics. The performance metrics declared by each worker node are obtained from the model update data. Based on the local performance metrics and the performance metrics declared by each worker node, the confidence score of each worker node is determined. Work nodes with confidence scores below a preset threshold are identified as proof-of-work anomalies to obtain the audit results.
8. The method according to claim 7, characterized in that, The performance metrics include loss value and accuracy; the confidence score is obtained by combining the relative error of loss and the absolute error of accuracy.
9. The method according to claim 1, characterized in that, The step of determining the associated working nodes of the suspected collaborative pair as having collaborative anomalies to obtain the audit results includes: The transitive closure algorithm is used to merge all suspected cooperative pairs into related cooperative groups. Work nodes belonging to the same collaboration group are uniformly identified as collaboration anomalies to obtain the audit results.
10. The method according to any one of claims 4 to 9, characterized in that, The process of updating the reputation value of each working node based on the audit results includes: Based on the anomaly type of each work node in the audit results, determine the deduction score for each work node; The updated reputation value is determined based on the historical reputation value of each working node, the deduction score, and the time decay coefficient.
11. The method according to any one of claims 1 to 9, characterized in that, Based on the updated reputation value, control the permissions of each working node to participate in model training and validation tasks, including: If the reputation value of the target worker node is greater than or equal to the first reputation threshold, the target worker node is determined as a trusted node, and the trusted node is configured to be allowed to participate normally in model training and audit node election. If the reputation value of the target working node is less than the first reputation threshold and greater than or equal to the second reputation threshold, the target working node is identified as a suspicious node, and the first weight of the model update data submitted by the target working node in the global model aggregation is reduced to the second weight. If the reputation value of the target worker node is less than the second reputation threshold, the target worker node is identified as a malicious node, and the malicious node is prohibited from participating in model training and audit node election.
12. The method according to any one of claims 1 to 9, characterized in that, The method further includes: Based on the updated reputation value and the status of each working node, a list of safe nodes for the current training round is generated, which includes node information that has not been identified as a malicious node. Once the list of secure nodes has been confirmed by multiple signatures from the auditing nodes, it is broadcast to all nodes.
13. A node permission determination device for distributed artificial intelligence training, characterized in that, The device includes: The dynamic election module is used to elect an audit node from among the nodes based on the current reputation value of each node, and to confirm the election result. The current reputation value is used to characterize the trustworthiness of the node. A multi-dimensional audit module is used to verify the credibility of model update data submitted by worker nodes by the confirmed audit nodes and generate audit results. The verification of the model update data submitted by worker nodes by the confirmed audit nodes and the generation of audit results includes: performing at least one of gradient anomaly detection, proof-of-work verification, and collaborative behavior analysis on the model update data to obtain the audit results; performing collaborative behavior analysis on the model update data to obtain the audit results includes: normalizing the gradient data in the model update data of each worker node to obtain the gradient direction vector of each worker node; determining a first similarity between every two gradient direction vectors; for each worker node, calculating the average similarity between the gradient direction vector of that worker node and the gradient direction vectors of all other worker nodes as the average similarity of that worker node; identifying worker node pairs where the first similarity is greater than a first threshold and the average similarity of both worker nodes is less than a second threshold as suspected collaborative pairs; and identifying the worker nodes associated with the suspected collaborative pairs as collaborative anomalies to obtain the audit results. The reputation management and policy execution module is used to update the reputation value of each working node based on the audit results, and to control the permissions of each working node to participate in model training and verification tasks based on the updated reputation value.
14. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 12.
15. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 12.
16. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Method for improving robustness of federal system
CN118095411A
Practical consensus algorithm for file resource sharing
CN121750195A
Intelligent traffic model credibility verification method based on zero-knowledge proof and reputation mechanism
CN121923826A
Privacy protection distributed collaborative auditing system and method based on block chain and federated map neural network
CN122174261A