Multi-party trusted cooperation method and system based on DAG block chain and federal learning
By dynamically selecting tip nodes on the DAG blockchain and combining freshness and data similarity for model aggregation, the inefficiency of traditional federated learning is solved, enabling efficient and reliable multi-party collaborative training, improving model convergence speed and accuracy, while ensuring data privacy.
Patent Information
- Application Number
- CN202511042703.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-28
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2045-07-28
AI Technical Summary
Traditional centralized data collaboration models suffer from low model training efficiency, resource waste, and insufficient privacy protection in federated learning. Existing federated learning methods based on DAG blockchains do not fully consider device heterogeneity, resulting in low model iteration efficiency and the risk of data leakage.
We adopt a multi-party trusted collaboration method based on DAG blockchain and federated learning. By dynamically selecting Tip nodes and combining freshness, accessibility and data distribution similarity for model aggregation, we use asymmetric encryption and hash commitment mechanisms to ensure data privacy and model trustworthiness, thereby achieving efficient collaboration and resource optimization.
It improved model training efficiency, optimized resource utilization, enhanced data privacy protection, increased the convergence speed and accuracy of the model in heterogeneous data environments, and achieved transparent traceability of the training path throughout the entire lifecycle.
Smart Images

Figure CN120880737A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the technical field of blockchain and privacy computing, and in particular to a multi-party trusted collaboration method and system based on DAG blockchain and federated learning. Background Technology
[0002] The statements in this section are merely background information relating to this disclosure and do not necessarily constitute prior art.
[0003] With the expansion of data transaction volume and the increasing demand for privacy protection, traditional centralized data collaboration models face severe challenges. In centralized federated learning (FL) frameworks, reliance on a single server to coordinate global model updates leads to problems such as low model training efficiency due to device asynchrony and decreased model accuracy due to data heterogeneity. While traditional blockchain technology can enhance data credibility through distributed ledgers, its chain-structure consensus mechanism (such as proof-of-work) suffers from high latency and low throughput, making it difficult to adapt to large-scale federated learning scenarios involving resource-constrained IoT devices. Taking medical data collaboration as an example, hospitals, pharmaceutical companies, and other parties need to share drug efficacy data while protecting patient privacy. However, existing platforms pose a risk of raw data leakage, and differences in the computing power of participating nodes during model updates can easily lead to slow global model convergence or even failure. Blockchain technology based on Directed Acyclic Graphs (DAGs) overcomes the performance bottleneck of traditional blockchains through parallel transaction processing mechanisms, and its asynchronous verification characteristics naturally align with the decentralized requirements of federated learning.
[0004] However, existing federated learning methods based on DAG blockchains often rely on random mechanisms for node selection, failing to adequately consider the impact of device heterogeneity on model iteration. For example, some schemes use fixed time windows to select participating nodes, leading to nodes with high computational latency contributing outdated model parameters and exacerbating the negative bias of non-independent and identically distributed (Non-IID) data. Furthermore, traditional blockchain federated learning frameworks require full storage of model parameters, resulting in wasted storage and communication resources, particularly in wireless network environments. To address these issues, there is an urgent need to construct a novel collaborative architecture that integrates dynamic node evaluation and lightweight verification mechanisms, enabling efficient collaboration and optimized resource allocation across heterogeneous devices while ensuring data privacy and model trustworthiness.
[0005] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0006] To address, or at least partially address, the aforementioned technical problems, this invention provides a multi-party trusted collaboration method and system based on DAG blockchain and federated learning.
[0007] This invention provides a multi-party trusted collaboration method based on DAG blockchain and federated learning, comprising: The task publisher initializes and publishes the global initial model in the DAG blockchain, and creates encrypted task metadata for collaborative training tasks, thus initiating the collaborative training task process; the training task metadata includes the task publisher identifier, initial global model summary information, and publication time. After the task training party retrieves the collaborative training task, it dynamically selects a Tip node from the DAG blockchain collaboration module based on the task requirements, freshness, accessibility, and data distribution similarity. It then uses the selected Tip node to pull the corresponding model, aggregates the model according to the set model aggregation strategy based on Tip selection rules to obtain a local training model, and executes the local training model training. After the task training party aggregates the model, it uploads the model aggregation path to the DAG blockchain collaboration module. The DAG blockchain collaboration module generates the corresponding hash commitment for structure verification and tamper-proof evidence storage. After the task trainer completes the training of the local training model, it encrypts the aggregated model parameters and model update metadata and uploads them to the DAG blockchain collaboration module. The model update metadata includes: trainer identifier, model accuracy, local feature signature and training round. The DAG blockchain collaboration module receives encrypted update data uploaded by each task training party, verifies the authenticity of the update data uploaded by each training party through a privacy protection mechanism, and executes Tip node updates using the update data after confirming that there has been no tampering. The task issuer terminates the training by issuing a training termination command based on preset global training termination conditions.
[0008] Furthermore, the task publisher and task trainer register their identities and access the system. Using an asymmetric encryption algorithm, a public-private key pair is generated for each of the registered and connected task publishers and task trainers. Each task publisher and task trainer has a public-private key pair, and the public key information is stored on the blockchain for subsequent identity authentication.
[0009] Furthermore, the freshness of a tip is calculated by the difference between the tip's timestamp and the current time, as well as the difference in global iteration rounds between the tip model and the current model. Based on the freshness, the tips in the optional tip set are assigned priority, with fresher tips having higher priority.
[0010] Furthermore, the task trainer calculates local feature signatures based on the intermediate layer features of the locally trained model, and compares them with the corresponding local feature signatures of the candidate Tip nodes using cosine similarity to evaluate the data distribution similarity of each Tip.
[0011] Furthermore, the calculation process for the local feature signature in the model update metadata is as follows: Task Training Method K i Having a private dataset , N i K represents the task training method. i The size of the private dataset, x h and y h These are the features and labels of the h-th data sample, respectively; Then the feature x of the h-th data sample of the task training side h The feature signature under the k-th feature extraction unit is: , In the above formula, zero(·) represents the number of zero elements in the statistical characteristic matrix; The k-th feature extraction unit in the task training... K i Local data D i The average value of the feature signatures is: , For task training K i The average feature signature of different feature extraction units used to extract features in its local training model forms a feature signature vector, which serves as the local feature signature. , where m is the number of feature extraction units selected in the model to construct the local feature signature.
[0012] Furthermore, model aggregation strategies based on tip selection rules include: Obtain the current global training state and determine the freshness index of each tip in the optional tip set by combining the training round information of its own model; Based on the intermediate layer features of the locally trained model, calculate the local feature signature and perform cosine similarity with the corresponding local feature signature of the candidate Tip node. The task trainer uses a breadth-first search algorithm to trace its most recent upload node in the DAG graph, identify the reachable Tip nodes of the current node, obtain the set of reachable Tips, and the set of remaining unreachable Tips; The task trainer sets a tip selection ratio coefficient λ based on the freshness, accessibility, and data distribution similarity of the tips, and selects the Top-N1 and Top-N2 tips from the accessible tip set and the unreachable tip set, respectively, for local model aggregation. The model trained on the k-th task in round t. The weighted aggregation formula is as follows: in, , in, For the model corresponding to the i1th reachable Tip node in round t-1, Let N1 = λN and N2 = (1-λ)N be the model corresponding to the i2th unreachable Tip node in the (t-1)th round, where N is the total number of selected Tips.
[0013] Furthermore, for the set of reachable tips, the model accuracy of each reachable tip node is directly calculated, and the models corresponding to the N1 tips with the highest accuracy are selected. For the set of unreachable tips, the local feature signature similarity of each unreachable tip node is calculated, and the models corresponding to the N2 tips with the highest similarity are selected.
[0014] Furthermore, after training is terminated, a complete and credible model training history is generated by backtracking based on the DAG structure. The final aggregated global model and complete DAG training path are then returned to the task publisher and each training party. Each party verifies the credible source and training process of the model based on the system's private key for deployment or application analysis.
[0015] This invention also provides a multi-party trusted collaboration system based on DAG blockchain and federated learning, comprising: a task publisher module, a task trainer module, and a DAG blockchain collaboration module, wherein the task publisher module and the task trainer module both interact with the DAG blockchain collaboration module; and implementing the multi-party trusted collaboration method based on DAG blockchain and federated learning as described in any one of claims 1-8.
[0016] Furthermore, the task publisher module includes a first identity registration unit, an initial model publishing unit, a training task management unit, and a DAG maintenance unit. The first identity registration module is responsible for the identity verification and registration of the task publisher, generating a public-private key pair, and storing the public key information on the blockchain for subsequent identity authentication. The initial model publishing unit is responsible for generating and publishing the global initial model to the DAG blockchain collaboration module to start the collaborative training process. The training task management unit is responsible for monitoring and managing the training tasks, including: model accuracy tracking, iteration count control, and issuing training termination commands. The DAG maintenance unit is responsible for monitoring the evolution of the overall DAG structure, ensuring the orderliness of tip generation and node reference relationships, and supporting subsequent integrity verification. The task training module includes: a second identity registration unit, a local training unit, a tip selection unit, a model upload unit, and a feature signature extraction unit. The second identity registration module is responsible for verifying and registering the trainer's identity, following the same process as the task publisher's identity registration. It generates a public-private key pair and stores the public key information on the blockchain for subsequent identity authentication, ensuring node trustworthiness. The local training unit performs model training based on local private data, protecting data privacy by ensuring data remains local. The tip selection unit dynamically selects the optimal tip node from the DAG blockchain collaboration module for model aggregation based on freshness, accessibility, and feature similarity. The model upload unit uploads the aggregated local training model and corresponding metadata to the DAG blockchain collaboration module, supporting asynchronous collaboration. The feature signature extraction unit extracts the feature distribution information of the local model and generates a local feature signature, used for subsequent tip selection and model data distribution similarity evaluation. The DAG blockchain collaboration module includes: a collaboration evidence storage unit, a tip selection support unit, a smart contract management unit, a hash verification unit, and a metadata management unit. The collaboration evidence storage unit is responsible for recording all uploaded model updates and model update metadata, forming a traceable and tamper-proof DAG structure. The tip selection support unit provides functions for querying tip freshness, accessibility, and feature similarity, supporting efficient tip node selection by the task training party. The smart contract management unit is responsible for maintaining the tip selection rules and feature similarity calculation smart contracts, ensuring the decision-making process is open and transparent. The hash verification unit, based on a block hash commitment mechanism, supports the verification of the DAG chain structure and model data integrity, preventing tampering. The metadata management unit manages the uploaded model update metadata.
[0017] The technical solutions provided in the embodiments of the present invention have the following advantages compared with the prior art: 1. This application utilizes the DAG blockchain structure to achieve efficient asynchronous collaboration in the federated learning process, avoiding the high energy consumption and latency issues caused by the PoW consensus mechanism in traditional blockchains, and significantly improving the model training efficiency and resource utilization of the system.
[0018] 2. This application emphasizes the traceability of the training process and the reliability of data updates. All model updates and metadata are recorded through a DAG structure and verified by hash commitment, which effectively prevents tampering and realizes transparent traceability of the training path throughout the entire lifecycle.
[0019] 3. This application provides a Tip selection mechanism for heterogeneous data environments. By combining multiple indicators such as Tip freshness, accessibility and feature similarity, it dynamically optimizes the source of aggregated models, improves the convergence speed and final accuracy of models in heterogeneous data scenarios, and enhances the robustness of the system under complex distributed data.
[0020] 4. This application introduces a similarity calculation strategy based on feature signature-based data distribution representation and smart contract assistance. It achieves cross-client data feature matching in a privacy-preserving manner, which not only protects data privacy but also further improves the model aggregation quality and training effect, thus balancing privacy protection and optimal performance. Attached Figure Description
[0021] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1 This is a schematic diagram of the structure of a multi-party trusted collaboration system based on DAG blockchain and federated learning, provided in an embodiment of the present invention.
[0024] Figure 2 A flowchart illustrating a multi-party trusted collaboration method based on DAG blockchain and federated learning, provided for embodiments of the present invention.
[0025] Figure 3 A flowchart of a model aggregation optimization method based on Tip selection rules provided in an embodiment of the present invention. Detailed Implementation
[0026] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0027] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0028] Example 1 like Figure 1 As shown, this invention provides a multi-party trusted collaboration system based on DAG blockchain and federated learning, including: a task publisher module, a task trainer module, and a DAG blockchain collaboration module. The task publisher module and the task trainer module both interact with the DAG blockchain collaboration module, thus implementing the multi-party trusted collaboration method based on DAG blockchain and federated learning of this application. In specific implementation, there can be multiple task publisher modules and task trainer modules as needed, and the DAG blockchain collaboration module serves as a unified global collaboration space.
[0029] The task publisher module is configured to at least implement the following: registration and access of the task publisher identity, publication of the initial global model, initiation and management of training tasks, maintenance of the DAG blockchain collaborative architecture, and publication of termination commands.
[0030] In specific implementation, the task publisher module includes a first identity registration unit, an initial model publishing unit, a training task management unit, and a DAG maintenance unit. The first identity registration module is responsible for the identity verification and registration of the task publisher, generating a public-private key pair, and storing the public key information on the blockchain for subsequent identity authentication. The initial model publishing unit is responsible for generating and publishing the global initial model to the DAG blockchain collaboration module to start the collaborative training process. The training task management unit is responsible for monitoring and managing the training tasks, including: model accuracy tracking, iteration count control, and issuing training termination commands. The DAG maintenance unit is responsible for monitoring the evolution of the overall DAG structure, ensuring the orderliness of tip generation and node reference relationships, and supporting subsequent integrity verification.
[0031] The task training module is configured to at least implement the following: registration and access of task training identity, privacy training of local data, model aggregation based on tip selection, model update and metadata upload, generation and storage of feature signatures, and maintenance of block verification paths.
[0032] The task training module specifically includes: a second identity registration unit, a local training unit, a tip selection unit, a model upload unit, and a feature signature extraction unit. The second identity registration module is responsible for verifying and registering the trainer's identity, following the same process as the task publisher's identity registration. It generates a public-private key pair and stores the public key information on the blockchain for subsequent identity authentication, ensuring node trustworthiness. The local training unit performs model training based on local private data, protecting data privacy by ensuring data remains local. The tip selection unit dynamically selects the optimal tip node from the DAG blockchain collaboration module for model aggregation based on freshness, accessibility, and feature similarity. The model upload unit uploads the aggregated local training model and corresponding metadata to the DAG blockchain collaboration module, supporting asynchronous collaboration. The feature signature extraction unit extracts the feature distribution information of the local model and generates a local feature signature, used for subsequent tip selection and model data distribution similarity evaluation.
[0033] The DAG blockchain collaboration module is configured to at least implement the following: on-chain storage of metadata for model updates during the training process, tip selection and management, notarization of training model update records, management of feature signatures and data distribution features of each model, execution of model aggregation rules, verification and traceability of DAG structural integrity, and deployment and execution of smart contracts.
[0034] The DAG blockchain collaboration module specifically includes: a collaboration evidence storage unit, a tip selection support unit, a smart contract management unit, a hash verification unit, and a metadata management unit. The collaboration evidence storage unit is responsible for recording all uploaded model updates and model update metadata, forming a traceable and tamper-proof DAG structure. The tip selection support unit provides functions for querying tip freshness, accessibility, and feature similarity, supporting efficient tip node selection by the task training party. The smart contract management unit is responsible for maintaining the tip selection rules and feature similarity calculation smart contracts, ensuring the decision-making process is open and transparent. The hash verification unit, based on a block hash commitment mechanism, supports the verification of the DAG chain structure and model data integrity, preventing tampering. The metadata management unit manages the uploaded model update metadata, including feature signatures, model accuracy, and training rounds, providing a basis for subsequent verification and collaborative optimization.
[0035] Example 2 This invention provides a multi-party trusted collaboration method based on DAG blockchain and federated learning, applied to a constructed multi-party trusted collaboration system based on DAG blockchain and federated learning, specifically including the following steps: In S100, the task publisher and task trainer register their identities and access the system using the task publisher module and task trainer module, respectively. The task publisher module and task trainer module use asymmetric encryption algorithms (such as RSA) to generate a public-private key pair for each registered and connected task publisher and task trainer. Each task publisher and task trainer possesses a public-private key pair and stores the public key information on the blockchain for subsequent identity authentication.
[0036] S200, the task publisher initializes and publishes the global initial model in the DAG blockchain collaboration module through the task publisher module, and creates the task metadata for the collaborative training task, thus initiating the collaborative training task process. The training task metadata includes the task publisher identifier, initial global model summary information, and publication time.
[0037] The task metadata format is defined as follows: MetaData=<PublisherID, InitialModelInfo, Timestamp> pk_platform Wherein, PublisherId is the identifier of the task publisher, InitialModelInfo is the initial model summary information, Timestamp is the publication time, and pk_platform represents the system's public key, which is used to encrypt the system's internal information. Here, the task metadata is encrypted to prevent information leakage. The training task metadata is uploaded to the DAG blockchain collaboration module after being encrypted with the system's public key.
[0038] Enables collaborative training data to be stored on the blockchain for retrieval by the task training participants.
[0039] Third, after retrieving the collaborative training task, the task trainer selects a Tip node from the DAG blockchain collaboration module according to the task requirements. The Tip node is a node that is not referenced in the DAG blockchain collaboration module. Specifically, the task trainer obtains the model update metadata of the Tip to obtain freshness, accessibility, and data distribution similarity. The task trainer dynamically selects the Tip node from the DAG blockchain collaboration module, pulls the corresponding model using the selected Tip node, and obtains the set of models to be aggregated through Tip selection. The model is then aggregated according to the set model aggregation strategy based on the Tip selection rules to obtain the local training model, and the local training model is trained. The task trainer uses local private data to perform model training, protecting data privacy by ensuring that the data does not leave the local machine.
[0040] In specific implementation, the model aggregation strategy based on Tip selection rules in this application includes: S201, the task trainer receives the optional tip set provided by the DAG blockchain collaboration module. Then, it obtains the current global training state and determines the freshness index of each tip in the optional tip set by combining its own model training round information. The freshness of the tip is calculated by the difference between the tip timestamp and the current time and the difference between the tip model and the current model's global iteration rounds. The tips in the optional tip set are given priority according to their freshness, with fresher tips having higher priority.
[0041] In S202, the task trainer calculates local feature signatures based on the intermediate layer features of the local training model, and compares them with the corresponding local feature signatures of the candidate Tip nodes using cosine similarity to evaluate the data distribution similarity of each Tip, thereby avoiding redundant calculations and accuracy loss.
[0042] Given a set K of task trainers, the i-th (meaning any) task trainer K i Having a private dataset , N i K represents the task training method. i The size of the private dataset, x h and y h These are the features and labels of the h-th data sample, respectively. In each local training round of the task trainer K... i In its private dataset D i Train a locally trained model w i During the forward propagation process, the local model w trained by the task trainer Ki... i The k-th feature extraction unit extracts features x from the given h-th data sample. h Extracting the feature matrix The feature matrix Let H be any intermediate layer feature, with dimensions H×W, where H and W represent the feature matrix, respectively. The height and width. Then the feature x of the h-th data sample of the task training side. h The feature signature under the k-th feature extraction unit is: , In the above formula, zero(·) represents the number of zero elements in the statistical characteristic matrix.
[0043] The k-th feature extraction unit in the task training... K i Local data D i The average value of the feature signatures is: , For task training K i The average feature signature of different feature extraction units used to extract features in its local training model forms a feature signature vector, which serves as the local feature signature. m is the number of feature extraction units selected in the model to construct local feature signatures. Each task trainer generates a feature signature vector that reflects the local data distribution, which is used to calculate the similarity with other clients.
[0044] Two task training methods K i and K j Formula for calculating the local signature similarity between two features: , in, and They are the task training side K i and K j Local signature features.
[0045] In S203, the task trainer uses a breadth-first search algorithm to trace its most recently uploaded node in the DAG graph, identify reachable Tip nodes for the current node, obtain a set of reachable Tips, and distinguish it from other Tip sets to obtain a set of reachable Tips and a set of unreachable Tips. Reachability reflects the similarity and inheritance of data distribution. Based on the similarity of data distribution and training contribution, historical Tip nodes are divided into high-similarity Tips (corresponding to reachable Tip nodes) and low-similarity Tips (corresponding to unreachable Tip nodes). Subsequently, different aggregation strategies are used for the models of reachable and unreachable Tip nodes to balance local optima and global convergence.
[0046] S204, the task trainer sets a tip selection ratio coefficient λ based on the freshness, reachability, and data distribution similarity of the tips. It then selects the Top-N1 and Top-N2 tips from the reachable and unreachable tip sets, respectively, for local model aggregation. Specifically, for the reachable tip set, the model accuracy of each reachable tip node is directly calculated, and the models corresponding to the N1 tips with the highest accuracy are selected. For the unreachable tip set, the local feature signature similarity to each unreachable tip node is calculated, and the models corresponding to the N2 tips with the highest similarity are selected. Here, N1 = λN, N2 = (1-λ)N, where N is the total number of selected tips, and the model of the k-th task trainer in round t is... The weighted aggregation formula is as follows: , in, For the model corresponding to the i1th reachable Tip node in round t-1, Let i be the model corresponding to the i2th unreachable Tip node in the (t-1)th round.
[0047] The task trainer aggregates the models corresponding to the selected tips, calculates the aggregated model parameters and uses them as the starting point for a new round of local training. At the same time, it extracts the local feature signature of the current training and uploads the corresponding updated metadata to the DAG blockchain collaboration module.
[0048] S300: After the task training party aggregates the model, it uploads the model aggregation path to the DAG blockchain collaboration module. The DAG blockchain collaboration module generates a corresponding hash commitment for structure verification and tamper-proof notarization. All tip selection processes are traceable and verifiable, ensuring the rationality of model updates and the credibility of training paths.
[0049] S400: After the task trainer completes the training of the local training model, it updates the model to the DAG blockchain collaboration module: extracting the local feature signature of the local training model, and encrypting the aggregated model parameters and model update metadata before uploading it to the DAG blockchain collaboration module. Specifically, the task trainer generates model update metadata containing the task trainer identifier, model accuracy, local feature signature, and training round, and uploads it to the DAG blockchain collaboration module after encrypting it with the platform's public key, forming a new Tip node; The format for model update metadata is defined as follows: UpdateData =<TrainerId, ModelAccuracy, CurrentEpoch,FeatureSignature, Timestamp> pk_platform Among them, TrainerId is the identifier of the task trainer, ModelAccuracy is the current model accuracy, CurrentEpoch is the training epoch, FeatureSignature is the feature signature of the locally trained model, and Timestamp in UpdateData is the model upload time.
[0050] After updating the local training model, any task trainer extracts the local feature signature of the local training model.
[0051] The S500 DAG blockchain collaboration module receives encrypted update data uploaded by each task training party. The DAG blockchain collaboration module verifies the authenticity of the update data uploaded by each training party through privacy protection mechanisms (such as feature signature comparison and hash verification) to ensure that the uploaded data has not been tampered with. Specifically, the DAG blockchain collaboration module verifies the uploaded feature signature and model update metadata through a smart contract mechanism.
[0052] S600, after confirming no tampering, uses the updated data to execute Tip node updates. The DAG blockchain collaboration module simultaneously updates the status and version control information of the Tip nodes on the chain, so that the updated Tip nodes can perform a new round of model aggregation based on Tip selection rules (combining Tip freshness, reachability and model accuracy), and maintain the orderly evolution of the DAG chain structure.
[0053] S700: Based on preset global training termination conditions (such as reaching the target accuracy or the maximum number of training rounds), the task issuer issues a training termination command. After termination, a complete and reliable model training history is generated by backtracking the DAG structure. The DAG blockchain collaboration module returns the finally aggregated global model and complete DAG training path to the task issuer and each training party. Each party verifies the reliable source of the model and the training process based on the system's private key, and further conducts deployment or application analysis.
[0054] The final global model and its corresponding training metadata are recorded in the DAG blockchain, ensuring the traceability, verifiability, and immutability of the model evolution process, and guaranteeing the credibility and privacy protection of the training results from multi-party collaboration.
[0055] In the embodiments provided by this invention, it should be understood that the disclosed structures and methods can be implemented in other ways. For example, the structural embodiments described above are merely illustrative. For instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection through some interfaces, structures, or units, and may be electrical, mechanical, or other forms.
[0056] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0057] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0058] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A multi-party trusted collaboration method based on DAG blockchain and federated learning, characterized in that, include: The task publisher initializes and publishes the global initial model in the DAG blockchain, and creates encrypted task metadata for collaborative training tasks, thus initiating the collaborative training task process; the training task metadata includes the task publisher identifier, initial global model summary information, and publication time. After the task trainer retrieves the collaborative training task, it dynamically selects a Tip node from the DAG blockchain collaboration module based on the task requirements, freshness, accessibility, and data distribution similarity. It then uses the selected Tip node to pull the corresponding model, aggregates the model according to the set model aggregation strategy based on Tip selection rules to obtain the local training model, and executes the local training model training. After the task training party aggregates the model, it uploads the model aggregation path to the DAG blockchain collaboration module. The DAG blockchain collaboration module generates the corresponding hash commitment for structure verification and tamper-proof evidence storage. After the task trainer completes the training of the local training model, it encrypts the aggregated model parameters and model update metadata and uploads them to the DAG blockchain collaboration module. The model update metadata includes: trainer identifier, model accuracy, local feature signature and training round. The DAG blockchain collaboration module receives encrypted update data uploaded by each task training party, verifies the authenticity of the update data uploaded by each training party through a privacy protection mechanism, and executes Tip node updates using the update data after confirming that there has been no tampering. The task issuer issues a training termination command to terminate the training based on preset global training termination conditions.
2. The multi-party trusted collaboration method based on DAG blockchain and federated learning according to claim 1, characterized in that, The task publisher and task trainer register their identities and access the system. An asymmetric encryption algorithm is used to generate a public-private key pair for each of the registered and connected task publishers and task trainers. Each task publisher and task trainer has a public-private key pair, and the public key information is stored on the blockchain for subsequent identity authentication.
3. The multi-party trusted collaboration method based on DAG blockchain and federated learning according to claim 1, characterized in that, Tip freshness is calculated by the difference between the Tip timestamp and the current time, as well as the difference between the Tip model and the current model's global iteration rounds. Tips in the optional Tip set are assigned priority based on freshness, with fresher Tips having higher priority.
4. The multi-party trusted collaboration method based on DAG blockchain and federated learning according to claim 1, characterized in that, The task trainer calculates local feature signatures based on the intermediate layer features of the local training model, and compares them with the corresponding local feature signatures of the candidate Tip nodes using cosine similarity to evaluate the data distribution similarity of each Tip.
5. The multi-party trusted collaboration method based on DAG blockchain and federated learning according to claim 1, characterized in that, The calculation process of local feature signatures in model update metadata is as follows: Task Training Method K i Having a private dataset , N i K represents the task training method. i The size of the private dataset, x h and y h These are the features and labels of the h-th data sample, respectively; Then the feature x of the h-th data sample of the task training side h The feature signature under the k-th feature extraction unit is: , In the above formula, zero(·) represents the number of zero elements in the statistical characteristic matrix; The k-th feature extraction unit in the task training... K i Local data D i The average value of the feature signatures is: , For task training K i The average feature signature of different feature extraction units used to extract features in its local training model forms a feature signature vector, which serves as the local feature signature. , where m is the number of feature extraction units selected in the model to construct the local feature signature.
6. The multi-party trusted collaboration method based on DAG blockchain and federated learning according to claim 1, characterized in that, Model aggregation strategies based on tip selection rules include: Obtain the current global training state and determine the freshness index of each tip in the optional tip set by combining the training round information of its own model; Based on the intermediate layer features of the locally trained model, calculate the local feature signature and perform cosine similarity with the corresponding local feature signature of the candidate Tip node. The task trainer uses a breadth-first search algorithm to trace its most recent upload node in the DAG graph, identify the reachable Tip nodes of the current node, obtain the set of reachable Tips, and the set of remaining unreachable Tips; The task trainer sets a tip selection ratio coefficient λ based on the freshness, accessibility, and data distribution similarity of the tips, and selects the Top-N1 and Top-N2 tips from the accessible tip set and the unreachable tip set, respectively, for local model aggregation. The model trained on the k-th task in round t. The weighted aggregation formula is as follows: in, , in, For the model corresponding to the i1th reachable Tip node in round t-1, For the model corresponding to the i2th unreachable Tip node in the (t-1)th round, N1 = λN, N2 = (1-λ)N, where N is the total number of selected Tips.
7. The multi-party trusted collaboration method based on DAG blockchain and federated learning according to claim 6, characterized in that, For the set of reachable tips, directly calculate the model accuracy of each reachable tip node, and select the models corresponding to the N1 tips nodes with the highest accuracy. For the set of unreachable tips, calculate the local feature signature similarity of each unreachable tip node, and select the models corresponding to the N2 tips nodes with the highest similarity.
8. The multi-party trusted collaboration method based on DAG blockchain and federated learning according to claim 1, characterized in that, After training is terminated, a complete and credible model training history is generated by backtracking based on the DAG structure. The final aggregated global model and complete DAG training path are returned to the task issuer and each training party. Each party verifies the credible source and training process of the model based on the system private key for deployment or application analysis.
9. A multi-party trusted collaboration system based on DAG blockchain and federated learning, characterized in that, It includes: a task publisher module, a task trainer module, and a DAG blockchain collaboration module, wherein the task publisher module and the task trainer module both interact with the DAG blockchain collaboration module; thereby implementing the multi-party trusted collaboration method based on DAG blockchain and federated learning as described in any one of claims 1-8.
10. The multi-party trusted collaboration architecture based on DAG blockchain and federated learning according to claim 9, characterized in that, The task publisher module includes a first identity registration unit, an initial model publishing unit, a training task management unit, and a DAG maintenance unit. The first identity registration module is responsible for the identity verification and registration of the task publisher, generating a public-private key pair, and storing the public key information on the blockchain for subsequent identity authentication. The initial model publishing unit is responsible for generating and publishing the global initial model to the DAG blockchain collaboration module, initiating the collaborative training process. The training task management unit is responsible for monitoring and managing training tasks, including: model accuracy tracking, iteration count control, and issuing training termination commands; the DAG maintenance unit is responsible for monitoring the evolution of the overall DAG structure, ensuring the orderliness of tip generation and node reference relationships, and supporting subsequent integrity verification. The task training module includes: a second identity registration unit, a local training unit, a tip selection unit, a model upload unit, and a feature signature extraction unit. The second identity registration module is responsible for verifying and registering the trainer's identity, following the same process as the task publisher's identity registration. It generates a public-private key pair and stores the public key information on the blockchain for subsequent identity authentication, ensuring node trustworthiness. The local training unit performs model training based on local private data, protecting data privacy by ensuring data remains local. The tip selection unit dynamically selects the optimal tip node from the DAG blockchain collaboration module for model aggregation based on freshness, accessibility, and feature similarity. The model upload unit uploads the aggregated local training model and corresponding metadata to the DAG blockchain collaboration module, supporting asynchronous collaboration. The feature signature extraction unit extracts the feature distribution information of the local model and generates a local feature signature, used for subsequent tip selection and model data distribution similarity evaluation. The DAG blockchain collaboration module includes: a collaboration evidence storage unit, a tip selection support unit, a smart contract management unit, a hash verification unit, and a metadata management unit. The collaboration evidence storage unit is responsible for recording all uploaded model updates and model update metadata, forming a traceable and tamper-proof DAG structure. The tip selection support unit provides functions for querying tip freshness, accessibility, and feature similarity, supporting efficient tip node selection by the task training party. The smart contract management unit is responsible for maintaining the tip selection rules and feature similarity calculation smart contracts, ensuring the decision-making process is open and transparent. The hash verification unit, based on a block hash commitment mechanism, supports the verification of the DAG chain structure and model data integrity, preventing tampering. The metadata management unit manages the uploaded model update metadata.
Citation Information
Patent Citations
Federal learning method and system based on block chain and trusted execution environment
CN113837761A
Federal learning method based on DAG block chain
CN115049071A
Federal learning privacy protection system and method based on blockchain without committee
CN118568783A
Clustering federal learning method under fog computing architecture
CN118780391A
Decentralized federated learning method and system based on personalized local differential privacy
CN118940857A
Cited By
Federal learning and block chain-based ultrasonic data verification method and system
CN121834864A