Data element privacy protection method based on zero-knowledge machine learning and on-chain verification

By adopting zero-knowledge machine learning and blockchain technology methods in the field of machine learning, the privacy protection problem of data during analysis and sharing is solved, and efficient privacy protection and secure model parameter updates are achieved.

CN120074838AActive Publication Date: 2025-05-30YUNNAN UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510229298.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-05-30
Estimated Expiration
2045-02-28

AI Technical Summary

Technical Problem

In the field of machine learning, data sharing and use often face the risk of privacy leakage, and traditional privacy protection methods cannot effectively solve the privacy protection problems of data during analysis and sharing.

Method used

Using data element privacy protection methods based on zero-knowledge machine learning and blockchain technology, the model calculation tasks are delegated to the off-chain oracle server for execution through ZK-SNARK technology, and a zero-knowledge machine learning model winding and parameter update algorithm are designed, and the model parameters that have undergone differential privacy processing are synchronized through the zero-knowledge gossip protocol.

Benefits of technology

It significantly reduces on-chain computing costs, effectively protects privacy, and ensures the security and traceability of model parameters, prevents security issues such as poisoning attacks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120074838A_ABST
    Figure CN120074838A_ABST
Patent Text Reader

Abstract

The invention discloses a data element privacy protection method based on zero-knowledge machine learning and on-chain verification. The method comprises the following steps: constructing a system model consisting of a decentralized application, a contract user, a data supplier, a distributed oracle framework and on-chain verification nodes; through a ZK-SNARK technology, a calculation-intensive task in the model is issued to an under-chain oracle machine for execution; based on a zero-knowledge machine learning model chaining and parameter updating algorithm, model parameters subjected to differential privacy processing are quickly synchronized to a plurality of oracle machine nodes through a zero-knowledge gossip protocol, and the updating process of the parameters is recorded in a block chain. Through the ZK-SNARK technology, the model calculation task is issued to the under-chain oracle server for execution, the on-chain calculation cost is remarkably reduced, and privacy is effectively protected.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of big data security technology, and specifically relates to a data element privacy protection method based on zero-knowledge machine learning and on-chain verification. Background Art

[0002] In the data-driven era, privacy protection has become particularly critical. With the widespread use of personal data, transaction information, and business data, how to ensure that these sensitive data are not leaked during analysis and sharing has become a major issue that needs to be addressed. Traditional privacy protection methods, such as data encryption, desensitization, and access control, were widely used in the early days, but as data needs became increasingly complex, the limitations of these methods gradually emerged. Especially in the field of machine learning, data sharing and use often face the risk of privacy leakage, while model training relies on a large amount of data, resulting in an increasingly acute contradiction between privacy protection and data availability. At the same time, the integrity and prediction accuracy of machine learning models are also increasingly attracting attention. How can model owners prove that their models are trained on a compliant basis? More importantly, how can they prove this while ensuring the privacy of the underlying dataset and model?

[0003] The current mainstream solution is zero-knowledge machine learning. Zero-knowledge proof technology can effectively verify the correctness of model predictions without leaking private information. Within the ZK-SNARK framework, the prover uses the given public input (i.e., the basic structure of the model) and private input data (i.e., the user data of the model). The verifier can use proof π to verify the integrity of model training or prediction without accessing private input data. However, in order to ensure the credibility and transparency of zero-knowledge proofs, especially in scenarios involving multiple parties, it becomes essential to rely on a decentralized, tamper-proof platform to maintain trust. At this time, blockchain technology provides an ideal solution. Blockchain can not only ensure the credibility of the operations of all parties through its decentralized and transparent characteristics, but also automatically verify and execute the zero-knowledge proof process through smart contracts to ensure that each step of the operation complies with the predetermined privacy protection specifications. With the help of blockchain, all training processes, model updates, and prediction verifications can be recorded openly and transparently on the chain, which not only enhances data privacy protection, but also improves the auditability and compliance of the entire system.

[0004] However, although blockchain provides a strong trust foundation and security mechanism for zero-knowledge proofs, in practical applications, blockchain technology itself also faces a series of challenges. First, the computing resources on the chain are insufficient and costly. Each node in the blockchain participates in verification and calculation, but the computing power of these nodes is relatively low, and each calculation incurs a computing cost. Second, the security of data transmission between the chain and off-chain. Data on the chain may be stolen or tampered with by certain malicious nodes, and sensitive information is prone to leakage during the transmission of off-chain data. Third, the scalability of smart contracts on the chain is insufficient. Smart contracts face limitations in computing and space resources, are unable to use machine learning models to complete complex business tasks, and the programmability and flexibility of programming languages are insufficient. Summary of the Invention

[0005] In view of this, the present invention proposes a data element privacy protection method based on zero-knowledge machine learning and on-chain verification. Through the ZK-SNARK technology, the model calculation tasks are offloaded to an off-chain oracle server for execution, significantly reducing the on-chain computing cost and effectively protecting privacy. And a zero-knowledge machine learning model on-chain and parameter update algorithm is designed. Through the zero-knowledge gossip protocol, the model parameters processed by differential privacy are quickly synchronized to multiple oracle nodes, and the update process of these parameters is recorded in the blockchain to ensure its security and traceability.

[0006] To achieve the above object, the present invention provides the following technical solutions:

[0007] A data element privacy protection method based on zero-knowledge machine learning and on-chain verification provided by the present invention includes:

[0008] Construct a system model consisting of a decentralized application, contract users, data providers, a distributed oracle architecture, and on-chain verification nodes;

[0009] Through the ZK-SNARK technology, offload the computationally intensive tasks in the model to an off-chain oracle for execution;

[0010] Based on the zero-knowledge machine learning model on-chain and parameter update algorithm, quickly synchronize the model parameters processed by differential privacy to multiple oracle nodes through the zero-knowledge gossip protocol, and record the update process of these parameters in the blockchain.

[0011] Preferably, in the system model:

[0012] The decentralized application is used to provide services for contract users in the form of smart contracts on the blockchain;

[0013] The contract user side includes a model trainer side and a general user side. The general user side uses private and real data as input in exchange for the decentralized services provided by the decentralized application. The model trainer side has the right to utilize the circuits in zero-knowledge proofs;

[0014] The data provider, independent of the blockchain and serving as the data source for user data and business data, verifies the authenticity of the data by signing the data for the user using its private key, and the signed data is verified through the corresponding public key;

[0015] The oracle, adopting a distributed architecture, is used to actively obtain data from off-chain data providers according to the requests of user contracts, generate proofs through ZK-SNARK, and then transmit them to the chain;

[0016] The on-chain verification node is used to verify the Proof and the model incoming to the smart contract.

[0017] Preferably, the zero-knowledge off-chain machine learning trusted data feeding process includes:

[0018] Starting from a request sent from the contract user side, a data feeding request is triggered by an arbitrary event request of the user contract to the oracle contract, then the oracle requests data from an external data source, generates a model proof in the oracle, and then obtains a result through model prediction. Finally, the user gets the returned result.

[0019] Preferably, the zero-knowledge off-chain machine learning application process mainly includes the following steps:

[0020] The contract user side submits a business request according to different functions provided by the centralized application by using the identity authentication information. The user's identity and business data are uploaded to the business contract, and the business contract forwards this information to the oracle on the blockchain after receiving the request;

[0021] Within the time specified by the business contract, the contract event trigger is automatically started and requests the oracle to obtain data;

[0022] The oracle requests the user's privacy data from the data source by carrying the user's identity authentication information through an HTTP request;

[0023] The data source verifies the authenticity of the user's identity according to the identity authentication information, signs the user's privacy data with its private key, and returns the signed data to the oracle;

[0024] The oracle verifies the integrity of the data, performs de-sensitization processing on the data according to the specific business contract model information and circuit information to obtain de-sensitized data, and then sends the business contract model information, circuit information, de-sensitized data, and the current model parameters to the model contract;

[0025] The model trainer sends a Drequest to the model contract to obtain model information, circuit information, model parameters, and desensitized data;

[0026] The model trainer uses the obtained data to train and get new model parameters, generates a public witness of the model through the public parameters and generates a proof, and then returns the training results and the generated data to the model contract of the blockchain by generating the verification key and the proof key of the model circuit;

[0027] The verification nodes of the blockchain regularly and actively verify the model proof transactions uploaded to the blockchain to determine the valid model parameters and fuse the valid model parameters to get new model parameters;

[0028] Repeat the steps of generating circuit proofs again to generate a new set of model proofs for the subsequent verification nodes to verify the model parameters;

[0029] After the verification passes, the machine learning model circuit parameters of this service are recognized by the blockchain and the oracle server for the business calls of ordinary users.

[0030] Preferably, after the oracle server updates the model parameters, the user's personal data will be converted into a private input witness of the model circuit and the corresponding zero-knowledge proof will be generated;

[0031] The verification nodes verify the zero-knowledge proof to ensure the credibility and accuracy of the model inference process;

[0032] After the verification nodes verify successfully, it is determined that the model inference result is correct, and the user's business request can continue to be processed.

[0033] Preferably, the process of the zero-knowledge machine learning model going on-chain and parameter updating algorithm includes:

[0034] Model initialization: For each node, call the Initialize Model() subroutine to initialize the model parameters;

[0035] Distributed training: In each training cycle, each node calls the Train Model() subroutine to update its model parameters;

[0036] Model synchronization and aggregation: Each node calls the Synchronize Model() subroutine to synchronize its model parameters and gets the final model parameters, and aggregates the final model parameters of all nodes into a global final model parameter.

[0037] Preferably, the model initialization specifically includes:

[0038] Generate initial model parameters: The central node or a pre-selected trusted node completes the initialization of the model parameters. During the initialization process, the initial parameters are generated randomly or set based on prior historical data and existing pre-trained models;

[0039] Generate model proofs: Nodes generate zero-knowledge proofs related to the initial model parameters;

[0040] Gossip protocol broadcast: Trusted nodes use the Gossip protocol to broadcast the generated initial model parameters, verification keys, and zero-knowledge proofs in a peer-to-peer manner, enabling the initial model parameters and their proofs to quickly cover the entire distributed network;

[0041] Verify model proofs: Each participating node uses the verification key to verify the received proofs after receiving the initial model parameters and zero-knowledge proofs, uses the model parameters that pass the verification to initialize the local model, rejects the model parameters corresponding to failed verification, and records relevant exceptions;

[0042] Initialization completion and consistency guarantee: When all nodes successfully verify and accept the initial model parameters, it is determined that the initialization process of the model is completed.

[0043] Preferably, the following steps are executed in each local training process corresponding to the distributed training process:

[0044] Obtain the model circuit structure and initial model parameters through the server nodes in the blockchain or oracle;

[0045] Use local data for model training, calculate the gradient of the current model parameters, and update the model parameters according to the calculated gradient and learning rate through the gradient descent algorithm;

[0046] Add noise to the updated model parameters to ensure differential privacy, generate a witness using the updated model parameters and private input, and generate a zero-knowledge proof based on the witness and public key, and the training is completed;

[0047] After training is completed, the node uploads the updated model parameters, zero-knowledge proofs, and verification keys to the blockchain.

[0048] Preferably, during the process of the zero-knowledge machine learning model going on-chain and parameter update algorithm, the following steps are used to perform the update work of the distributed model parameters:

[0049] Each node sends its updated model parameters, zero-knowledge proofs, verification keys, and witnesses to the network through the oracle contract and spreads them among multiple nodes using the Gossip protocol;

[0050] The receiving node asynchronously receives model parameters and zero-knowledge proofs from other nodes, and verifies the received model parameters and zero-knowledge proofs to ensure the legitimacy and reliability of the calculation results;

[0051] If the validation passes, the model parameters are added to a specific set, and if the validation fails, they are discarded;

[0052] After each round of synchronization, the node calculates new global model parameters through an aggregation algorithm based on the model parameters in a specific set;

[0053] The optimization of global model parameters is completed through multiple rounds of iterations. In each round of iteration, each node repeats the broadcast, verification and aggregation steps to gradually enhance the convergence and accuracy of the model and achieve consistency across the entire network.

[0054] When all scheduled iterations are completed, the model parameters of the entire network reach a converged state and form the final global model, ensuring that each node has a consistent and optimized model;

[0055] After the distributed model training synchronization is completed, the system will record the final aggregated model parameters and their corresponding zero-knowledge proof on the blockchain to realize the evidence storage function.

[0056] The present invention has achieved at least the following beneficial effects:

[0057] 1. Through ZK-SNARK technology, the model calculation tasks are delegated to the off-chain oracle server for execution, which significantly reduces the on-chain computing cost and effectively protects privacy.

[0058] 2. Designed a zero-knowledge machine learning model chain-up and parameter update algorithm, which quickly synchronizes the model parameters processed with differential privacy to multiple oracle nodes through the zero-knowledge gossip protocol, and records the update process of these parameters in the blockchain to ensure its security and traceability, and ensure that in the distributed machine learning training of multi-node oracles, the correctness of the model parameters can be verified without leaking training data, preventing security issues such as poisoning attacks.

[0059] 3. Through distributed Oracle technology, ensure that on-chain DApps remain efficient and scalable when handling complex tasks.

[0060] 4. Through zero-knowledge proof and differential privacy technology, the problem of privacy leakage of user data during the interaction between on-chain and off-chain is solved, providing a safe and reliable machine learning solution for privacy-sensitive fields such as medical care and finance.

[0061] Other advantages, objectives, and features of the present invention will be elaborated in the subsequent description, and to some extent, they are obvious to those skilled in the art, or those skilled in the art can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] In order to make the objectives, technical solutions, and beneficial effects of the present invention clearer, the present invention provides the following drawings for illustration:

[0063] Figure 1 It is the architecture diagram of the zero-knowledge machine learning trusted data feed model in the embodiment of the present invention;

[0064] Figure 2 It is the timing diagram of the zero-knowledge machine learning data feed in the embodiment of the present invention;

[0065] Figure 3 It is the architecture diagram of the zero-knowledge machine learning model on-chain and parameter update in the embodiment of the present invention;

[0066] Figure 4 It is the pseudocode schematic diagram of Algorithm 1 in the embodiment of the present invention;

[0067] Figure 5 It is the pseudocode schematic diagram of Algorithm 2 in the embodiment of the present invention;

[0068] Figure 6 It is the pseudocode schematic diagram of Algorithm 3 in the embodiment of the present invention;

[0069] Figure 7 It is the pseudocode schematic diagram of Algorithm 4 in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0070] The following describes the preferred embodiments of the present invention with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.

[0071] The present invention provides a data element privacy protection method based on zero-knowledge machine learning and on-chain verification, following the principle of "off-chain calculation, on-chain verification". Through the ZK-SNARK technology, the model calculation tasks are offloaded to an off-chain oracle server for execution, significantly reducing the on-chain calculation cost and effectively protecting privacy. In addition, to ensure that the correctness of model parameters can be verified without leaking training data and preventing security issues such as poisoning attacks in the distributed machine learning training of multi-node oracles, we designed a zero-knowledge machine learning model on-chain and parameter update algorithm. Through the zero-knowledge gossip protocol, the differentially private processed model parameters are quickly synchronized to multiple oracle nodes, and the update process of these parameters is recorded in the blockchain to ensure its security and traceability.

[0072] By combining ZK-SNARK and distributed Oracle technologies, it provides an efficient and secure solution for the integration of blockchain and machine learning. First, through the "off-chain calculation, on-chain verification" framework, the computationally intensive tasks are moved to off-chain execution, greatly alleviating the computational resource pressure on the blockchain; through the distributed Oracle technology, it ensures that the on-chain DApp remains efficient and scalable when dealing with complex tasks; through zero-knowledge proof and differential privacy technologies, it solves the privacy leakage problem of user data in the on-chain and off-chain interaction process, providing a secure and reliable machine learning solution for privacy-sensitive fields such as healthcare and finance. The research results of this paper demonstrate the collaborative potential of blockchain technology and machine learning in practical applications, expand the application boundary of blockchain in the field of intelligent decision-making, and provide a technical foundation for realizing more powerful smart contracts and decentralized applications.

[0073] The overall architecture of the system model of the solution is as Figure 1 shown, consisting of 8 entities, namely on-chain verification nodes, decentralized applications (Dapps), contract users, model trainers, oracles, government authorities, and non-authoritative institutions. Among them,

[0074] 1) Decentralized application: Provides services in the form of smart contracts on the blockchain.

[0075] 2) Contract users: Contract users are divided into model trainers and ordinary users. Ordinary users can directly access the decentralized services provided by the Dapp, which requires his private and real data as input. For privacy considerations, Dapp users hope to maintain the privacy of their data while enjoying decentralized services. Like a typical blockchain, each DU has one or more public / private key pairs. Model trainers can use circuits in zero-knowledge proofs.

[0076] 3) Data Providers: Data providers include government authorities and non-authorities. Among them, government agencies refer to authoritative institutions such as hospitals and courts, which can provide highly professional and sensitive privacy data; non-authoritative institutions refer to institutions such as factories and weather forecast providers, which can provide non-sensitive data such as business data and guarantee the authenticity of the data. Data providers are independent of the blockchain and are usually the data sources that generate user data and business data. It can verify data by signing data for users using its private key. The DA also knows when the data is generated and with whom the data is associated, but it will never disclose this data to anyone. The signed data can be verified by everyone using the corresponding public key pk a For example, a medical examination report is generated and signed by a trusted hospital and is associated with a specific individual. The hospital plays the role of the DA in this case and is trusted not to disclose the content of the report to others.

[0077] 4) Oracles: Bridges that connect on-chain and off-chain, giving the blockchain the ability to actively obtain data off-chain. It can actively obtain data from data providers and transfer it on-chain. It can ensure the security of data during transmission, but cannot guarantee the reliability of the data source data. The oracle server receives requests from on-chain user contracts, actively obtains data from data providers, generates proofs for the data of data providers through ZK-SNARK, and then returns it on-chain to achieve data privacy.

[0078] 5) On-chain Verification Nodes: Verification nodes can be any node on the blockchain. Verifiers can verify through verification keys and ZKP. On-chain nodes can verify the Proof and models of incoming smart contracts.

[0079] In a specific embodiment, the system model designs a decentralized insurance contract platform based on the blockchain, mainly providing flight insurance services for airline passengers. The platform combines multiple roles, such as contract users (passengers), model trainers (flight company engineers), data providers (authoritative and non-authoritative institutions), and oracles (obtaining off-chain data and providing privacy protection). Specifically as follows:

[0080] a) Decentralized Application: Insurance Contract Dapp

[0081] The Dapp provides flight insurance services for users, including functions such as insurance purchase, compensation cost prediction, and flight risk analysis. Users can conveniently view the risk assessment of their flights through the Dapp and decide whether to purchase insurance based on the assessment results. The Dapp uses smart contracts to execute all insurance transactions, including the signing of insurance and the setting of compensation trigger conditions. To ensure the privacy of user data, the Dapp adopts encryption or zero-knowledge proof technology to protect user data. Even if the user's private data is used for model training, it will not be leaked to any unauthorized third party. Data such as the personal information, flight situation, and health status provided by the user will be encrypted or processed through zero-knowledge proof when input into the model to ensure the privacy of the data. Throughout the process, the specific data content of the user is always protected to avoid any leakage.

[0082] b) Contract user: Passenger

[0083] The contract user in this role refers to the air passenger who purchases insurance. On the platform, passengers obtain risk assessments and insurance quotes by entering data such as flight information, personal health status, and travel plans. Since the privacy of passenger data is of utmost importance, they hope to ensure the security and privacy of their personal data while enjoying decentralized insurance services. Passengers authenticate their identities through public and private keys to ensure that only authorized users can access their private data and transaction records. After selecting an insurance product, passengers provide the required data (such as personal health status, flight information, etc.) and complete insurance purchases and claims requests on the blockchain through smart contracts, without relying on any central entity. During the data transmission and processing process, the passengers' data always remains encrypted to ensure effective protection of privacy.

[0084] c) Model trainer: Flight company engineer

[0085] As the model trainer, flight company employees are responsible for collecting anonymized data of passengers (such as flight information, historical insurance compensation records, etc.) and using this data to train machine learning models for predicting flight risks and insurance compensation amounts. The model trainer verifies the correctness of model training and prediction through the blockchain platform with the help of zero-knowledge proof (ZKP) to ensure that the model output is based on real and fair data. The flight company trains risk assessment and compensation prediction models based on the anonymized data provided by passengers (such as flight delay records, historical accident data, etc.). After the model training is completed, the flight company generates zero-knowledge proofs to ensure that the model is trained based on compliant and real data, without disclosing the private information of passengers. The verified training model proofs will be uploaded and verified through the blockchain smart contract to ensure the transparency and compliance of the entire process.

[0086] d) Data provider: Authoritative and non-authoritative institutions

[0087] Data providers can be divided into authoritative institutions and non-authoritative institutions. Authoritative institutions, such as hospitals and health data providers, are responsible for providing sensitive data related to passengers' health conditions; while non-authoritative institutions, such as weather forecasting companies and flight data providers, provide real-time information related to flights and weather forecasts, etc. Regardless of the type of institution, the core responsibility of data providers is to ensure the authenticity and accuracy of the data, and use private keys to sign the provided data to verify the legality and credibility of the data source. Specifically, authoritative institutions (such as hospitals) provide passengers' health data (such as whether there is a medical history, whether fit to fly, etc.), and sign these data with private keys to ensure the reliability and accuracy of the data. The hospital does not disclose the specific health conditions of passengers, but provides encrypted certification documents to ensure the protection of passengers' privacy. Non-authoritative institutions (such as weather forecasting companies) provide real-time weather information (such as the weather on the flight day, whether there is severe weather, etc.), and these data are also signed and verified to ensure the credibility of their sources. After the data is signed, it is transmitted to the blockchain through an oracle for passengers and model trainers to use. These encrypted and verified data will be used for risk assessment and insurance compensation calculation to ensure that the model makes accurate predictions based on real and credible data while protecting passengers' privacy.

[0088] e) Oracle: The bridge between on-chain and off-chain data

[0089] As the bridge between on-chain and off-chain data, the oracle is responsible for actively obtaining data from external data providers (such as weather forecasting companies, flight companies, etc.) and transmitting it to the blockchain. It provides external data support for the blockchain system, promoting model prediction and the execution of smart contracts. The oracle ensures the security of data during transmission and guarantees data privacy through zero-knowledge proofs (ZK-SNARK), but it cannot fully verify the reliability of the data source. Specifically, after receiving a request from a smart contract, the oracle actively obtains information such as weather and flights off-chain and encrypts the data. Subsequently, the oracle generates zero-knowledge proofs to ensure the privacy and security of the encrypted data before uploading it to the blockchain. Although the oracle can ensure the security of the data transmission process, since it cannot fully verify the accuracy of the data source, it needs to establish a trust mechanism with trusted data providers to ensure the reliability of the data.

[0090] f) On-chain verification nodes: Verifiers in the blockchain network

[0091] Verification nodes are important participants in the blockchain network, responsible for verifying model training results, zero-knowledge proofs, and data authenticity. Each node can ensure the accuracy and compliance of contract execution by checking the parameters of smart contracts, model update proofs, and data signatures. When a verification node receives information such as smart contracts, model training results, and data signatures on the chain, it uses verification keys and zero-knowledge proofs for verification. The verification node confirms whether the model training and data provision processes comply with regulations and ensures that all operations are carried out in a transparent and immutable environment. If the verification is successful, the node records the transaction and continues to execute operations on the chain; if the verification fails, it triggers the exception handling mechanism of the contract to ensure the security and reliability of the system.

[0092] In a specific embodiment, the business process of the system includes:

[0093] A. System initialization

[0094] 1. Parameter initialization: Set global model parameters {θ, info model , info circuit , pk a , sk a}, where θ is the weight parameter of the model, used to represent the training status of the model. info model includes the input, output format of the model, model type (e.g., regression model, classification model, etc.), and other specific information related to the model architecture. info circuit is the configuration related to the zero-knowledge proof circuit, specifically including: the number and type of public inputs (such as the structure information of the model), the number and type of private inputs (such as user data or training data), the circuit proof logic (i.e., how the model verifies the correctness of inputs and outputs), and other information. pk a , sk a are the signature public key and private key generated by the data source. These global parameters {θ, info model , info circuit , pk a} will be broadcast and stored in the consortium blockchain and stored and managed through smart contracts. These parameters not only provide shared basic information for the participants of the system but also ensure the trust and transparency of all parties. On this basis, all participants (such as model trainers, data providers, oracles, etc.) can access and use these global parameters to execute related operations (such as data verification, model training and verification, zero-knowledge proof generation, etc.) through smart contracts. In addition, the system should also ensure the security of these parameters to prevent malicious tampering and data leakage.

[0095] When the system is initialized, the smart contract defines and deploys some basic rules and protocols to ensure that participants can perform operations according to the preset contract terms. For example, data providers can provide encrypted data according to the contract requirements, while model trainers use the public model information and circuit configurations for training and generate zero-knowledge proofs. All these operations will be recorded on the blockchain to ensure the transparency and immutability of each step.

[0096] 2. Circuit Initialization: Circuit initialization is a crucial step in the zero-knowledge proof system. It ensures that the circuit can correctly verify the validity of the computation process and generate appropriate proofs and verification keys. This process involves setting a set of parameters {λ, ξ, pk, vk, C} related to the circuit and is executed as follows:

[0097] Security Parameter λ: First, a security parameter λ needs to be set. This parameter is usually a positive integer representing the security level required by the zero-knowledge proof system. The security parameter determines the circuit complexity and the strength of the encryption algorithm used. A larger λ generally provides higher security but also increases the computational cost.

[0098] Circuit Setup Setup(λ) → ξ: Through the security parameter λ, a set of basic public parameters ξ are generated on a bilinear group (or other appropriate mathematical structure). These parameters include the keys and algorithms required for constructing zero-knowledge proofs and provide the basis for subsequent circuit compilation and proof generation. The public parameter ξ is the information shared by all parties involved in the circuit.

[0099] Circuit Compile Compile(ξ, Info circuit ) → C: Next, the generated public parameter ξ and the detailed configuration of the circuit (i.e., Info circuit , including public inputs, private inputs, proof logic, etc.) are used to compile the circuit. In this step, by performing the compilation process, the circuit is transformed from an abstract description into an executable proof logic, resulting in the circuit description C, which contains the specific rules for how to process input data and verify the correctness of the computation.

[0100] Key Generation GenKey(C) → (pk, vk): After the circuit is compiled, keys for generating proofs and verification need to be generated. Through the key generation algorithm GenKey(C), based on the compiled circuit C, a proof key pk and a verification key vk are generated. The proof key pk is used to generate zero-knowledge proofs, ensuring that the prover can prove the validity of a certain computational process without exposing private data. The verification key vk is used by the verifier to check the validity of the zero-knowledge proof and ensure the correctness of the calculation. The proof key vk is usually held by the circuit builder or model trainer, responsible for providing the necessary key information when generating zero-knowledge proofs. The verification key pk is usually held by the verifier (such as the verification nodes in a blockchain network) and is used to verify the validity of the zero-knowledge proof. This key can also be made public so that anyone can verify the authenticity of the calculation.

[0101] Circuit Parameters and Key Storage: Once these keys and circuit descriptions are generated, relevant information (such as ξ, pk, vk, and C) is usually stored in a blockchain or other decentralized storage systems. For participants (such as model trainers, data providers, etc.), being able to access this information is crucial. The public parameters of the circuit and the verification key will ensure that all parties can perform appropriate verification at different stages without exposing sensitive data.

[0102] B. Trusted Data Feeding Process for Machine Learning under Zero-Knowledge Off-Chain

[0103] In the trusted data feeding process for machine learning under zero-knowledge off-chain here, the circuit model has been initialized and trained, and the model's inference task can be directly completed. This process refers to the entire process starting from when the user makes a request. A data feeding request is triggered by a user contract to a certain event request to the oracle contract, then to the oracle server to request data from an external data source, generate a model proof in the oracle, and then obtain the result through model prediction. Finally, the user gets the returned result.

[0104] As Figure 2 shown, the application process of machine learning under zero-knowledge off-chain mainly includes the following steps:

[0105] 1. User Request: When the user accesses a decentralized application (Dapp), they submit a business request according to the different functions provided by the Dapp. In the request, the user needs to carry identity authentication information, which may include personal identity identification IDAuth i , authentication information, and other privacy data related to the business (such as health data, flight information, etc.). At this step, the user's identity and business data are uploaded to the business contract. After receiving the request, the business contract forwards this information to the oracle server on the blockchain. Here, although the user's privacy data is transmitted, it will not be directly made public, but is processed in a decentralized manner to ensure that privacy is not leaked.

[0106] 2. Oracle Obtains Data: Within the time specified by the business contract, the contract event trigger will automatically start and request the oracle server to obtain data. The oracle requests the user's privacy data from the data source by carrying the user's authentication information IDAuth_i in an HTTP request. The data source may include authoritative institutions such as hospitals and flight companies or non-authoritative institutions such as weather forecast companies. The data source will verify the authenticity of the user's identity and sign the user's privacy data with the private key sk a for the user's privacy data. The signed data sig = Sig(IDAuth i , Hash(data), sk a ), data, and pk a are returned to the oracle server. The oracle server verifies the integrity of the data and performs data desensitization on the data according to the specific business contract model information (info model ) and circuit information (info circuit ) to obtain the desensitized data m. Subsequently, the oracle sends info model , info circuit , m, and the current server's model parameters θ 1 to the model contract. At this time, the data is still in an encrypted or desensitized state, ensuring that the user's privacy is not exposed.

[0107] 3. Off-chain Model Training and Upload: The model trainer sends a Drequest to the model contract to obtain the model information info model , circuit information info circuit , model parameters θ 1 , and desensitized data m. The public data m is converted into the private witness of the model circuit θ 1 and m and other data are used as input parameters for Algorithm 1. The model trainer obtains the new model parameters θ 2 by training with the data, generates the public witness x of the model through the public parameters ξ, θ 1 and θ 2 , generates the proof π 2 , and then generates the verification key and proof key of the model circuit and returns the result to the model contract of the blockchain through . The verification nodes of the blockchain will periodically and actively verify the model proof transactions uploaded to the blockchain. If the result is true, it means that the model inference is correct, the model parameters θ 2 are valid, and the model parameters are fused to obtain the new model parameters θ, and then the previous circuit proof generation steps are repeated to generate a new set of model proofs π 2, for the subsequent verification nodes to verify the model parameters. After successful verification, the machine learning model circuit parameters of this service are recognized by the blockchain and the oracle server and can be used for the service calls of ordinary users.

[0108] 4. Verification and business process: When the oracle server updates the model parameters, the user's personal data will be converted into the private input witness of the model circuit and the corresponding zero-knowledge proof will be generated. The proof will pass through for verification to ensure the credibility and accuracy of the model inference process. After the verification node verifies successfully, the model inference result is considered correct, and the user's business request can continue to be processed. For example, assume this is a prediction service for flight delay insurance. The oracle server inputs the real-time information of the flight and the user's health condition data into the trained model, and the model returns the probability of flight delay and the corresponding insurance compensation amount. If the model inference is successful, the contract will continue to execute the compensation process and pay the corresponding amount to the user. Similarly, for the agricultural insurance compensation caused by bad weather, the user's data will be used for model inference, and the compensation amount will be finally generated.

[0109] All these operations are automatically executed through smart contracts to ensure the decentralization and transparency of the business process. The business contract uses blockchain and zero-knowledge proof technologies to ensure that every step is credible, data privacy is protected, and the model inference process is tamper-proof.

[0110] In a specific embodiment, in the centralized oracle architecture, since all training data, computing resources, and model updates are managed and executed by a single central node, the model parameter update and verification operations during the entire training process are concentrated within this node, without considering the synchronization problem between nodes, nor the need for additional algorithms to coordinate the parameter updates of multiple nodes. After the model training is completed, the central node will directly update the parameters and store the updated model.

[0111] However, in the distributed oracle architecture, since different nodes have their own independent data sets and each node conducts independent training locally, the final model parameters need to be synchronized through an effective parameter update algorithm and merged into a globally consistent model through parameter aggregation operations. The local training results and updated model parameters of each node must be ensured to be legal and consistent through corresponding mechanisms to avoid malicious nodes from tampering with or wrongly updating and affecting the global model.

[0112] To ensure the security of the model update process and data privacy protection, parameter update and aggregation operations in a distributed environment not only require designing effective algorithms to handle parameter synchronization among nodes but also need to incorporate privacy protection mechanisms such as zero-knowledge proofs to ensure that the update process of each node is both legal and does not expose sensitive data. In addition, due to different training progress among distributed nodes, some nodes may fail to update their model parameters in a timely manner due to network latency or failures. Therefore, the entire system also needs to have a certain degree of fault tolerance and consistency guarantee.

[0113] Therefore, to achieve the above goals, a zero-knowledge machine learning model on-chain and parameter update algorithm is proposed to ensure the security, efficiency, and consistency of the model training and update process, referring to Figure 3 . This algorithm has the following core features and innovations:

[0114] 1. Ensure the correctness and privacy protection of model parameters

[0115] By introducing zero-knowledge proof technology, this algorithm ensures that in a multi-node distributed oracle, each update of the model parameters undergoes strict verification, thus guaranteeing the correctness of the training process. Throughout the process, zero-knowledge proof can verify the legitimacy of the training results without exposing the specific content of the training data or model parameters. This design not only protects the privacy of participating nodes but also effectively prevents security risks such as poisoning attacks by malicious nodes, ensuring the credibility of the system.

[0116] 2. Prevent poisoning attacks and malicious updates

[0117] By generating and verifying zero-knowledge proofs during the parameter update process, this algorithm can identify and reject incorrect updates or deliberately tampered model parameters from malicious nodes, fundamentally eliminating the threat posed by poisoning attacks to the global model. After each node completes local training, the submitted model parameters will automatically generate corresponding zero-knowledge proofs, and other nodes will verify the proofs before synchronizing the parameters, thereby ensuring that only legitimate updates can affect the global model.

[0118] 3. Distributed synchronization and gossip protocol

[0119] This algorithm realizes fast distributed synchronization of model parameters through the gossip protocol. In a multi-node environment, each node can gradually synchronize the model parameters generated by local training and the corresponding zero-knowledge proofs to other nodes in a peer-to-peer manner, thus avoiding the single-point failure problem existing in traditional centralized architectures. In addition, by performing differential privacy processing on the model parameters, the privacy protection ability of the system during the distributed synchronization process is further enhanced.

[0120] 4. On-chain record and traceability

[0121] To further enhance security and transparency, the algorithm records the model parameter update process and its corresponding zero-knowledge proofs on the blockchain. Through the immutability and traceability of the blockchain, the update history of all nodes can be securely stored and verified. This design not only ensures the complete transparency of the parameter synchronization process in a distributed environment but also provides technical support for subsequent possible auditing or traceability requirements.

[0122] 5. Global Model Consistency and Security Assurance

[0123] After parameter synchronization is completed, each node performs a parameter aggregation operation (such as weighted average or other methods) based on the collected model parameters and their zero-knowledge proofs to generate globally consistent model parameters. The combination of the update history recorded on the blockchain and the zero-knowledge proof verification mechanism ensures the security, consistency, and legality of the global model parameters under multi-node collaboration.

[0124] This algorithm consists of the following three sub-processes: model initialization, local training of the model by the model trainer and uploading the model, and synchronization, update, and aggregation of model parameters.

[0125] Algorithm 1 Main Function ZKML Training and Synchronization Algorithm introduces the entire process (refer to Figure 4 ):

[0126] 1) Model Initialization: For each node i, call the Initialize Model() subroutine to initialize the model parameters M i .

[0127] 2) Distributed Training: In each training cycle t, each node i calls the Train Model() subroutine to update its model parameters M i .

[0128] Model Synchronization and Aggregation: Each node i calls the Synchronize Model() subroutine to synchronize its model parameters and obtains the final model parameters The final model parameters of all nodes are aggregated into a global final model parameter M*.

[0129] In a specific embodiment, the distributed ZKML model initialization process includes:

[0130] Model initialization is the first step in the training of a distributed zero-knowledge machine learning (ZKML) system. Its main goal is to ensure that all participating nodes in the network can synchronously have consistent and trustworthy initial model parameters, laying a foundation for the subsequent distributed training process. The specific process is as shown in Algorithm 2 (refer to Figure 5 ):

[0131] Generate initial model parameters: The initialization of model parameters is the responsibility of the central node or a pre-selected trusted node. These initial parameters M can be generated randomly, or set based on some prior knowledge (such as historical data) or existing pre-trained models. The generation process of the initial model needs to ensure that it meets the basic requirements of subsequent training and has sufficient robustness.

[0132] Generate model proofs: To ensure the correctness and privacy of model parameters, nodes will generate zero-knowledge proofs related to the initial model parameters. Specifically, nodes will use private inputs and public inputs (private data E, zero-knowledge circuit initialization parameters ξ, and initialization model parameters M) to generate a witness w, and then use this witness and the public key pk to generate a zero-knowledge proof π. The specific definition of the circuit is to calculate the loss value of the model through the model parameters M and private data E. If the requirements are met, the model proof is valid.

[0133] Broadcast via the Gossip protocol: To ensure that all nodes can receive consistent initial model parameters, trusted nodes use the Gossip protocol to broadcast the following: the generated initial model parameters M, verification key vk, and zero-knowledge proof π. The Gossip protocol enables the initial model parameters and their proofs to quickly cover the entire distributed network through a peer-to-peer propagation method, while enhancing the fault tolerance of the propagation.

[0134] Verify model proofs: After each participating node receives the initial model parameters M and zero-knowledge proof π, it performs the following operations:

[0135] A. Verify the zero-knowledge proof: The node uses the verification key vk to verify the received proof π. The verification process depends on the public circuit structure and model parameters M, without accessing the private data E. The verification key vk will be used to verify the zero-knowledge proof. In this process, participants only know the public dataset, the structure of the circuit, and the public witness (model parameters M), and can know whether the model parameters are reliable without knowing the specific values of the private data of real users.

[0136] B. Accept or reject model parameters: The model parameters that pass the verification are accepted and used to initialize the local model M i ; if the verification fails, the node will reject the parameter and record the relevant exception.

[0137] Initialization completion and consistency guarantee: When all nodes successfully verify and accept the initial model parameters M, the model initialization process is completed. Through the combination of zero-knowledge proofs and the Gossip protocol, the security, consistency, and traceability of the initialized model parameters are ensured. Each node can be convinced of the credibility and legality of the initial model parameters without accessing the private data of others.

[0138] In a specific embodiment, in a distributed zero - knowledge machine learning (ZKML) system, local training is an important process where each participating node independently updates the model parameters based on its own data. This process aims to ensure the credibility and privacy of the training process through zero - knowledge proof technology, while preventing data leakage through differential privacy technology. As Figure 3 shown, during the local training process, the model trainer interacts with the blockchain and oracle nodes to obtain the necessary information. After completing the model training, the generated proof and updated model parameters are uploaded to the blockchain for the whole network to verify and synchronize.

[0139] The specific details are as in Algorithm 3, referring to Figure 6 shown:

[0140] 1. Obtain circuit and model information

[0141] The model trainer obtains the model circuit structure and initial model parameters through server nodes in the blockchain or oracle. These information define the computational processes required during training, including the definition of the loss function, the gradient calculation method, and the construction of the zero - knowledge circuit, etc.

[0142] 2. Local gradient calculation and update

[0143] Model Trainer i uses local data E i to perform model training and obtain the gradient of the current model parameter M i . The gradient calculation can be expressed as

[0144]

[0145] where represents the loss function, represents the current model parameter. Then, according to the calculated gradient and the learning rate η, the model parameter M i is updated. Specifically, it is This process is called the gradient descent step. The learning rate η is a hyperparameter.

[0146] 3. Differential privacy processing

[0147] To protect data privacy, the updated model parameters need to add noise N i to ensure differential privacy. The generation method of noise N i is to follow a Gaussian distribution:

[0148]

[0149] where Δ is the sensitivity and ∈ is the privacy budget. The model parameter after adding noise is expressed as:

[0150] M′ i = M i + N i

[0151] The purpose of adding noise is to make it difficult for external attackers to infer the specific content of the original data E even if they obtain the model parameters M′ i , thus achieving privacy protection. i

[0152] 4. Generate zero - knowledge proof

[0153] The model trainer uses the updated model parameters M′ i and the private input ξ to generate a witness w i :

[0154] w i = GenWitness(M′ i , ξ)

[0155] Based on the witness w i and the public key pk, generate a zero - knowledge proof π i :

[0156] π i = Prove(M′ i , pk, w i )

[0157] This proof ensures the correctness of the training process and parameter updates, while hiding the content of private data.

[0158] 5. Model proof upload and verification

[0159] After training is completed, the node uploads the following information to the blockchain: the updated model parameters M′ i (public input), the zero - knowledge proof π i , and the verification key vk. After the blockchain verification node receives these data, it uses the verification key vk to verify the zero - knowledge proof π i :

[0160] A. Verification passed: It indicates that the node's calculation is correct and trustworthy, and the updated model parameters are recorded and synchronized to the oracle node.

[0161] B. Verification failed: The update of this node is rejected to prevent untrusted data from contaminating the global model.

[0162] 6. Interaction and synchronization

[0163] Through the above process, the updated model parameters of all local nodes and their proofs are verified and recorded on the blockchain. The verified parameters are synchronized to the distributed network through the oracle contract, ensuring that all nodes in the network can perform further global model aggregation based on the trusted local update results.

[0164] In a specific embodiment, the distributed ZKML model parameter update is a key step to achieve global model consistency and optimization, mainly completed through repeated iterations of broadcasting, verification, and aggregation, ensuring that the system improves computational efficiency and security while protecting data privacy.

[0165] As Figure 3 shown, the model parameter M′ i , the proof π i , the verification key vk, and the witness w i are sent to the oracle nodes through the oracle contract and propagated to all nodes through the Gossip protocol. T1, T2, and T3 in the figure are the rounds of Gossip propagation, and each propagation in the figure is these parameters. In distributed machine learning, model synchronization is an iterative process, and the consistency of global model parameters is achieved through multiple rounds of broadcasting, receiving, verification, and aggregation.

[0166] Specifically, as shown in Algorithm 4, refer to Figure 7 shown:

[0167] 1. Broadcast of parameters

[0168] Each node will send its updated model parameter M′ i , the zero-knowledge proof π i , the verification key vk, and the witness w i to the network through the oracle contract. And it is propagated among multiple nodes using the Gossip protocol. The diffusion of information is completed in turn according to the rounds T1, T2, T3 in Figure 3 . The nodes for each broadcast are randomly selected by the system, while ensuring that all nodes in the network have the opportunity to receive this information. This randomized mechanism can not only effectively avoid network isolation problems but also significantly reduce the risk of single-point failures, thereby improving the robustness and reliability of the entire system.

[0169] 2. Receiving and verifying model parameters

[0170] A. Asynchronous reception: Node P will asynchronously receive the model parameters and zero-knowledge proof π from other nodes. The number and time of reception vary depending on the network conditions. This flexibility improves the fault tolerance of the system, allowing the system to operate normally in a dynamic and uncertain network environment.

[0171] B. Zero - Knowledge Verification: Each node strictly verifies the received model parameters and zero - knowledge proofs to ensure the legality and reliability of the calculation results. The specific steps of the verification process are as follows:

[0172] IsVali|d j ←Verify(w j ,vk,π j )

[0173] Verify the model parameters against the set circuit conditions through the witness w i , verification key vk, and proof π j . If the verification passes, the model parameter M j ′ will be added to the set R; if the verification fails, the parameter will be discarded, thus defending against malicious attacks.

[0174] 3. Aggregation of Parameters

[0175] After each round of synchronization, the nodes will calculate new global model parameters through an aggregation algorithm. Common aggregation methods include:

[0176] Simple Averaging: The model parameters of all nodes participate in the aggregation with the same weight. The formula is as follows:

[0177]

[0178] This method is simple to calculate and is suitable for scenarios where the node data volume and computing power are similar.

[0179] Weighted Averaging: Weights w j are assigned to the model parameters according to the characteristics of each node (such as local data volume or computing power) for weighted calculation:

[0180]

[0181] where: ω j is the weight of the j - th node, usually calculated based on the local data volume or computing power of the node, and the weight ω j >0. is the normalization factor of the weights, used to ensure the correctness of the calculation results.

[0182] Weighted averaging can more effectively reflect the differences in data distribution and computing resources, thereby improving the efficiency and fairness of model training. The aggregated global model parameter M can not only more comprehensively synthesize the calculation results of each node but also reduce the impact of abnormal nodes on the overall model, ensuring the accuracy and stability of the global model. In multiple rounds of iteration, the continuous optimization of parameter aggregation makes the global model gradually converge, ultimately achieving high - quality distributed machine learning.

[0183] 4. Multi - Round Iterative Synchronization

[0184] The optimization of global model parameters is completed synchronously through N rounds of iterations. In each round of iteration, each node repeats the broadcast, verification and aggregation steps to gradually enhance the convergence and accuracy of the model and achieve consistency across the entire network. The specific description is as follows:

[0185] A. Multi-round broadcast: In each round of iteration, each node will update the current model parameter M′ i And zero-knowledge proof π i The information is then propagated to other randomly selected neighboring nodes through the Gossip protocol. This approach ensures that each node has a high probability of receiving model updates from all other nodes, thus avoiding information islands.

[0186] B. Step-by-step verification: In each round, the node independently verifies the received model parameters and zero-knowledge proofs, and only accepts model parameters that have passed the verification. This mechanism effectively resists attacks by malicious nodes and prevents incorrect parameters from contaminating the global model.

[0187] C. Iterative aggregation: After each round of synchronization is completed, each node aggregates the model parameters that have passed the verification in the current round to generate new global model parameters. This process gradually integrates the calculation results of each node, so that the global model gradually converges.

[0188] D. Improved convergence: Through N rounds of synchronous iterations, the global model parameters in the system will continue to approach the optimal solution. In a distributed environment, multiple rounds of iterations can also effectively alleviate the problems caused by network delays or node asynchrony, ensuring the accuracy and consistency of the final results.

[0189] When all scheduled iteration rounds are completed, the model parameters of the entire network reach a converged state and form the final global model, ensuring that each node has a consistent and optimized model.

[0190] 5. Evidence of final results

[0191] After the distributed model training synchronization is completed, the system records the final aggregated model parameter M and its corresponding zero-knowledge proof π on the blockchain to realize the evidence storage function. This process not only improves the transparency and trust of the distributed system, but also provides security and traceability for subsequent verification and auditing. The specific description is as follows:

[0192] A. Model parameter storage: The final global model parameter M is aggregated and verified, and then uploaded to the blockchain for storage. The tamper-proof nature of the blockchain ensures that the recorded parameters are complete and reliable, preventing parameter contamination caused by malicious modification or system failure.

[0193] B. Zero-Knowledge Proof Archiving: The zero-knowledge proof π uploaded together with the model parameters is used to prove the legitimacy of the parameters and the correctness of the calculation process. The blockchain verification nodes can verify π through the public verification key vk to ensure that the final model meets the set circuit conditions without accessing the private data of the nodes.

[0194] C. Security and Traceability: The archived model parameters and proofs provide an open and trustworthy basis for subsequent use. Whether it is the audit of the operation process of the distributed system or the verification of the model results, archiving provides a reliable data source, enhancing the security and transparency of the system.

[0195] D. Multi-Party Verification: After the archived information is stored on the blockchain, any authorized participating party can read the relevant records and verify the validity of the model parameters. This mechanism not only enhances the openness of the system but also further improves the trust level in cross-institutional collaboration.

[0196] Through the archiving function of the blockchain, every step of the distributed model training can be strictly recorded and verified, providing important support for building a secure, transparent, and efficient distributed system.

[0197] Finally, it should be noted that the above preferred embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail through the above preferred embodiments, those skilled in the art should understand that various changes can be made in form and details without departing from the scope defined by the claims of the present invention.

Claims

1. A data element privacy protection method based on zero-knowledge machine learning and on-chain verification, characterized in that: include: Build a system model consisting of decentralized applications, contract users, data providers, distributed oracle architecture, and on-chain verification nodes; Through ZK-SNARK technology, the computationally intensive tasks in the model are delegated to off-chain oracles for execution; Based on the zero-knowledge machine learning model chain and parameter update algorithm, the model parameters that have undergone differential privacy processing are quickly synchronized to multiple oracle nodes through the zero-knowledge gossip protocol, and the update process of these parameters is recorded in the blockchain.

2. According to claim 1, a data element privacy protection method based on zero-knowledge machine learning and on-chain verification is characterized in that: In the system model: Decentralized applications are used to provide services to contract users in the form of smart contracts on the blockchain; Contract users, including model trainers and ordinary users. Ordinary users use private and real data as input in exchange for decentralized services provided by decentralized applications. Model trainers have the right to use the circuits in zero-knowledge proofs. Data providers, which are independent of the blockchain and serve as the data source for user data and business data, verify the authenticity of data by signing data for users using their private keys. The signed data is then verified using the corresponding public key; The oracle uses a distributed architecture to actively obtain data from off-chain data providers based on user contract requests and transmits it to the chain after generating proofs through ZK-SNARK; On-chain verification nodes are used to verify the Proof and model passed into the smart contract.

3. According to claim 1, a data element privacy protection method based on zero-knowledge machine learning and on-chain verification is characterized in that: The trusted data feeding process of zero-knowledge off-chain machine learning includes: Starting from the request issued by the contract user, a data feed request is triggered by the user contract to request any event to the oracle contract, and then the oracle requests data from the external data source, generates a model proof in the oracle, and then predicts the result of the model, and finally the user gets the return result.

4. According to claim 3, a data element privacy protection method based on zero-knowledge machine learning and on-chain verification is characterized in that: The zero-knowledge chain machine learning application process mainly includes the following steps: The contract user uses the identity authentication information to submit a business request based on the different functions provided by the centralized application. The user's identity and business data are uploaded to the business contract. After receiving the request, the business contract forwards this information to the oracle on the blockchain. Within the time specified in the business contract, the contract event trigger is automatically started and requests the oracle to obtain data; The oracle requests the user's private data from the data source via HTTP requests carrying the user's identity authentication information; The data source verifies the authenticity of the user's identity based on the identity authentication information, signs the user's private data with the private key, and returns the signed data to the oracle; The oracle verifies the integrity of the data and desensitizes the data according to the specific business contract model information and circuit information to obtain desensitized data. Then, the business contract model information, circuit information, desensitized data and current model parameters are sent to the model contract. The model trainer sends a Drequest to the model contract to obtain model information, circuit information, model parameters, and desensitized data; The model trainer obtains new model parameters through data training, generates a public witness of the model and a proof through public parameters, and then returns the training results and generated data to the model contract on the blockchain by generating the verification key and proof key of the model circuit. The verification nodes of the blockchain regularly and proactively verify the model proof transactions uploaded to the blockchain to determine the valid model parameters, and merge the valid model parameters to obtain new model parameters; Repeat the steps of generating circuit proofs again to generate a new set of model proofs for subsequent verification nodes to verify model parameters; Once the verification is passed, the circuit parameters of the business's machine learning model are recognized by the blockchain and oracle server and can be used by ordinary users for business calls.

5. According to claim 4, a data element privacy protection method based on zero-knowledge machine learning and on-chain verification is characterized in that: When the oracle server updates the model parameters, the user’s personal data is converted into a private input witness of the model circuit and a corresponding zero-knowledge proof is generated; The verification node verifies the zero-knowledge proof to ensure the credibility and accuracy of the model reasoning process; After the verification node is successfully verified, it is determined that the model inference result is correct and the user's business request can continue to be processed.

6. According to claim 1, a data element privacy protection method based on zero-knowledge machine learning and on-chain verification is characterized in that: The zero-knowledge machine learning model on-chain and parameter update algorithm process includes: Model initialization: For each node, call the Initialize Model() subroutine to initialize the model parameters; Distributed training: In each training cycle, each node calls the Train Model() subroutine to update its model parameters; Model synchronization and aggregation: Each node calls the Synchronize Model() subroutine to synchronize its model parameters and obtain the final model parameters, aggregating the final model parameters of all nodes into a global final model parameter.

7. A data element privacy protection method based on zero-knowledge machine learning and on-chain verification according to claim 6, characterized in that: Model initialization specifically includes: Generate initial model parameters: The model parameters are initialized by the central node or a pre-selected trusted node. During the initialization process, the initial parameters are generated randomly or set based on prior historical data and existing pre-trained models. Generate model proof: Generate zero-knowledge proof related to initial model parameters through nodes; Gossip protocol broadcast: Trusted nodes use the Gossip protocol to broadcast the generated initial model parameters, verification keys, and zero-knowledge proofs in a point-to-point manner, so that the initial model parameters and their proofs can quickly cover the entire distributed network; Verify model proof: After receiving the initial model parameters and zero-knowledge proof, each participating node uses the verification key to verify the received proof, uses the verified model parameters to initialize the local model, refuses to accept the corresponding model parameters that fail the verification, and records the relevant exceptions; Initialization completion and consistency assurance: When all nodes successfully verify and accept the initial model parameters, the model initialization process is determined to be completed.

8. A data element privacy protection method based on zero-knowledge machine learning and on-chain verification according to claim 6, characterized in that: The following steps are performed in each local training process corresponding to the distributed training process: Obtain the model circuit structure and initial model parameters through the server node in the blockchain or oracle; Use local data to train the model, calculate the gradient of the current model parameters, and update the model parameters according to the calculated gradient and learning rate through the gradient descent algorithm; Add noise to the updated model parameters to ensure differential privacy, generate a witness using the updated model parameters and private input, generate a zero-knowledge proof based on the witness and public key, and the training is complete; After training is completed, the node uploads the updated model parameters, zero-knowledge proof, and verification key to the blockchain.

9. A data element privacy protection method based on zero-knowledge machine learning and on-chain verification according to claim 6, characterized in that: During the zero-knowledge machine learning model on-chain and parameter update algorithm process, the distributed model parameter update is performed through the following steps: Each node sends its updated model parameters, zero-knowledge proof, verification key, and witness to the network through the oracle contract, and uses the Gossip protocol to propagate among multiple nodes; The receiving node asynchronously receives model parameters and zero-knowledge proofs from other nodes, and verifies the received model parameters and zero-knowledge proofs to ensure the legitimacy and reliability of the calculation results; If the validation passes, the model parameters are added to a specific set, and if the validation fails, they are discarded; After each round of synchronization, the node calculates new global model parameters through an aggregation algorithm based on the model parameters in a specific set; The optimization of global model parameters is completed through multiple rounds of iterations. In each round of iteration, each node repeats the broadcast, verification and aggregation steps to gradually enhance the convergence and accuracy of the model and achieve consistency across the entire network. When all scheduled iterations are completed, the model parameters of the entire network reach a converged state and form the final global model, ensuring that each node has a consistent and optimized model; After the distributed model training synchronization is completed, the system will record the final aggregated model parameters and their corresponding zero-knowledge proof on the blockchain to realize the evidence storage function.

Citation Information

Patent Citations

  • Validation of measurement data sets using oracle consensus

    CN113950679A

  • SNARK-based non-interactive public verifiable system and calculation method in block chain scene

    CN116545603A

  • System and method for decentralized data management and dynamic verification, valuation, and monetization of data queries

    US12155781B1