A data trusted computing framework based on alliance block chain
By using a data trust computing framework based on consortium blockchain, the issues of protecting the ownership of data providers and supervising data demanders during the data computing process are resolved, thus achieving data security and integrity, reshaping trust relationships in the data computing process, and preventing data loss of control and leakage.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-28
- Publication Date
- 2026-03-24
AI Technical Summary
Existing technologies neglect the protection of the data provider's ownership of the original data during the data computation process, and do not consider the regulatory measures after the data requester obtains the data, resulting in damage to the interests of the data provider and the existence of data loss of control and potential risks.
By adopting a data trust computing framework based on consortium blockchain, data is standardized and registered by dividing it into three parts: identifier, feature data, and data entity. A data trust computing mechanism and traceability mechanism are designed to realize data processing as a service, ensuring that data providers can provide data analysis results without data requesters having access to the original data, and ensuring the security and integrity of data flow through on-chain traceability.
It ensures data security and control for data providers, prevents data leaks, rebuilds trust relationships in the data computation process, alleviates data barriers and loss of control issues, and ensures that data requesters obtain analysis results while protecting the interests of data providers.
Smart Images

Figure CN119885267B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of big data computing technology, and in particular to a trusted data computing framework based on consortium blockchain. Background Technology
[0002] With the advent of the big data era, how to achieve secure and reliable data computing on the open, dynamic, and uncontrollable internet has become an urgent problem to be solved. Currently, most centralized data trading platforms operate on a Data-as-a-Service (DaaS) model, where raw data is stored on a third-party data trading platform, and data buyers obtain the seller's entire raw dataset through the third-party platform. Traditional data hosting trading models are shown in the attached... Figure 1 As shown, this typically involves three parties: the data provider, the traditional data platform (i.e., some intermediary), and the data requester. Specifically, the data provider sends the dataset to a trusted data exchange platform and sets an appropriate selling price. The data requester selects data products of interest and places an order online, similar to an online supermarket. After receiving payment from the data requester, the traditional data platform transmits the purchased data to the data requester and pays the data provider (after deducting management fees or commissions). The data requester usually obtains the entire dataset from the traditional data platform. However, the data requester may only need the results of data analysis to inform data-driven decision-making, rather than accessing the complete dataset. For example, the data requester may not need access to all the revenue data for a specific quarter of a year from the data provider, but may only need a summary of certain statistics from that quarter's revenue data. Moreover, transmitting the entire dataset to the data requester poses a risk of data falsification, malicious collection and resale of user data resources by the platform, especially if data oversight is weak. Furthermore, detecting whether user data has been tampered with or resold by the platform, and balancing the allocation of responsibilities and rights in the data sharing process, is extremely difficult.
[0003] Blockchain technology, with its decentralized and immutable characteristics, provides a technological approach for building trusted data computing mechanisms. Simultaneously, blockchain can clearly define the ownership of data resources, including data ownership and privacy protection rights, incentivizing data providers to become members of the blockchain network. However, traditional public blockchains such as Bitcoin and Ethereum suffer from low throughput and poor privacy, limiting their application in the big data market. This has led to increasing attention on the development of consortium blockchains. Hyperledger Fabric, as a representative of enterprise-grade consortium blockchains, features a highly modular and configurable architecture, effectively meeting the business needs of current big data scenarios.
[0004] Current research primarily focuses on building data computing platforms without third-party intervention using consortium blockchains, aiming to enable efficient data resource flow between various data platforms and systems. Simultaneously, a unified trust system is jointly constructed by the data computing participants, protecting their privacy while ensuring the data provider's control over the data. However, due to the highly redundant storage method of blockchain, as the number of blockchain nodes increases, data replication and storage costs rise simultaneously. This leads to significant waste of storage resources and intolerable performance degradation in practical applications. Considering the pressure on on-chain storage and system throughput, some researchers are optimizing on-chain data storage in off-chain repository systems by combining new standardized protocols such as IPFS and DOIP, thereby better achieving interconnection, interoperability, and interoperability of heterogeneous, cross-domain, and cross-owner data.
[0005] However, the above solutions still have certain limitations. Existing research generally focuses on ensuring the integrity and security of data delivery / transmission, neglecting the protection of the data provider's ownership of the original data, and failing to consider regulatory measures after the data requester obtains the data. If dishonest data requesters resell the data resources needed for computing, it will greatly affect the interests of data providers and their motivation to share data.
[0006] To address this, this invention proposes a data-trusted computing framework based on consortium blockchain. First, a data architecture paradigm based on consortium blockchain is defined, standardizing the data registration process by dividing data into three parts: identifiers, feature data, and data entities. This provides a unified access method for heterogeneous, multi-source, and multi-domain data resources, efficiently integrating data resources. Simultaneously, a data-trusted computing mechanism is designed. This transforms data hosting / exchange as a service into data processing as a service. Instead of hosting the original data with a third party for data requesters to analyze, the data provider provides the results of data resource analysis and processing without the data requester accessing the original data, achieving "usable but invisible" data to address trust challenges in the data computing process. Furthermore, this invention designs a data-trusted traceability mechanism, ensuring the security and integrity of data flow between data requesters, computing nodes, and data providers by leaving traces on the data computing process on the blockchain. Summary of the Invention
[0007] This invention proposes a trusted data computing framework based on consortium blockchain, transforming the traditional data hosting / exchange-as-a-service model into a data processing-as-a-service model. Instead of hosting raw data with a third party for data requesters, the data provider offers the results of data resource analysis and processing without the data requester having access to the raw data, thus achieving "usable but invisible" data. This avoids the leakage of raw data from the data provider while not affecting the data requester's access to the data analysis results, rebuilding the trust relationship between the two parties in the data computing process. It can alleviate many problems inherent in traditional data computing platforms, such as data silos, data loss of control, and data ownership issues.
[0008] The present invention adopts the following technical solution.
[0009] A data-trusted computing framework based on consortium blockchains is proposed. The entities within the Data Trusted Computing Framework (DTCF) include data requesters, data providers, consensus nodes, and certificate authorities; such as... Figure 2 As shown, the computing framework is structured into a data registration layer, a data trust layer, a data consensus layer, a data processing layer, and a data delivery layer.
[0010] The data registration layer adopts a data architecture paradigm based on consortium blockchain, which divides data into identifiers, feature data, and data entities to standardize the data registration process, providing a unified access method for heterogeneous, heterogeneous, and heterogeneous data resources, and efficiently integrating data resources.
[0011] The data trust layer uses a data trust computing mechanism to enable data providers to provide data resource analysis and processing results to data requesters without them having access to the original data.
[0012] The data consensus layer includes a data trust traceability mechanism, which ensures the security and integrity of data flow between data requesters, computing nodes, and data providers by leaving traces on the data computation process chain.
[0013] The data computation process of the trusted data computing framework includes the following steps;
[0014] Step S1: The data provider transforms the data resources according to the data architecture paradigm and registers them on the blockchain;
[0015] Step S2: The data requester discovers data through the registry query module, finds interested data providers and suitable computing nodes to match the demand, and then sends a data request carrying purchase pricing information to the consortium blockchain network.
[0016] Step S3: The consortium blockchain network sends the request to the corresponding computing node. The computing node encapsulates the data request to obtain data resources from the data provider. Then, it performs data analysis and processing locally on the node. Only the execution result is sent to the data requester, not the original data.
[0017] Step S4: The computing nodes that meet the computing requirements will store the data computing result storage record containing the signatures of the data computing participants in the consortium blockchain network of the trusted data computing framework.
[0018] Step S5: Automatically pay through the smart contract module; evaluate the satisfaction level of this data calculation process.
[0019] The data registration layer is responsible for transforming data architecture rules and standardizing the registration process for heterogeneous data resources from different sources and regions; the data trust layer is used to ensure the security and trustworthiness of data transmission and record the hash digest of the calculation results to prevent tampering; the data consensus layer is used to ensure consensus on the distributed ledger data on the chain; the data processing layer is responsible for controlling the process of data calculation; and the data delivery layer is used to support the trusted delivery and evaluation of the final data processing results.
[0020] Let DP i For data providers, DR i For data demanders, CN i For computing nodes, To compute the state information of the nodes, For data resources, UID i A unique identifier for data resources, UI i For the identification of data resources, CD i For the characteristic data of data resources, DE i For data entities of data resources, CBDAO i For data architecture objects, UUID i The following are the key components of the data computation process: RC (Registry Contract), DRMC (Data Request Management Contract), RISC (Result Information Storage Contract), DEMC (Data Evaluation Management Contract), and SK. i For private key, PK i For public key, Cert i For digital certificates, K i For symmetric keys, Sig() is the short signature function, and dataProof... i For data calculation and recording, ω is the satisfaction evaluation index;
[0021] Data Provider: A data provider is an individual user or organization that owns data resources. They access this framework through FabricSDK, publish their data resources, and sell processed data resources to suitable data providers through demand matching.
[0022] Data demanders are potential customers of data resources. They access this framework through the Fabric SDK and discover data from all network nodes by querying the registry to match suitable data providers to meet their data computing needs.
[0023] Compute Node (CN): A device or node used to meet the computing needs of different data technology scenarios, assuming that it is secure, reliable, and capable of performing relevant computing tasks correctly and independently.
[0024] Consensus Nodes: Consensus nodes consist of all nodes that have joined the trusted computing framework for data, and together they maintain consensus on the distributed ledger data on the chain, ensuring the security and consistency of the consortium blockchain network.
[0025] The trusted data computing framework uses a Certificate Authority (CA): an authoritative and trusted certification authority responsible for authorizing and issuing exclusive public-private key pairs to data computing participants accessing the framework. Data computing participants will use these keys to verify their identity and encrypt data during communication within the consortium blockchain network.
[0026] In step S1, the data architecture paradigm CBDA (Consortium Blockchain-based Data Architecture) is used to provide a unified way to persistently identify, describe, locate, and manage heterogeneous data resources from different sources and regions. The CBDA transformation rules will register data providers (DPs). j Data resources It is divided into identifiers, feature data, and data entities; among which the identifier (UI) i It is an identifier for data resources, used to persistently and uniquely identify each data resource, and is described using the following definition 1:
[0027] Definition 1.UI i It is a data resource In CBDA, the identifier portion is represented by the following quadruple:
[0028]
[0029] UID i It is a data resource The unique identifier, tStamp is the generated timestamp, PK j It is the data provider DP j SK's public key j It is the data provider DP j The private key, M is a cryptographic string defined by the data provider, and Sig() is a short signature function. It is the data provider DP jUse its private key SK j The signature information for string M;
[0030] Feature Data CD i This refers to metadata information of data resources, used for retrieving and discovering corresponding data resources. Its specific meaning is defined in Definition 2 below:
[0031] Definition 2.CD i It is a data resource In CBDA, the feature data portion, i.e., metadata, is represented as the following quadruple:
[0032]
[0033] Meta i It is a data resource metadata, It is the data provider DP j Use its private key SK j Data resources UID i Signature information, apiInfo i It is a data resource The API call information;
[0034] Data Entity DE i It is a mapping of data resources after they have been processed by the registry storage module of the on-chain registry contract RC, rather than the original data, as shown in Definition 3 below:
[0035] Definition 3.DE i It is a data resource In CBDA, the data entity portion is represented as the following quintuple:
[0036]
[0037] Among them AD i Represents data resources The access address, dl represents the valid expiration time of the data resource, AS is the status indicator indicating whether the data resource is available, K DTSP This represents the proposed symmetric key. Through the symmetric key K DTSP Data resources The signature information of the access address;
[0038] Data Architecture Objects (CBDAO) i This refers to the final formatted storage form of data resources within the proposed framework, the specific meaning of which is defined in Definition 4 below:
[0039] Definition 4.CBDAO i It is a data resource After CBDA rule transformation, the on-chain representation is shown as the following triple:
[0040] CBDAO i =(UI) i CD i DE i ) Formula 4;
[0041] UI i CD i DE i These are data resources The identifier, feature data, and data entity parts after CBDA rule transformation.
[0042] The CBDA data transformation process is as follows: The proposed framework allocates a UI definition 1 for each data resource. i This serves as the unique identifier for its corresponding CBDA object; simultaneously, it constructs the CD defined in Definition 2. i and with preset mapping functions and UI i One-to-one correspondence; finally, encapsulate the data entity DE defined in definition 3 in a standardized format. i It also provides an interface for receiving data requests to the outside world;
[0043] When a data provider is ready to publish its data resources, it performs standardized registration of the data resources according to the data transformation process described above; the data registration process uses Algorithm 1.
[0044] Algorithm 1 is expressed in pseudocode as follows:
[0045]
[0046] The initialization parameters and functionalities required in Algorithm 1 are as follows:
[0047] `registerCheck` is used to identify whether the data registration requirements are met; `verifyId()` is used to verify the authentication information of the role; `decrypt()` is used for decryption; `matchStyle()` is used to verify the compliance of the data resource format; and `DataHashList`... i Collection records data resources The data hash value generated by all computational operations is used by the `dataCheck()` function to verify whether there is duplicate registration, and the `register()` function is used to register CBDAO. i Registering on the blockchain involves using the `update()` function to update the state of the distributed ledger on the chain. `registerProperty` is the data resource. The final registration credential on the blockchain is obtained through the getInfo() function;
[0048] The data registration process is as follows: First, the registry contract (RC) verifies the role identity and parses out the data schema object CBDAO to be registered using the verifyId() and decrypt() functions. i Then, the `matchStyle()` and `dataCheck()` functions are used to check the data format and whether there are duplicate registrations. If the above conditions are met, CBDAO will be processed. i Register and return the credentials for this data registration.
[0049] The DTCF data calculation process includes the following security features:
[0050] Identity authentication: Identity authentication means that the key identity information of the data requester should not be forged by malicious participants, thereby bypassing the authentication and authorization of the consortium blockchain network to carry out illegal data calculations. It is necessary to ensure the privacy and security of the consortium blockchain network.
[0051] Data autonomy: Whether data is shared should be entirely controlled by the data provider, meaning the data provider can decide whether to share its data resources; Data security and integrity: The data provider's data is encrypted during sharing, and no one other than the computing node can access the requested original data resources; Data integrity is protected during the data computing process, meaning the data should not be tampered with during the data computing process;
[0052] Security against dishonest data providers: Dishonest data providers are those who provide invalid data that is irrelevant to the intended data request, or those who send additional data to gain more benefits; the trusted data computing framework can identify and resist dishonest data providers;
[0053] Regarding the security of dishonest computing nodes: Dishonest computing nodes are data nodes that forge data requests to obtain data resources from data providers or fail to process data according to predetermined requirements. The trusted data computing framework can strictly monitor this phenomenon, protect the interests of data computing participants, and prevent insider theft.
[0054] For the security of dishonest data requesters: For dishonest data requesters who send false data requests to data providers by forging the identity information of other participants, or who deny and refuse to pay for completed data computation requests, the trusted data computing framework can preserve key data computation records or take reliable security measures to protect the interests and reputation of data providers.
[0055] The smart contract module of the Data Trusted Computing Framework (DTCF) includes functions such as data registration, demand matching, data computation, and satisfaction evaluation, which are presented in the form of a Registry Contract (RC), a Data Request Management Contract (DRMC), a Result Information Storage Contract (RISC), and a Data Evaluation Management Contract (DEMC), respectively.
[0056] RC: RC stores and manages all registered data resources on the blockchain through the registry storage module, providing a unified access method for data resources from different sources; at the same time, RC's registry query module provides a query interface for data requesters to match suitable data resources.
[0057] DRMC: DRMC's request information storage module is responsible for storing and managing all data requests during the data computation process, including data computation requests initiated by data requesters, data computation requests processed by the proposed framework, and data acquisition requests initiated by computing nodes; DRMC's prepayment storage module requires data requesters to prepay to data providers and computing nodes after demand matching.
[0058] RISC: RISC stores and manages key evidence information during the data computation process through a data computation evidence storage module, including the data provider's data response information and the trusted credentials of the final data computation results on the blockchain by the computing nodes; the result matching module statistically analyzes the data processing results of the computing nodes, and the result with the most identical results is the final result; the payment module is responsible for issuing payments to the transaction addresses of the data provider and the computing nodes based on the final consistent result data.
[0059] DEMC: DEMC stores the identity information of all data computation participants through the user information storage module; the evaluation storage module is used to store the satisfaction evaluation of the data requester on the data provider and computing node after each data computation process is completed; the evaluation query module can help the data requester better select the appropriate data provider and computing node for subsequent demand matching.
[0060] In step S2, the data discovery method specifically involves the following: Within the trusted data computing framework, the data requester initiates a query request to the registry contract (RC) of the consortium blockchain by constructing query parameters. The registry contract (RC) processes the request and returns the query results to the data requester. This method is based on the query function of the blockchain, such as... Figure 5As shown. A data query parameter includes a query keyword description, the public key information of the data requester, and the identifier information of the data requester, represented as Request=<Request_Keyinfo,Pk,Uid)> The triple; the data query result contains a complete JSON format query result information, timestamp, public key information of the data requester, and identifier information of the data requester, represented as Reply=<Reply_info,Pk,Uid)> The triplet.
[0061] In step S2, the specific method for demand matching is as follows: user registration information is divided into two parts: user private information and user public information; as the unique identity proof of the participants in data calculation, it is described using the following definitions 11 to 12:
[0062] Definition 11. User private information δS k This is the private identity information returned after the data computing participants register, specifically represented by the following seven-tuple:
[0063]
[0064] Among them, U i This is the user's identity identifier. `tStamp` represents the user's registration time, `M` represents a random string entered by the user, used to encrypt and generate the identity key, and `PK`... k and SK k This indicates that the CA generates a public / private key pair for the user. k This indicates the user's identity status information. This represents the user's contract account address;
[0065] Definition 12. User Public Information δP k This refers to the publicly available identity information of data computation participants, which is stored on the blockchain after registration. Specifically, it is represented by the following four-tuple:
[0066]
[0067] Wherein, δP k It is user private information δS k The anonymized version hides the relevant confidential identity information;
[0068] Data Demand Side (DR) k The complete requirement matching process is as follows: Data Requester (DR) A Obtain user identity information as defined in Definition 11 through the registration process provided by the proposed framework. Consortium blockchain networks define user public information as 12. Synchronize to the on-chain data evaluation management contract DEMC; after successful registration, the data requester DRk After accessing the consortium blockchain network through a proxy client, data discovery is performed using the RC registry query module to find the required data resources. Definition 4 on-chain information DCAO i In addition, data requesters also need to select suitable computing nodes and deploy data analytics code to process data resources. The evaluation criteria are based on the evaluation information recorded in the Data Evaluation Management Contract (DEMC), thereby selecting suitable data providers and computing nodes. The requesting party can further query the evidence records of specific data computation processes on the consortium blockchain to verify the authenticity of the information. Once both parties agree on the data computation service and payment, the data requesting party (DR)... k The prepayment is sent to the RISC results information storage contract and deployed as an executable binary on the selected compute nodes with anti-counterfeiting features.
[0069] In step S3, let the data requester (DR) be... k Data Provider (DP) j Data resources The selected compute node is CN. l The data calculation process is then described by the following definitions 5 to 10;
[0070] Definition 5. The calling procedure for data processing code, Train_Process, is the deployment parameter of the data requester for its selected compute node, represented as the following six-tuple:
[0071]
[0072] Where tStamp represents the deployment time of the data processing code, Desc represents the calling process description of the data processing code, checksum represents the hash checksum of the data processing code, and prePay represents the prepayment for the data calculation process;
[0073] Definition 6. The data processing code Process_Code can perform data analysis and processing on computational data to meet the data computation needs of the data requester, specifically represented by the following four-tuple:
[0074] Process_Code=(PType, PMod, PDoc, PFunc) Formula 6;
[0075] Among them, PType is the data processing code type, PMod is the structured module of the data processing code, PDoc is the technical document of the data processing code, and PFunc is the functional division of the data processing code.
[0076] Definition 7. Data Calculation Request Information: Train_Request is a data calculation request initiated by the data requester after requirement matching, represented as the following four-tuple:
[0077]
[0078] Where tStamp is the initiation time of the data computation request, and the computation node selected for demand matching is represented as... Any nodeInfo is represented as nodeInfo = (UUID) i ,nodeParameters), where nodeName represents the identifier name of the compute node, nodeParameters represents the computing power parameters of the compute node; Train_Info represents the parameter information of the selected data resources for demand matching;
[0079] Definition 8. Data computation parameter information Train_Info is the specific computation parameter information sent to the selected computing node. It can identify the identities of both parties in the data computation process and locate the required data resources, represented as the following six-tuple:
[0080]
[0081] The location information of the data resources is represented as follows: Any dataInfo is represented as dataInfo = (UUID) i The data resource name is defined as `dataName`, `dataParameters` as the location parameters for the method address of the data resource, and `dataProcess` as the calling process parameters of the data resource. `Desc` is the plaintext description of the data resource, and `dl` represents the valid deadline for the data computation request. Definition 9. Data Acquisition Request Information: `Data_Request` is a request from the computing node to the data provider to obtain the data resources required for data computation, specifically represented by the following seven-tuple:
[0082]
[0083] Where tStamp represents the time the data retrieval request was initiated, and Desc represents the basic descriptive information of the data retrieval request. This indicates a request for encrypted data computation bearing the signature of the data requester.
[0084] Definition 10. Data Evaluation Information (Train_Evaluation) is the evaluation and execution information of the data requester on the data provider and the computation nodes participating in the process after each data computation, represented as the following 5-tuple:
[0085]
[0086] Where tStamp is the time when the data evaluation was initiated, ω DS This represents the satisfaction rating of the data requester towards the data provider, ω. node This indicates the satisfaction rating of data requesters with the computing nodes. It is the hash value of the encrypted result data with the signature of the computing node;
[0087] In step S3, Figure 6 It demonstrates a detailed workflow for data computation, in which...<PK,SK> The public and private key pairs of the participants in the data computation are represented by CN, and the set of individual computing nodes is represented by CN. B =CN B1 ...CN Bi , ..., CN Bn For i ∈ n, the following computation nodes are represented by CN. B Generally speaking, Sig() represents a short signature function; the specific steps are as follows:
[0088] Step A1, Data Requester (DR) A After matching the requirements, use its private key SK A The data request Train_Request for this data calculation is encrypted to obtain the encrypted value. Send to the Data Request Management Contract (DRMC);
[0089] Step A2: The consortium blockchain network receives the encrypted data computation request. Then, verify whether it is registered and use its public key to PK. A Parse the data calculation parameter information Train_Info (Definition 8) and identify the selected computing node CN for this data calculation. B And on this basis, further construct the parameters of the data calculation request. The binary tuple includes the following: the first part is the parsed data calculation parameter information Train_Info, and the second part is the data requester DB obtained in step A1. A Encryption request with signature Then it is sent to all selected compute nodes CN in this data calculation process. B ;
[0090] Step A3, compute node CN BBased on Train_Info, the data provider DP who participated in this data calculation was identified. C The address message is used to construct and encrypt data to obtain the request Data_Request in definition 9. Send to the corresponding data provider DP C ;
[0091] Step A4, Data Provider DP C Data requesters (DRs) from consortium blockchain networks A and computing node CN B public key PK A PK B And decrypt and verify Verify and confirm the data requester (DR) A and computing node CN B After identifying the user's identity, the system obtains the data retrieval request (Data_Request) and locates the specific data resource (Train_Data) based on the dataInfo. Then, it further constructs a data response based on this information. The binary tuple includes the following: the first part is the encrypted data resource. The second part is the encrypted data resource hash value.
[0092] Step A5, compute node CN B Using registered data provider DP C public key PK C and its own private key SK B Verify and decrypt The data resource Train_Data and its anti-counterfeiting hash value Train_dataHash are obtained through parsing.
[0093] Step A6, compute node CN B After verifying the data resource Train_Data, the pre-deployed data processing code is used to analyze and process the data to obtain the result data Train_Result, and then the data requester DR is used. A public key PK A Encryption to obtain PK A (Train_Result) and its anti-counterfeiting hash digest
[0094] Step A7, compute node CN B Using SK B The encrypted data result PK obtained in step 6 A Sign (Train_Result) to obtain And return it to the data requester's DB via remote procedure call.A ;
[0095] Step A8: To ensure the transparency and traceability of the data computation process, the computation node CN... B Returning the data processing results to the data requester, DR. A Simultaneously, it is necessary to construct a trusted data computation credential Train_Proof; Train_Proof consists of four parts, namely, the computation node CN B The process of calling the signed data processing code Includes compute node CN B The hash value of the signed data processing result Includes Data Requester (DR) A Signature calculation request This data retrieval request includes the signatures of all participants in this data calculation. Next, compute node CN B The Train_Proof is published to the consortium blockchain network in the form of a transaction, and the hash value of the data calculation result is stored in DataHashList. i In the process, the RISC result information storage contract verifies the validity of the results and sends them to the correct compute node CN. B and data provider DP C Payments are made through the transaction account; if data participants have a dispute over the data calculation process, they can resolve the dispute through the on-chain trusted data calculation certificate;
[0096] Step A9, Data Requester (DR) A After completing this data calculation, a satisfaction evaluation is conducted, and the data evaluation information Train_Evaluation (Definition 10) is submitted to the data evaluation management contract DEMC for on-chain storage via a transaction.
[0097] In step S5, the hash value of the data processing result is used as the sole basis for calculating the satisfaction evaluation of the data, which avoids malicious data requesters from making multiple invalid and abnormal evaluations.
[0098] During evaluation, `matchCheck` is used to indicate whether the evaluation requirements are met, `check()` is used to check the execution status of all computation nodes in the current data computation process to indicate whether the task has been completed, and `evaluate()` evaluates the current data computation based on the satisfaction evaluation index ω, as shown in Definition 13 below:
[0099] Definition 13. The satisfaction evaluation index ω is a comprehensive evaluation score of the data providers and calculation nodes involved in the data calculation, represented by the following triplet:
[0100] ω=(Accuracy,Score,Efficiency) Formula 13;
[0101] Here, Accuracy is the score evaluating the accuracy of the data results, Score is the subjective evaluation score from the data requester regarding the data calculation, and Efficiency is the score evaluating the time taken for data delivery. For a single data calculation, to avoid unrealistic or malicious evaluations from the data requester, the satisfaction evaluation index is kept within a controllable range, i.e., ω. i =(Accuracy) i Score i Efficiency i )∈((-ξ,ξ),(-ξ,ξ),(-ξ,ξ)),ξ is the threshold to prevent malicious data requesters from making abnormal evaluations;
[0102] The satisfaction evaluation process is described using Algorithm 2, specifically as follows: The data evaluation management contract DEMC first uses the verifyId() function to check the data provider's identifier and public key to verify whether its identity meets the requirements for data evaluation. Then, it checks whether all computation nodes have completed their execution. If both conditions are met, the analysis() function will parse the hash value of the resulting data and store it in the DataHashList. i The data is matched within the DataHashList. If a match is found, the `evaluate()` function can be used to evaluate the satisfaction level. The data requester will objectively evaluate the data calculation process based on the satisfaction evaluation index ω in Definition 13. After the evaluation, the results will be retrieved from the DataHashList. i Delete the hash value record of this data entry to ensure the accuracy and uniqueness of the evaluation.
[0103] The satisfaction evaluation part of Algorithm 2 is expressed in pseudocode as follows:
[0104]
[0105] The trusted data computing framework includes a trusted data traceability mechanism based on consortium blockchains. It designs and builds a proposed traceability mechanism through the Hyperledger Fabric basic framework to manage the entire data computing process, aiming to support the characteristics of auditable data use, traceable data source, and accountable data computing.
[0106] The RISC contract, corresponding to the data traceability mechanism, is responsible for persistently storing records of the entire data computation process. It ensures transparency and traceability through dual safeguards of timestamps and key signatures. Algorithm 2 sets initial parameters and functions to meet the requirements; among them, `recordCheck` identifies whether the requirements for computation record notarization have been met, and `Sig`...j (dataProof i ) is data resources The data is used to calculate voucher information. The `analysis()` function is used to parse and extract the corresponding calculated data hash value from the voucher information, and the `isExist()` function is used to verify the uniqueness of the data hash value.
[0107] The data traceability process, as described in Algorithm 2, involves the Result Information Storage Contract RISC verifying the identity of the notary public and reviewing the format of the computation result record to be notified using the verifyId() and matchStyle() functions. Subsequently, the Result Information Storage Contract RISC uses the analysis() and isExist() functions to parse the Sig... j (dataProof i The corresponding data hash value dataHash i It also checks for duplicate notarization. If it is the first notarization, it adds the DTCF signature to the notarization information and stores it in the DataHashList. i Finally, the update() function writes this data calculation record and returns the certificate information indicating successful data calculation record storage.
[0108] The data tracing part of Algorithm 2 is represented in pseudocode as follows:
[0109]
[0110] The security strategy of the trusted computing framework is as follows:
[0111] 1) Identity Authentication: Before a User, a participant in data computation, accesses the consortium blockchain network, identity authentication is required. The CA (Certificate Authority) generates a unique public-private key pair {PK} for the User based on the randomly entered string M. user SK user} and digital certificate Cert user To access the consortium blockchain network, all registered participants' public user information δP user User private information δS will be stored in the consortium blockchain network. user The resources are then kept by the participating parties themselves; only authorized participating parties can obtain access to resources in the Hyperledger Fabric network, thereby maintaining the security and privacy characteristics of the consortium blockchain.
[0112] 2) Data Autonomy: The data provider (DP) retrieves encrypted data based on the data acquisition request sent by the computing node (CN). The specific data calculation parameters determine whether to respond to a data acquisition request, greatly avoiding unfair and unreasonable data calculation tasks. Unlike traditional data exchange / hosting-as-a-service models, by building a trusted data computing mechanism, the computing node (CN) needs to correctly forward the encrypted request from the data requester (DR). Only by obtaining data resources can other data requesters (DRs) and compute nodes (CNs) be prevented from bypassing the data provider to perform data computations. Simultaneously, data analysis and processing operations are executed on the compute node (CN), and only the hash value of the final result data is uploaded. When the evidence is stored on the blockchain, the data requester (DR) cannot access the original data.
[0113] 3) Data security and integrity: To ensure the security of the data delivery process, the original data resources will be protected by the public key PK of the computing node CN. CN The private key SK of the data provider DP DP Encryption is performed to ensure that other participants cannot obtain the plaintext of the data; at the same time, the flow of data resources always relies on digital signature and hash digest technology to ensure that every step of the data computation is clearly recorded and proven, and is finally reflected through the data computation trusted certificate Train_Proof, thereby realizing the monitoring and management of the entire data computation process, and thus ensuring the integrity of the computation result data.
[0114] 4) Security for dishonest data providers: Dishonest data providers (DPs) may provide irrelevant data or even fail to deliver data at all. DPs need to sign and verify the original data they provide and its hash digest. and And construct data response In response to data acquisition requests, the Data Requester (DR) verifies the authenticity of the acquired data resources based on this hash digest to prevent the data provider from denying or repudiating in the event of a dispute. In addition, when matching requests, the DR obtains relevant evaluation information ω of the data provider (DP) from the Data Evaluation Management Contract (DEMC) for further screening.
[0115] 5) Security for Dishonest Computing Nodes: Dishonest computing nodes (CNs) may forge Data Requests (Data_Request) to obtain data resources from data providers (DPs). DPs can verify the authenticity of these requests through the Data Request Management Contract (DRMC). Only requests matching the signature of the data requester (DR) will be delivered to the computing node (CN), proving that the request was indeed initiated by the corresponding data requester (DR). This prevents CNs from maliciously obtaining data resources. To prevent CNs from processing data improperly to save computing power while still receiving expected rewards, CNs must upload a Trusted Data Computation Certificate (Train_Proof) to the blockchain after delivering the data processing results to receive the expected reward. This certificate contains the signature value of the hash digest of the resulting data from the CN. If a compute node (CN) commits malicious acts, it can be traced and punitive measures can be taken to effectively prevent compute nodes (CNs) from stealing from their own premises.
[0116] 6) Security for dishonest data requesters: Dishonest data requesters (DRs) may not only forge identities to send false data computation requests, but may also deny or repudiate the data computations. When processing requests, the consortium blockchain network will use the DR's public key (PK). DR Identity authentication effectively prevents identity forgery. Meanwhile, to protect the reputation and interests of the data provider (DP), the RISC result information storage contract in the consortium blockchain network records the trusted credential information Train_Proof, which includes the signatures of all participants in each data calculation, and provides data traceability functionality, thereby reducing the potential risk of dishonest data requesters (DRs) denying or repudiating during the data calculation process.
[0117] This invention proposes a data trusted computing scheme based on consortium blockchain. First, it defines a data architecture paradigm based on consortium blockchain, which divides data into three parts: identifier, descriptive data, and data entity to standardize the data registration process, providing a unified access method for heterogeneous, heterogeneous, and heterogeneous data resources, and efficiently integrating data resources.
[0118] Secondly, this invention designs a trusted data computing mechanism that transforms data hosting / exchange as a service into data processing as a service, thereby achieving "usable but invisible" data to address trust challenges in the data computing process.
[0119] Finally, this invention designs a data trust traceability mechanism, which ensures the security and integrity of data flow between data requesters, computing nodes, and data providers by leaving traces on the data computation process chain.
[0120] This invention transforms the traditional data hosting / exchange-as-a-service model into a data processing-as-a-service model. Instead of hosting raw data with a third party for the data requester, the data provider offers the results of data resource analysis and processing without the data requester having access to the raw data, thus achieving "usable but invisible" data. This avoids the leakage of the data provider's raw data while not affecting the data requester's access to the data analysis results. It rebuilds the trust relationship between the two parties in the data computing process and alleviates many problems inherent in traditional data computing platforms, such as data silos, data loss of control, and data ownership issues.
[0121] This invention proposes a trusted data computing framework based on consortium blockchain. First, it defines a data architecture paradigm based on consortium blockchain, standardizing the data registration process by dividing data into three parts: identifiers, feature data, and data entities. This provides a unified access method for heterogeneous, multi-source, and multi-domain data resources, efficiently integrating them. Simultaneously, it designs a trusted data computing mechanism by transforming data hosting / exchange as a service into data processing as a service. Instead of hosting the original data with a third party for data requesters to analyze, the data provider provides the results of data resource analysis and processing without the data requester accessing the original data, achieving "usable but invisible" data to address trust challenges in the data computing process. Furthermore, this invention designs a trusted data traceability mechanism, ensuring the security and integrity of data flow between data requesters, computing nodes, and data providers by leaving traces of the data computing process on the blockchain. Attached Figure Description
[0122] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments:
[0123] Appendix Figure 1 This is a schematic diagram of a data custody transaction model based on traditional technology;
[0124] Appendix Figure 2 This is a schematic diagram providing an overview of the data trust computing framework of this invention;
[0125] Appendix Figure 3 This is a schematic diagram of the data conversion model of the present invention (taking the medical device industry bidding data dataset as an example);
[0126] Appendix Figure 4 This is a schematic diagram of the module design of the smart contract in this invention;
[0127] Appendix Figure 5 This is a schematic diagram illustrating the principle of data discovery in this invention;
[0128] Appendix Figure 6 This is a schematic diagram of the data trust calculation process of the present invention;
[0129] Appendix Figure 7 This is a schematic diagram showing the results of performance testing of the method described in this invention.
[0130] Appendix Figure 8 This is another schematic diagram showing the results of performance testing of the method described in this invention. Detailed Implementation
[0131] As shown in the figure, in this example, the trusted data computing framework is based on the Go language and uses the Hyperledger Fabric v2.3 framework to build a consortium blockchain network. Three machines are used to simulate the data demander, data provider, and computing nodes respectively, to simulate a real data computing scenario. The network configuration includes 3 Order nodes and 3 Org organizations. Each Org is configured with 4 Peer nodes. The nodes use the Raft consensus mechanism, and all the smart contracts required by the proposed framework are deployed on all Peer nodes, including the Registry Contract (RC), Data Request Management Contract (DRMC), Result Information Storage Contract (RISC), and Data Evaluation Management Contract (DEMC). Detailed equipment configuration and required testing tools can be found in the Equipment Detailed Configuration Table - Testing Tools Table.
[0132]
[0133] Detailed Equipment Configuration
[0134]
[0135] Testing tools
[0136] The proposed framework's smart contracts underwent read / write performance stress tests by simulating 20-500 virtual client agents. The Hyperledger Foundation's Caliper performance benchmark tool and the Fabric-Nodejs-SDK interaction tool were used. The average execution time of the proposed smart contracts was measured by generating a fixed 10 transaction requests per second. The performance stress test consisted of five rounds, with the number of concurrent virtual clients gradually increasing from 20 to 500 in each round. Figure 7-8 As shown, as the system throughput gradually increases, the time consumed by read and write operations will also gradually increase, eventually stabilizing within an acceptable fluctuation range.
[0137] The trusted computing framework in this example allows different data providers to register their data and different data computing demanders to discover their data and find the corresponding data providers and computing nodes. The main achievement of this invention is to design a data architecture paradigm based on consortium blockchain, which divides data into three parts: identifier, descriptive data, and data entity, and realizes the standardized registration of metadata information on the blockchain.
[0138] The data computation process includes three key steps: demand matching, data computation, and satisfaction evaluation. First, in the demand matching phase, the data computation requester can discover data and find the corresponding data provider and computation node through a registry supported by the proposed data architecture. Second, through the proposed data computation process, the data computation requester can receive and parse the output results of the reliable data computation processing. Finally, the data computation requester can provide feedback on the data computation process through the satisfaction evaluation mechanism.
[0139] The data trust traceability mechanism can persistently store the entire data computation process record through smart contracts, and at the same time, it can ensure that the data computation process is transparent and traceable through the dual protection of timestamps and key signatures.
Claims
1. A data-trusted computing framework based on consortium blockchain, characterized in that: The entities in the Data Trusted Computing Framework (DTCF) include data requesters, data providers, consensus nodes, and certificate authorities; the framework's structure is divided into a data registration layer, a data trust layer, a data consensus layer, a data processing layer, and a data delivery layer. The data registration layer adopts a data architecture paradigm based on consortium blockchain, which divides data into identifiers, feature data, and data entities to standardize the data registration process, providing a unified access method for heterogeneous, heterogeneous, and heterogeneous data resources, and integrating data resources. The data trust layer uses a data trust computing mechanism to enable data providers to provide data resource analysis and processing results to data requesters without them having access to the original data. The data consensus layer includes a data trust traceability mechanism, which ensures the security and integrity of data flow between data requesters, computing nodes, and data providers by leaving traces on the data computation process chain. In step S1, the Data Architecture Paradigm CBDA is used to provide a unified way to persistently identify, describe, locate, and manage heterogeneous data resources from different sources and regions. In the CBDA transformation rules, the registered data provider DP will be... j Data resources It is divided into identifiers, feature data, and data entities; among which UI i It is an identifier for data resources, used to persistently and uniquely identify each data resource, and is described using the following definition 1: Definition 1.UI i It is a data resource In CBDA, the identifier portion is represented by the following quadruple: UID i It is a data resource The unique identifier, tStamp is the generated timestamp, PK j It is the data provider DP j SK's public key j It is the data provider DP j The private key, M is a cryptographic string defined by the data provider, and Sig() is a short signature function. It is the data provider DP j Use its private key SK j The signature information for string M; characteristic data CD i is metadata information of data resource, used for retrieving and discovering corresponding data resource, and its specific meaning is shown in definition 2 below: Definition 2.CD i It is a data resource In CBDA, the feature data portion, i.e., metadata, is represented as the following quadruple: Meta i It is a data resource metadata, It is the data provider DP j Use its private key SK j Data resources UID i Signature information, apiInfo i It is a data resource The API call information; Data Entity DE i is the mapping of the data resource after being processed by the registry storage module of the on-chain registry contract RC, rather than the original data, as shown in the following definition 3: Definition 3.DE i It is a data resource In CBDA, the data entity portion is represented as the following quintuple: Among them AD i Represents data resources The access address, dl represents the valid expiration time of the data resource, AS is the status indicator indicating whether the data resource is available, K DTSP This represents the proposed symmetric key. Through the symmetric key K DTSP Data resources The signature information of the access address; Data Architecture Object CBDAO i is the final formatted storage form of the data resources within the proposed framework, the specific meaning of which is given in definition 4 below: Definition 4.CBDAO i It is a data resource After CBDA rule transformation, the on-chain representation is shown as the following triple: CBDAO i = (UI i , CD i , DE i ) Equation 4; UI i CD i DE i These are data resources The identifier, feature data, and data entity parts after CBDA rule transformation; the CBDA data transformation process is as follows: The proposed framework allocates a UI definition 1 for each data resource. i This serves as the unique identifier for its corresponding CBDA object; simultaneously, it constructs the CD defined in Definition 2. i and with preset mapping functions and UI i One-to-one correspondence; finally, encapsulate the data entity DE defined in definition 3 in a standardized format. i It also provides an interface for receiving data requests to the outside world; When a data provider is ready to publish its data resources, it shall register the data resources in a standardized manner according to the above data transformation process.
2. The data trust computing framework based on consortium blockchain according to claim 1, characterized in that: The data computation process of the trusted data computing framework includes the following steps; Step S1: The data provider transforms the data resources according to the data architecture paradigm and registers them on the blockchain; Step S2: The data requester discovers data through the registry query module, finds interested data providers and suitable computing nodes to match the demand, and then sends a data request carrying purchase pricing information to the consortium blockchain network. Step S3: The consortium blockchain network sends the request to the corresponding computing node. The computing node encapsulates the data request to obtain data resources from the data provider. Then, it performs data analysis and processing locally on the node. Only the execution result is sent to the data requester, not the original data. Step S4: The computing nodes that meet the computing requirements will store the data computing result storage record containing the signatures of the data computing participants in the consortium blockchain network of the trusted data computing framework. Step S5: Automatic payment via smart contract module; evaluate the satisfaction level of this data calculation process.
3. The data trust computing framework based on consortium blockchain according to claim 2, characterized in that: The data registration layer is responsible for converting data architecture rules and standardizing the registration process for heterogeneous data resources from different sources and regions. The data trust layer is used to ensure the security and trustworthiness of data transmission and to record the hash digest of the calculation results to prevent tampering; the data consensus layer is used to ensure the consensus of the distributed ledger data on the chain; and the data processing layer is responsible for the process control of the data calculation process. The data delivery layer is used to support the reliable delivery and evaluation of the final data processing results.
4. The data trust computing framework based on consortium blockchain according to claim 3, characterized in that: Let DP i For data providers, DR i For data demanders, CN i For computing nodes, To compute the state information of the nodes, For data resources, UID i A unique identifier for data resources, UI i For the identification of data resources, CD i For the characteristic data of data resources, DE i For data entities of data resources, CBDAO i For data architecture objects, UUID i The following are the key components of the data computation process: RC (Registry Contract), DRMC (Data Request Management Contract), RISC (Result Information Storage Contract), DEMC (Data Evaluation Management Contract), and SK. i For private key, PK i For public key, Cert i For digital certificates, K i For symmetric keys, Sig() is the short signature function, and dataProof... i For data calculation and recording, ω is the satisfaction evaluation index; Data Provider: A data provider is an individual user or organization that owns data resources. They access this framework through the Fabric SDK, publish their data resources, and sell the processed data resources to suitable data providers through demand matching. Data demanders are potential customers of data resources. They access this framework through the Fabric SDK and discover data from all network nodes by querying the registry to match suitable data providers to meet their data computing needs. Compute Node (CN): A device or node used to meet the computing needs of different data technology scenarios, assuming that it is secure, reliable, and capable of executing relevant computing tasks correctly and independently. Consensus Nodes: Consensus nodes consist of all nodes that have joined the trusted computing framework for data, and together they maintain consensus on the distributed ledger data on the chain, ensuring the security and consistency of the consortium blockchain network. The trusted data computing framework uses a Certificate Authority (CA): an authoritative and trusted certification authority responsible for authorizing and issuing exclusive public-private key pairs to data computing participants accessing the framework. Data computing participants will use these keys to verify their identity and encrypt data during communication within the consortium blockchain network.
5. The data trust computing framework based on consortium blockchain according to claim 4, characterized in that: The DTCF data calculation process includes the following security features: Identity authentication: Identity authentication means that the key identity information of the data requester should not be forged by malicious participants, thereby bypassing the authentication and authorization of the consortium blockchain network to carry out illegal data calculations. It is necessary to ensure the privacy and security of the consortium blockchain network. Data autonomy: Whether data is shared should be entirely controlled by the data provider, meaning the data provider can decide whether to share its data resources; Data security and integrity: The data provided by the data provider is encrypted during sharing, and no one other than the computing node can access the requested raw data resource; This ensures data integrity during the data computation process, meaning the data should not be tampered with during computation; it also addresses security for dishonest data providers: dishonest data providers are those who provide invalid data that is irrelevant to the intended data request, or those who send additional data to gain more revenue. A trusted computing framework can identify and resist dishonest data providers; Security concerns regarding dishonest computing nodes: Dishonest computing nodes are data nodes that forge data requests to obtain data resources from data providers or fail to process data in accordance with predetermined requirements; For the security of dishonest data requesters: For dishonest data requesters who send false data requests to data providers by forging the identity information of other participants, or who deny and refuse to pay for completed data computation requests, the trusted data computing framework can preserve key data computation records or take reliable security measures to protect the interests and reputation of data providers. The smart contract module of the Data Trusted Computing Framework (DTCF) includes functions such as data registration, demand matching, data computation, and satisfaction evaluation, which are presented in the form of a Registry Contract (RC), a Data Request Management Contract (DRMC), a Result Information Storage Contract (RISC), and a Data Evaluation Management Contract (DEMC), respectively. RC: RC stores and manages all registered on-chain data resources through the registry storage module, providing a unified access method for data resources from different sources; at the same time, RC's registry query module provides a query interface for data requesters to match suitable data resources. DRMC: The request information storage module of DRMC is responsible for storing and managing all data requests during the data computation process, including data computation requests initiated by data requesters, data computation requests processed by the proposed framework, and data acquisition requests initiated by computing nodes. DRMC's prepaid storage module requires data requesters to prepay data providers and computing nodes after demand matching; RISC: RISC stores and manages key evidence information during the data computation process through a data computation evidence storage module, including the data provider's data response information and the trusted credentials of the final data computation results on the blockchain by the computing nodes; the result matching module statistically analyzes the data analysis and processing results of the computing nodes, and the result with the most identical results is the final result; The payment module is responsible for issuing payments to the transaction addresses of data providers and computing nodes based on the final consistent result data; DEMC: DEMC stores the identity information of all data calculation participants through a user information storage module; The evaluation storage module is used to store the satisfaction evaluations of the data requester on the data provider and computing nodes after each data calculation process is completed; The evaluation query module helps data requesters better select suitable data providers and computing nodes for subsequent demand matching.
6. The data trust computing framework based on consortium blockchain according to claim 1, characterized in that: In step S2, the data discovery method is as follows: In the trusted computing framework, the data requester initiates a query request to the registry contract (RC) of the consortium blockchain by constructing query parameters. The registry contract (RC) processes the request and returns the query results to the data requester. This method is based on the query function of the blockchain. A data query parameter includes a query keyword description, the public key information of the data requester, and the identification information of the data requester, represented as Request =<Request_Keyinfo,Pk,Uid)> The triple; the data query result contains a complete JSON format query result information, timestamp, public key information of the data requester, and identifier information of the data requester, represented as Reply=<Reply_info,Pk,Uid)> The triplet; In step S2, the specific method for demand matching is as follows: user registration information is divided into two parts: user private information and user public information; As the sole identity proof for participants in data computation, it is described using the following definitions 11 to 12: Definition 11. User private information δS k is the private identity information returned after the data computing participant registers, which is specifically represented as a seven-tuple: Among them, U i This is the user's identity identifier. `tStamp` represents the user's registration time, `M` represents a random string entered by the user, used to encrypt and generate the identity key, and `PK`... k and SK k This indicates that the CA generates a public / private key pair for the user. k This indicates the user's identity status information. This represents the user's contract account address; Definition 12. User public information δP k is the public identity information of the data computing participant registered and stored on the chain, which is specifically represented as a four-tuple: where δP k is a desensitized version of the user's private information δS k , concealing the relevant confidential identity information; Data Demand Side (DR) k The complete requirement matching process is as follows: Data Requester (DR) A Obtain user identity information as defined in Definition 11 through the registration process provided by the proposed framework. Consortium blockchain networks define user public information as 12. Synchronize to the on-chain data evaluation management contract DEMC; after successful registration, the data requester DR k After accessing the consortium blockchain network through a proxy client, data discovery is performed using the RC registry query module to find the required data resources. Definition 4 on-chain information DCAO i In addition, data requesters also need to select suitable computing nodes and deploy data analytics code to process data resources. The evaluation criteria are based on the evaluation information recorded in the Data Evaluation Management Contract (DEMC) to select suitable data providers and computing nodes. The requesting party further queries the evidence records of specific data computation processes on the consortium blockchain to verify the authenticity of the information. Once both parties agree on the data computation service and payment, the data requesting party (DR)... k The prepayment is sent to the RISC results information storage contract and deployed as an executable binary on the selected compute nodes with anti-counterfeiting features.
7. A data trust computing framework based on consortium blockchain according to claim 6, characterized in that: In step S3, let the data requester (DR) be... k Data Provider (DP) j Data resources The selected compute node is CN. l The data calculation process is then described by the following definitions 5 to 10; Definition 5. The calling procedure for data processing code, Train_Process, is the deployment parameter of the data requester for its selected compute node, represented as the following six-tuple: Where tStamp represents the deployment time of the data processing code, Desc represents the calling process description of the data processing code, checksum represents the hash checksum of the data processing code, and prePay represents the prepayment for the data calculation process; Definition 6. The data processing code Process_Code can perform data analysis and processing on computational data to meet the data computation needs of the data requester, specifically represented by the following four-tuple: Process_Code=(PType,PMod,PDoc,PFunc) Formula 6; Among them, PType is the data processing code type, PMod is the structured module of the data processing code, PDoc is the technical document of the data processing code, and PFunc is the functional division of the data processing code. Definition 7. Data Calculation Request Information: Train_Request is a data calculation request initiated by the data requester after requirement matching, represented as the following four-tuple: Where tStamp is the initiation time of the data computation request, and the computation node selected for demand matching is represented as... Any nodeInfo is represented as nodeInfo = (UUID) i , nodeParameters), where nodeName represents the identifier name of the compute node, nodeParameters represents the computing power parameters of the compute node; Train_Info represents the parameter information of the selected data resources for demand matching; Definition 8. Data computation parameter information Train_Info is the specific computation parameter information sent to the selected computing node, identifying the identities of both parties in the data computation process and locating the required data resources, represented as the following six-tuple: The location information of the data resources is represented as follows: Any dataInfo is represented as dataInfo = (UUID i , dataName, dataParameters, dataProcess), where dataName represents the data resource name, dataParameters represents the positioning parameters of the method address of the data resource, and dataProcess represents the calling process parameters of the data resource; Desc is the plaintext description information of the data resource, and dl represents the effective deadline of the data calculation request; Definition 9. Data acquisition request information Data_Request is the request of the computing node to the data provider to acquire the data resource required for data calculation, which is specifically represented as the following seven-tuple: Where tStamp represents the time the data retrieval request was initiated, and Desc represents the basic descriptive information of the data retrieval request. This indicates a request for encrypted data computation bearing the signature of the data requester. Definition 10. Data Evaluation Information (Train_Evaluation) is the evaluation and execution information of the data requester on the data provider and the computation nodes participating in the process after each data computation, represented as the following 5-tuple: Where tStamp is the time when the data evaluation was initiated, ω DS This represents the satisfaction rating of the data requester towards the data provider, ω. node This indicates the satisfaction rating of data requesters with the computing nodes. It is the hash value of the encrypted result data with the signature of the computing node; In step S3, the detailed workflow of data computation is as follows: using <PK, SK> to represent the public-private key pair of the data computation participant, the set of individual computing nodes is represented as CN B = CN B1 ,... CN Bi ,..., CN Bn , i∈n, the following computing nodes are collectively referred to as CN B , and Sig() represents a short signature function; the specific steps are as follows: Step A1, Data Requester (DR) A After matching the requirements, use its private key SK A The data request Train_Request for this data calculation is encrypted to obtain the encrypted value. Send to the Data Request Management Contract (DRMC); Step A2: The consortium blockchain network receives the encrypted data computation request. Then, verify whether it is registered and use its public key to PK. A Parse the data calculation parameter information Train_Info and identify the selected computing node CN for this data calculation. B And on this basis, further construct the parameters of the data calculation request. The binary tuple includes the following: the first part is the parsed data calculation parameter information Train_Info, and the second part is the data requester DB obtained in step A1. A Encryption request with signature Then it is sent to all selected compute nodes CN in this data calculation process. B ; Step A3, compute node CN B Based on Train_Info, the data provider DP who participated in this data calculation was identified. C The address message is used to construct and encrypt data to obtain the request Data_Request in definition 9. Send to the corresponding data provider DP C ; Step A4, Data Provider DP C Data requesters (DRs) from consortium blockchain networks A and computing node CN B public key PK A PK B And decrypt and verify Verify and confirm the data requester (DR) A and computing node CN B After identifying the user's identity, the system obtains the data retrieval request (Data_Request) and locates the specific data resource (Train_Data) based on the dataInfo. Then, it further constructs a data response based on this information. The binary tuple includes the following: the first part is the encrypted data resource. The second part is the encrypted data resource hash value. Step A5, compute node CN B Using registered data provider DP C public key PK C and its own private key SK B Verify and decrypt The data resource Train_Data and its anti-counterfeiting hash value Train_dataHash are obtained through parsing. Step A6, compute node CN B After verifying the data resource Train_Data, the pre-deployed data processing code is used to analyze and process the data to obtain the result data Train_Result, and then the data requester DR is used. A public key PK A Encryption to obtain PK A (Train_Result) and its anti-counterfeiting hash digest Step A7, compute node CN B Using SK B The encrypted data result PK obtained in step 6 A Sign (Train_Result) to obtain And return it to the data requester's DB via remote procedure call. A ; Step A8: To ensure the transparency and traceability of the data computation process, the computation node CN... B Returning the data processing results to the data requester, DR. A Simultaneously, it is necessary to construct a trusted data computation credential Train_Proof; Train_Proof consists of four parts, namely, the computation node CN B The process of calling the signed data processing code Includes compute node CN B The hash value of the signed data processing result Includes Data Requester (DR) A Signature calculation request This data retrieval request includes the signatures of all participants in this data calculation. Next, compute node CN B The Train_Proof is published to the consortium blockchain network in the form of a transaction, and the hash value of the data calculation result is stored in DataHashList. i In the process, the RISC result information storage contract verifies the validity of the results and sends them to the correct compute node CN. B and data provider DP C Payments are made through the transaction account; if data participants have a dispute over the data calculation process, they can resolve the dispute through the on-chain trusted data calculation certificate; Step A9, Data Requester (DR) A After completing this data calculation, a satisfaction evaluation is conducted, and the data evaluation information Train_Evaluation defined in 10 is submitted to the data evaluation management contract DEMC for on-chain storage in the form of a transaction.
8. A data trust computing framework based on consortium blockchain according to claim 7, characterized in that: In step S5, the hash value of the data processing result is used as the sole basis for calculating the satisfaction evaluation of the data, which avoids malicious data requesters from making multiple invalid and abnormal evaluations. During evaluation, `matchCheck` is used to indicate whether the evaluation requirements are met, `check()` is used to check the execution status of all computation nodes in the current data computation process to indicate whether the task has been completed, and `evaluate()` evaluates the current data computation based on the satisfaction evaluation index ω, as shown in Definition 13 below: Definition 13. The satisfaction evaluation index ω is a comprehensive evaluation score of the data providers and calculation nodes involved in the data calculation, represented by the following triplet: ω=(Accuracy,Score,Efficiency) Formula 13; Among them, Accuracy is the evaluation score for the accuracy of the data results, Score is the subjective evaluation score of the data requester for this data calculation, and Efficiency is the evaluation score for the time taken to deliver the data. For a single data calculation, to avoid unrealistic and malicious evaluations from the data requester, the satisfaction evaluation index is kept within a controllable range, i.e., ω. i =(Accuracy) i Score i Efficiency i )∈((-ξ,ξ),(-ξ,ξ),(-ξ,ξ)),ξ is the threshold to prevent malicious data requesters from making abnormal evaluations; The satisfaction evaluation process is described using Algorithm 2, specifically as follows: The data evaluation management contract DEMC first uses the verifyId() function to check the data provider's identifier and public key to verify whether its identity meets the requirements for data evaluation. Then, it checks whether all computing nodes have completed their execution. If the checks and verifications pass, the analysis() function parses the hash value of the resulting data and stores it in the DataHashList. i The data is matched within the DataHashList. If a match is found, the `evaluate()` function can be used to evaluate the satisfaction level. The data requester will objectively evaluate the data calculation process based on the satisfaction evaluation index ω in Definition 13. After the evaluation, the results will be retrieved from the DataHashList. i Delete the hash value record of this data entry to ensure the accuracy and uniqueness of the evaluation.
9. A data trust computing framework based on consortium blockchain according to claim 7, characterized in that: The trusted data computing framework includes a trusted data traceability mechanism based on consortium blockchains. It designs and builds a proposed traceability mechanism through the Hyperledger Fabric basic framework to manage the entire data computing process, aiming to support the characteristics of auditable data use, traceable data source, and accountable data computing. The RISC contract, corresponding to the data traceability mechanism, is responsible for persistently storing records of the entire data computation process. It ensures transparency and traceability through dual safeguards of timestamps and key signatures. Algorithm 2 sets initial parameters and functions to meet the requirements; among them, `recordCheck` identifies whether the requirements for computation record notarization have been met, and `Sig`... j (dataProof i ) is data resources The data is used to calculate voucher information. The `analysis()` function is used to parse and extract the corresponding calculated data hash value from the voucher information, and the `isExist()` function is used to verify the uniqueness of the data hash value. The data traceability process, as described in Algorithm 2, involves the Result Information Storage Contract RISC verifying the identity of the notary public and reviewing the format of the computation result record to be notified using the verifyId() and matchStyle() functions. Subsequently, the Result Information Storage Contract RISC uses the analysis() and isExist() functions to parse the Sig... j (dataProof i The corresponding data hash value dataHash i It also checks for duplicate notarization. If it is the first notarization, it adds the DTCF signature to the notarization information and stores it in the DataHashList. i In the middle; finally, the update() function is used to write this data calculation record and return the certificate information indicating that the data calculation record has been successfully stored; The security strategy of the trusted computing framework is as follows: 1) Identity Authentication: Before a User, a participant in data computation, accesses the consortium blockchain network, identity authentication is required. The CA (Certificate Authority) generates a unique public-private key pair {PK} for the User based on the randomly entered string M. user SK user } and digital certificate Cert user To access the consortium blockchain network, all registered participants' public user information δP user User private information δS will be stored in the consortium blockchain network. user The resources are then kept by the participating parties themselves; only authorized participating parties can obtain access to resources in the Hyperledger Fabric network, thereby maintaining the security and privacy characteristics of the consortium blockchain. 2) Data Autonomy: The data provider (DP) retrieves encrypted data based on the data acquisition request sent by the computing node (CN). The specific data calculation parameters determine whether to respond to a data acquisition request, greatly avoiding unfair and unreasonable data calculation tasks. Unlike traditional data exchange / hosting-as-a-service models, by building a trusted data computing mechanism, the computing node (CN) needs to correctly forward the encrypted request from the data requester (DR). Only by obtaining data resources can other data requesters (DRs) and compute nodes (CNs) be prevented from bypassing the data provider to perform data computations. Simultaneously, data analysis and processing operations are executed on the compute node (CN), and only the hash value of the final result data is uploaded. When the evidence is stored on the blockchain, the data requester (DR) cannot access the original data. 3) Data security and integrity: To ensure the security of the data delivery process, the original data resources will be protected by the public key PK of the computing node CN. CN The private key SK of the data provider DP DP Encryption is performed to ensure that other participants cannot obtain the plaintext of the data; at the same time, the flow of data resources always relies on digital signature and hash digest technology to ensure that every step of the data computation is clearly recorded and proven, and is finally reflected through the data computation trusted certificate Train_Proof, thereby realizing the monitoring and management of the entire data computation process, and thus ensuring the integrity of the computation result data. 4) Security for dishonest data providers: Dishonest data providers (DPs) may provide irrelevant data or even fail to deliver data at all; DPs need to sign and verify the original data they provide and its hash digest. and And construct data response In response to data acquisition requests, the Data Requester (DR) verifies the authenticity of the acquired data resources based on this hash digest to prevent the data provider from denying or repudiating in the event of a dispute. In addition, when matching requests, the DR obtains relevant evaluation information ω of the data provider (DP) from the Data Evaluation Management Contract (DEMC) for further screening. 5) Security for dishonest compute nodes: Dishonest compute nodes (CNs) can forge Data Requests (Data_Request) to obtain data resources from data providers (DPs). Data providers (DPs) can verify the authenticity of these requests through the Data Request Management Contract (DRMC). Only requests matching the signature of the data requester (DR) will be delivered to the compute node (CN), proving that the request was indeed initiated by the corresponding data requester (DR). This prevents compute nodes (CNs) from maliciously obtaining data resources. To prevent compute nodes (CNs) from processing data improperly to save computing power and obtain expected rewards, compute nodes (CNs) must upload a Trusted Data Proof (Train_Proof) to the blockchain after delivering the data processing results to receive the expected rewards. This proof contains the signature value of the hash digest of the result data from the compute node (CN). If a compute node (CN) commits malicious acts, it can be traced and punitive measures can be taken to effectively prevent compute nodes (CNs) from stealing from their own premises. 6) Security against dishonest data requesters: Dishonest data requesters (DRs) can not only forge identities to send false data computation requests, but also deny and repudiate the data computations; when processing requests, the consortium blockchain network will use the DR's public key (PK). DR Identity authentication effectively prevents identity forgery. Meanwhile, to protect the reputation and interests of the data provider (DP), the RISC result information storage contract in the consortium blockchain network records the trusted credential information Train_Proof, which includes the signatures of all participants in each data calculation, and provides data traceability functionality, thereby reducing the potential risk of dishonest data requesters (DRs) denying or repudiating during the data calculation process.