Machine learning multi-dimensional data credibility prediction method based on block chain deterministic execution

By introducing a hierarchical random number generation mechanism and hardware error statistical verification on the blockchain, the problem of inconsistent calculation results caused by randomness and hardware heterogeneity in machine learning algorithms in blockchain networks is solved, and efficient and reliable multidimensional data prediction is achieved.

CN120850356APending Publication Date: 2025-10-28BOYA ZHENGLIAN (BEIJING) TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510916469.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-03
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

In existing technologies, machine learning algorithms, when executed in blockchain networks, result in inconsistent computational results due to randomness. Furthermore, floating-point errors caused by hardware heterogeneity exacerbate the inconsistency of computational results, leading to unreliable predictions. In addition, traditional solutions are inefficient and waste computational resources significantly.

Method used

By employing a hierarchical random number generation mechanism and hardware error statistics verification, it is ensured that blockchain nodes execute based on the same random number sequence. A verifiable random function is used to generate a private random number sequence, and a Merkle tree verification mechanism is used to record floating-point operation logs, ensuring reliable evidence storage of the prediction process.

Benefits of technology

It significantly improves the consistency of calculation results, reduces the consumption of computing resources, realizes reliable prediction of multidimensional data, and solves the core pain point of unreliable prediction results in traditional solutions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120850356A_ABST
    Figure CN120850356A_ABST
Patent Text Reader

Abstract

The invention discloses a machine learning multi-dimensional data credibility prediction method based on block chain deterministic execution, and relates to the technical field of cross of block chains and artificial intelligence. By constructing a hierarchical random number generation mechanism, a random source in execution of a machine learning algorithm is eliminated from a bottom layer, so that when different block chain nodes are executed based on the same random number sequence, the consistency of calculation results is remarkably improved, and the problem of prediction result deviation caused by random seed difference is solved. A hardware error statistical verification system is introduced, floating point errors caused by heterogeneous hardware are controlled within an acceptable range through floating point operation log records and chi-square distribution verification, calculation result inconsistency caused by hardware architecture differences is avoided, and calculation resource consumption is greatly reduced. Besides, for multi-dimension and multi-source heterogeneous characteristics of multi-dimensional data, through an RLP coding serialization model and a Merkel tree verification mechanism, full-link credible evidence storage in a prediction process is realized, and a core pain point that a prediction result is not credible in a traditional scheme is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of interdisciplinary technology of blockchain and artificial intelligence, specifically involving a reliable prediction method for multidimensional data based on deterministic execution of blockchain using machine learning. Background Technology

[0002] Machine learning (ML) is a core branch of artificial intelligence (AI), dedicated to enabling computer systems to automatically learn and improve performance through data and algorithms. Its core idea is to extract patterns from historical data (training models) and then make predictions or decisions based on new data. Traditional machine learning models that rely on single-machine or centralized server deployments face challenges in application, including data reliability issues (data sharing across multiple departments is prone to tampering, and input randomness affects model consistency) and consensus problems with computational results (inconsistent results due to random model operations across multiple nodes).

[0003] Existing solutions combining machine learning and blockchain suffer from several drawbacks, including the contradiction between blockchain consensus mechanisms and the randomness of machine learning, the exacerbation of inconsistent results due to hardware heterogeneity, and the lack of adaptation mechanisms designed for the multi-dimensional characteristics of data. Specifically, in blockchain technology, the consensus mechanism requires all nodes to reach strict agreement on the computation results, but the execution of machine learning algorithms often relies on randomness, leading to a fundamental contradiction. Randomness in machine learning manifests itself in three main ways: random number selection during the input phase, random source generation during the computation process, and randomization errors introduced by hardware floating-point operations. This randomness can cause different nodes to produce different computation results when executing the same machine learning algorithm, thus undermining the consensus mechanism of the blockchain network. Currently, pre-compiled contracts on the blockchain only support deterministic computation and cannot handle the randomness dependency problem of machine learning algorithms. Furthermore, due to differences in hardware architecture among different nodes, the distribution of floating-point errors in GPU parallel computing is also inconsistent, further exacerbating the inconsistency of computation results. Traditional solutions require all nodes to repeatedly execute the entire machine learning computation process for verification, which is not only inefficient but also results in a huge waste of computing resources. Therefore, how to maintain the randomness of machine learning algorithms while ensuring consensus among nodes in a blockchain network has become a pressing technical problem. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention provides a reliable prediction method for multidimensional data in machine learning based on blockchain deterministic execution. By employing a hierarchical random number generation mechanism and hardware error statistical verification, it solves the practical problems of inconsistent consensus caused by the randomness of machine learning computation in a decentralized environment, which in turn leads to unreliable data prediction results.

[0005] The technical solution of this invention is as follows:

[0006] The first aspect of this invention provides a reliable prediction method for multidimensional data based on blockchain deterministic execution using machine learning, comprising the following steps:

[0007] Users submit machine learning models to the blockchain. After the machine learning model is uploaded to the blockchain, the blockchain returns a unique identifier, model_id, to the user.

[0008] The caller initiates a machine learning model call through a pre-compiled contract interface based on the unique identifier model_id of the machine learning model. The block-producing node on the blockchain loads and executes the machine learning model on the blockchain to make predictions, obtains the prediction results, and generates a new block.

[0009] Block producers on the blockchain broadcast the block information of new blocks; the block information is the information stored in the block header and block body of the new block;

[0010] Verification nodes on the blockchain receive block information and verify it.

[0011] After successful verification, the machine learning model is reused to perform predictions and obtain the prediction results.

[0012] The two predictions are compared; if they match, a consensus is reached; otherwise, a fork occurs.

[0013] Furthermore, the user submits the machine learning model to the blockchain. After the machine learning model is uploaded to the blockchain, the blockchain returns a unique identifier, model_id, to the user, specifically:

[0014] A1: First, encode the machine learning model into a binary data stream;

[0015] A2: Generate metadata and use the ModelUpload interface of the smart contract to upload the metadata and the encoded binary data stream to the blockchain;

[0016] A3: Generate a unique identifier, model_id, for the machine learning model based on the encoded binary data stream.

[0017] Furthermore, the step of generating the unique identifier model_id of the machine learning model based on the encoded binary data stream specifically involves:

[0018] The binary data stream is packaged using RLP encoding, and then fragmented and stored on the blockchain, with the hash value used as the model_id. The model_id = H(RLP(ModelData)), where ModelData represents the binary data stream, RLP represents RLP encoding, and H represents a cryptographic hash function; or the binary data stream is stored in IPFS, and the CID is recorded on the chain as the model_id.

[0019] Furthermore, the caller initiates a machine learning model call through a pre-compiled contract interface based on the unique identifier model_id of the machine learning model. The block-producing node loads and executes the machine learning model on the blockchain to make predictions, obtains the prediction results, and generates a new block, specifically as follows:

[0020] B1: Block-producing nodes automatically load machine learning models on the blockchain using model_id;

[0021] B2: The block-producing node uses the previous block hash prevBlockHash and the caller transaction counter nonce under the caller address callerAddress to generate a global random seed seed = H(prevBlockHash, callerAddress, nonce);

[0022] B3: Encapsulate the global random seed and the user-input call parameters together into an input' = (seed, params), where params are the user-input call parameters; the user-input call parameters are the parameters input to the called machine learning model;

[0023] B4: Use the block-producing node's private key sk to execute a verifiable random function VRF to generate a private random number sequence R. private and its proof π, (R private ,π)=VRF sk (seed), VRF sk This indicates that a verifiable random function (VRF) is executed based on the block-producing node's private key (sk).

[0024] B5: Based on the call parameters in the input', the block-producing node executes a machine learning model to make predictions, obtains the prediction results, and simultaneously inserts a log hook to record the input value x for each floating-point operation during the prediction process. in (i) and output value x out (i) And construct a floating-point operation log sequence Log = {(x in (i) ,x out (i) )}k i=1 , where i represents the index of the number of floating-point operations, and k represents the total number of floating-point operations;

[0025] B6: Using a private random number sequence R private Construct a Merkle tree from the floating-point operation log sequence Log, and then obtain the root hash RootHash of the Merkle tree = MerkleTreeCreate(H(R private ),H(Log)),MerkleTreeCreate means creating a Merkle tree;

[0026] B7: The block producer puts the Merkle tree root hash into the block header, and packages the predicted output, the private random number sequence and its proof, and the floating-point operation log sequence into the block body. The block header and block body are then combined to obtain a new block.

[0027] Furthermore, the verification node receives block information and verifies the block information, specifically as follows:

[0028] C1: After receiving the block information, the verification node verifies the integrity of the block information BlockData by verifying the Merkle Tree Verify and the blockchain root hash BlockHash of the new block; the blockchain root hash BlockHash is the hash value of the block header of the new block.

[0029] The verification method is as follows:

[0030] MerkleTreeVerify(BlockHash,BlockData,H(R private ),H(Log))

[0031] C2: Verify the correctness of the global random seed using the previous block hash PreBlockHash and the caller transaction counter nonce under the caller address callerAddress;

[0032] assert(H(prevBlockHash,callerAddress,nonce)==seed) (3)

[0033] Here, assert represents the assertion method for judging the execution result. If the parameter is False, the verification fails; if the parameter is True, the verification passes.

[0034] C3: Verify the correctness of the private random number sequence using the public key of the block-producing node;

[0035] assert(VerifyVRF(pubkey,seed,random_bytes,proof)==true) (4)

[0036] The VerifyVRF function is used to verify whether the random number sequence output by the verifiable random function is valid and has not been tampered with. pubkey represents the public key of the block-producing node, random_bytes represents the random number sequence generated by VRF, and proof represents the VRF integrity proof.

[0037] C4: Calculate the standardized error of floating-point numbers in the floating-point operation log sequence;

[0038] C5: Normalizes the standardization error of floating-point numbers;

[0039] C6: Verify, through chi-square test, whether the standardized error of floating-point operations after normalization follows a uniform distribution. If it satisfies... If the verification passes, it will pass; otherwise, it will fail.

[0040]

[0041] Where, χ 2 It is the chi-square statistic; O e Here, E is the actual observed frequency in the e-th error interval, and E is the theoretical frequency with a value of [value missing]. n is the total number of samples, and α is the significance level. Let α be the confidence interval for parameter α.

[0042] Furthermore, the reusing of the machine learning model to perform predictions specifically involves: if random operations are encountered during the prediction process using the machine learning model, random numbers are sequentially selected from a private random number sequence for randomization; if floating-point operations are encountered during the prediction process, the output value is used directly without calculation by querying the input and output in the floating-point operation log sequence.

[0043] A second aspect of the present invention provides an electronic device, comprising: a processor, a memory, and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor communicates with the memory via the bus, and when the machine-readable instructions are executed by the processor, the steps of the machine-readable multidimensional data reliable prediction method based on blockchain deterministic execution are performed.

[0044] A third aspect of the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, performs the steps of the described machine learning multidimensional data reliable prediction method based on blockchain deterministic execution.

[0045] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0046] By constructing a hierarchical random number generation mechanism (combining a global random seed with a VRF-verifiable private random number sequence), the source of randomness in machine learning algorithm execution is eliminated from the ground up. This significantly improves the consistency of calculation results when different blockchain nodes execute based on the same random number sequence, completely resolving the prediction result deviation problem caused by differences in random seeds in traditional solutions. Simultaneously, an innovative hardware error statistical verification system is introduced. Through floating-point operation log recording and chi-square distribution testing, floating-point errors caused by heterogeneous hardware are controlled within an acceptable range, avoiding inconsistencies in calculation results caused by differences in hardware architecture. Compared to the traditional method of repeated calculation verification across all nodes, this significantly reduces computational resource consumption. Furthermore, addressing the multi-dimensional and multi-source heterogeneous characteristics of multi-dimensional data, an RLP encoding serialization model and Merkle tree verification mechanism are used to achieve end-to-end reliable evidence storage in the prediction process, resolving the core pain point of unreliable prediction results in traditional solutions. Attached Figure Description

[0047] Figure 1 This is a flowchart illustrating the execution of the on-chain machine learning algorithm according to an embodiment of the present invention.

[0048] Figure 2 This is a diagram illustrating the protocol process by which consensus nodes reach a consensus on a machine learning algorithm according to an embodiment of the present invention.

[0049] Figure 3 This is a schematic diagram illustrating how the computational verification data is organized into a Merkle tree according to an embodiment of the present invention. Detailed Implementation

[0050] The present invention will now be described in detail with reference to the accompanying drawings and embodiments.

[0051] This embodiment provides a machine learning method for reliable prediction of multidimensional data based on blockchain deterministic execution, such as... Figure 1 and Figure 2 As shown, it includes the following steps:

[0052] Step 1: The user submits the machine learning model to the blockchain. After the machine learning model is uploaded to the blockchain, the blockchain returns a unique identifier (model_id) to the user. The machine learning model includes decision tree models and ensemble models.

[0053] Step 1.1: First, encode the machine learning model into a binary data stream;

[0054] The decision tree model is encoded into a binary data stream using a standardized serialization format:

[0055] For each non-leaf node N in the decision tree modelq Its splitting rule can be expressed as a conditional judgment:

[0056]

[0057] Among them, split(N) q ) represents a non-leaf node N. q The splitting rule, N q For non-leaf nodes, q is the index of the non-leaf node, and feature(N) q ) represents the feature selected for non-leaf nodes, threshold(N) q ) is the split threshold; leaf node L j Then record the category label or regression value, where j is the number of the leaf node; finally, the decision tree model is encoded as a binary data stream including all node parameters;

[0058] Encode the ensemble model (such as a random forest) as a binary data stream: Set the ensemble model to consist of M base models. Each basic model is independently encoded as Encode(F) m The serialized model data of the ensemble model, i.e., the binary data stream representation, is as follows:

[0059]

[0060] Where Encode (Ensemble) is the serialized model data of the ensemble model, F m The base model is defined as follows: m is the base model number, M is the number of base models, Ensemble is the type identifier of the base model (e.g., "Decision Tree", "GBDT", etc.), and Encode(F) is the base model type identifier. m ) is based on the model F m Independent encoding, w m Based on model F m The weights are defined by the aggregation function, such as a weighted average. Or voting strategy, x m The parameters of the basic model are configured (such as decision tree depth, leaf node threshold, etc.); during encoding, the encoding, weights and fusion rules of each basic model need to be recorded to form complete serialized model data;

[0061] Step 1.2: Generate metadata and use the ModelUpload interface of the smart contract to upload the metadata and the encoded binary data stream to the blockchain;

[0062] In this embodiment, metadata (additional information), such as the number of trees in the random forest, the depth of the decision tree, and the model combination method of ensemble learning, are used for subsequent model parsing and computation. In this embodiment, the ModelUpload interface of the smart contract is: function ModelUpload(bytes memory modelData, bytes memory metadata) public returns(bytes32).

[0063] Step 1.3: Generate a unique identifier, model_id, for the machine learning model based on the encoded binary data stream;

[0064] In this embodiment, RLP encoding can be used to package the binary data stream and generate a 256-bit hash: the binary data stream can be fragmented and stored on the blockchain and the hash value can be used as model_id, or the binary data stream can be stored on IPFS and the CID can be recorded on the chain as model_id. model_id is a unique identifier for storing machine learning models on the blockchain. It is generated by a hash algorithm, model_id = H(RLP(ModelData)), where ModelData represents the binary data stream, RLP represents RLP encoding, and H represents a cryptographic hash function.

[0065] Step 2: The caller initiates a machine learning model call through the pre-compiled contract interface based on the unique identifier model_id of the machine learning model. The block-producing node loads and executes the machine learning model on the blockchain to predict multi-dimensional data, obtains the prediction results, and generates a new block.

[0066] In this embodiment, when the blockchain network starts, a pre-compiled contract for a specific machine learning algorithm is added to the blockchain node: Predict(model_id, blockheader, input) → (output, proof). Predict represents the prediction method of the machine learning algorithm in the pre-compiled contract, blockheader represents the block header information of the current block, input represents the caller's input, output is the algorithm execution result (prediction result), and proof contains verifiable evidence such as VRF proof, floating-point operation logs, and intermediate state hashes. The contract address is fixed at 0x000...ML01, supporting on-chain model loading, parsing, and computation. Simultaneously, adapted pre-compiled contract modules are designed for decision trees and ensemble learning algorithms to ensure their normal operation on the chain.

[0067] {

[0068] "opcode":0xML01, / / Predefined machine learning opcode

[0069] "model_id":"0x8921...a3f",

[0070] "input":[params]

[0071] }

[0072] Step 2.1: The block-producing node automatically loads the machine learning model on the blockchain using the model_id;

[0073] Step 2.2: The block-producing node uses the hash of the previous block, prevBlockHash, and the caller transaction counter nonce under the caller address, callerAddress, to generate a global random seed seed = H(prevBlockHash, callerAddress, nonce);

[0074] In this embodiment, prevBlockHash: previous block hash (256 bits), callerAddress: caller's Ethereum address (160 bits), and nonce: caller's transaction counter (64-bit unsigned integer).

[0075] Step 2.3: Encapsulate the global random seed and the user-input call parameters together into an input' = (seed, params), where params are the call parameters; the user-input call parameters are the parameters input to the called machine learning model;

[0076] Step 2.4: Use the block-producing node's private key sk to execute the verifiable random function VRF to generate a private random number sequence R. private and its proof π, (R private ,π)=VRF sk (seed) ensures that the random numbers are publicly verifiable and unpredictable; VRF sk This indicates that a verifiable random function (VRF) is executed based on the block-producing node's private key (sk).

[0077] In this embodiment, a VRF on the elliptic curve secp256k1 is used, and a verifiable random function (VRF) is executed to output a 512-bit random byte stream and a 96-byte zero-knowledge proof.

[0078] Step 2.5: Based on the calling parameters in the input', the block-producing node executes the machine learning model to predict the multidimensional data, obtains the prediction results, and inserts a log hook to record the input value x for each floating-point operation during the prediction process. in (i) and output value x out (i)And construct a floating-point operation log sequence Log = {(x in (i) ,x out (i) )} k i=1 , where i represents the index of the number of floating-point operations, and k represents the total number of floating-point operations;

[0079] The execution process of the machine learning model in this embodiment is as follows:

[0080] For random forests: During model inference, log hooks are inserted to record the node access order, feature comparison results, path selection, and other operations for each tree traversal process; at the same time, input and output operations such as matrix multiplication (if there are feature vector and weight matrix calculations), convolution operations (if there is image-related processing), activation functions (ReLU, Sigmoid, etc., if nonlinear transformations are involved), Dropout layer random mask (if there is random drop operation), and floating-point operations are recorded.

[0081] For ensemble learning (such as decision trees): Record the above operations of each basic model in the inference process according to the model combination method, and record the input and output of the fusion process of the output results of each basic model (such as weighted summation, voting, etc.).

[0082] Step 2.6: Use the private random number sequence R private Constructing a log sequence of floating-point operations as follows Figure 3 The Merkle tree shown is used to obtain the root hash of the Merkle tree, RootHash = MerkleTreeCreate(H(R private ),H(Log)),MerkleTreeCreate means creating a Merkle tree;

[0083] Step 2.7: The block producer puts the Merkle tree root hash RootHash into the block header, and packs the prediction output, private random number sequence and its proof, and floating-point operation log sequence into the block body. The block header and block body are then combined to obtain a new block.

[0084] Step 3: The block producer broadcasts the block information of the new block; the block information is the information stored in the block header and block body of the new block;

[0085] Step 4: The validator node receives the block information and verifies it;

[0086] Verification nodes ensure result correctness through triple verification: 1) checking VRF random numbers and Merkle proofs; 2) verifying the floating-point error distribution; 3) re-executing the algorithm using the same random numbers, directly calling the floating-point results in the log to guarantee deterministic recalculation. Non-block-producing consensus nodes verify the execution results of the machine learning algorithm as follows:

[0087] Step 4.1: After receiving the block information, the verification node (non-block producing node) verifies the integrity of the block information BlockData through Merkle Tree Verify and the blockchain root hash BlockHash of the new block; the blockchain root hash BlockHash is the hash value of the block header of the new block.

[0088] The verification method is as follows:

[0089] MerkleTreeVerify(BlockHash,BlockData,H(R private ),H(Log))

[0090] Step 4.2: Verify the correctness of the global random seed using the previous block hash PreBlockHash and the caller transaction counter nonce under the caller address callerAddress;

[0091] assert(H(prevBlockHash,callerAddress,nonce)==seed) (3)

[0092] Here, assert represents the assertion method for judging the execution result. If the parameter is False, the verification fails; if the parameter is True, the verification passes.

[0093] Step 4.3: Verify the correctness of the private random number sequence using the public key of the block-producing node;

[0094] assert(VerifyVRF(pubkey,seed,random_bytes,proof)==true) (4)

[0095] The VerifyVRF function verifies whether the random number sequence output by the Verify Random Function (VRF) is valid and has not been tampered with. pubkey represents the public key of the block-producing node, random_bytes represents the random number sequence generated by the VRF, and the integrity of this random number sequence will be verified here; proof represents the VRF integrity proof, ensuring that the random numbers are generated by a legitimate private key and have not been tampered with.

[0096] Step 4.4: Calculate the standardized error of the floating-point numbers in the floating-point operation log sequence;

[0097] In this embodiment, single-precision floating-point numbers are used as an example. The method for calculating the standardized error is as follows:

[0098]

[0099] Where, Δ i Let x be the standardized error of the i-th floating-point operation. theoretical It is the theoretical value, x actual It is the actual value. ulp(x) is a function used to calculate the last digit of the single-precision floating-point number x, which is used to measure the precision of the single-precision floating-point number. floor indicates the floor operation.

[0100] Step 4.5: Normalize the standardization error of the floating-point number to the (0,1) interval;

[0101]

[0102] Where, Δ i′ is the normalized error of the i-th floating-point operation after normalization, min(Δ) is the minimum value among all original normalized errors, and max(Δ) is the maximum value among all original normalized errors;

[0103] Step 4.6: Verify through chi-square test whether the standardized error of floating-point operations after normalization follows a uniform distribution. If it satisfies... If the verification passes, it will pass; otherwise, it will fail.

[0104]

[0105] Where, χ 2 It is the chi-square statistic, used to test whether floating-point errors follow a uniform distribution; O e Here, E is the actual observed frequency in the e-th error interval, and E is the theoretical frequency with a value of [value missing]. n is the total number of samples, and α = 0.05 is the significance level. Let α be the confidence interval.

[0106] Decision Tree: Discretize the error values ​​in the decision tree inference process into 100 intervals (0.00-0.01,...,0.99-1.00), perform a chi-square test, and if the rejection region is χ²... 2 >χ 2 If the value _{0.05,99} falls within the rejection field, the verification will fail.

[0107] Ensemble learning (e.g., random forest): Based on the combination of base models, the error values ​​of each base model are discretized and subjected to a chi-square test. If the error test of any base model fails, the validation fails. For example, a random forest consists of multiple trees, potentially resulting in multiple sets of floating-point operation results. The error value of each tree is discretized into 100 intervals (0.00-0.01,...,0.99-1.00), and then a chi-square test is performed on each tree, with the rejection region being χ². 2 >χ 2 If the error check of any tree falls within the rejection region, the verification fails.

[0108] Step 5: After the global random seed, private random number sequence, and floating-point operation log sequence have been verified, the input is reconstructed using the global random seed and the user-input call parameters. The machine learning model is then used again to perform predictions based on the private random number sequence and floating-point operation log sequence to obtain the prediction results.

[0109] After the above verification is passed, the calculation begins. During the execution process, if randomization operations (such as dropout) are required, random numbers are selected sequentially from a private random number sequence for randomization. When floating-point operations are encountered, the input and output of the floating-point operation log sequence are queried, and the output value is used directly without calculation. Depending on the model type, the corresponding calculation logic is called for inference, such as random forest traversing multiple trees and aggregating the results, decision tree deducing step by step from the root node, and ensemble learning fusing the outputs of various basic models.

[0110] Step 6: Compare the two prediction results. If they match, the consensus is successful; otherwise, a fork occurs.

[0111] This embodiment also provides an electronic device, including: a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor communicates with the memory via the bus. When the machine-readable instructions are executed by the processor, the steps of the machine-readable multidimensional data reliable prediction method based on blockchain deterministic execution are performed.

[0112] This embodiment also provides a computer-readable storage medium storing a computer program that, when executed by a processor, performs the steps of the machine learning multidimensional data reliable prediction method based on blockchain deterministic execution.

[0113] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope defined by the present invention.

Claims

1. A machine learning method for reliable prediction of multidimensional data based on blockchain deterministic execution, characterized in that, Includes the following steps: Users submit machine learning models to the blockchain. After the machine learning model is uploaded to the blockchain, the blockchain returns a unique identifier, model_id, to the user. The caller initiates a machine learning model call through a pre-compiled contract interface based on the unique identifier model_id of the machine learning model. The block-producing node on the blockchain loads and executes the machine learning model on the blockchain to make predictions, obtains the prediction results, and generates a new block. Block producers on the blockchain broadcast the block information of new blocks; the block information is the information stored in the block header and block body of the new block; Verification nodes on the blockchain receive block information and verify it. After successful verification, the machine learning model is reused to perform predictions and obtain the prediction results. The two predictions are compared; if they match, a consensus is reached; otherwise, a fork occurs.

2. The machine learning multidimensional data reliable prediction method based on blockchain deterministic execution according to claim 1, characterized in that, The user submits the machine learning model to the blockchain. After the machine learning model is uploaded to the blockchain, the blockchain returns a unique identifier, model_id, to the user. A1: First, encode the machine learning model into a binary data stream; A2: Generate metadata and use the ModelUpload interface of the smart contract to upload the metadata and the encoded binary data stream to the blockchain; A3: Generate a unique identifier, model_id, for the machine learning model based on the encoded binary data stream.

3. The machine learning multidimensional data reliable prediction method based on blockchain deterministic execution according to claim 2, characterized in that, The process of generating a unique identifier, model_id, for the machine learning model based on the encoded binary data stream is as follows: The binary data stream is packaged using RLP encoding, and then fragmented and stored on the blockchain, with the hash value used as the model_id. The model_id = H(RLP(ModelData)), where ModelData represents the binary data stream, RLP represents RLP encoding, and H represents a cryptographic hash function; or the binary data stream is stored in IPFS, and the CID is recorded on the chain as the model_id.

4. The machine learning multidimensional data reliable prediction method based on blockchain deterministic execution according to claim 1, characterized in that, The caller initiates a machine learning model call through a pre-compiled contract interface based on the unique identifier `model_id` of the machine learning model. The block-producing node loads and executes the machine learning model on the blockchain to make predictions, obtains the prediction results, and generates a new block. Specifically: B1: Block-producing nodes automatically load machine learning models on the blockchain using model_id; B2: The block-producing node uses the previous block hash prevBlockHash and the caller transaction counter nonce under the caller address callerAddress to generate a global random seed seed = H(prevBlockHash, callerAddress, nonce); B3: Encapsulate the global random seed and the user-input call parameters together into an input' = (seed, params), where params are the user-input call parameters; the user-input call parameters are the parameters input to the called machine learning model; B4: Use the block-producing node's private key sk to execute a verifiable random function VRF to generate a private random number sequence R. private and its proof π, (R private ,π)=VRF sk (seed), VRF sk This indicates that a verifiable random function (VRF) is executed based on the block-producing node's private key (sk). B5: Based on the call parameters in the input', the block-producing node executes a machine learning model to make predictions, obtains the prediction results, and simultaneously inserts a log hook to record the input value x for each floating-point operation during the prediction process. in (i) and output value x out (i) And construct a floating-point operation log sequence Log = {(x in (i) ,x out (i) )} k i=1 , where i represents the index of the number of floating-point operations, and k represents the total number of floating-point operations; B6: Using a private random number sequence R private Construct a Merkle tree from the floating-point operation log sequence Log, and then obtain the root hash RootHash of the Merkle tree = MerkleTreeCreate(H(R private ),H(Log)),MerkleTreeCreate means creating a Merkle tree; B7: The block producer puts the Merkle tree root hash into the block header, and packages the predicted output, the private random number sequence and its proof, and the floating-point operation log sequence into the block body. The block header and block body are then combined to obtain a new block.

5. The machine learning multidimensional data reliable prediction method based on blockchain deterministic execution according to claim 1, characterized in that, The verification node receives block information and verifies the block information, specifically as follows: C1: After receiving the block information, the verification node verifies the integrity of the block information BlockData by verifying the Merkle Tree Verify and the blockchain root hash BlockHash of the new block; the blockchain root hash BlockHash is the hash value of the block header of the new block. The verification method is as follows: MerkleTreeVerify(BlockHash,BlockData,H(R private ),H(Log)) C2: Verify the correctness of the global random seed using the previous block hash PreBlockHash and the caller transaction counter nonce under the caller address callerAddress; assert(H(prevBlockHash,callerAddress,nonce)==seed) (3) Here, assert represents the assertion method for judging the execution result. If the parameter is False, the verification fails; if the parameter is True, the verification passes. C3: Verify the correctness of the private random number sequence using the public key of the block-producing node; assert(VerifyVRF(pubkey,seed,random_bytes,proof)==true) (4) The VerifyVRF function is used to verify whether the random number sequence output by the verifiable random function is valid and has not been tampered with. pubkey represents the public key of the block-producing node, random_bytes represents the random number sequence generated by VRF, and proof represents the VRF integrity proof. C4: Calculate the standardized error of floating-point numbers in the floating-point operation log sequence; C5: Normalizes the standardization error of floating-point numbers; C6: Verify, through chi-square test, whether the standardized error of floating-point operations after normalization follows a uniform distribution. If it satisfies... If the verification passes, it will pass; otherwise, it will fail. Where, χ 2 It is the chi-square statistic; O e Here, E is the actual observed frequency in the e-th error interval, and E is the theoretical frequency with a value of [value missing]. n is the total number of samples, and α is the significance level. Let α be the confidence interval for parameter α.

6. The machine learning multidimensional data reliable prediction method based on blockchain deterministic execution according to claim 1, characterized in that, The reusing of the machine learning model to perform predictions specifically involves: if random operations are encountered during the prediction process using the machine learning model, random numbers are sequentially selected from the private random number sequence for randomization; if floating-point operations are encountered during the prediction process, the output value is used directly without calculation by querying the input and output in the floating-point operation log sequence.

7. An electronic device, characterized in that, include: The device includes a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor communicates with the memory via the bus. When the machine-readable instructions are executed by the processor, they perform the steps of the machine-learning multidimensional data reliable prediction method based on blockchain deterministic execution as described in any one of claims 1-6.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, performs the steps of the machine learning multidimensional data reliable prediction method based on blockchain deterministic execution as described in any one of claims 1-6.