A blockchain-based artificial intelligence model training method and system
Through smart contracts encrypting data and using blockchain storage, combined with decentralized computing resources and consensus mechanisms, the problems of data privacy and transparency in artificial intelligence model training are solved, and a safe and efficient model training process and transparent update mechanism are realized.
Patent Information
- Application Number
- CN202510389947.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-03-31
AI Technical Summary
During the existing artificial intelligence model training process, the problems of data privacy, transparency of the training process and credibility of the results have not been effectively solved, and relying on centralized computing resources has problems such as data leakage risks and uneven allocation of computing resources.
Training data is automatically collected and encrypted through smart contracts, and it is tamper-free to store using blockchain. Data access permissions are verified according to blockchain smart contract rules, decentralized computing resources are used for training, and the training process is recorded in real time through blockchain. A consensus mechanism is used to monitor abnormal behaviors during the training process, optimize training parameters, and ensure the stability and transparency of model training.
It improves the security, efficiency and credibility of artificial intelligence model training, ensures data privacy and transparency of training processes, and realizes traceability of model updates.
Smart Images

Figure CN119892377B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence, and particularly to an artificial intelligence model training method and system based on blockchain. Background Art
[0002] With the rapid development of artificial intelligence technology, model training has achieved remarkable results in multiple fields. However, in the process of artificial intelligence model training, issues such as data privacy, transparency of the training process, and credibility of results are still important factors restricting technological progress. Traditional artificial intelligence training methods usually rely on centralized computing resources. Although this method is efficient, it also has problems such as data leakage risk, non-traceability of the process, and uneven distribution of computing resources.
[0003] As a decentralized distributed ledger technology, blockchain technology has advantages such as immutability, transparency, and strong security. It can solve problems such as data privacy protection, transparency, and traceability, and automate the management of data access, computing resource scheduling, and model updates through smart contracts. However, the current application of blockchain technology in artificial intelligence model training is still in its infancy. How to effectively combine blockchain with the artificial intelligence training process to improve the efficiency and reliability of training remains a technical problem to be solved urgently.
[0004] Therefore, developing an artificial intelligence model training method based on blockchain to ensure data security, transparency of the training process, and traceability of the model through the decentralized characteristics of blockchain has important technical significance and broad application prospects. Summary of the Invention
[0005] The present invention provides an artificial intelligence model training method based on blockchain, including the following steps:
[0006] Automatically collect and encrypt training data through a smart contract to ensure data privacy and security, and store it immutably through blockchain;
[0007] Automatically verify and determine the access rights of data according to the rules of the blockchain smart contract to ensure that only authorized computing nodes can access the training data;
[0008] Transmit the verified data through the blockchain network to distributed computing nodes, use decentralized computing resources to train the artificial intelligence model, and record the intermediate results in real time during the training process through blockchain;
[0009] Monitor and correct abnormal behaviors during the training process through the consensus mechanism of blockchain nodes, and optimize training parameters to ensure the stability and efficiency of model training;
[0010] After the training is completed, the final model and its related information are stored through the blockchain, and the model version is released through a smart contract to ensure the traceability and transparency of model updates.
[0011] A method for training an artificial intelligence model based on blockchain as described above, wherein the training data is collected and encrypted, specifically including automatically obtaining raw data from multiple data sources, and after encrypting the raw data through an encryption algorithm, storing it in the blockchain to ensure the privacy and security of the data; the multiple data sources include data collected by sensors, real-time monitoring data, user input data, images, videos, or data from other sensor networks.
[0012] A method for training an artificial intelligence model based on blockchain as described above, wherein according to the rules of the blockchain smart contract, the access rights of the data are automatically verified and determined, specifically by authenticating the identity of the computing nodes according to a preset permission control model, and automatically confirming whether the nodes have the authorization to read the training data. If a node fails the verification, its access request is rejected.
[0013] A method for training an artificial intelligence model based on blockchain as described above, wherein decentralized computing resources are used to train the artificial intelligence model, specifically by executing the model training task through decentralized computing resources and recording the generated intermediate results in real time through the blockchain during the training process.
[0014] A method for training an artificial intelligence model based on blockchain as described above, wherein when monitoring abnormal behaviors during the training process through the consensus mechanism of blockchain nodes, it further includes comparing the error between the calculation results of each node and the global model during the training process, and dynamically adjusting the influence of the node on the global training result based on the error value.
[0015] A method for training an artificial intelligence model based on blockchain as described above, wherein the training parameters are optimized, specifically including dynamically adjusting the learning rate of the computing nodes to adapt to the data volume, computing load, and historical training performance of different nodes.
[0016] A method for training an artificial intelligence model based on blockchain as described above, wherein the model version is released through a smart contract, specifically including recording and publishing the model parameters, performance metrics, and summary information of the training data after the training is completed through the smart contract to the blockchain, ensuring that the release and update of each version have a clear timestamp and can be traced back to the source.
[0017] The present invention also provides a system for training an artificial intelligence model based on blockchain, including:
[0018] A data collection module, which is used to automatically collect and encrypt training data through a smart contract, ensure the privacy and security of the data, and perform tamper-proof storage through the blockchain;
[0019] A smart contract module, which is used to automatically verify and determine the access rights of data according to the rules of the blockchain smart contract, ensuring that only authorized computing nodes can access the training data;
[0020] A computing resource scheduling module, which is used to transmit the verified data to distributed computing nodes through the blockchain network, utilize decentralized computing resources to train the artificial intelligence model, and record the intermediate results during the training process in real time through the blockchain;
[0021] A blockchain verification module, which is used to monitor and correct abnormal behaviors during the training process through the consensus mechanism of blockchain nodes, and optimize the training parameters to ensure the stability and efficiency of model training;
[0022] A model release module, which is used to store the final model and its related information through the blockchain after the training is completed, and release the model version through the smart contract to ensure the traceability and transparency of model updates.
[0023] The beneficial effects achieved by the present invention are as follows: The present invention not only solves the problems of security, transparency, and efficiency faced in traditional artificial intelligence training, but also improves the overall performance of artificial intelligence model training through innovative blockchain applications. Description of the Drawings
[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present invention. For those of ordinary skill in the art, other drawings can also be obtained according to these drawings.
[0025] Figure 1 It is a flowchart of a method for training an artificial intelligence model based on blockchain provided in Embodiment 1 of the present application;
[0026] Figure 2 It is a schematic diagram of a data annotation system based on big data management provided in Embodiment 2 of the present application. Detailed Embodiments
[0027] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0028] Example 1
[0029] As Figure 1 shown, Example 1 of this application provides an artificial intelligence model training method based on blockchain, including the following steps:
[0030] Step S10: Automatically collect and encrypt training data through a smart contract to ensure the privacy and security of the data, and store it immutably through the blockchain;
[0031] First, the system establishes connections with multiple data sources and automatically collects the required training data. This data includes data collected by sensors, real-time monitoring data, user input data, images, videos, or other data from sensor networks. The data sources can be devices, sensors, or cloud services, etc. Through interfaces, the smart contract automatically controls the data collection process to ensure the real-time, integrity, and legality of data collection.
[0032] After the data collection is completed, the system encrypts the data through an encryption algorithm to prevent the data from being accessed or tampered with by unauthorized entities during storage and transmission. The system uses modern encryption technologies, such as symmetric encryption algorithms (e.g., AES algorithm) or asymmetric encryption algorithms (e.g., RSA algorithm), and selects appropriate encryption methods according to the different natures of the data. The smart contract automatically generates, manages, and distributes keys to ensure the secure exchange and storage of keys. The encryption process is fully controlled by the smart contract to ensure that each encryption link meets the established security standards, and the encryption keys are stored on the blockchain to further prevent key leakage and tampering.
[0033] The encrypted data will be uploaded to the blockchain for storage. In this process, the system generates a unique hash value for each data segment to ensure the integrity and immutability of the data. By storing the data hash value on the blockchain, any attempt to modify the data will be detected by the blockchain network because the change in the hash value means a change in the data. As a decentralized distributed ledger, the blockchain ensures the immutability and transparency of data storage. Each stored data generates a block, and these blocks are linked to form a chain to ensure the security of data storage.
[0034] Throughout the process, the smart contract automatically executes operations such as data collection, encryption, storage, and verification. The smart contract manages the data collection frequency, encryption method, and storage path according to predetermined rules to ensure the efficiency and consistency of the data collection and encryption processes. The management of encryption keys is also automatically executed by the smart contract, and the system can flexibly manage encryption keys according to the different natures of the data to ensure the efficiency and security during the encryption process.
[0035] During the data collection, encryption, and storage processes, the smart contract also performs data integrity checks to ensure that all collected data is not lost or tampered with. After each data collection, the smart contract calculates the hash value of the data and compares it with the previous data to ensure that the newly collected data is consistent with the original data. If any data anomalies are detected, the smart contract will automatically trigger an alarm and suspend further processing of the data to ensure data reliability.
[0036] Step S20: Automatically verify and determine the access rights of the data according to the rules of the blockchain smart contract to ensure that only authorized computing nodes can access the training data;
[0037] In step S20, the system automatically executes the verification and management of data access rights through the rules of the blockchain smart contract to ensure that only authorized computing nodes can access the encrypted training data stored on the blockchain.
[0038] When a computing node needs to access the training data, the smart contract authenticates the node requesting access. Each computing node has a unique identifier on the blockchain, which is usually based on a public key, digital signature, or other forms of identity credentials. This identity information has been previously recorded in the blockchain, and the smart contract matches the node information stored in the blockchain with the current requesting node's identity to confirm whether the node is a legitimate computing node. If the verification fails, the node's access request will be rejected.
[0039] Once the authentication is successful, the smart contract continues to check the node's access rights. The access rights for each dataset are preset according to the sensitivity of the data, the role of the computing node, and its historical reputation. The smart contract dynamically calculates the access rights of the node based on the node's historical behavior and its reputation score. The reputation score formula generates a weighted score by combining factors such as the node's historical behavior, access frequency, and consumed computing resources. The following is the formula for dynamically calculating the node's reputation score: , where is the final reputation score of the computing node, reflecting the overall reputation of the node and determining whether the node can access a specific dataset; is the weight coefficient related to the i-th behavior score is the reputation score of the node in its i-th past behavior. Each behavior (such as data request, computing task execution, etc.) generates a score. represents the reputation value of the i-th behavior; n is the number of all historical behaviors of the node in the system, representing the total number of the node's behaviors; is the adjustment factor used to adjust the influence of the average behavior reputation of all nodes on the final reputation score; is the kth score in the node's historical behavior; m is the number of historical scores, reflecting the node's behavior frequency and behavior pattern; It is a measure of the access frequency of a node, reflecting how frequently a node requests access to data; is the frequency adjustment factor, which is used to adjust the final reputation score according to the access frequency of the node; It is a measure of the maximum access time of a node in the system, and is used to measure the degree of consumption of system resources by the node; is the average access time of the node, which indicates the time consumed by the node to access data each time; It is a time adjustment factor used to adjust the impact of a node on consuming system resources. The higher the reputation score, the greater the possibility that the node is authorized to access data.
[0040] After calculating the final reputation score according to the rules of the smart contract, the smart contract will compare the node's score with the preset access rights threshold. If the node's reputation score exceeds the set threshold, the smart contract will grant the node access to the training data. If the node's score is below the threshold, the access request will be denied, and the denial will be automatically recorded in the blockchain, ensuring that all access decisions are transparent and traceable.
[0041] Smart contracts also monitor the permission management process in real time. Whenever a node requests access to data, the smart contract automatically records all relevant information about the request, including the node's identity, the timestamp of the access, the type of data set requested for access, and the result of granting access rights. This information will be encrypted and stored on the blockchain to ensure complete transparency of data access and can be audited and traced back at any time.
[0042] If a node is found to frequently initiate abnormal access requests during system operation, the smart contract will automatically adjust the node's permissions to reduce its chances of accessing sensitive data. Smart contracts can also analyze node behavior in real time and dynamically adjust access permissions based on the node's behavior history to ensure the security and fairness of the system.
[0043] Step S30: The verified data is transmitted to the distributed computing nodes through the blockchain network, and the AI model is trained using decentralized computing resources, and the intermediate results of the training process are recorded in real time through the blockchain;
[0044] In step S30, the verified and authorized computing nodes automatically obtain the encrypted training data stored in the blockchain through smart contracts. All transmission processes are carried out through the blockchain network to ensure that the data is not leaked or tampered with during transmission. Each computing node starts to process the training task in parallel based on its local computing resources and conducts model training on the local dataset. In distributed training, the training results of all computing nodes, such as the calculated gradients, loss function values, and model update parameters, will be recorded in real-time on the blockchain to ensure the transparency and immutability of the training process.
[0045] During the training process, the gradient update mechanism used by the computing nodes is based on the gradient descent method. Each computing node i calculates the gradient of the loss function with respect to its local dataset , and updates the local model parameters according to the gradient. For the parameter update of each node, the calculation formula is: , where represents the model parameters of the i-th node at the t-th iteration; is the learning rate; is the gradient calculated by the i-th node at the t-th iteration. The local update of each node is only based on its local dataset, while the global model update is based on the weighted average of the calculation results of all nodes. The global model update formula is calculated by weighting the gradients of the nodes to ensure that the contribution of each node in the global training is reasonable. The specific global update formula is as follows: , where represents the parameters of the global model at the (t + 1)-th iteration; are the parameters of the global model at the t-th iteration; is the weight of the i-th node; is the gradient calculated by the i-th node. Through this weighted average method, the gradient information of all nodes is aggregated into the global model, thus realizing the collaborative update of each computing node in distributed training.
[0046] To ensure the stability and accuracy of the training process, the system also introduces a dynamic adjustment mechanism, especially the adjustment of the adaptive learning rate. The learning rate of each node will be dynamically adjusted according to its historical performance, data quality, and training efficiency. The specific adjustment formula is: , where is the adaptive learning rate of the i-th node at the (t + 1)-th iteration; is the learning rate of the i-th node at the t-th iteration; is the adjustment factor; is the amount of local data processed by the i-th node; It is the data volume of the node with the largest data volume among all nodes. This formula ensures that the learning rate of each node is dynamically adjusted according to the amount of data it processes. If a node has a large amount of local data, the learning rate is low; conversely, it is high. This adjustment ensures that the update pace of each node matches its computing resources, avoiding instability in the training process caused by unbalanced computing loads.
[0047] During the entire training process, intermediate results of all computing nodes, such as gradients, loss functions, model parameters, etc., will be recorded in the blockchain in real time by the smart contract. Each time a node completes a training, the smart contract will encrypt and store the calculation results and training status on the blockchain, ensuring that the data is tamper-proof and always traceable. The update and participation behaviors of each node can be audited through the blockchain, ensuring the transparency and fairness of each step and avoiding the impact of any potential malicious behavior or calculation errors.
[0048] To ensure the convergence of the global model, the smart contract will also dynamically adjust the training strategy according to the performance of each node during the training process. If the gradient deviation of a certain node is large or the error during its training process is too high, the smart contract will automatically trigger an adjustment mechanism to reduce the weight of this node in the global training process or perform other optimizations. Through this mechanism, the system ensures that the training process always converges in the correct direction and avoids the improper behavior of any single node affecting the training effect of the entire system.
[0049] Step S40: Monitor and correct abnormal behaviors during the training process through the consensus mechanism of blockchain nodes, and optimize training parameters to ensure the stability and efficiency of model training;
[0050] In step S40, the system monitors and corrects abnormal behaviors during the training process of the artificial intelligence model in real time through the consensus mechanism of blockchain nodes, and automatically optimizes problems such as possible model instability or low training efficiency during the training process. The introduction of the consensus mechanism not only ensures the fairness and transparency of the training process, but also guarantees the efficiency and consistency of model training through a decentralized verification mechanism.
[0051] During each round of training, all computing nodes update their parameters according to the local dataset and the global model and calculate the gradients of the loss function. These calculation results are uploaded to the blockchain in real time through the smart contract and verified and recorded by blockchain nodes. However, due to factors such as data imbalance, node performance differences, and possible network delays in the distributed training process, some nodes may produce inconsistent results or abnormal behaviors during the training process. These abnormal behaviors may affect the update of the global model and even slow down or prevent the convergence of the model training.
[0052] To solve this problem, through its built-in consensus mechanism, the blockchain verifies the calculation results of each node in real time. The consensus mechanism automatically identifies possible outliers by comparing the gradient information submitted by all nodes. If the calculation result of a certain node significantly deviates from the results of other nodes, the system will automatically mark the result of this node as "abnormal" and identify and warn it through a smart contract. At this time, the smart contract will correct the calculation result of this node according to the preset rules, eliminate the abnormal data, and record relevant information through the blockchain to ensure the transparency and auditability of the training process.
[0053] After the abnormal behavior is identified and marked, the smart contract will automatically adjust the training parameters or learning rate of this node to avoid the adverse impact of this node on the global model. For example, the smart contract can automatically reduce the weight of this node in the global model update according to the error margin of the node, or adjust its learning rate to make its calculation result contribute more smoothly in the subsequent training process. Through this dynamic adjustment mechanism, the system can optimize the training process in real time, ensure the consistency between the training effect of each node and the global model, and ensure the stability and efficiency during the training process.
[0054] The smart contract will also automatically adjust other hyperparameters in the global model training process, such as the learning rate, batch size, etc., to cope with the training progress and error conditions of different nodes. By analyzing the training performance of each node, the smart contract can adaptively optimize these hyperparameters to ensure the maximization of the model training efficiency. For example, if the training process of some nodes is relatively slow, the smart contract can automatically increase the learning rate of these nodes, thereby accelerating their convergence speed and avoiding the "bottleneck" phenomenon during the training process.
[0055] Step S50: After the training is completed, store the final model and its related information through the blockchain, and release the model version through the smart contract to ensure the traceability and transparency of the model update.
[0056] After the model training is completed, the smart contract will automatically summarize all the information during the training process, including but not limited to the final parameters of the model, the source of the training data, the hyperparameter configuration during the training process (such as the learning rate, batch size, etc.), the convergence of the loss function, and the performance of the model on the training set and the validation set. All this information will be stored on the blockchain through an encryption algorithm to ensure the security and immutability of the data. The decentralized feature of the blockchain enables this information to be verified and stored on multiple nodes, thus avoiding the risk of single-point failure and ensuring the transparency and credibility of the data.
[0057] Subsequently, the smart contract will generate a new model version according to the training process and release this version to the blockchain. Each released model version contains the following content:
[0058] The final parameters of the model, i.e., the model parameters obtained after the training process;
[0059] Summary information of the dataset used in the training process;
[0060] Hyperparameter configuration during the model training process (such as learning rate, batch size, etc.);
[0061] Performance metrics during the training process (such as loss function value, accuracy, etc.);
[0062] The generation timestamp of the model version and the relevant blockchain transaction hash value.
[0063] Through the automated control of the smart contract, the model version number and relevant information will be updated in real time, and a new record will be generated on the blockchain. Each newly released version contains a unique identifier and is linked to the hash value of the previous version, forming an ordered version control chain to ensure that the update of each version can be traced and verified.
[0064] To ensure that each newly released model version has been strictly optimized and updated, the smart contract will evaluate each model update according to comprehensive optimization goals. Specifically, the optimization goals include multiple factors such as training loss, validation loss, model stability, and node contribution. The optimization process of the model version can be described by the following comprehensive loss function: , where, are the final model parameters obtained after training; are the model parameters of the previous version, used to calculate version consistency; and are the training dataset and the test dataset respectively, used to train and validate the model; is the loss function of the i-th computing node on the training set; is the loss function of the j-th computing node on the test set; N and M represent the total number of computing nodes participating in the training and the total number of computing nodes used for validation respectively; is the weight factor used to adjust each loss; and are the weights used to weight the training loss and the validation loss respectively; is a small constant used to ensure the stability of the calculation; is the Hessian matrix of the i-th node, used to capture the change of the gradient; represents the gradient of the i-th node during the training process, measuring the direction and magnitude of the model parameter update; is the second-order gradient of the i-th node, measuring the curvature of the loss function and optimizing the gradient update during the model training process; Denotes the trace of the matrix, calculates the sum of the diagonal elements of the matrix, and is used to adjust the pace of model updates to ensure the stability of the training process. This loss function comprehensively considers multiple factors such as training error, validation error, model stability, and node contributions. By adjusting the weight factors, the smart contract can flexibly balance different optimization objectives to ensure that each new version of the model is rigorously optimized and meets various performance requirements.
[0065] After the model is released, all model updates and version information will be recorded through the blockchain. After each model update, relevant information such as version number, training loss, validation loss, and model parameters will be permanently stored through blockchain transactions to ensure the transparency and immutability of the update process. The decentralized feature of the blockchain ensures the immutability of model versions, meaning that even the model developers cannot modify the released versions. The release of each version is linked to the blockchain blocks, enabling subsequent version updates to be traced back to the previous version and ensuring the transparency of each update.
[0066] Embodiment 2
[0067] As Figure 2 shown, Embodiment 2 of the present application provides an artificial intelligence model training system based on the blockchain, including:
[0068] A data acquisition module 21, configured to automatically acquire and encrypt training data through a smart contract, ensure the privacy and security of the data, and perform tamper-proof storage through the blockchain;
[0069] The data acquisition module is responsible for automatically acquiring and encrypting training data through a smart contract to ensure the privacy and security of the data and perform tamper-proof storage through the blockchain. This module aims to obtain the required training data from multiple data sources (such as sensors, external databases, etc.) in an automated manner and encrypt the data to prevent it from being stolen or tampered with during transmission. The main functions and workflow of the data acquisition module are as follows:
[0070] Collect training data from multiple data sources. These data sources can be sensors, external APIs, historical data storage, real-time data streams, etc. The module automatically selects data sources and collects data according to the rules defined in the smart contract to ensure the integrity and timeliness of the data.
[0071] Before the collected data is encrypted, the data acquisition module will perform preliminary preprocessing on the data. These preprocessing steps include but are not limited to data cleaning (removing duplicate or invalid data), data formatting (unifying the data structure), data normalization (standardizing or normalizing numerical data), etc.
[0072] All the collected training data is encrypted by an encryption algorithm to ensure the privacy and security of the data during transmission. The encryption method can be symmetric encryption or asymmetric encryption, and the specific algorithm is selected according to actual requirements. The encrypted data cannot be accessed or tampered with by unauthorized entities.
[0073] The data collection module uploads the encrypted data to the blockchain through a smart contract. During the data upload process, the data is stored on the blockchain in an encrypted form to ensure the immutability of the data. Each data upload generates a block, recording the hash value of the data, the upload time, and other relevant information. Through the decentralized feature of the blockchain, the persistence and transparency of the data are ensured.
[0074] Before uploading the data to the blockchain, the data collection module manages the data access permissions through a smart contract. The smart contract determines which nodes or users can access specific data sets based on the identity and permission rules of the nodes. Unauthorized computing nodes will not be able to access sensitive data, ensuring the privacy of the data.
[0075] Each time data is collected and uploaded, the data collection module generates the hash value of the data and compares it with the original data to ensure that no changes or damages have occurred to the data during transmission.
[0076] The data collection module also has a real-time monitoring function, which can track each step of the data collection process and record the detailed logs of each data collection, processing, and upload. The smart contract manages these logs and provides an audit function for data access to ensure the transparency of the entire process.
[0077] The smart contract module 22 is used to automatically verify and determine the data access permissions according to the rules of the blockchain smart contract, ensuring that only authorized computing nodes can access the training data;
[0078] The smart contract module is responsible for automatically verifying and determining the data access permissions to ensure that only authorized computing nodes can access the training data. The main task of this module is to verify the identity and permissions of computing nodes through a smart contract based on pre-set rules and decide whether to allow a node to access specific training data. The main functions and work processes of the smart contract module are as follows:
[0079] The smart contract module manages the access permissions of computing nodes to ensure that each node follows the pre-determined rules when accessing the training data. The smart contract automatically assigns access permissions based on information such as the identity, task, and computing power of each node. The setting of access permissions can be based on factors such as the historical reputation score of the node, the node type, and the urgency of the task.
[0080] At each data access request, the smart contract module first authenticates the computing node of the request. The verification process includes checking whether the identity information of the node meets the predetermined conditions. The identity information usually includes the public key, signature, historical behavior records, etc. of the node, ensuring that only legal and authorized nodes can initiate data access requests. Once the node identity is verified, the smart contract module decides whether the node has the right to access the data according to the pre-set permission rules. These permission rules include: the historical performance of the node (e.g., reputation score, accuracy of task execution, etc.); the frequency and type of data access by the node (some nodes may only be allowed to access specific types of data); the computing power and task priority of the node (nodes with stronger computing power or higher task priority may be given priority to obtain data access rights).
[0081] After the permission verification and decision are completed, the smart contract module will perform corresponding operations according to the permission rules. If the node has the right to access the data, the smart contract will automatically generate an access permission credential and allow the node to access the corresponding training data; if the node fails the verification or does not have sufficient permissions, the smart contract will reject its data access request. In addition, all access operations and permission verification processes will be recorded on the blockchain to ensure the transparency and traceability of each access.
[0082] The smart contract module records the detailed information of each data access in real time through the blockchain, including the access time, access node, data type accessed, and operations performed. This information will be used for subsequent auditing and tracing to ensure that each data access can be verified and traced. Through the immutability of the blockchain, all access records can prevent malicious tampering and data leakage.
[0083] The smart contract module supports the dynamic permission adjustment function, which adjusts the access permissions in real time according to the historical behavior of the node and system requirements. For example, if the performance of a certain node does not meet expectations, the smart contract can automatically reduce the access permissions of the node, or in an emergency, increase the permissions of some nodes to ensure the smooth progress of the training task; the smart contract module also supports multi-level access control, allowing computing nodes at different levels to access data with different sensitivities. For example, some nodes may only be able to access low-sensitivity data, while nodes that require higher permissions can access critical data sets. Multi-level access control ensures fine-grained protection of data.
[0084] The computing resource scheduling module 23 is used to transmit the verified data to the distributed computing nodes through the blockchain network, utilize the decentralized computing resources for the training of the artificial intelligence model, and record the intermediate results in real time during the training process through the blockchain;
[0085] The computing resource scheduling module is responsible for transmitting the data verified by the smart contract to the distributed computing nodes for artificial intelligence model training. The core task of this module is to intelligently schedule data based on factors such as the load, computing power, and task priority of the computing nodes, ensuring the optimal use of distributed computing resources and improving the efficiency of the training process. The main functions and working processes of the computing resource scheduling module are as follows:
[0086] After receiving the data verified by the smart contract module, the computing resource scheduling module is responsible for transmitting this data to the appropriate distributed computing nodes. The scheduling module selects a suitable node for data allocation according to the current state, computing power, and task type of the node. This module ensures that the verified data can reach the appropriate computing nodes in a timely manner, avoiding overloading between nodes or idling of computing resources.
[0087] The computing resource scheduling module continuously monitors the load conditions of each computing node to ensure that data and computing tasks can be evenly distributed among the nodes. This module dynamically adjusts the data allocation strategy according to factors such as real-time computing load, CPU and memory usage of the nodes, and network bandwidth. If the computing load of some nodes is high or the resources are approaching saturation, the scheduling module will avoid allocating new data or training tasks to these nodes and instead select idle or low-load nodes.
[0088] According to the computing power, available resources of the nodes, and requirements of the training tasks, the computing resource scheduling module uses intelligent scheduling algorithms (such as load balancing algorithms, priority scheduling, etc.) for data allocation. By optimizing the use of computing resources, the scheduling module can accelerate the training process, avoid resource waste, and improve the overall efficiency of the system. The scheduling strategy also evaluates the nodes based on their historical performance (such as task completion status, computing efficiency), so as to dynamically select the most suitable nodes to execute computing tasks.
[0089] The computing resource scheduling module automatically adjusts the computing resource allocation according to the priority of the training tasks. Some high-priority tasks may need to be processed immediately, while low-priority tasks can be executed later. Task priority is usually evaluated based on factors such as the urgency of the task, the complexity of the training goal, and resource requirements. The scheduling module ensures that critical tasks are given priority by adjusting the execution order of different tasks, thereby enhancing the overall efficiency of the training process.
[0090] The computing resource scheduling module is not only responsible for data transmission but also ensures the synchronization of task scheduling and data transmission. During the training process, data is transmitted from the blockchain to the computing nodes and is synchronized with the computing tasks. The scheduling module ensures that each node can quickly start the training task and return the training results to the system through interaction with the training task management system.
[0091] To cope with training tasks of different scales and sudden computing demands, the computing resource scheduling module can automatically expand computing resources. When the system detects that the computing demand of a training task exceeds the current resource capacity, the scheduling module can automatically call more computing nodes for resource expansion to ensure that the training task can be completed in a timely manner. In addition, the module also has adaptive scheduling capabilities and can dynamically adjust the resource allocation strategy according to changes in computing load and task priority adjustments to adapt to changes during the training process.
[0092] The computing resource scheduling module continuously monitors the training progress of each node and provides real-time feedback. Through real-time monitoring of computing nodes, the scheduling module can detect potential training bottlenecks or performance degradation problems. Once abnormal node performance or task delays are detected, the module will reallocate computing tasks or adjust data transmission strategies to ensure the smooth progress of the training process.
[0093] The blockchain verification module 24 is used to monitor and correct abnormal behaviors during the training process through the consensus mechanism of blockchain nodes, and optimize training parameters to ensure the stability and efficiency of model training;
[0094] The main function of the blockchain verification module is to record the intermediate results during the training process in real time and ensure the immutability of data. Through the combination of smart contracts and blockchain, this module ensures that the intermediate calculation results (such as gradients, loss function values, model parameters, etc.) during each training process can be encrypted and stored through the blockchain, and prevent any data tampering behavior from occurring, thereby ensuring the transparency of the training process and the credibility of data. The main functions and working processes of the blockchain verification module are as follows:
[0095] The blockchain verification module will collect and record the intermediate results calculated by each computing node in real time during each training process. These intermediate results include but are not limited to model parameters, gradients, loss function values, etc. During each calculation update, the module records these intermediate results in encrypted form on the blockchain through smart contracts to ensure data integrity and confidentiality. During the training process, all intermediate results will be encrypted using encryption algorithms to ensure that only authorized entities can read this data. Then, a unique identifier (hash value) of this data is generated through the hash algorithm, and the hash value and related data are stored on the blockchain. Each piece of data stored is verified by the smart contract to ensure the immutability of the data.
[0096] The decentralized and immutable characteristics of blockchain technology make it impossible to modify or delete the data stored each time after recording. Each intermediate result is connected to the previous block on the blockchain through a hash value, forming a complete chain. Even if a node attempts to modify the intermediate result of a certain calculation, the tampering behavior of this node will be immediately detected because the modified data will cause the hash value on the blockchain to not match, resulting in a verification failure.
[0097] The blockchain verification module is not only responsible for storing intermediate results, but also allows all participants (including computing nodes and system administrators) to view the intermediate results stored on the blockchain. Through publicly transparent data records, all steps, calculation results, and node performances during the training process can be audited and verified. This enables the system to effectively prevent any malicious behavior and data tampering, enhancing the credibility of the system.
[0098] The blockchain verification module ensures the correctness and consistency of the calculation results of all computing nodes through the consensus mechanism of the blockchain. The calculation results submitted by each node will be verified through the consensus mechanism of the blockchain, and only the calculation results that have passed the consensus can be recorded on the blockchain. If the calculation result of a certain node is significantly different from that of other nodes, the system will automatically mark the result of this node as abnormal and perform correction or recalculation to avoid the wrong result affecting the global model.
[0099] The blockchain verification module also has the function of detecting abnormal behaviors. During the training process, if the calculation result of a certain node significantly deviates from the calculation results of other nodes, the module will mark it as abnormal through a smart contract and generate relevant reports. All abnormal behaviors and their handling processes will be recorded on the blockchain to ensure the transparency of the training process and provide a basis for subsequent auditing and analysis.
[0100] The blockchain verification module creates a new block for each training step and stores the intermediate results and their version information on the blockchain. Through smart contract management, the system can achieve version control of the training process. Each training update and parameter adjustment will generate a new version to ensure that the historical records of all versions can be traced and verified.
[0101] The model release module 25 is used to store the final model and its related information through the blockchain after the training is completed, and release the model version through a smart contract to ensure the traceability and transparency of model updates.
[0102] The model release module is responsible for storing the final trained model through the blockchain and releasing version information. This module ensures that each trained final model can be effectively stored and releases new model versions through smart contracts. By leveraging the immutability feature of the blockchain, the model release module ensures that each version of the model has a complete historical record that can be traced and verified, thus guaranteeing the transparency, traceability, and credibility of model updates. The main functions and workflow of the model release module are as follows:
[0103] The model release module encrypts the final model parameters (including weights, biases, training hyperparameters, etc.) after training and stores this data on the blockchain through a smart contract. Blockchain storage ensures that the model cannot be tampered with during storage and transmission, and only authorized users and computing nodes can read it. The model storage process includes: encrypting the final model parameters; generating a model storage record through a smart contract; using the decentralized storage technology of the blockchain to ensure the security and integrity of each model version.
[0104] After each training is completed, the model release module generates new model version information and publishes it on the blockchain. Each version information includes: version number: a number that uniquely identifies each model version; version timestamp: records the specific time of model release; relevant performance metrics: including training accuracy, validation error, loss function value, etc.; model metadata: such as a summary of the training dataset and the hyperparameter configuration used during training.
[0105] By publishing this information through a smart contract, the system can generate a hash value for each version, ensuring the immutability and transparency of the version record.
[0106] The model release module performs version control on each updated model version through a smart contract. When a new model version is released, the system automatically records its version number and related information and associates it with the previous version. This version control mechanism ensures that the historical update process of the model is fully traceable. Each model release forms an immutable blockchain record that can provide full audit and verification.
[0107] The model release module utilizes the transparency feature of the blockchain to ensure that each released model version can be publicly verified. Any person or system can query the blockchain to verify the version history of the model and its related performance metrics. This makes the model release process completely transparent, eliminates any possible concealment or tampering behavior, and enhances the fairness and credibility of the model.
[0108] The model release module automatically releases each model version through a smart contract. Whenever the training is completed, the smart contract will automatically trigger the release operation, including: encrypting and storing the model parameters; recording the model version and related training information; releasing the model version information and forming an ordered chain with other versions on the blockchain.
[0109] At each model release, the model release module generates a checksum (hash value) for that version and stores it through the blockchain. These hash values will be stored on the blockchain together with information such as the model's training data and hyperparameter configurations. By verifying the hash values, the system can verify the integrity of the stored model, ensuring that the released model is consistent with the intermediate results during the training process and has not been tampered with.
[0110] Through blockchain storage and version management, the model release module provides a solid foundation for the optimization and iteration of subsequent versions. The release of each new version can be traced back to the previous version, facilitating subsequent analysis and improvement. Through this mechanism of continuous update and version management, the system can continuously optimize the model while maintaining the transparency and reliability of each iteration.
[0111] Corresponding to the above embodiments, an embodiment of the present invention provides a computer storage medium, including: at least one memory and at least one processor;
[0112] The memory is used to store one or more program instructions;
[0113] The processor is used to run one or more program instructions to execute an artificial intelligence model training method based on the blockchain.
[0114] Corresponding to the above embodiments, an embodiment of the present invention provides a computer-readable storage medium. The computer storage medium contains one or more program instructions, and the one or more program instructions are used to be executed by a processor to execute an artificial intelligence model training method based on the blockchain.
[0115] An embodiment disclosed by the present invention provides a computer-readable storage medium. Computer program instructions are stored in the computer-readable storage medium. When the computer program instructions run on a computer, the computer is caused to execute the above artificial intelligence model training method based on the blockchain.
[0116] In an embodiment of the present invention, the processor may be an integrated circuit chip with signal processing capabilities. The processor may be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0117] It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present invention may be directly embodied as being executed and completed by a hardware decoding processor, or executed and completed by a combination of hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, register, etc. The processor reads the information in the storage medium and combines its hardware to complete the steps of the above method.
[0118] The storage medium may be a memory, for example, it may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories.
[0119] Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory.
[0120] The volatile memory may be a Random Access Memory (RAM) which serves as an external cache. By way of example but not limitation, many forms of RAM are available, such as Static RAM (SRAM), Dynamic RAM (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced SDRAM (ESDRAM), Synchlink DRAM (SLDRAM), and Direct Rambus RAM (DRRAM).
[0121] The storage media described in the embodiments of the present invention are intended to include, but are not limited to, these and any other suitable types of memory.
[0122] Those skilled in the art should be aware that, in one or more of the above examples, the functions described in the present invention can be implemented by a combination of hardware and software. When applying software, the corresponding functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium. The computer-readable medium includes computer storage media and communication media, where the communication media includes any medium that facilitates the transfer of a computer program from one place to another. The storage media can be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0123] The above specific embodiments have further elaborated on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for training an artificial intelligence model based on blockchain, characterized in that, It includes the following steps: Automatically collect and encrypt training data through a smart contract to ensure the privacy and security of the data, and store it immutably through the blockchain; Automatically verify and determine the access rights of the data according to the rules of the blockchain smart contract to ensure that only authorized computing nodes can access the training data; When a computing node needs to access the training data, the smart contract authenticates the node requesting access; Each computing node has a unique identifier on the blockchain. The smart contract matches the node information stored in the blockchain with the current requested node identity to confirm whether the node is a legitimate computing node. If the verification fails, the node's access request is rejected; Once the authentication is successful, the smart contract continues to check the access rights of the node; the access rights for each dataset are preset according to the sensitivity of the data, the role of the computing node, and its historical reputation; the smart contract dynamically calculates the access rights of the node based on the node's historical behavior and its reputation score; Reputation score formula: , where is the final reputation score of the computing node, reflecting the overall reputation of the node and determining whether the node can access a specific dataset; is related to the th behavior score associated weight coefficient; is the reputation score of the node in the th past behavior; is the number of all historical behaviors of the node in the system, representing the total number of node behaviors; is the adjustment factor; is the kth score in the node's historical behaviors; m is the number of historical scores; is the access frequency metric of the node, reflecting how frequently the node requests access to data; is the frequency adjustment factor used to adjust the final reputation score according to the node's access frequency; is the metric of the maximum access duration of the node in the system, used to measure the degree of consumption of system resources by the node; is the average access duration of the node, representing the time consumption each time the node accesses data; is the duration adjustment factor used to adjust the impact of the node's consumption of system resources; the higher the reputation score, the greater the likelihood that the node is authorized to access data. Transmit the verified data through the blockchain network to the distributed computing nodes, use the decentralized computing resources to train the artificial intelligence model, and record the intermediate results during the training process in real time through the blockchain; Monitor and correct abnormal behaviors during the training process through the consensus mechanism of the blockchain nodes, and optimize the training parameters to ensure the stability and efficiency of model training; After the training is completed, store the final model and its related information through the blockchain, and release the model version through the smart contract to ensure the traceability and transparency of model updates.
2. The artificial intelligence model training method based on blockchain according to claim 1, wherein Collect and encrypt training data, specifically including automatically obtaining raw data from multiple data sources, encrypting the raw data through an encryption algorithm, and storing it in the blockchain to ensure the privacy and security of the data; the multiple data sources include data collected by sensors, real-time monitoring data, user input data, images, videos, or data from other sensor networks.
3. The artificial intelligence model training method based on blockchain according to claim 1, wherein, Automatically verify and determine the access rights of the data according to the rules of the blockchain smart contract, specifically by authenticating the computing node according to the preset access control model and automatically confirming whether the node has the authorization to read the training data. If the node fails the verification, its access request is rejected.
4. The method according to claim 1, wherein Use decentralized computing resources to train the artificial intelligence model, specifically by executing the model training task through decentralized computing resources and recording the generated intermediate results in real time through the blockchain during the training process.
5. The artificial intelligence model training method based on blockchain according to claim 1, wherein, When monitoring abnormal behaviors during the training process through the consensus mechanism of the blockchain nodes, it further includes comparing the error between the calculation result of each node and the global model during the training process, and dynamically adjusting the influence of the node on the global training result based on the error value.
6. The method for training an artificial intelligence model based on a blockchain according to claim 1, wherein, Optimize the training parameters, specifically including dynamically adjusting the learning rate of the computing node to adapt to the data volume, computing load, and historical training performance of different nodes.
7. The artificial intelligence model training method based on blockchain according to claim 1, wherein, Release model versions through smart contracts, specifically including recording and publishing the model parameters, performance metrics, and summary information of the training data after training through smart contracts to the blockchain, ensuring that the release and update of each version have a clear timestamp and can be traced back to the source.
8. An artificial intelligence model training system based on blockchain, characterized in that, Including: A data collection module for automatically collecting and encrypting training data through smart contracts, ensuring the privacy and security of the data, and storing it immutably through the blockchain; A smart contract module for automatically verifying and determining the access rights of data according to the rules of blockchain smart contracts, ensuring that only authorized computing nodes can access the training data; When a computing node needs to access the training data, the smart contract authenticates the node requesting access; Each computing node has a unique identifier on the blockchain. The smart contract matches the node information stored in the blockchain with the identity of the currently requested node to confirm whether the node is a legitimate computing node. If the verification fails, the access request of the node is rejected; Once the authentication is successful, the smart contract continues to check the access rights of the node; the access rights of each dataset are preset according to the sensitivity of the data, the role of the computing node, and its historical reputation; the smart contract dynamically calculates the access rights of the node based on the historical behavior of the node and its reputation score; Reputation score formula: , where is the final reputation score of the computing node, reflecting the overall reputation of the node and determining whether the node can access a specific data set; is related to the th behavior score associated weight coefficient; is the reputation score of the node in the th behavior in the past; is the number of all historical behaviors of the node in the system, representing the total number of node behaviors; is the adjustment factor; is the kth score in the node's historical behaviors; m is the number of historical scores; is the access frequency metric of the node, reflecting the frequency of the node's requests to access data; is the frequency adjustment factor, used to adjust the final reputation score according to the access frequency of the node; is the metric of the maximum access duration of the node in the system, used to measure the degree of consumption of system resources by the node; is the average access duration of the node, indicating the time consumption of the node each time it accesses data; is the duration adjustment factor, used to adjust the impact of the node's consumption of system resources; the higher the reputation score, the greater the likelihood that the node will be authorized to access data; A computing resource scheduling module for transmitting the verified data through the blockchain network to distributed computing nodes, training artificial intelligence models using decentralized computing resources, and recording the intermediate results during the training process in real time through the blockchain; A blockchain verification module for monitoring and correcting abnormal behaviors during the training process through the consensus mechanism of blockchain nodes, and optimizing the training parameters to ensure the stability and efficiency of model training; A model release module for storing the final model and its related information through the blockchain after training, and releasing model versions through smart contracts to ensure the traceability and transparency of model updates.
Citation Information
Patent Citations
Federal learning defecation vehicle node detection method based on block chain information interaction
CN118520341A