Data tracing and tracking method for block chain based on artificial intelligence technology
By introducing artificial intelligence methods into blockchain technology, data is collected, stored, processed and transmitted securely, the security and accuracy problems of blockchain data traceability and tracking are solved, and efficient and secure data management and risk detection are achieved.
Patent Information
- Application Number
- CN202510446001.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-25
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the prior art, blockchain data traceability and tracking capabilities have limitations, affecting the security and accuracy of data access and transmission.
Data is collected through the blockchain node API interface and distributed crawling technology, combined with hash checksum digital signature verification, and large-scale data storage and parallel processing are used to use MongoDB and MapReduce to identify transaction patterns. Neo4j traces the data path and realizes asynchronous communication through RESTful API and Kafka. OAuth2.0 and SSL/TLS are used to ensure data security.
It improves the accuracy and efficiency of blockchain data traceability and tracking, ensures the security of data access and transmission, supports real-time monitoring and risk warning by financial institutions and regulatory agencies, and improves the data management efficiency and security of blockchain technology.
Smart Images

Figure CN120372584A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of data science and blockchain technology, and specifically to a method for data traceability and tracking of blockchain based on artificial intelligence technology. Background Art
[0002] With the continuous advancement of the global digitalization process, blockchain, as a decentralized, highly secure, and transparent distributed ledger technology, is gradually integrating into multiple industries such as finance, supply chain management, healthcare, Internet of Things, and energy, demonstrating strong application potential. In terms of technological innovation, blockchain technology is constantly optimizing performance, enhancing security and scalability to meet the needs of different application scenarios. At the same time, the interoperability between different blockchains is also strengthening, promoting the formation of a unified and efficient blockchain ecosystem. In terms of the policy environment, many governments attach great importance to the development of blockchain technology and have introduced a series of encouraging policies, providing a good external environment for the innovation and application of blockchain. In terms of industrial applications, the application of blockchain in the financial field has taken initial shape and is gradually expanding to other industries, promoting industrial upgrading and digital transformation. In terms of international cooperation, the cross-border application of blockchain technology is increasing, and international cooperation in the blockchain field is strengthening, which helps to promote the formulation of technical standards and the healthy development of the global blockchain industry. In summary, blockchain technology is expected to play an important role in more fields in the future, bringing profound changes to economic and social development.
[0003] However, in the existing technology, there are certain limitations in the data traceability and tracking capabilities of blockchain, which affect the security of data during access and transmission. Therefore, how to improve the accuracy of data traceability and tracking and ensure the security of data access and transmission is the problem we need to solve. For this reason, a method for data traceability and tracking of blockchain based on artificial intelligence technology is proposed. Summary of the Invention
[0004] In view of the deficiencies of the existing technology, the present invention provides a method for data traceability and tracking of blockchain based on artificial intelligence technology, which solves the problems in the above background art.
[0005] To achieve the above objectives, the present invention is realized through the following technical solutions: A method for data traceability and tracking of blockchain based on artificial intelligence technology, which consists of the following steps:
[0006] S1. Collect transaction list data and smart contract event log data through the blockchain node API interface and distributed crawler technology;
[0007] S2. Prevent malicious data injection through hash verification and digital signature verification;
[0008] S3. Introduce the MongoDB document database and the MapReduce distributed computing framework to achieve the storage of large-scale data sets and accelerate parallel processing tasks under a large amount of data;
[0009] S4. Identify transaction patterns and detect abnormal behaviors through the gradient boosting decision tree algorithm and clustering analysis;
[0010] S5. Trace the data paths of transaction list data and smart contract event log data through the graph algorithms in the Neo4j graph database;
[0011] S6. Use the RESTful API as the service endpoint and the message queue system Kafka to achieve asynchronous communication and client access;
[0012] S7. Implement security control during data access and transmission through the OAuth2.0 authorization protocol and the SSL / TLS encrypted transport layer.
[0013] Furthermore, in the above-mentioned S1, the process of collecting transaction list data and smart contract event log data through the blockchain node API interface and the distributed crawler technology includes:
[0014] Use the API library to set API access permissions and interact with the blockchain nodes on the blockchain platform, deploy the API server to the local server, implement the OAuth2.0 security protocol for user authentication, set the API request rate limit, enable logging and real-time monitoring;
[0015] Obtain transaction list data and smart contract event log data through API calls, provide interfaces for querying historical transaction records and event logs, simultaneously crawl data from different blockchain nodes, allocate tasks through the load balancing mechanism to evenly distribute the workload of each node, and adopt a fault tolerance mechanism to achieve automatic retry and error recovery functions;
[0016] Check and remove duplicate transaction list data and smart contract event log data, unify data in different formats into a standardized format, adjust the time stamps to make the time references of the collected data consistent, and for fields with missing values, use mean filling;
[0017] Distinguish different transaction types, where the transaction types include contract deployment and contract call, record the sender address, receiver address participating in the transaction, and the contract address when involving smart contracts, construct time series features based on the time stamp of each transaction, extract the transaction ID, and record the block height where the transaction is located.
[0018] Furthermore, in the above-mentioned S2, the process of preventing malicious data injection through hash verification includes:
[0019] When each block is generated, corresponding hash values are calculated for the transaction list data and the smart contract event log data. When retrieving data from a blockchain node, the hash value of the retrieved data is recalculated and compared with the hash value recorded in the block header. If the two calculated hash values match, it indicates that the data has not been tampered with. If the two calculated hash values do not match, the data has been modified, and it is rejected and marked as suspicious.
[0020] Further, in S2, the process of preventing malicious data injection through the digital signature verification mechanism includes:
[0021] For each transaction and smart contract operation, the sender uses a private key to digitally sign the transaction data. The API server verifies the validity of the digital signature using the sender's public key. If the digital signature verification passes, it proves that the transaction was initiated by a legitimate user and has not been tampered with. If the verification fails, it indicates a potential malicious behavior, and the transaction is rejected and a security alert is triggered.
[0022] Further, in S3, the process of introducing the MongoDB document database and the MapReduce distributed computing framework to achieve large-scale dataset storage and accelerate parallel processing tasks under a large amount of data includes:
[0023] The MongoDB document database is introduced to store and standardize the transaction list data and smart contract event log data retrieved from the blockchain node. Through the indexing mechanism of MongoDB, historical transaction records and event logs are queried and retrieved, and data backup is performed using the replication set and sharding techniques of MongoDB;
[0024] The process of batch data processing and in-depth analysis based on the MapReduce distributed computing framework includes a mapping stage and a reduction stage. In the mapping stage, the input transaction list data and smart contract event log data are split into small pieces and distributed to different computing nodes for independent processing. The sender address, recipient address, and contract address are recorded, and the hash value is verified. In the reduction stage, the intermediate results generated in the mapping stage are collected, reorganized in the form of key-value pairs, and passed to the reduction function for summarization and merging to obtain the total transaction volume statistics, frequent interaction analysis, and abnormal behavior detection analysis results.
[0025] Further, in S4, the process of identifying transaction patterns through the gradient boosting decision tree algorithm includes:
[0026] Based on the transaction list data and smart contract event log data, key features are extracted as the training set. The key features include transaction type, sender address, recipient address, contract address, timestamp of each transaction, and block height, and a feature set is comprehensively obtained;
[0027] Introduce a constant model to make initial predictions for all samples and obtain initial prediction values. The initial prediction values are the average values of the target variables in the training set. For each training sample, calculate the residual between its true value and the current model's prediction value. Based on the calculated residuals, train a decision tree to minimize the residuals and fit the residual values. Add the trained decision tree to the existing model, and sum the initial prediction value and the decision tree's prediction value for the residuals to calculate the new prediction value, and multiply by the learning rate to control the impact degree of each update;
[0028] Repeat the steps from calculating the residuals to updating the model until the model performance no longer improves. Each iteration gradually improves the overall model by introducing new decision trees, obtaining a GBDT ensemble model composed of different decision trees, and integrating the prediction results of different decision trees to achieve the recognition of trading patterns.
[0029] Further, in S4, the process of detecting abnormal behaviors through cluster analysis includes:
[0030] Convert the key features in the feature set into numerical feature vectors. Define the maximum distance within the cluster for samples and the parameter of the minimum number of samples required for the cluster according to the data characteristics and requirements. Starting from an unvisited data point, search for neighboring points within the maximum distance within the cluster for samples centered on this data point. If the number of neighboring points reaches the minimum number of samples required for the cluster, a new cluster is formed. If the number of neighboring points does not reach the minimum number of samples required for the cluster, this point is regarded as a noise point;
[0031] For each new cluster, continue to search for neighboring points within the maximum distance within the cluster for samples of its member points and add the found neighboring points to the cluster. Repeat the process of expanding the cluster until no new neighboring points are added. After completing the clustering, mark the data points that are not assigned to any cluster as abnormal points, representing abnormal behaviors, and regard each cluster in the obtained set of clusters as a trading pattern.
[0032] Further, in S5, the process of tracing the data paths of the transaction list data and the smart contract event log data through the graph algorithms in the Neo4j graph database includes:
[0033] Import the numerical feature vectors converted from the transaction list data and the smart contract event log data into the Neo4j graph database. In the Neo4j graph database, regard transactions and smart contract events as nodes, and represent the interaction relationships between nodes as edges. The interaction relationships include transfer behaviors between the sender address and the receiver address, contract deployment, and contract invocation. Attach the timestamp, block height, and transaction ID attributes to the corresponding nodes and edges to record the time series features and location information of each transaction;
[0034] Using the graph algorithm library provided by Neo4j, perform shortest path search, connected component analysis, and PageRank calculation operations. When tracing the fund flow path of a transaction, define the starting transaction node and the final receiving transaction node, apply the shortest path algorithm to determine all paths between the two, and use connected component analysis to identify closed transaction loops to detect potential loop structures;
[0035] Analyze the path of smart contract event log data by constructing a call relationship network between smart contract addresses. When a contract call event occurs, a directed edge is established between the contract address nodes accordingly, thereby forming a directed graph. Use graph algorithms to trace the actual execution path from the initial contract deployment to subsequent calls, and identify the external accounts that trigger the contract calls and the mutual influence relationships between each contract call.
[0036] Further, in S6, the process of using RESTful API as a service endpoint and implementing asynchronous communication and client access through the message queue system Kafka includes:
[0037] Deploy a RESTful API interface as a service endpoint. This API interface is used to receive HTTP requests from the client. The HTTP requests include querying historical transaction records and obtaining smart contract event operations. After the API server processes the received requests, it interacts with the local MongoDB document database, retrieves data from it, and returns the results to the client in JSON format;
[0038] Introduce the message queue system Kafka for asynchronous communication. When new transactions and smart contract events occur, the API server stores the generated transaction list data and smart contract event log data in the local MongoDB document database B, and serializes them as messages and publishes them to Kafka topics. Each topic corresponds to different types of transaction confirmation, contract deployment, and call event streams;
[0039] The application program of the client receives real-time updates by subscribing to the topics of interest. The Kafka consumer pulls messages from the topics and processes them according to the business logic. The background worker process consumes the corresponding messages, executes MapReduce task allocation and graph algorithm calculation operations, saves the processing results back to the local MongoDB document database, and during network failures and system maintenance, temporarily stores the messages that have not been processed in time in Kafka and continues to transmit them after the connection is restored.
[0040] Further, in S7, the process of implementing security control during data access and transmission through the OAuth2.0 authorization protocol and the SSL / TLS encrypted transport layer includes:
[0041] The OAuth2.0 authorization protocol is adopted to manage the access rights of API endpoints. The client application initiates an authentication request to the OAuth2.0 authorization server, providing the client ID and secret key. The authorization server verifies the client's identity and sends a temporary authorization code to the client. The client uses the temporary authorization code to exchange for an access token and carries the access token as an HTTP header to request the RESTful API interface;
[0042] The API server receives the request containing the access token, internally verifies the validity, scope, and expiration status of the access token, authorizes the client to execute the API call. For API requests involving sensitive operations, the access token is refreshed to extend the validity period of the access rights, and the access token is rotated regularly;
[0043] At the network communication level, the SSL / TLS encryption transport layer protocol is implemented to encrypt the data transmission between the client and the API server, as well as between the API server and the MongoDB document database and the Kafka message queue system. The SSL / TLS protocol combines symmetric encryption algorithms and asymmetric encryption algorithms. At the initial stage of establishing a connection, the handshake protocol is used to negotiate and share the secret keys of both parties, and the secret keys are used to encrypt and decrypt data during data transmission, and the digital certificates are used to verify the true identities of the communicating parties.
[0044] The present invention provides a method for data traceability and tracking of blockchain based on artificial intelligence technology. It has the following beneficial effects:
[0045] Through the blockchain node API interface and distributed crawler technology in step S1, this method can simultaneously capture transaction list data and smart contract event log data from multiple nodes, ensuring data integrity and accuracy. Multi-source parallel acquisition not only speeds up data acquisition but also enhances the coverage of the entire blockchain network. In S2, the hash verification and digital signature verification mechanisms prevent malicious data injection. The hash value is recalculated each time data is obtained and compared with the block header record to ensure that the data has not been tampered with. The sender uses a private key for signature, and the API server uses a public key for verification to confirm the authenticity and legality of the transaction, enhancing system security. S3 introduces the MongoDB document database and the MapReduce distributed computing framework, optimizing large-scale data storage and processing performance. MongoDB provides efficient unstructured data storage, and MapReduce accelerates parallel processing tasks under large data volumes, shortening processing time and resource consumption, especially suitable for complex scenarios such as time series feature construction and frequent interaction analysis. S4 applies the GBDT algorithm and clustering analysis to achieve accurate transaction pattern recognition and abnormal behavior detection. By deeply mining key features, not only can normal transaction patterns be discovered, but potential risks or non-compliant operations can also be timely warned, providing strong support for financial institutions and regulatory agencies. S5 utilizes the Neo4j graph database and its graph algorithm library to trace the complex association paths between transaction list data and smart contract event logs. Intuitively presenting the fund flow and contract call link has important value for anti-money laundering investigations, contract audits, etc., helping users understand complex transaction relationships. S6 deploys a RESTful API as a service endpoint and combines it with a Kafka message queue system to build an efficient and reliable asynchronous communication platform. The client application receives real-time updates, and the background worker process seamlessly processes messages, improving the system's response speed and message delivery consistency, suitable for application scenarios that require instant feedback. In stage S7, through the OAuth2.0 authorization protocol and the SSL / TLS encrypted transport layer, user permissions are strictly managed and data security is protected. Only authenticated legitimate users can access specific resources, and all communications are encrypted with high strength to avoid the risk of unauthorized information leakage or tampering. In summary, this method comprehensively applies advanced AI algorithms, distributed computing, and graph analysis, greatly improving the efficiency, accuracy, and security of blockchain data management and promoting the development of blockchain technology to a higher level. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 It is a flowchart of a method for data traceability and tracking of a blockchain based on artificial intelligence technology according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0047] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0048] As Figure 1 shown, the present invention provides a technical solution: a method for data traceability and tracking of a blockchain based on artificial intelligence technology, which consists of the following steps:
[0049] S1. Collect transaction list data and smart contract event log data through the blockchain node API interface and distributed crawler technology;
[0050] S2. Prevent malicious data injection through hash verification and digital signature verification;
[0051] S3. Introduce the MongoDB document database and the MapReduce distributed computing framework to achieve large-scale dataset storage and accelerate parallel processing tasks under a large amount of data;
[0052] S4. Identify transaction patterns and detect abnormal behaviors through the gradient boosting decision tree algorithm and clustering analysis;
[0053] S5. Trace the paths of transaction list data and smart contract event log data through the graph algorithms in the graph database Neo4j;
[0054] S6. Use the RESTful API as a service endpoint and the message queue system Kafka to achieve asynchronous communication and client access;
[0055] S7. Implement security control during the data access and transmission process through the OAuth2.0 authorization protocol and the SSL / TLS encrypted transport layer.
[0056] In the above S1, the process of collecting transaction list data and smart contract event log data through the blockchain node API interface and distributed crawler technology includes:
[0057] Use the API library to set API access permissions and interact with blockchain nodes on the blockchain platform, deploy the API server to the local server, implement the OAuth2.0 security protocol for user authentication, set the API request rate limit, and enable logging and real-time monitoring;
[0058] Obtain transaction list data and smart contract event log data through API calls, provide interfaces for querying historical transaction records and event logs, simultaneously capture data from different blockchain nodes, allocate tasks through a load balancing mechanism to evenly distribute the workload of each node, and adopt a fault tolerance mechanism: implement automatic retry and error recovery functions;
[0059] Check and remove duplicate transaction list data and smart contract event log data, unify data in different formats into a standardized format, adjust the timestamp to make the time reference for the collected data consistent, and for fields with missing values, fill them with the mean value;
[0060] Distinguish different transaction types, where the transaction types include contract deployment and contract call, record the sender address, recipient address participating in the transaction, and the contract address when involving a smart contract, construct time series features based on the timestamp of each transaction, extract the transaction ID, and record the block height where the transaction is located.
[0061] In step S2, the process of preventing malicious data injection through hash verification includes:
[0062] When each block is generated, calculate the corresponding hash value for the transaction list data and smart contract event log data. When obtaining data from the blockchain node, recalculate the hash value of the obtained data and compare it with the hash value recorded in the block header. If the two calculated hash values match, it indicates that the data has not been tampered with. If the two calculated hash values do not match, the data has been modified, reject the reception and mark it as suspicious.
[0063] In step S2, the process of preventing malicious data injection through a digital signature verification mechanism includes:
[0064] For each transaction and smart contract operation, the sender uses the private key to digitally sign the transaction data. The API server verifies the validity of the digital signature using the sender's public key. If the digital signature verification passes, it proves that the transaction is initiated by a legitimate user and has not been tampered with. If the verification fails, it indicates potential malicious behavior, reject processing the transaction and trigger a security alert.
[0065] In step S3, the process of introducing the MongoDB document database and the MapReduce distributed computing framework to achieve large-scale dataset storage and accelerate parallel processing tasks under a large amount of data includes:
[0066] Introduce the MongoDB document database to store and standardize the transaction list data and smart contract event log data obtained from the blockchain node. Through the indexing mechanism of MongoDB, query and retrieve historical transaction records and event logs, and use the replication set and sharding technologies of MongoDB for data backup;
[0067] The process of batch data processing and in-depth analysis based on the MapReduce distributed computing framework includes a mapping stage and a reduction stage. In the mapping stage, the input transaction list data and smart contract event log data are split into small pieces and allocated to different computing nodes for independent processing, recording the sender address, recipient address, and contract address, and verifying the hash value. In the reduction stage, the intermediate results generated in the mapping stage are collected, reorganized in the form of key-value pairs, and passed to the reduction function for summarization and merging to obtain the total transaction volume statistics, frequent interaction analysis, and abnormal behavior detection analysis results.
[0068] In S4, the process of identifying transaction patterns through the gradient boosting decision tree algorithm includes:
[0069] Based on the transaction list data and smart contract event log data, key features are extracted as the training set. The key features include transaction type, sender address, recipient address, contract address, timestamp and block height of each transaction, and a feature set is comprehensively obtained;
[0070] A constant model is introduced to make an initial prediction for all samples to obtain the initial prediction value, which is the average value of the target variable in the training set. For each training sample, the residual between its true value and the current model prediction value is calculated. Based on the calculated residual, a decision tree is trained to minimize the residual and fit the residual value. The trained decision tree is added to the existing model, and the initial prediction value and the prediction value of the decision tree for the residual are summed to calculate the new prediction value, which is multiplied by the learning rate to control the impact degree of each update;
[0071] Repeat the steps from calculating the residual to updating the model until the model performance no longer improves. Each iteration gradually improves the overall model by introducing a new decision tree to obtain a GBDT ensemble model composed of different decision trees, and comprehensively combines the prediction results of different decision trees to achieve the identification of transaction patterns.
[0072] In S4, the process of detecting abnormal behavior through clustering analysis includes:
[0073] The key features in the feature set are converted into numerical feature vectors. According to the data characteristics and requirements, parameters such as the maximum distance within the cluster for samples and the minimum number of samples required for the cluster are defined. Starting from an unvisited data point, neighboring points are searched within the maximum distance range of the samples within the cluster. If the number of neighboring points reaches the minimum number of samples required for the cluster, a new cluster is formed. If the number of neighboring points does not reach the minimum number of samples required for the cluster, the point is regarded as a noise point;
[0074] For each new cluster, continue to search for neighboring points within the maximum distance of the in-cluster samples of its member points, and add the found neighboring points to the cluster. Repeat the process of expanding the cluster until no new neighboring points are added. After completing the clustering, mark the data points that are not assigned to any cluster as outliers, representing abnormal behaviors, and consider each cluster in the resulting set of clusters as a transaction pattern.
[0075] In the step S5, the process of tracing the data paths of the transaction list data and the smart contract event log data through the graph algorithms in the graph database Neo4j includes:
[0076] Import the numerical feature vectors converted from the transaction list data and the smart contract event log data into the Neo4j graph database. In the Neo4j graph database, regard transactions and smart contract events as nodes, and represent the interaction relationships between the nodes as edges. The interaction relationships include the transfer behavior between the sender address and the recipient address, contract deployment, and contract invocation. Attach the timestamp, block height, and transaction ID attributes to the corresponding nodes and edges to record the time series features and location information of each transaction.
[0077] Utilize the graph algorithm library provided by Neo4j to perform shortest path search, connected component analysis, and PageRank calculation operations. When tracing the fund flow path of a transaction, define the starting transaction node and the final receiving transaction node, apply the shortest path algorithm to determine all paths between them, and use connected component analysis to identify closed transaction loops to detect potential loop structures.
[0078] Analyze the path of the smart contract event log data by constructing a call relationship network between smart contract addresses. When a contract invocation event occurs, establish a directed edge between the contract address nodes accordingly, thereby forming a directed graph, and use graph algorithms to trace the actual execution path from the initial contract deployment to subsequent invocations, and identify the external accounts that trigger the contract invocations and the mutual influence relationships between each contract invocation.
[0079] In the step S6, the process of using the RESTful API as the service endpoint and implementing asynchronous communication and client access through the message queue system Kafka includes:
[0080] Deploy the RESTful API interface as the service endpoint. This API interface is used to receive HTTP requests from the client. The HTTP requests include querying historical transaction records and obtaining smart contract event operations. After the API server processes the received requests, it interacts with the local MongoDB document database, retrieves data from it, and returns the results to the client in JSON format.
[0081] The message queue system Kafka is introduced for asynchronous communication. When new transactions and smart contract events are generated, the API server stores the generated transaction list data and smart contract event log data in the local MongoD document database B, and serializes them into messages and publishes them to Kafka topics. Each topic corresponds to different types of transaction confirmation, contract deployment and call event streams.
[0082] Client applications receive real-time updates by subscribing to topics of interest. Kafka consumers pull messages from topics and process them according to business logic. Background worker processes consume corresponding messages, perform MapReduce task allocation and graph algorithm calculation operations, and save the processing results back to the local MongoDB document database. During network failures and system maintenance, messages that are not processed in time are temporarily stored in Kafka and continue to be transmitted after the connection is restored.
[0083] In S7, the process of implementing security control in data access and transmission through the OAuth2.0 authorization protocol and the SSL / TLS encryption transmission layer includes:
[0084] The OAuth2.0 authorization protocol is used to manage API endpoint access rights. The client application initiates an authentication request to the OAuth2.0 authorization server, providing the client ID and key. The authorization server verifies the client's identity and sends a temporary authorization code to the client. The client uses the temporary authorization code to exchange for an access token and carries the access token as an HTTP header to request the RESTful API interface.
[0085] The API server receives a request containing an access token, internally verifies the validity, scope, and expiration of the access token, authorizes the client to execute the API call, and for API requests involving sensitive operations, extends the validity period of access rights by refreshing the access token and rotates the access token regularly;
[0086] The SSL / TLS encrypted transport layer protocol is implemented at the network communication level to encrypt data transmission from the client to the API server and between the API server and the MongoDB document database and Kafka message queue system. The SSL / TLS protocol combines symmetric encryption algorithms with asymmetric encryption algorithms. At the beginning of establishing a connection, the handshake protocol is used to negotiate and share the secret keys of both parties. The secret keys are used for encryption and decryption during data transmission, and the true identity of the communicating parties is verified through digital certificates.
[0087] It should be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device comprising said element.
[0088] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for data traceability and tracking of a blockchain based on artificial intelligence technology, characterized in that, Including the following steps: S1. Collect transaction list data and smart contract event log data through the blockchain node API interface and distributed crawler technology; S2. Prevent malicious data injection through hash verification and digital signature verification; S3. Introduce the MongoDB document database and the MapReduce distributed computing framework to achieve large-scale dataset storage and accelerate parallel processing tasks under large amounts of data; S4. Identify transaction patterns and detect abnormal behaviors through gradient boosting decision tree algorithms and clustering analysis; S5. Trace the paths of transaction list data and smart contract event log data through graph algorithms in the Neo4j graph database; S6. Use the RESTful API as a service endpoint and the Kafka message queue system to achieve asynchronous communication and client access; S7. Achieve security control during data access and transmission through the OAuth2.0 authorization protocol and the SSL / TLS encrypted transport layer.
2. The method for data traceability and tracking of a blockchain based on artificial intelligence technology according to claim 1, wherein: In S1, the process of collecting transaction list data and smart contract event log data through the blockchain node API interface and distributed crawler technology includes: Use the API library to set API access permissions and interact with blockchain nodes on the blockchain platform, deploy the API server to the local server, implement the OAuth2.0 security protocol for user authentication, set API request rate limits, and enable logging and real-time monitoring; Obtain transaction list data and smart contract event log data through API calls, provide interfaces for querying historical transaction records and event logs, simultaneously crawl data from different blockchain nodes, distribute tasks through a load balancing mechanism to evenly distribute the workload of each node, and adopt a fault tolerance mechanism to achieve automatic retry and error recovery functions; Check and remove duplicate transaction list data and smart contract event log data, unify data in different formats into a standardized format, adjust timestamps to make the time references of the collected data consistent, and for fields with missing values, use mean filling; Distinguish different transaction types, where the transaction types include contract deployment and contract call, record the sender address, recipient address participating in the transaction, and the contract address when involving smart contracts, construct time series features based on the timestamp of each transaction, extract the transaction ID, and record the block height where the transaction is located.
3. The method for data traceability and tracking of the blockchain based on artificial intelligence technology according to claim 2, characterized in that: In S2, the process of preventing malicious data injection through hash verification includes: When each block is generated, calculate the corresponding hash values for the transaction list data and smart contract event log data. When obtaining data from the blockchain node, recalculate the hash value of the obtained data and compare it with the hash value recorded in the block header. If the two calculated hash values match, it means the data has not been tampered with. If the two calculated hash values do not match, the data has been modified, reject it and mark it as suspicious.
4. The method for data traceability and tracking of a blockchain based on artificial intelligence technology according to claim 3, wherein: In S2, the process of preventing malicious data injection through the digital signature verification mechanism includes: For each transaction and smart contract operation, the sender uses the private key to digitally sign the transaction data. The API server uses the sender's public key to verify the validity of the digital signature. If the digital signature verification passes, it proves that the transaction is initiated by a legitimate user and has not been tampered with. If the verification fails, it indicates potential malicious behavior, and the transaction is rejected for processing and a security alert is triggered.
5. The method for data traceability and tracking of a blockchain based on artificial intelligence technology according to claim 4, wherein: In S3, the process of introducing the MongoDB document database and the MapReduce distributed computing framework to achieve large-scale dataset storage and accelerate parallel processing tasks under a large amount of data includes: Introduce the MongoDB document database to store and standardize the transaction list data and smart contract event log data obtained from blockchain nodes. Through the indexing mechanism of MongoDB, query and retrieve historical transaction records and event logs, and use the replication set and sharding technologies of MongoDB for data backup; The process of batch data processing and in-depth analysis based on the MapReduce distributed computing framework includes a mapping phase and a reduction phase. In the mapping phase, the input transaction list data and smart contract event log data are split into small pieces and distributed to different computing nodes for independent processing, recording the sender address, recipient address, and contract address, and verifying the hash value. In the reduction phase, collect the intermediate results generated in the mapping phase, reorganize them in the form of key-value pairs and pass them to the reduction function for summarization and merging to obtain the total transaction volume statistics, frequent interaction analysis, and abnormal behavior detection analysis results.
6. The method for data traceability and tracking of a blockchain based on artificial intelligence technology according to claim 5, characterized in that: In S4, the process of identifying transaction patterns through the gradient boosting decision tree algorithm includes: Based on the transaction list data and smart contract event log data, extract key features as the training set. The key features include transaction type, sender address, recipient address, contract address, timestamp and block height of each transaction, and comprehensively obtain the feature set; Introduce a constant model to make an initial prediction for all samples to obtain the initial prediction value. The initial prediction value is the average value of the target variable in the training set. For each training sample, calculate the residual between its true value and the current model prediction value. Based on the calculated residual, train a decision tree to minimize the residual and fit the residual value. Add the trained decision tree to the existing model, and sum the initial prediction value and the prediction value of the decision tree for the residual to calculate the new prediction value, and multiply by the learning rate to control the impact degree of each update; Repeat the steps from calculating the residual to updating the model until the model performance no longer improves. Each iteration gradually improves the overall model by introducing new decision trees to obtain a GBDT ensemble model composed of different decision trees, and comprehensively integrate the prediction results of different decision trees to achieve the identification of transaction patterns.
7. The method for data traceability and tracking of blockchain based on artificial intelligence technology according to claim 6, wherein: In S4, the process of detecting abnormal behavior through clustering analysis includes: Convert the key features in the feature set into a numerical feature vector. Define the maximum distance between samples within a cluster and the minimum number of samples required for a cluster according to the data characteristics and requirements. Starting from an unvisited data point, use this data point as the center and search for neighboring points within the maximum distance of the samples within the cluster. If the number of neighboring points reaches the minimum number of samples required for the cluster, a new cluster is formed. If the number of neighboring points does not reach the minimum number of samples required for the cluster, this point is regarded as a noise point; For each new cluster, continue to search for neighboring points within the maximum distance of the samples within the cluster of its member points and add the found neighboring points to the cluster. Repeat the process of expanding the cluster until no new neighboring points are added. After completing the clustering, mark the data points not assigned to any cluster as outliers, representing abnormal behaviors, and regard each cluster in the obtained set of clusters as a transaction pattern.
8. The method for data traceability and tracking of a blockchain based on artificial intelligence technology according to claim 7, wherein: In S5, the process of tracing the data paths of the transaction list data and the smart contract event log data through the graph algorithms in the graph database Neo4j includes: Import the numerical feature vectors into which the transaction list data and the smart contract event log data have been converted into the Neo4j graph database. In the Neo4j graph database, regard transactions and smart contract events as nodes, and represent the interaction relationships between nodes as edges. The interaction relationships include transfer behaviors between the sender address and the recipient address, contract deployment, and contract invocation. Attach the timestamp, block height, and transaction ID attributes to the corresponding nodes and edges to record the time series characteristics and location information of each transaction; Utilize the graph algorithm library provided by Neo4j to perform shortest path search, connected component analysis, and PageRank calculation operations. When tracing the fund flow path of a transaction, define the starting transaction node and the final receiving transaction node, apply the shortest path algorithm to determine all paths between the two, and use connected component analysis to identify closed transaction loops to detect potential loop structures; Analyze the path of the smart contract event log data by constructing a call relationship network between smart contract addresses. When a contract call event occurs, establish a directed edge between the contract address nodes accordingly, thereby forming a directed graph. Use graph algorithms to trace the actual execution path from the initial contract deployment to subsequent calls, and identify the external accounts that trigger the contract calls and the mutual influence relationships between each contract call.
9. The method for data traceability and tracking of blockchain based on artificial intelligence technology according to claim 8, wherein: In S6, the process of using RESTful API as a service endpoint and implementing asynchronous communication and client access through the message queue system Kafka includes: Deploy the RESTful API interface as a service endpoint. This API interface is used to receive HTTP requests from clients. The HTTP requests include querying historical transaction records and obtaining smart contract event operations. After the API server processes the received requests, it interacts with the local MongoDB document database, retrieves data from it, and returns the results to the clients in JSON format; Introduce the Kafka message queue system for asynchronous communication. When new transactions and smart contract events occur, the API server stores the generated transaction list data and smart contract event log data in the local MongoDB document database B, and serializes them into messages to be published to Kafka topics. Each topic corresponds to different types of transaction confirmations, contract deployment, and call event streams; The client application receives real-time updates by subscribing to the topics of interest. The Kafka consumer pulls messages from the topics and processes them according to the business logic. The background worker process consumes the corresponding messages and executes MapReduce task allocation and graph algorithm calculation operations, saving the processing results back to the local MongoDB document database. During network failures and system maintenance, the messages that have not been processed in time are temporarily stored in Kafka and continue to be transmitted after the connection is restored.
10. The method for data traceability and tracking of blockchain based on artificial intelligence technology according to claim 9, characterized in that: In step S7, the process of implementing security control during data access and transmission through the OAuth2.0 authorization protocol and the SSL / TLS encrypted transport layer includes: Use the OAuth2.0 authorization protocol to manage API endpoint access permissions. The client application initiates an authentication request to the OAuth2.0 authorization server, providing the client ID and secret key. The authorization server verifies the client's identity and sends a temporary authorization code to the client. The client exchanges the temporary authorization code for an access token and carries the access token as an HTTP header to request the RESTful API interface; The API server receives the request containing the access token, internally verifies the validity, scope, and expiration status of the access token, authorizes the client to execute the API call. For API requests involving sensitive operations, the access token is refreshed to extend the validity period of the access permission, and the access token is rotated regularly; Implement the SSL / TLS encrypted transport layer protocol at the network communication level to encrypt the data transmission between the client and the API server, as well as between the API server and the MongoDB document database and the Kafka message queue system. The SSL / TLS protocol combines symmetric encryption algorithms and asymmetric encryption algorithms, negotiates and shares the secret keys of both parties through the handshake protocol at the initial stage of establishing a connection, uses the secret keys to encrypt and decrypt data during data transmission, and verifies the true identity of the communicating parties through digital certificates.