A big data analysis method and system based on blockchain
By introducing the Merkle tree hash structure and hash chain mechanism into the blockchain system, combining the distributed graph computing framework GraphX and graph mining algorithm, the problem of low efficiency and insufficient correlation mining capabilities of the blockchain system in financial big data analysis is solved, and efficient data analysis and risk model discovery is achieved.
Patent Information
- Application Number
- CN202411798405.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-09
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-12-09
AI Technical Summary
The existing blockchain systems have problems in financial big data analysis, such as low data retrieval efficiency, large communication overhead, heavy computing burden, and lack of effective mining and insight into the relationship between transaction subjects.
By introducing the Merkle tree hash structure and hash chain mechanism, a data integrity proof is constructed, and a distributed graph computing framework GraphX is used to graphically model transaction data, and data correlation analysis is achieved in combination with a graph mining algorithm.
It improves data mining efficiency, ensures the credibility and immutability of data, and can effectively discover abnormal behaviors and risk patterns in the transaction network.
Smart Images

Figure CN119739786B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing, and in particular to a big data analysis method and system based on blockchain. Background Art
[0002] With the rapid development and digital transformation of the financial industry, financial transaction data has shown an explosive growth trend. How to effectively analyze and mine the value contained in massive financial data has become a major challenge facing financial institutions. Traditional financial data analysis methods are mainly based on relational database and data warehouse technology, using SQL query and aggregation analysis to process data. However, these methods have problems such as low analysis efficiency and limited data association mining capabilities when processing massive, heterogeneous, and dynamic financial big data, making it difficult to meet real-time and intelligent business needs.
[0003] In recent years, blockchain technology, as a decentralized, tamper-proof, and highly reliable distributed ledger technology, has provided new ideas for solving the pain points in financial big data analysis. Blockchain uses cryptographic principles to ensure the integrity and immutability of data by storing data in the form of blocks in a distributed network, providing strong support for the trusted storage and sharing of financial data. At the same time, each node in the blockchain network saves a complete copy of the data, and the data is synchronized and agreed upon in real time between the nodes, creating conditions for cross-institutional financial data correlation analysis.
[0004] However, the existing blockchain system still has limitations in the analysis of financial big data. On the one hand, the data stored on the blockchain lacks an effective index and query mechanism, and the data retrieval efficiency is low; on the other hand, the data synchronization and consensus mechanism between blockchain nodes brings a large communication overhead and computing burden, which makes it difficult to support the real-time analysis of massive financial data. In addition, the existing blockchain data analysis methods mainly focus on statistical analysis of a single dimension, lack of mining and insight into the relationship between transaction entities, and cannot effectively discover abnormal behaviors and risk patterns in the transaction network. Summary of the invention
[0005] In response to the problem of low efficiency in financial data association mining and analysis in the prior art, the present application provides a big data analysis method and system based on blockchain. By introducing the Merkle tree hash structure and hash chain mechanism, blockchain technology is used to construct data integrity proof to ensure the credibility of the data. At the same time, the distributed graph computing framework GraphX is used to graphically model the transaction data, and the graph mining algorithm is combined to realize data association analysis, thereby improving the efficiency of data mining.
[0006] The purpose of this application is achieved through the following technical solutions.
[0007] One aspect of the present application provides a big data analysis method based on blockchain, including: obtaining transaction data; preprocessing the obtained transaction data to obtain structured data; using a hash function to encode and map the structured data to generate a unique identifier for the corresponding data; according to the Merkle tree hash structure, taking the unique identifier as the leaf node, iteratively generating the Merkle tree root hash value from bottom to top, and using the root hash value as a proof token for data integrity verification; linking and storing the proof tokens through a hash chain mechanism, and recording the hash value of each token on the chain of the next block; using a hash-based indexing method, using the hash value of the proof token to map the storage location of the structured data on the blockchain to generate a blockchain index; using a distributed file system IPFS, uploading the structured data and the corresponding proof token to the blockchain node; using a distributed graph computing framework GraphX to perform graph modeling on the stored structured data, taking the transaction subject as a graph node, and the transaction behavior as a directed edge, to construct a topological structure diagram of the association relationship between the transaction subjects; according to the topological structure diagram of the association relationship, using a graph mining algorithm, obtaining the associated transaction path of the transaction data, and outputting it as a data analysis result.
[0008] Furthermore, a hash function is used to encode and map the structured data to generate a unique identifier for the corresponding data, including: performing field extraction on the acquired transaction data to extract the transaction amount, transaction time and transaction subject; converting the transaction amount data into a floating point format using a regular expression matching method to obtain standardized transaction amount data; converting the transaction time data into a standard date and time format using a date and time parsing function to obtain standardized transaction time data; parsing the transaction subject data using a string matching method to extract the subject identification substring, and mapping the substring to an integer code to obtain an encoded transaction subject ID; concatenating the standardized transaction amount, standardized transaction time and transaction subject ID to obtain structured data; and performing hash calculation using a hash algorithm SHA-256 based on the structured data to obtain a hash value as the unique identifier for the corresponding transaction data.
[0009] Furthermore, according to the Merkle tree hash structure, with the unique identifier as the leaf node, the Merkle tree root hash value is iteratively generated from bottom to top, and the root hash value is used as a proof token for data integrity verification, including: arranging the unique identifier in a predefined order as the hash value of the Merkle tree child node; pairing the hash values of two adjacent leaf nodes, generating the corresponding parent node hash value by joint hash calculation, and arranging the parent node hash values in a predefined order as the input of the next round of hash calculation; repeating the pairing and hash calculation steps until the hash value of the Merkle tree root node is generated, and using the hash value of the root node as the Merkle proof of the entire transaction data; combining the root hash value of the Merkle proof with the height and timestamp of the current block, and performing a hash operation, and using the obtained hash value as the transaction integrity verification token of the current block; in the next block, embedding the hash value of the transaction integrity verification token of the previous block into the block header, and performing a hash operation with the current block header information to generate a new transaction integrity verification token.
[0010] Furthermore, the proof tokens are linked and stored through a hash chain mechanism, and the hash value of each token is recorded on the chain of the next block, including: parsing the transaction integrity verification token, extracting the root hash value of the Merkle proof, and forming a key-value pair with the root hash value and the unique identifier of the corresponding transaction data; through the hash index method, the root hash value in the key-value pair is used to map the storage location of the transaction data on the blockchain to build a blockchain index; using the content-addressed storage model of IPFS, the structured data and the corresponding proof tokens are uploaded to the blockchain node group; the hash value of the unique identifier of the structured data is used as the content-addressed identifier of the corresponding data in the IPFS network.
[0011] Furthermore, the distributed graph computing framework GraphX is used to perform graph modeling on the stored structured data, with the transaction subject as the graph node and the transaction behavior as the directed edge, to construct a topological structure diagram of the association relationship between the transaction subjects, including: obtaining multiple transaction data within a preset time range from the blockchain index, and obtaining the corresponding structured transaction data through the content addressing mechanism of IPFS according to the unique identifier of the transaction data; performing graphical analysis on the obtained structured transaction data to extract the transaction subject ID in each transaction; using a hash algorithm to perform a hash operation on the transaction subject ID to obtain the hash value as the identity fingerprint of the transaction subject, and using the identity fingerprint as the topological structure diagram of the association relationship. Nodes; perform graphical analysis on the acquired structured transaction data to extract the transaction amount, transaction time and transaction direction of each transaction; use the time window division algorithm to divide the transaction time, and mark the transaction time with serial numbers according to the time window in which the transaction occurs; use the amount interval discretization method to group and encode the transaction amount, and map the transaction amount according to the preset amount interval; use the serial numbered transaction time and the mapped transaction amount as the edges of the association relationship topology structure diagram; set the transaction frequency as the additional weight of the edge, and count the number of transactions between the two target transaction entities within the preset time range; construct the association relationship topology structure diagram based on the nodes, edges and additional weights.
[0012] Furthermore, the transaction frequency is set as an additional weight of the edge, and the number of transactions between two target transaction entities within a preset time range is counted, including: traversing all transaction data within the preset time range, using the MapReduce distributed computing framework to process the transaction data in parallel, and counting the number of transactions between each pair of transaction entities; wherein, the Map stage takes the transaction data as input, extracts the transaction entity pairs and the corresponding transaction times, and the Reduce stage takes the transaction entity pairs as keys, summarizes the transaction times output by each Map task, and obtains the total transaction times for each pair of transaction entities; in the Reduce stage, the combiner function is used to perform parallel processing on the transaction data at the Map end. The subject pairs are merged and calculated to obtain the cumulative transaction frequency between the node pairs. The intermediate results of the node pair transaction frequency output in the Reduce stage are stored in a column storage database. Each node pair represents two transaction subjects, and the frequency value is used as an additional attribute of the edge. For all the node pair transaction frequencies stored in the column storage database, the K node pairs with the highest current transaction frequency are maintained through the maximum heap data structure, and the transaction frequency is normalized by linear interpolation to generate a list of frequent transaction partners between the transaction subjects. According to the node pairs in the frequent transaction partner list and the corresponding normalized transaction frequency, the number of transactions between the two target transaction subjects within the preset time range is obtained.
[0013] Furthermore, according to the association relationship topology structure diagram, the graph mining algorithm is used to obtain the association transaction path of the transaction data as the output of the data analysis result, including: using the Dijkstra shortest path algorithm to obtain the shortest transaction path between any two target nodes in the association relationship topology structure diagram, arranging the obtained shortest transaction path in the order of the transaction occurrence, and constructing the shortest transaction sequence of the target node. Specifically, the association relationship topology structure diagram is converted into an adjacency matrix representation, the nodes in the association relationship topology structure diagram are used as the rows and columns of the adjacency matrix, and the transaction time on the edge of the association relationship topology structure diagram is used as the element value of the corresponding position in the adjacency matrix to obtain the adjacency matrix representing the association relationship topology structure diagram; in the adjacency matrix, the element value with a transaction time of 0 is replaced by a maximum value greater than all transaction times, and the elements in the matrix less than 0 are replaced by 0 to obtain a revised adjacency matrix; the revised adjacency matrix is used as the input of the Dijkstra algorithm, and the shortest path between any two target nodes in the adjacency matrix is calculated by the Dijkstra algorithm to obtain A shortest path result including a node sequence and a path length on the shortest path; according to the node sequence in the shortest path result, corresponding nodes and edges are extracted from the association relationship topology structure diagram to construct a shortest path subgraph; the edges in the shortest path subgraph are traversed, and the edges in the shortest path subgraph are sorted in ascending order of transaction time according to the transaction time on the edges to obtain a sorted shortest path edge sequence; according to the sorted shortest path edge sequence, the source node and the target node in the shortest path edge sequence are extracted, and the shortest transaction sequence between the target nodes is constructed according to the order of the shortest path edges; the constructed shortest transaction sequence of the target node is used as input data for a subsequent transaction pattern matching algorithm.
[0014] An AC automaton pattern matching algorithm is adopted, and a preset frequent transaction pattern is used as a matching template to perform a matching search on the shortest transaction sequence of a target node to obtain a target transaction fragment that meets the matching template; specifically, each transaction in the preset frequent transaction pattern is taken as a character to construct a pattern string set of the AC automaton, and a Go-to function and a Failure function of the AC automaton are constructed based on the pattern string set; each transaction in the shortest transaction sequence of the target node is taken as a character to construct a target text string of the AC automaton; using the constructed AC automaton, a multi-pattern matching search is performed on the target text string to obtain the termination position information of all frequent transaction patterns appearing in the target text string; according to the termination position information of the frequent transaction pattern, a backtracking algorithm is adopted to start from the termination position and backtrack the matching path along the Failure function of the AC automaton until the starting position of the matching path, to obtain the starting position information of all frequent transaction patterns ending with the termination position, and to construct a position interval set of the frequent transaction pattern; overlapping position intervals in the position interval set are merged to obtain a position interval set of the merged target transaction fragment; according to the position interval set of the target transaction fragment, the target transaction is extracted from the shortest transaction sequence of the target node.
[0015] Calculate the time difference between every two adjacent transactions in the target transaction segment respectively, count the number of transactions with a time difference less than a threshold, and take the number of transactions with a time difference less than the threshold as the time density of the target transaction segment; calculate the transaction amount of each transaction in the target transaction segment respectively, count the number of transactions with a transaction amount greater than the threshold, and take the number of transactions with a transaction amount greater than the threshold as the transaction activity of the target transaction segment; calculate the confidence of the target transaction segment based on the time density and transaction activity of the target transaction segment; define the target transaction segment with a confidence lower than the threshold as a suspicious transaction segment; extract the nodes and edges contained in each suspicious transaction segment in the association relationship topological structure diagram, and construct a suspicious transaction subgraph.
[0016] The Louvain community mining algorithm is used to conduct community mining on the constructed suspicious transaction subgraph, and the community mining results are output as data analysis results. Specifically, the suspicious transaction subgraph is converted into a weighted adjacency matrix representation, and the elements in the adjacency matrix represent the sum of the transaction amounts between the corresponding nodes in the suspicious transaction subgraph, and the initial community division is obtained, in which each node constitutes an initial community; based on the adjacency matrix, the following sub-steps are iteratively performed until the community division no longer changes: traverse each node in the suspicious transaction subgraph, calculate the change in modularity after each node joins the community where its neighboring nodes are located, and select the community that can maximize the modularity improvement as the target community of the node; according to the target community of each node, all nodes are re-divided into corresponding communities, and the adjacency matrix representation of the community is updated according to the division results; traverse each community in the adjacency matrix, merge the sub-communities connected within the community into one community, and update the adjacency matrix of the merged community Representation; based on the adjacency matrix representation of the community, calculate the internal characteristics of each community, including: counting the number of nodes within the community; counting the number of edge connections between nodes within the community, and calculating the edge density between nodes within the community; counting the transaction amounts of nodes within the community, and calculating the sum, mean, variance and other statistical characteristics of the transaction amounts within the community; traversing each community, calculating the ratio of the sum of the transaction amounts of nodes within the community to the total of all transaction amounts within the community, and marking the community with a ratio greater than a preset threshold as a suspicious transaction community; outputting the information of each suspicious transaction community as the result of the suspicious transaction pattern analysis, including: identification information of all nodes within the suspicious transaction community; identification information and weight information of all edges within the suspicious transaction community; statistical characteristics related to the transaction amounts of nodes within the suspicious transaction community.
[0017] Furthermore, the time difference between every two adjacent transactions in the target transaction segment is calculated using the following formula: ,in, Indicates the time difference between the i-th transaction and the i+1-th transaction in the target transaction segment; represents the transaction time of the i-th transaction in the target transaction segment; Indicates the transaction time of the i+1th transaction in the target transaction segment.
[0018] Furthermore, the confidence of the target transaction segment is calculated according to the time density and transaction activity of the target transaction segment, using the following formula: , where Confidence represents the confidence of the target transaction fragment, and its value range is [0, 1]; Indicates the temporal closeness of the target transaction segment; Indicates the transaction activity of the target transaction segment; ,in, Indicates that the time difference is less than the preset time threshold The number of transactions, N represents the total number of transactions in the target transaction segment; ,in, Indicates that the transaction amount is greater than the preset amount threshold Number of transactions; Indicates the transaction amount of the i-th transaction in the target transaction segment.
[0019] Another aspect of the present application provides a blockchain-based big data analysis system for executing a blockchain-based big data analysis method of the present application.
[0020] Compared with the prior art, the advantages of this application are:
[0021] By introducing the Merkle tree hash structure, the unique identifier of the transaction data is used as the leaf node, and the Merkle tree root hash value is iteratively generated from the bottom up, and the root hash value is used as a proof token for data integrity verification. The proof token is linked and stored using the hash chain mechanism, and the hash value of each token is recorded on the chain of the next block. A hash-based indexing method is used to map the storage location of structured data on the blockchain using the hash value of the proof token to generate a blockchain index. The integrity and immutability of the transaction data are guaranteed to prevent the data from being forged or tampered with.
[0022] The distributed graph computing framework GraphX is used to perform graph modeling on the stored structured transaction data. The transaction subjects are taken as graph nodes and the transaction behaviors are taken as directed edges to construct a topological structure diagram of the association relationship between the transaction subjects. On this basis, the transaction frequency is introduced as an additional weight of the edge. The transaction frequency of node pairs is counted and stored through the MapReduce distributed computing framework and column storage database. The maximum heap data structure is used to maintain the K node pairs with the highest transaction frequency, and a list of frequent transaction partners between transaction subjects is generated. Efficient graphical association modeling and frequent pattern mining are performed on massive transaction data, which significantly improves the efficiency of association mining and analysis of blockchain transaction data.
[0023] Based on the transaction subject association topology, the shortest transaction path between any two target nodes is obtained through the Dijkstra shortest path algorithm, and the shortest transaction sequence of the target node is constructed. The AC automaton pattern matching algorithm is used to match the shortest transaction sequence with the preset frequent transaction pattern as the matching template to obtain the target transaction fragment that meets the matching template. By analyzing the time density and transaction activity of the target transaction fragment, the confidence of the transaction fragment is calculated, and the transaction fragment with a confidence lower than the threshold is defined as a suspicious transaction fragment, and then the suspicious transaction subgraph is extracted. Suspicious transaction patterns in blockchain transaction data can be quickly discovered. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 is an exemplary flow chart of a big data analysis method based on blockchain as shown in this application;
[0025] Figure 2 is an exemplary flow chart of generating a unique identifier according to the present application;
[0026] Figure 3 is an exemplary flow chart of constructing an association relationship topology diagram according to the present application;
[0027] Figure 4 is an exemplary flow chart for generating data analysis results according to the present application. DETAILED DESCRIPTION
[0028] The method and system provided in the embodiments of the present application are described in detail below with reference to the accompanying drawings.
[0029] like Figure 1 As shown, the transaction data is obtained; specifically, the application communicates with the RPC interface provided by the blockchain node and calls the corresponding application programming interface (API) to obtain the required transaction data. As the core component of the blockchain network, the blockchain node maintains a complete transaction database and provides a series of standardized API interfaces to the outside world, allowing external applications to access and query transaction data through RPC calls. Common blockchain platforms all provide a rich RPC interface that supports obtaining multiple data types such as block information, transaction information, and account balances. The block information of the specified block height can be obtained through the eth_getBlockByNumber method in the JSON-RPC interface, which contains detailed data of all transactions in the block. The transaction data returned by the RPC interface is usually encoded in JSON format, and each transaction record contains multiple key fields, such as transaction hash, block height, transaction timestamp, transaction amount, transaction subject, etc. These fields provide key information such as the unique identification, time attribute, quantity attribute, and relationship attribute of the transaction, and are important inputs for subsequent data analysis. Transaction data includes fields such as transaction hash, block hash, block height, transaction initiator address, transaction recipient address, transaction amount, transaction timestamp, etc., providing the original data basis for subsequent data preprocessing and analysis.
[0030] like Figure 2As shown, the acquired transaction data is preprocessed to obtain structured data; the structured data is encoded and mapped using a hash function to generate a unique identifier for the corresponding data; the acquired transaction data is subjected to field extraction to extract the transaction amount, transaction time, and transaction subject; converting the transaction amount data into a standard floating point format is one of the key links in data normalization. In order to achieve the format conversion and standardized representation of the transaction amount, this application adopts the technical means of regular expression matching. Regular expressions are a powerful tool for matching string patterns, which can describe the structure and content of the target string by defining a series of rules and special characters. When processing transaction amount data, we need to extract the numerical part from the original string representation and convert it into a floating point format for subsequent numerical calculations and statistical analysis. Specifically, this application uses the regular expression / d+ ( / . / d+)? to match the numerical part in the transaction amount field. The meaning of this regular expression is as follows: / d+ matches one or more consecutive numeric characters to extract the integer part of the amount; ( / . / d+)? matches the decimal point and one or more numeric characters after it to extract the decimal part of the amount, where? indicates that this part is optional, that is, the amount can be an integer or a decimal. Through this regular expression, we can accurately extract the numerical part from the original transaction amount string and ignore other irrelevant characters. For the following transaction amount strings: "10.5ETH"; "0.02BTC"; "1,000,000USD"; using the regular expression / d+ ( / . / d+)? for matching, the numerical parts of 10.5, 0.02 and 1,000,000 can be extracted respectively. After extracting the numerical part, we also need to convert it to a standard floating point format for subsequent numerical calculations and comparisons. Different programming languages and environments provide different functions or methods for converting strings to floating point numbers, such as the float() function in Python and the Double.parseDouble() method in Java. These functions or methods can automatically handle special formats in numeric strings, such as thousands separators, decimal points, etc., and convert them into corresponding floating-point values.
[0031] For the extracted numerical string 1,000,000, the float (*) function is used for conversion to obtain the floating point number 1000000.0; for the numerical string 0.02, the floating point number 0.02 is obtained after conversion. Through regular expression matching and floating point number conversion, this application realizes the standardized representation of transaction amount data, uniformly converts the original string format into a standard floating point format, eliminates the differences and ambiguities in the data format, and provides a consistent data basis for subsequent numerical statistics, comparison, sorting and other operations. The standardized transaction amount data can directly participate in numerical calculations, such as summing, averaging, maximum and minimum values, etc., which improves the efficiency and accuracy of data processing.
[0032] Normalizing transaction time data is another important task in the data preprocessing stage. In order to convert transaction time data into a standard date and time format for time series analysis and comparison, this application adopts the technical means of date and time parsing functions. Transaction time data in blockchain networks are usually recorded in the form of timestamps, that is, the number of seconds or milliseconds from a fixed time point (such as Unix epoch time January 1, 1970 00:00:00 UTC) to the time when the transaction occurred. However, this timestamp representation is not friendly to human reading and understanding, nor is it convenient for direct comparison and calculation of time attributes. Therefore, we need to convert the timestamp to a standard date and time format, such as "YYYY-MM-DDHH:MM:SS", to improve the readability and processability of time data. In order to realize the conversion of timestamps to date and time formats, this application utilizes date and time parsing functions provided by programming languages or environments. These functions accept timestamps as input and convert the timestamps into corresponding date and time string representations according to the specified time zone and format string.
[0033] Take the Unix timestamp as an example, it represents the number of seconds from 00:00:00 UTC on January 1, 1970 to the time when the transaction occurred. Given a Unix timestamp, such as 1622736000, we can use the date and time parsing function to convert it to a standard date and time format. Different programming languages provide different date and time parsing functions, such as datetime.from time stamp() in Python, java.util.Date and java.text.SimpleDateFormat in Java, etc. Taking Python's datetime.from time stamp() function as an example, the Unix timestamp can be converted to a date and time string in the following way: timestamp = 1622736000; dt = date time.from timestamp (timestamp); formatted_time = dt.strftime ("%Y-%m-%d %H: %M: %S"); In the above code, the datetime.from timestamp (*) function accepts a Unix timestamp as input and returns the corresponding date and time object dt. Then, through the strftime(*) method, according to the specified format string "%Y-%m-%d%H:%M:%S", the date and time object is converted to a standard date and time string representation such as "2021-06-0312:00:00".
[0034] Through the date and time parsing function, this application realizes the standardized representation of transaction time data, converts the original timestamp into a standard date and time format, and improves the readability and understandability of time data. Standardized transaction time data can easily compare, sort, group and other operations of time attributes, such as aggregating transaction amounts by date, counting transaction frequencies by time period, etc., which provides a basis for time series analysis and trend prediction. In order to extract the unique identifier of the transaction subject and convert it into an integer code that is easy to store and calculate, this application adopts string matching and mapping technical means. The transaction subject in the blockchain network is usually represented in the form of a string, such as a wallet address, account address, etc. These address strings are usually long and contain multiple characters such as letters and numbers, which are not convenient for direct comparison and calculation. In order to improve the efficiency and consistency of data processing, we need to extract a relatively short and unique identification substring from the original address string and map it to an integer code as the unique ID of the transaction subject.
[0035] Specifically, this application uses a string matching method to parse and extract transaction subject data. According to the address format characteristics of different blockchain networks, we can predefine a set of matching rules to identify and extract key parts in the address string. For example, the account address is a 20-byte (40 hexadecimal characters) string starting with "0x", such as "0x742d35Cc6634C0532925a3b844Bc454e4438f44e". For this address format, we can use the following matching rules: match the string starting with "0x"; extract the first 20 characters after "0x" as the unique identification substring; through this set of matching rules, we can extract from the original address string. A 20-character substring of the form "742d35Cc6634C0532925" as the unique identifier of the transaction subject. After extracting the unique identification substring, we also need to map it to an integer code for subsequent storage, comparison and calculation. A commonly used mapping method is hash mapping, which uses a hash function to map a string to an integer value of fixed length. Common hash functions include MD5, SHA-1, SHA-256, etc., and the appropriate hash function can be selected according to the length of the identification substring and the probability of collision. For the extracted identification substring "742d35Cc6634C0532925", we can use the SHA-256 hash function to hash it and obtain a 256-bit hash value. Then, we can convert the hash value into an integer and take its lower 32 bits as the integer encoding ID of the transaction subject. Through hash mapping, we can map different identification substrings to unique integer encodings, ensuring the uniqueness and consistency of the transaction subject ID. Compared with string identification, integer encoding has the advantages of small storage space and high comparison efficiency, and is more suitable for large-scale data processing and analysis scenarios.
[0036] By adopting the technical means of string matching and hash mapping, this application realizes the parsing and encoding of transaction subject data, converts the original address string into a unique integer encoding ID, and improves the efficiency and consistency of data processing. The encoded transaction subject ID can be conveniently used for data mining tasks such as association analysis and graph calculation, such as building a relationship network between transaction subjects, calculating the centrality index of the subject, etc., to provide data support for in-depth understanding of the behavior pattern and risk characteristics of the transaction subject. The standardized transaction amount, transaction time and transaction subject ID are spliced to form structured data, and the SHA-256 hash algorithm is used to generate a unique identifier. The previous data preprocessing step has converted the transaction amount into a standard floating point format, the transaction time into a standard date and time format, and the transaction subject is mapped to a unique integer encoding ID. Now, we need to splice these standardized data elements to form a structured data record for subsequent storage, query and analysis. Specifically, this application adopts the string splicing method to splice the standardized transaction amount, transaction time and transaction subject ID in a fixed order and format. We can define the following concatenation format: transaction amount, transaction time, transaction subject ID; where the transaction amount is in floating point format, the transaction time is a date and time string in the format of "YYYY-MM-DDHH:MM:SS", and the transaction subject ID is an integer code. These data elements are separated by commas to form a structured string record. For a normalized transaction data, the transaction amount is 1000.00, the transaction time is "2023-05-2014:30:00", and the transaction subject ID is 42. The concatenated structured data is: 1000.00, 2023-05-2014:30:00, 42. By concatenating the normalized data elements in a fixed format, we get a structured data record that contains the key information of the transaction, which is convenient for subsequent storage and processing.
[0037] In order to ensure the uniqueness and integrity of the data, we also need to generate a unique identifier for each structured data. This application uses the SHA-256 hash algorithm to achieve this goal. SHA-256 is a secure cryptographic hash function that can map data of any length to a fixed-length (256-bit) hash value. Using the SHA-256 hash function, we can perform hash calculations on the spliced structured data to obtain a 256-bit hash value as the unique identifier of the transaction data. For the above-mentioned spliced structured data "1000.00, 2023-05-2014:30:00, 42", after SHA-256 hash calculation, the hash value obtained is: a7f5b7c9d1f2e3d4c5b6a7c8d9e0f1a2b3c4d5e6f7a8b9c0d1e2f3a4b5c6d7. This hash value can be used as a unique identifier for the transaction data for data indexing, querying and verification. The introduction of hash identifiers ensures the uniqueness and integrity of data, avoids data duplication and tampering, and improves the reliability of data processing.
[0038] The Merkle tree hash structure and hash chain mechanism are used to implement the integrity verification and storage optimization of transaction data. By constructing the Merkle tree and hash chain, a transaction integrity verification token is generated, and combined with the block header information of the blockchain, the verifiability and non-tamperability of the transaction data are achieved. Specifically, the present application first constructs the leaf nodes of the Merkle tree based on the unique identifier of the transaction data generated previously. The unique identifier of each transaction data serves as a leaf node, representing the hash value of the transaction data. In order to ensure the balance and consistency of the Merkle tree, the leaf nodes need to be arranged in a predefined order and sorted according to the lexicographic order of the unique identifier or the size of the value. There are 8 unique identifiers for the following transaction data: Tx1: a7f5b7c9d1f2e3d4c5b6a7c8d9e0f1a2b3c4d5e6f7a8b9c0d1e2f3a4b5c6d7; Tx2: b8c6d4e2f1a9b7c5d3e1f9a7b5c3d1e9f7a5b3c1d9e7f5a3b1c 9d7e5f3a1b9; Tx3: c9d5e3f1a7b5c3d1e9f7a5b3c1d9e7f5a3b1c9d7e5f3a1b9c7d5e 3f1a7b5c3; Tx4: d1e9f7a5b3c1d9e7f5a3b1c9d7e5f3a1b9c7d5e3f1a7b5c3d1e9f7a 5b3c1d9; Tx5: e7f5a3b1c9d7e5f3a1b9c7d5e3f1a7b5c3d1e9f7a5b3c1d9e7f5a3b1c 9d7e5; Tx6: f3a1b9c7d5e3f1a7b5c3d1e9f7a5b3c1d9e7f5a3b1c9d7e5f3a1b9c7d5e 3f1; Tx7: a5b3c1d9e7f5a3b1c9d7e5f3a1b9c7d5e3f1a7b5c3d1e9f7a5b3c1d9e7f5a3; Tx8: b1c9d7e5f3a1b9c7d5e3f1a7b5c3d1e9f7a5b3c1d9e7f5a3. These unique identifiers are used as leaf nodes of the Merkle tree and arranged in a predefined order to obtain the initial leaf node sequence of the Merkle tree: Next, the hash values of two adjacent leaf nodes are paired, and the hash value of the corresponding parent node is generated by joint hash calculation. Joint hash calculation is to concatenate the hash values of the two child nodes and perform hash calculation again to obtain the hash value of the parent node. and , the hash value of its parent node is: Hash(Tx1+Tx2)=Hash("a7f5b7c9d1f2e3d4c5b6a7c8d9e0f1a2b3c4d5e6f7a8b9c0d1e2f3a4b5c6d7"+"b8c6d4e2f1a9b7c5d3e1f9a7b5c3d1e9f7a5b3c1d9e7f5a3b1c9d7e5f3a1b9"); where "+" represents the string concatenation operation, and Hash(*) represents the hash function, such as SHA-256.
[0039] Repeat this pairing and hashing process to generate the hash value of the parent node of the next layer until the hash value of the root node of the Merkle tree is generated. In each round of pairing and hashing, the parent node hash values need to be arranged in a predefined order as the input for the next round of calculation. , after three rounds of pairing and hashing , , in the first round, pairing and hashing are performed to generate the parent node hash: ; Second round of hash calculation: ; Finally generate the root node hash value of the Merkle tree: Through this layer-by-layer pairing and hash calculation method, a large amount of transaction data is compressed into a root hash value of a fixed length, which serves as the Merkle proof of the entire transaction data set.
[0040] In order to further enhance the security and verifiability of transaction data, this application combines the root hash value of the Merkle proof with the height and timestamp of the current block, and performs a hash operation to obtain a transaction integrity verification token for the current block. This technical solution utilizes the block header information of the blockchain to associate transaction data with a specific block, prevent transaction data from being moved or reused between different blocks, and improve the credibility of transaction data. Specifically, after generating the root hash value of the Merkle proof, this application combines the root hash value with the height and timestamp of the current block. The block height indicates the sequence number or position of the block in the blockchain, and the timestamp indicates the time point when the block was generated. Combining the root hash value with these two block attributes can uniquely identify the location and time of the transaction data set corresponding to the Merkle proof in the blockchain. The current block height is 1000, the timestamp is 2023-05-2014:30:00, and the root hash value of the Merkle proof is: RootHash = "a7f5b7c9d1f2e3d4c5b6a7c8d9e0f1a2b3c4d5e6f7a8b9c0d1e2f3a4b5c6d7". The root hash value, block height and timestamp are concatenated in a predefined format to obtain a combined string: CombinedString = "a7f5b7c9d1f2e3d4c5b6a7c8d9e0f1a2b3c4d5e6f7a8b9c0d1e2f3a4b5c6d7, 1000, 2023-05-2014:30:00". Perform a hash operation on the combined string to obtain the transaction integrity verification token of the current block: TransactionIntegrity Token = Hash("a7f5b7c9d1f2e3d4c5b6a7c8d9e0f1a2b3c4d5e6f7a8b9c0d1e2f3a4b5c6d7, 1000, 2023-05-2014:30:00").
[0041] In order to further enhance the security and traceability of transaction data, this application uses a hash chain mechanism to link and store transaction integrity verification tokens. In the chain structure of the blockchain, the hash value of the transaction integrity verification token of each block is recorded in the block header of the next block to form a continuous hash chain. For the transaction integrity verification token of block 1000, its hash value is recorded in the block header of block 1001: Block1001 Header: {... "PrevBlockHash": "...", "TransactionIntegrityTokenHash": "Hash (TransactionIntegrityToken of Block 1000)", ...}. Through this hash chain mechanism, the transaction integrity verification token forms a continuous chain with the block header information of the blockchain. Any tampering with the transaction data or the transaction integrity verification token will cause the hash chain to break, which can be easily detected. At the same time, by tracing the hash chain, the historical authenticity and integrity of the transaction data can be verified, and the traceability of the transaction data can be achieved.
[0042] In order to optimize the storage and retrieval efficiency of transaction data, a hash-based indexing method and a distributed file system IPFS technical solution are adopted. By using the hash value of the transaction integrity verification token to index the structured transaction data and storing the transaction data in the IPFS network, efficient data retrieval and distributed storage are achieved. Specifically, after generating the transaction integrity verification token, this application extracts the root hash value of the Merkle proof contained in the token by parsing the token. The root hash value, as a unique identifier of the integrity of the transaction data, can form a key-value pair with the unique identifier of the corresponding transaction data.
[0043] For a certain transaction data, its unique identifier is "Tx1", and the corresponding Merkle proof root hash value is: RootHash = "a7f5b7c9d1f2e3d4c5b6a7c8d9e0f1a2b3c4d5e6f7a8b9c0d1e2f3a4b5c6d7". The unique identifier and the root hash value form a key-value pair: Key-ValuePair: ("Tx1", "a7f5b7c9d1f2e3d4c5b6a7c8d9e0f1a2b3c4d5e6f7a8b9c0d1e2f3a4b5c6d7"). Using the root hash value in the key-value pair, the storage location of the transaction data on the blockchain is mapped through the hash index method. The hash index method is a data indexing technology based on hash functions. By performing hash operations on key values, the hash values are mapped to the corresponding storage locations or data blocks. Perform a hash operation on the root hash value to obtain a hash index value: HashIndex=Hash("a7f5b7c9d1f2e3d4c5b6a7c8d9e0f1a2b3c4d5e6f7a8b9c0d1e2f3a4b5c6d7")%TotalStorage Blocks; where TotalStorageBlocks represents the total number of storage blocks on the blockchain. The hash index value can be used to quickly locate the storage location or data block of the transaction data on the blockchain. Through this hash-based indexing method, the present application constructs an efficient blockchain index. The blockchain index associates the unique identifier of the transaction data with its corresponding Merkle proof root hash value, and maps it to the storage location on the blockchain through a hash operation. When querying transaction data, the corresponding root hash value can be quickly retrieved through the unique identifier, and then the hash index method can be used to locate the storage location of the transaction data on the blockchain to achieve efficient data retrieval.
[0044] In order to further improve the availability and distributed storage capabilities of transaction data, this application uses the distributed file system IPFS to store structured transaction data and corresponding proof tokens. IPFS is a distributed file system based on content addressing. By using the hash value of the file content as the unique identifier of the file, efficient data distribution and redundant storage are achieved. In this application, the hash value of the unique identifier of the structured transaction data is used as the content addressing identifier (ContentIdentifier, CID) of the corresponding data in the IPFS network. For transaction data with a unique identifier of "Tx1", its hash value is calculated: CID=Hash("Tx1"); the transaction data and the corresponding proof token are uploaded to the IPFS network, and the CID is associated with the transaction data. The nodes in the IPFS network will store and distribute data based on the CID to achieve distributed storage and redundant backup of transaction data. When retrieving data, the corresponding transaction data can be quickly located and obtained in the IPFS network through the CID. The content addressing mechanism of IPFS ensures the integrity and consistency of the data, and at the same time improves the availability and fault tolerance of the data through data distribution and redundant storage between nodes.
[0045] like Figure 4As shown, in order to conduct in-depth analysis of transaction data, the distributed graph computing framework GraphX and the MapReduce distributed computing framework are used to perform graph modeling and frequent transaction partner analysis on the stored structured transaction data. By constructing a topological structure diagram of the association relationship between transaction entities and calculating the transaction frequency between transaction entities, a list of frequent transaction partners is generated to provide data support for subsequent abnormal transaction detection and risk assessment. Specifically, after obtaining multiple transaction data within a preset time range, this application uses the content addressing mechanism of IPFS to obtain the corresponding structured transaction data based on the unique identifier (CID) of the transaction data. The content addressing mechanism of IPFS maps the data content to a unique CID based on the hash value of the data content. Each CID is a unique identifier for the data content, and the corresponding data can be quickly located and retrieved in the IPFS network through the CID. The CID of a transaction data is "QmXyz...". Through the content addressing mechanism of IPFS, the corresponding structured transaction data can be obtained from the IPFS network according to the CID: CID: QmXyz... RetrievedData: {"transactionID": "Tx123", "sender": "Alice", "recipient": "Bob", "amount": 100, "timestamp": "2023-06-01T10:30:00Z", "direction": "outbound"}. After obtaining the structured transaction data, this application will analyze it graphically and extract the key information in each transaction, including the transaction subject ID, transaction amount, transaction time and transaction direction. These key information will be used to construct a topological structure diagram of the association relationship between transaction subjects.
[0046] In order to uniquely identify the transaction subject in the association relationship topology diagram, the present application uses a hash algorithm to perform a hash operation on the transaction subject ID to obtain a hash value as the identity fingerprint of the transaction subject. The hash algorithm is an algorithm that maps data of any length to a hash value of a fixed length, and has the characteristics of uniqueness and irreversibility. By performing a hash operation on the transaction subject ID, a fixed-length, unique identity fingerprint can be generated to identify the transaction subject. For the transaction subject ID "Alice", the SHA-256 hash algorithm is used for hash operation: SHA-256 ("Alice") = "1abc...xyz"; the obtained hash value "1abc...xyz" is used as Alice's identity fingerprint, which is used to represent Alice, the transaction subject, in the association relationship topology diagram. In the association relationship topology diagram, the identity fingerprint is used as a node to represent the unique identification of the transaction subject. Through the identity fingerprint, the same transaction subject in different transaction data can be associated to build a complete transaction subject relationship network.
[0047] When constructing the topological structure diagram of the association relationship between transaction entities, this application uses the time window partitioning algorithm and the amount interval discretization method to process the transaction time and transaction amount respectively, and converts the continuous time and amount data into discrete representations as the edge attributes of the topological structure diagram of the association relationship. For transaction time, this application uses the time window partitioning algorithm for processing. The time window partitioning algorithm divides a continuous time range into time windows of fixed size, and each time window represents a fixed time interval. By marking the time windows with serial numbers, continuous timestamps can be mapped to discrete time window numbers. The preset time window size is 1 hour. For a period of time ranging from 2023-06-01 00:00:00 to 2023-06-02 23:59:59, it can be divided into 48 time windows, each of which corresponds to a time interval of one hour: TimeWindow1: 2023-06-01 00:00:00~2023-06-01 00:59:59; TimeWindow2: 2023-06-01 01:00:00~2023-06-01 01:59:59; ...; TimeWindow48: 2023-06-02 23:00:00~2023-06-02 23:59:59. Each time window is marked with a sequence number, TimeWindow1 is numbered 1, TimeWindow2 is numbered 2, and so on. Mapping the transaction timestamp to the corresponding time window number can convert continuous time data into discrete time window numbers. For a transaction with a transaction timestamp of "2023-06-0110:30:00", it falls in TimeWindow11, so it is mapped to time window number 11.
[0048] For the transaction amount, this application uses the amount interval discretization method for processing. The amount interval discretization method groups and encodes the transaction amount according to the preset amount interval, and maps the continuous transaction amount to a discrete amount interval. The preset amount intervals are as follows: Interval1: 0~999; Interval2: 1000~4999; Interval3: 5000~9999; Interval4: 10000~49999; Interval5: 50000andabove. Encode each amount interval, Interval1 is encoded as 1, Interval2 is encoded as 2, and so on. By mapping the transaction amount to the corresponding amount interval code, the continuous transaction amount can be converted into a discrete amount interval representation. For a transaction with a transaction amount of 3500, it falls in Interval2, so it is mapped to the amount interval code 2. Through the time window partitioning algorithm and the amount interval discretization method, this application converts the continuous transaction time and transaction amount into a discrete representation, and uses the time window sequence number and the amount interval code as the edge attributes of the association relationship topology diagram respectively. This discretization process can simplify the representation of transaction data, reduce the dimension of data, and retain the time and amount characteristics of transaction behavior. In the topological structure diagram of the association relationship, the transaction time after the serial number mark and the transaction amount after the mapping are used as the attributes of the edge to represent the transaction behavior between the transaction subjects. Through the time window serial number and amount interval coding, the time distribution and amount distribution characteristics of the transaction behavior can be characterized, providing important information for subsequent transaction pattern analysis and anomaly detection.
[0049] In order to characterize the frequency of transactions between transaction entities, this application sets the transaction frequency as an additional weight for the edge in the association relationship topology diagram. The transaction frequency indicates the number of transactions between two transaction entities within a preset time range. By calculating the transaction frequency, the transaction activity and closeness between transaction entities can be quantified. In order to efficiently calculate the transaction frequency, this application uses the MapReduce distributed computing framework to parallel process all transaction data within a preset time range. MapReduce is a parallel computing model based on the divide-and-conquer idea. It can divide large-scale data sets into multiple small data blocks, and process them in parallel on multiple computing nodes, and finally merge the results to obtain the final calculation result. In the Map phase of MapReduce, transaction data is used as input to process each transaction data. Specifically, for each transaction data, the transaction subject pair information, that is, the initiator and recipient of the transaction, is extracted, and a key-value pair is generated with the transaction subject pair as the key and the number of transactions as the value. For a transaction data, the transaction subject A and the transaction subject B are extracted to generate a key-value pair: , indicating that a transaction occurred between transaction subject A and transaction subject B. The output result of the Map phase is a set of key-value pairs, where the key is the transaction subject pair and the value is the number of transactions that occurred for the transaction subject pair. In the Reduce phase, the key-value pairs output in the Map phase are grouped and aggregated using the transaction subject pair as the key. Specifically, for each transaction subject pair, the values of all the same keys output in the Map phase are accumulated to obtain the total number of transactions for the transaction subject pair within the preset time range.
[0050] For the transaction subject pair (A, B), the Reduce phase receives all the values of the same key output by the Map phase: (A, B) -> (1, 1, 1, ...), indicating that multiple transactions occurred between transaction subjects A and B in different Map tasks. The Reduce phase accumulates these values and obtains , indicating that a total of 10 transactions occurred between transaction subject A and transaction subject B within the preset time range. In order to optimize the computational efficiency of the Reduce phase, this application introduces a combiner function to locally summarize the transaction subject pairs at the Map end. , The combiner function is similar to the Reduce function, but it merges the values of the same key locally in the Map task, reducing the amount of data that needs to be transferred to the Reduce task.
[0051] In the local Map task, for the transaction subject pair (A, B), the combiner function locally summarizes the values of the same key (1, 1, 1) output by the Map task, and obtains (A, B) -> 3, which means that in this Map task, 3 transactions occurred between transaction subject A and transaction subject B. Through the local summary of the combiner function, the amount of data transmitted to the Reduce task can be significantly reduced, and the overall computing efficiency can be improved. The result output by the Reduce stage is the total transaction frequency of each transaction subject pair, that is, the key-value pair of (transaction subject pair, transaction frequency). In order to efficiently store and query these results, this application uses a column storage database for storage. Column storage databases are suitable for storing sparse and structured data. By organizing data by columns, efficient data compression and query performance can be achieved. In a column storage database, each transaction subject pair is a row, and the transaction frequency is the corresponding column value. Through this storage method, the transaction frequency can be easily queried, sorted, and aggregated, supporting rapid retrieval of the transaction frequency between transaction subjects.
[0052] In order to generate a list of frequent trading partners, the present application analyzes the transaction frequencies of all node pairs stored in the column storage database. The column storage database stores all transaction subject pairs and their corresponding transaction frequencies output by the Reduce phase. By analyzing these data, the trading partners with the highest transaction frequency can be identified. Specifically, the present application uses a maximum heap data structure to maintain the K node pairs with the highest current transaction frequency. A maximum heap is a binary tree structure in which the value of each node is greater than or equal to the value of its child node. Through the maximum heap, the top K largest elements, that is, the K transaction subject pairs with the highest transaction frequency, can be efficiently maintained.
[0053] The column storage database stores the following node pair transaction frequency data: If you want to maintain the top three node pairs with the highest transaction frequency (K=3), you can use the max heap data structure. Initially, the max heap is empty. Insert the node pairs and their transaction frequencies into the max heap one by one: , the maximum heap is ;insert , the maximum heap is [ ];insert , the maximum heap is [ ;insert , the maximum heap is [ ] (smaller than the top element of the heap, not inserted); insert , the maximum heap is [ ]. Finally, the maximum heap contains the top three node pairs with the highest transaction frequency: . In order to facilitate comparison and sorting, this application normalizes the transaction frequency and uses linear interpolation to map the transaction frequency to the range of [0, 1]. Normalization can eliminate the magnitude difference of transaction frequency between different node pairs, making the transaction frequency comparable. For the node pair (A, B), its transaction frequency is 10, the minimum transaction frequency is 5, and the maximum transaction frequency is 20. The normalized transaction frequency is: . The normalized transaction frequency is 0.333, indicating that among all node pairs, the transaction frequency of (A, B) is at a medium level. By maintaining the top K node pairs with the highest transaction frequency through the maximum heap and normalizing the transaction frequency, a list of frequent transaction partners can be obtained. The list of frequent transaction partners reflects the closeness and transaction activity between transaction entities, and provides important data support for subsequent abnormal transaction detection and risk assessment. The list of frequent transaction partners may contain the following contents: .
[0054] According to the list of frequent trading partners, the transaction frequency between two target trading entities can be quickly queried. For trading entity B and trading entity D, by querying the list of frequent trading partners, the normalized transaction frequency between them can be obtained as 0.667, indicating that they have a high transaction frequency within the preset time range. Finally, in order to intuitively present the association relationship and transaction behavior pattern between trading entities, this application uses the distributed graph computing framework GraphX to construct an association relationship topology structure graph. Based on the Spark distributed computing framework, GraphX provides a wealth of graph algorithms and operations, which can efficiently process large-scale graph data. When constructing the association relationship topology structure graph, the transaction subject is used as the node of the graph, the transaction behavior is used as the edge of the graph, and the transaction frequency is used as the additional weight of the edge. Through the graph algorithms and operations of GraphX, various analyses and calculations can be performed on the association relationship topology structure graph: calculate the degree centrality of the node to identify the key trading entity; calculate the shortest path between the nodes, and analyze the association path between the trading entities; perform community detection to find closely related groups of trading entities; calculate the PageRank value of the node to evaluate the importance and influence of the trading entity. The relationship topology diagram constructed by GraphX can intuitively display the relationship and transaction behavior patterns between transaction entities, laying the foundation for in-depth analysis of transaction data.
[0055] S5, according to the association relationship topology structure graph, use the graph mining algorithm to obtain the associated transaction path of the transaction data as the data analysis result output. Specifically, obtain the shortest transaction path between any two target nodes in the association relationship topology structure graph. The representation of the association relationship topology structure graph is as follows: there are n nodes in the association relationship topology structure graph, numbered from 1 to n. The adjacency matrix AdjMatrix[n][n] is used to represent the association relationship topology structure graph, where AdjMatrix[i][j] represents the transaction time on the edge from node i to node j. If there is no directly connected edge between node i and node j, AdjMatrix[i][j] is initialized to 0 or negative infinity (indicating unreachable). The adjacency list is represented by an array AdjList[n] of length n, where AdjList[i] stores a linked list or array containing nodes adjacent to node i and the transaction time on the corresponding edge.
[0056] Construction of the adjacency matrix: Initialize the adjacency matrix AdjMatrix[n][n] and set all elements to 0 or negative infinity. Traverse all edges Edge(i, j, time) in the association relationship topology graph, where i and j represent the start and end nodes of the edge, respectively, and time represents the transaction time on the edge. For each edge Edge(i, j, time), update the value of AdjMatrix[i][j] to time, which represents the transaction time on the edge from node i to node j. Modification of the adjacency matrix: Define a variable MaxTime to represent the maximum value of all transaction times. Traverse all elements of the adjacency matrix AdjMatrix[n][n]: If the value of AdjMatrix[i][j] is 0, it means that there is an edge with a transaction time of 0 between node i and node j, and replace it with MaxTime. If the value of AdjMatrix[i][j] is less than 0, it means that there is no direct connection between node i and node j, and replace it with 0. For example, there is a relationship topology graph with 4 nodes, the nodes are numbered from 1 to 4, and the edge information is as follows: Edge (1, 2, 5): the edge from node 1 to node 2, the transaction time is 5. Edge (1, 3, 3): the edge from node 1 to node 3, the transaction time is 3. Edge (2, 4, 7): the edge from node 2 to node 4, the transaction time is 7. Edge (3, 4, 0): the edge from node 3 to node 4, the transaction time is 0. Construct the adjacency matrix AdjMatrix[4][4]: Initialize all elements to 0 or negative infinity. Update the matrix elements according to the edge information: AdjMatrix[1][2]=5; AdjMatrix[1][3]=3; AdjMatrix[2][4]=7; AdjMatrix[3][4]=0; MaxTime is 10, and modify the adjacency matrix: the value of AdjMatrix[3][4] is 0, replace it with MaxTime, that is, AdjMatrix[3][4]=10. The final adjacency matrix is:
[0057] By representing and modifying the adjacency matrix, the association relationship topology graph can be converted into a data structure that is convenient for shortest path calculation, and special cases can be handled, such as edges with transaction time of 0 and node pairs without direct connection. Dijkstra shortest path algorithm, input data: modified adjacency matrix AdjMatrix[n][n], where n is the number of nodes. Source node number source and target node number target. Initialize data structure: create a distance array dist[n] to store the shortest distance from the source node to other nodes. Initialize dist[source] to 0, indicating that the distance from the source node to itself is 0. Initialize dist[i] (i≠source) to the value of AdjMatrix[source][i], indicating the initial distance from the source node to other nodes. Create a predecessor node array prev[n] to store the predecessor node of each node, with an initial value of -1, indicating an unknown predecessor. Create a minimum priority queue pq to store the node number and the corresponding distance value. Algorithm execution process: Add the source node source to the minimum priority queue pq with a distance value of 0. When pq is not empty, repeat the following steps: Take out the node u with the smallest distance from pq. If u is the target node target, the algorithm ends and jumps to step 4. Traverse all nodes v directly connected to node u: calculate the distance from the source node to node v through node u, that is, dist[u]+AdjMatrix[u][v]. If the distance is less than the current value of dist[v], update dist[v] to the distance and update prev[v] to u. If node v is not in pq, add it to pq with the distance value updated to dist[v]. If node v is already in pq, update its distance value in pq to the updated dist[v].
[0058] Construction of the shortest path: backtrack the shortest path through the predecessor node array prev. Starting from the target node target, keep searching for its predecessor nodes until the source node source. Store the found nodes in an array path in the order from the source node to the target node. The array path is the sequence of nodes on the shortest path from the source node to the target node. The length of the shortest path is the value of dist[target]. For example: Taking the above-mentioned corrected adjacency matrix as an example, the source node is 1 and the target node is 4. Initialize the data structure: dist[4]=[0,∞,∞,∞]; prev[4]=[-1,-1,-1,-1]; pq=[(1,0)]. Algorithm execution process: Take out node 1, update the distance of the nodes connected to node 1: dist[2]=min(∞,0+5)=5, prev[2]=1; dist[3]=min(∞,0+3)=3, prev[3]=1; pq=[(3,3),(2,5)]; Take out node 3, update the distance of the nodes connected to node 3: dist[4]=min(∞,3+10)=13, prev[4]=3; pq=[(2,5),(4,13)]; Take out node 2, update the distance of the nodes connected to node 2: dist[4]=min(13,5+7)=12, prev[4]=2, pq=[(4,12)], take out node 4, which is the target node, and the algorithm ends.
[0059] Construction of the shortest path: path = [4, 2, 1], the shortest path length is dist[4] = 12, therefore, the shortest transaction path from node 1 to node 4 is , the shortest path length is 12. The Dijkstra algorithm can be used to effectively calculate the shortest transaction path between any two target nodes in the association relationship topology diagram. The algorithm maintains the distance array and the predecessor node array, continuously updates the shortest distance between nodes, and uses the minimum priority queue to select the node with the shortest distance for expansion, and finally obtains the shortest path and path length from the source node to the target node.
[0060] Construction of the shortest path subgraph: According to the shortest path node sequence path[k] obtained by the Dijkstra algorithm, where k is the number of nodes on the shortest path. Create a new graph data structure SubGraph to store the shortest path subgraph. Traverse the shortest path node sequence path: For each pair of adjacent nodes path[i] and path[i + 1] (0 ≤ i < k - 1) in the sequence: Search for the edge edge(path[i], path[i + 1]) connecting path[i] and path[i + 1] and its transaction time time in the original associated relationship topology graph. Add the nodes path[i], path[i + 1] and the edge edge(path[i], path[i + 1]) to the shortest path subgraph SubGraph. In SubGraph, set the transaction time of the edge between nodes path[i] and path[i + 1] to time. The finally obtained SubGraph is the shortest path subgraph, which contains all the nodes and edges on the shortest path.
[0061] Sorting of the shortest path edge sequence: Create an edge sequence array edgeList to store all edges in the shortest path subgraph SubGraph. Traverse all edges in SubGraph: For each edge edge(u, v), add it to edgeList and record the edge's starting node u, ending node v, and transaction time time. Sort edgeList using an appropriate sorting algorithm, such as quick sort or merge sort. The comparison condition for sorting is the edge's transaction time time, which is arranged in ascending order. The edgeList obtained after sorting is the shortest path edge sequence, which is arranged in ascending order of transaction time. Construction of the shortest transaction sequence: According to the sorted shortest path edge sequence edgeList, extract the source node source and the target node target. Create a shortest transaction sequence array tradeSequence to store the nodes and edges of the shortest transaction sequence. Initialize tradeSequence and add the source node source to the start of the sequence. Traverse the sorted shortest path edge sequence edgeList: For each edge edge(u, v), add its ending node v to tradeSequence. In tradeSequence, the edge between node u and node v represents a transaction, and the transaction time is the transaction time time of edge (u, v). The final tradeSequence is the shortest transaction sequence, which contains the nodes and edges on the shortest path from the source node to the target node, and is arranged in the order of transaction time. For example: Taking the above shortest path as an example, the shortest path node sequence is path=[1, 2, 4]. Construction of the shortest path subgraph: Create SubGraph. Add nodes 1, 2 and edge edge (1, 2) to SubGraph, and the transaction time is 5. Add nodes 2, 4 and edge edge (2, 4) to SubGraph, and the transaction time is 7. Sorting of the shortest path edge sequence: Initialize edgeList=[edge (1, 2), edge (2, 4)]. Sort edgeList by transaction time to get edgeList=[edge (1, 2), edge (2, 4)].
[0062] Construction of the shortest transaction sequence: the source node is 1 and the target node is 4. Initialize tradeSequence=[1]. Traverse edgeList: edge (1, 2), add node 2 to tradeSequence, and get tradeSequence=[1, 2]. Edge edge (2, 4), add node 4 to tradeSequence, and get tradeSequence=[1, 2, 4]. The final shortest transaction sequence is tradeSequence=[1, 2, 4], which represents the shortest transaction path from node 1 to node 4, arranged in the order of transaction time. According to the shortest path node sequence obtained by the Dijkstra algorithm, the shortest path subgraph is constructed, and the shortest path edge sequence is sorted to finally obtain the shortest transaction sequence. The shortest path subgraph contains all nodes and edges on the shortest path. The shortest path edge sequence is arranged from small to large according to the transaction time. The shortest transaction sequence shows the shortest transaction path from the source node to the target node in the order of transaction time. Perform pattern matching search on the shortest transaction sequence of the target node, preset frequent transaction patterns and construction of AC automaton: Suppose there are m preset frequent transaction patterns, which are . Each frequent trading pattern As a string, where each transaction is a character. Construct a set of pattern strings for AC automata Based on the pattern string set P, construct the Go-to function goto and Failure function fail of the AC automaton: goto(state, character) indicates the next state to be transferred to after receiving the character character from state state. fail(state) indicates the next state to be transferred to after the state state fails to match. Processing of the shortest transaction sequence of the target node: Let the shortest transaction sequence of the target node be sequence, which contains n transactions and is expressed as . Treat the shortest transaction sequence sequence as a target text string , where each transaction As a character.
[0063] AC automaton multi-pattern matching search: Using the constructed AC automaton, the target text string Perform multi-pattern matching search. Assume that the initial state of the AC automaton is , the current state is current_state, initialized to . Traverse the target text string Each character in : In the AC automaton, from the current state Start, receive characters , transfer to the next state .if If it does not exist, the next state to be transferred is found according to the Failure function fail, until a legal state is found or the initial state is returned. . Update current status for .if is a pattern string The termination status of the pattern string is recorded in the target text string. Frequent transaction pattern location interval extraction: For each frequent transaction pattern that appears in the target text string , according to its end position , use the backtracking algorithm to find its starting position : Initialize the current state is the pattern string The final state. From the final position Start by backtracking along the Failure function fail of the AC automaton until you reach the pattern string The starting state or initial state . Record the position corresponding to each state passed during the backtracking process and construct a position sequence . Pattern string Starting position for For each frequent trading pattern , construct its position interval , indicating the range of positions where the pattern can appear in the target text string.
[0064] Extraction of target transaction segments: Extract the location intervals of all frequent transaction patterns Merge to get the merged position interval set For each merged position interval : Extract subsequence from the shortest transaction sequence sequence of the target node . Take the subsequence subsequence as a target transaction fragment and add it to the target transaction fragment set middle.
[0065] The multi-pattern matching search for the shortest transaction sequence of the target node is realized by using the Aho-Corasick automaton. First, the preset frequent transaction patterns are constructed into the set of pattern strings of the Aho-Corasick automaton, and the corresponding Go-to function and Failure function are constructed. Then, the shortest transaction sequence of the target node is regarded as the target text string, and the Aho-Corasick automaton is used to perform multi-pattern matching search on it to obtain the termination positions of each frequent transaction pattern in the target text string. Next, the backtracking algorithm is adopted. Starting from the termination position, the matching path is backtracked along the Failure function of the Aho-Corasick automaton to find the starting positions of each frequent transaction pattern, and the set of position intervals of the frequent transaction patterns is constructed. Finally, the position intervals are merged, and the target transaction segments are extracted from the shortest transaction sequence according to the merged position intervals.
[0066] Calculating the time tightness and transaction activity of the target transaction segment, extraction of the target transaction segment: According to the set of target transaction segments obtained by pattern matching search , select a target transaction segment fragment for analysis. Let the target transaction segment fragment contain N transactions, denoted as . Each transaction contains the transaction time and the transaction amount two attributes. Calculation of time tightness: Let the preset time threshold be TimeThreshold. When the time difference between two adjacent transactions is less than this threshold, these two transactions are considered to be time-tight. Initialize the time tightness TimeDensity to 0, and initialize the number of transactions with a time difference less than the threshold to 0. Traverse each pair of adjacent transactions and in the target transaction segment fragment, where 1 ≤ i < N: Calculate the time difference between the two transactions. If the time difference is less than the time threshold TimeThreshold, then increment by 1, indicating that a pair of time-tight transactions is found. According to the calculation formula of time tightness, calculate the time tightness of the target transaction segment.
[0067] Calculation of transaction activity: Let the preset amount threshold be AmountThreshold. When the transaction amount is greater than this threshold, this transaction is considered to be active. Initialize the transaction activity TransActivity to 0, and initialize the number of transactions with a transaction amount greater than the threshold to 0. Traverse each transaction in the target transaction segment fragment, where 1 ≤ i ≤ N: Obtain the transaction amount If the transaction amount If the amount is greater than the amount threshold AmountThreshold, Add 1 to indicate that an active transaction is found. Calculate the transaction activity of the target transaction segment according to the transaction activity calculation formula . Result output: Output the time density TimeDensity and transaction activity TransActivity of the target transaction fragment. Time density TimeDensity indicates the proportion of transactions that are dense in time within the target transaction fragment. The value range is [0, 1]. The larger the value, the denser the transactions are in time. Transaction activity TransActivity indicates the proportion of active transactions within the target transaction fragment. The value range is [0, 1]. The larger the value, the more active the transactions are. For example: The target transaction fragment contains 5 transactions, and the transaction time and transaction amount are as follows: : ; : ; : ; : ; : ; Set the time threshold TimeThreshold to 10 seconds and the amount threshold AmountThreshold to 1000.
[0068] Calculate time closeness: time difference between adjacent transactions: ; ; ; The number of transactions with a time difference less than the threshold , time tightness . Calculate transaction activity: the number of transactions with transaction amounts greater than the threshold , transaction activity Therefore, the time density of the target transaction segment is 0.5, and the transaction activity is 0.6. Quantitatively evaluate the density and activity of transactions in terms of time and amount. Time density reflects the concentration of transactions in time, and transaction activity reflects the activity of transactions in terms of amount.
[0069] Calculate the confidence of the target transaction fragment and identify suspicious transaction fragments. Calculate the confidence of the target transaction fragment: For each target transaction fragment , where 1≤i≤M (M is the total number of target transaction fragments), calculate its confidence . Based on the time density calculated previously and transaction activity , use the confidence calculation formula to calculate the confidence: Confidence The value range of is [0, 1]. The larger the value, the higher the suspicious degree of the target transaction segment. Identification of suspicious transaction segments: Let the preset confidence threshold be ConfidenceThreshold, which means that the target transaction segment with confidence lower than the threshold is regarded as a suspicious transaction segment. Traverse each target transaction segment : If the confidence If the value is less than the confidence threshold ConfidenceThreshold, the target transaction segment Mark suspicious transaction segments. Add to Suspicious Transaction Snippet Collection Result output: Output each target transaction fragment Confidence Output suspicious transaction fragment set , which contains all target transaction fragments whose confidence is lower than the threshold.
[0070] In this implementation, there are 3 target transaction segments, and their time density and transaction activity are as follows: : , ; : , ; : , ; Set the confidence threshold ConfidenceThreshold to 0.5. Calculate the confidence of each target transaction segment: ; ; .
[0071] Identify suspicious transaction fragments: Confidence ,therefore Transaction fragments marked as suspicious. Confidence ,therefore Transaction fragments marked as suspicious. Confidence ,therefore Marked as suspicious transaction fragments. Finally, the collection of suspicious transaction fragments is obtained. The confidence of the target transaction fragment is calculated based on its time density and transaction activity, and suspicious transaction fragments are identified based on the preset confidence threshold. The confidence comprehensively considers the density and activity of the transaction in terms of time and amount. The smaller the value, the more suspicious the transaction fragment is. By setting an appropriate confidence threshold, transaction fragments with low confidence can be screened out as suspicious transaction fragments for focus and further analysis.
[0072] Perform community mining on suspicious transaction fragments and construct suspicious transaction subgraphs: For each suspicious transaction fragment , where 1≤i≤K (K is the total number of suspicious transaction fragments), extract the nodes and edges it contains. The extracted nodes and edges are constructed into a suspicious transaction subgraph , expressed as ,in Represents a collection of nodes. Represents the edge set. Construction of the adjacency matrix: For each suspicious transaction subgraph , convert it into a weighted adjacency matrix representation . Adjacency Matrix The dimension is ,in Represents a suspicious transaction subgraph The number of nodes in the adjacency matrix Elements in Representation Node and nodes The sum of the transaction amounts between , where e represents the connection node and The edge, Represents the transaction amount of edge e.
[0073] Community mining: Louvain community mining algorithm is used to mine each suspicious transaction subgraph. Perform community mining. Initialize each node Belong to different societies , that is, each node forms a community. Iterate the following sub-steps until the community division does not change: traverse the suspicious transaction subgraph Each node in :Compute node The change in modularity after joining the community of its neighbor node , where k represents the neighbor node The community number you are in. Choose the community that can improve the modularity the most As a node The target community is , ; The node Divide into target groups Update the adjacency matrix representation of the community , merge nodes in the same community into a supernode , the edge weight between super nodes is the sum of the edge weights between nodes within the community. Traversing the adjacency matrix Each community in :If the community If there are connected sub-communities inside, these sub-communities will be merged into a new community . Update the merged community adjacency matrix representation .
[0074] Community feature calculation: For each community , calculate its internal features: number of nodes: community The number of nodes contained inside is expressed as . Edge density: community The ratio of the actual number of edges between internal nodes to the maximum possible number of edges is expressed as ,in Indicates the community The number of edges that actually exist inside. The total transaction amount: community The sum of transaction amounts between all internal nodes is expressed as ,in Indicates the community Internal Node and The transaction amount between. The average transaction amount: community The average transaction amount between all internal nodes is expressed as . Variance of transaction amount: Community The mean of the sum of squares of the differences between the transaction amounts and the mean between all internal nodes is expressed as .
[0075] Identification of suspicious transaction groups: Set the preset transaction amount ratio threshold as ThresholdRatio. Traverse each group :Computing Society The sum of transaction amounts of internal nodes Accounting Society The total amount of all internal transactions The ratio of sum(G_i) If the ratio If it is greater than the preset threshold ThresholdRatio, the community Mark as a suspicious transaction community. Result output: Output the information of each suspicious transaction community, including: community number and identification information of internal nodes. Identification information and weight information of the edges within the community. Statistical features related to the transaction amount of the nodes within the community, such as sum, mean, variance, etc.
[0076] For example, there is a suspicious transaction subgraph G, which contains 5 nodes and 7 edges. The node identifiers are , the transaction amount of the edge is as follows: : , amount=1000; : , amount=800; : , amount=1200; : , amount=600; : , amount=1500; : , amount=900; : , amount=1100. The adjacency matrix A is expressed as:
[0077] After community mining, the following two communities may be obtained: : Contains nodes , the internal edge is . Societies : Contains nodes , the internal edge is . Societies Features: Number of nodes: , edge density: ; Total transaction amount: ; Average transaction amount: ; Variance of transaction amount: .
[0078] Societies Features: Number of nodes: ; Edge density: ; Total transaction amount: ; Average transaction amount: ; Variance of transaction amount: ; The preset transaction amount ratio threshold ThresholdRatio is 0.4, which can calculate the transaction amount ratio of the cada community: ; Therefore, the community was marked as a suspicious transaction community, and the community Not marked as a suspicious transaction community. Perform community mining on suspicious transaction fragments to identify suspicious transaction communities. The Louvain community mining algorithm achieves community division by optimizing modularity, and can effectively discover closely related node communities. By calculating the characteristics and transaction amount ratio within the community, suspicious communities with a high proportion of transaction amounts can be further screened out as key focus objects.
[0079] According to the association relationship topology structure diagram, the graph mining algorithm is used to obtain the associated transaction path of the transaction data as the output of the data analysis result. The Dijkstra shortest path algorithm is used to obtain the shortest transaction path between any two target nodes in the association relationship topology structure diagram, and the obtained shortest transaction path is arranged in the order of the transaction occurrence to construct the shortest transaction sequence of the target node. Specifically, the association relationship topology structure diagram is converted into an adjacency matrix representation, the nodes in the association relationship topology structure diagram are used as the rows and columns of the adjacency matrix, and the transaction time on the edge of the association relationship topology structure diagram is used as the element value of the corresponding position in the adjacency matrix to obtain the adjacency matrix representing the association relationship topology structure diagram; in the adjacency matrix, the element value with a transaction time of 0 is replaced by a maximum value greater than all transaction times, and the elements in the matrix less than 0 are replaced by 0 to obtain the corrected adjacency matrix; the corrected adjacency matrix is used as the input of the Dijkstra algorithm, and the Dijkstra algorithm is used to calculate the shortest path between any two target nodes in the adjacency matrix to obtain. A shortest path result including a node sequence and a path length on the shortest path; according to the node sequence in the shortest path result, corresponding nodes and edges are extracted from the association relationship topology structure diagram to construct a shortest path subgraph; the edges in the shortest path subgraph are traversed, and the edges in the shortest path subgraph are sorted in ascending order of transaction time according to the transaction time on the edges to obtain a sorted shortest path edge sequence; according to the sorted shortest path edge sequence, the source node and the target node in the shortest path edge sequence are extracted, and the shortest transaction sequence between the target nodes is constructed according to the order of the shortest path edges; the constructed shortest transaction sequence of the target node is used as input data for a subsequent transaction pattern matching algorithm.
[0080] An AC automaton pattern matching algorithm is adopted, and a preset frequent transaction pattern is used as a matching template to perform a matching search on the shortest transaction sequence of a target node to obtain a target transaction fragment that meets the matching template; specifically, each transaction in the preset frequent transaction pattern is taken as a character to construct a pattern string set of the AC automaton, and a Go-to function and a Failure function of the AC automaton are constructed based on the pattern string set; each transaction in the shortest transaction sequence of the target node is taken as a character to construct a target text string of the AC automaton; using the constructed AC automaton, a multi-pattern matching search is performed on the target text string to obtain the termination position information of all frequent transaction patterns appearing in the target text string; according to the termination position information of the frequent transaction pattern, a backtracking algorithm is adopted to start from the termination position and backtrack the matching path along the Failure function of the AC automaton until the starting position of the matching path, to obtain the starting position information of all frequent transaction patterns ending with the termination position, and to construct a position interval set of the frequent transaction pattern; overlapping position intervals in the position interval set are merged to obtain a position interval set of the merged target transaction fragment; according to the position interval set of the target transaction fragment, the target transaction is extracted from the shortest transaction sequence of the target node.
[0081] Calculate the time difference between every two adjacent transactions in the target transaction segment, count the number of transactions with a time difference less than a threshold, and use the number of transactions with a time difference less than the threshold as the time density of the target transaction segment; calculate the transaction amount of each transaction in the target transaction segment, count the number of transactions with a transaction amount greater than the threshold, and use the number of transactions with a transaction amount greater than the threshold as the transaction activity of the target transaction segment; ,in, Indicates the time difference between the i-th transaction and the i+1-th transaction in the target transaction segment; represents the transaction time of the i-th transaction in the target transaction segment; Indicates the transaction time of the i+1th transaction in the target transaction segment. Calculate the confidence of the target transaction segment based on the time density and transaction activity of the target transaction segment; , where Confidence represents the confidence of the target transaction fragment, and its value range is [0, 1]; Indicates the temporal closeness of the target transaction fragment; Indicates the transaction activity of the target transaction segment; ,in, Indicates that the time difference is less than the preset time threshold The number of transactions, N represents the total number of transactions in the target transaction segment; ,in, Indicates that the transaction amount is greater than the preset amount threshold Number of transactions; Indicates the transaction amount of the i-th transaction in the target transaction segment. Define the target transaction segment with a confidence level lower than the threshold as a suspicious transaction segment; extract the nodes and edges contained in each suspicious transaction segment in the association relationship topology graph to construct a suspicious transaction subgraph.
[0082] The Louvain community mining algorithm is used to conduct community mining on the constructed suspicious transaction subgraph, and the community mining results are output as data analysis results. Specifically, the suspicious transaction subgraph is converted into a weighted adjacency matrix representation, and the elements in the adjacency matrix represent the sum of the transaction amounts between the corresponding nodes in the suspicious transaction subgraph, and the initial community division is obtained, in which each node constitutes an initial community; based on the adjacency matrix, the following sub-steps are iteratively performed until the community division no longer changes: traverse each node in the suspicious transaction subgraph, calculate the change in modularity after each node joins the community where its neighboring nodes are located, and select the community that can maximize the modularity improvement as the target community of the node; according to the target community of each node, all nodes are re-divided into corresponding communities, and the adjacency matrix representation of the community is updated according to the division results; traverse each community in the adjacency matrix, merge the sub-communities connected within the community into one community, and update the adjacency matrix of the merged community Representation; based on the adjacency matrix representation of the community, calculate the internal characteristics of each community, including: counting the number of nodes within the community; counting the number of edge connections between nodes within the community, and calculating the edge density between nodes within the community; counting the transaction amounts of nodes within the community, and calculating the sum, mean, variance and other statistical characteristics of the transaction amounts within the community; traversing each community, calculating the ratio of the sum of the transaction amounts of nodes within the community to the total of all transaction amounts within the community, and marking the community with a ratio greater than a preset threshold as a suspicious transaction community; outputting the information of each suspicious transaction community as the result of the suspicious transaction pattern analysis, including: identification information of all nodes within the suspicious transaction community; identification information and weight information of all edges within the suspicious transaction community; statistical characteristics related to the transaction amounts of nodes within the suspicious transaction community.
Claims
1. A big data analysis method based on blockchain, characterized in that: include: Get transaction data; Preprocess the acquired transaction data to obtain structured data; Use hash functions to encode and map structured data to generate a unique identifier for the corresponding data; According to the Merkle tree hash structure, the unique identifier is used as the leaf node, and the Merkle tree root hash value is iteratively generated from bottom to top. The root hash value is used as a proof token for data integrity verification, including: Arrange the unique identifiers in a predefined order as the hash values of the Merkle tree child nodes; Pair the hash values of two adjacent leaf nodes, generate the corresponding parent node hash value through joint hash calculation, and arrange the parent node hash values in a predefined order as the input for the next round of hash calculation; Repeat the pairing and hash calculation steps until the hash value of the Merkle tree root node is generated, and the hash value of the root node is used as the Merkle proof of the entire transaction data; Combine the root hash value of the Merkle proof with the height and timestamp of the current block, perform a hash operation, and use the resulting hash value as the transaction integrity verification token of the current block; In the next block, the hash value of the transaction integrity verification token of the previous block is embedded in the block header, and hashed with the current block header information to generate a new transaction integrity verification token; The proof tokens are linked and stored through the hash chain mechanism, and the hash value of each token is recorded on the chain of the next block; A hash-based indexing method is used to map the storage location of structured data on the blockchain using the hash value of the proof token to generate a blockchain index; Using the distributed file system IPFS, the structured data and the corresponding proof tokens are uploaded to the blockchain node; The distributed graph computing framework GraphX is used to perform graph modeling on the stored structured data. The transaction subjects are used as graph nodes, and the transaction behaviors are used as directed edges to construct a topological structure diagram of the association relationship between the transaction subjects. According to the association relationship topology diagram, the graph mining algorithm is used to obtain the associated transaction path of the transaction data and output it as the data analysis result.
2. The big data analysis method based on blockchain according to claim 1 is characterized in that: Preprocess the acquired transaction data to obtain structured data; Use hash functions to encode and map structured data to generate unique identifiers for the corresponding data, including: Perform field extraction on the acquired transaction data to extract the transaction amount, transaction time and transaction subject; The transaction amount data is converted into a floating point format using a regular expression matching method to obtain standardized transaction amount data; Use the date and time parsing function to convert the transaction time data into the standard date and time format to obtain the standardized transaction time data; The transaction subject data is parsed using a string matching method, the subject identification substring is extracted, and the substring is mapped to an integer code to obtain the encoded transaction subject ID; The standardized transaction amount, standardized transaction time and transaction subject ID are concatenated to obtain structured data; According to the structured data, hash calculation is performed using the hash algorithm SHA-256 to obtain the hash value as the unique identifier of the corresponding transaction data.
3. The big data analysis method based on blockchain according to claim 1 is characterized in that: The proof token is linked and stored through a hash chain mechanism, including: Parse the transaction integrity verification token, extract the root hash value of the Merkle proof, and form a key-value pair with the root hash value and the unique identifier of the corresponding transaction data; Through the hash index method, the root hash value in the key-value pair is used to map the storage location of the transaction data on the blockchain to build a blockchain index; Using IPFS’s content-addressed storage model, structured data and corresponding proof tokens are uploaded to the blockchain node group; The hash value of the unique identifier of the structured data is used as the content addressing identifier of the corresponding data in the IPFS network.
4. The big data analysis method based on blockchain according to claim 1 is characterized in that: Construct a topological structure diagram of the relationship between transaction entities, including: Obtain multiple transaction data within a preset time range from the blockchain index, and obtain the corresponding structured transaction data through the content addressing mechanism of IPFS based on the unique identifier of the transaction data; Perform graphic analysis on the acquired structured transaction data and extract the transaction subject ID in each transaction; A hash algorithm is used to perform a hash operation on the transaction subject ID, and the hash value is used as the identity fingerprint of the transaction subject, and the identity fingerprint is used as a node of the association relationship topology diagram; Perform graphical analysis on the acquired structured transaction data to extract the transaction amount, transaction time and transaction direction of each transaction; The time window division algorithm is used to divide the transaction time, and the transaction time is marked with a serial number according to the time window in which the transaction occurs; The transaction amount is grouped and coded using the amount interval discretization method, and the transaction amount is mapped according to the preset amount interval; The transaction time marked with the serial number, as well as the mapped transaction amount and transaction direction, are used as the edges of the association relationship topology structure graph; Set transaction frequency as the additional weight of the edge to count the number of transactions between two target transaction entities within the preset time range; Construct a topological structure diagram of the association relationship based on nodes, edges and additional weights.
5. The big data analysis method based on blockchain according to claim 4 is characterized in that: Set transaction frequency as an additional weight for the edge, including: Traverse all transaction data within the preset time range, use the MapReduce distributed computing framework to process the transaction data in parallel, and count the number of transactions between each pair of transaction entities; the Map stage uses the transaction data as input to extract the transaction entity pairs and the corresponding transaction times, and the Reduce stage uses the transaction entity pairs as keys to summarize the transaction times output by each Map task to obtain the total transaction times for each pair of transaction entities; In the Reduce phase, the combiner function is used to merge and calculate the transaction subject pairs on the Map side to obtain the cumulative transaction frequency between the node pairs. A column storage database is used to store the intermediate results of the node pair transaction frequency output in the Reduce phase. Each node pair represents two transaction entities, and the frequency value is used as an additional attribute of the edge. The transaction frequencies of all node pairs stored in the column-type storage database are maintained through the maximum heap data structure to maintain the K node pairs with the highest current transaction frequencies. The transaction frequencies are normalized using linear interpolation to generate a list of frequent transaction partners between transaction entities. According to the node pairs in the frequent transaction partner list and the corresponding normalized transaction frequency, the number of transactions between two target transaction entities within a preset time range is obtained.
6. The big data analysis method based on blockchain according to claim 5 is characterized in that: Using graph mining algorithms, we can obtain the associated transaction paths of transaction data and output them as data analysis results, including: The Dijkstra shortest path algorithm is used to obtain the shortest transaction path between any two target nodes in the association relationship topology graph, and the obtained shortest transaction paths are arranged in the order of the transaction occurrence to construct the shortest transaction sequence of the target node; Adopting AC automaton pattern matching algorithm, taking preset frequent transaction patterns as matching templates, matching search is performed on the shortest transaction sequence of the target node to obtain the target transaction fragments that meet the matching templates; Calculate the time difference between every two adjacent transactions in the target transaction segment, count the number of transactions with a time difference less than a threshold, and take the number of transactions with a time difference less than the threshold as the time density of the target transaction segment; Calculate the transaction amount of each transaction in the target transaction segment respectively, count the number of transactions with transaction amounts greater than the threshold, and use the number of transactions with transaction amounts greater than the threshold as the transaction activity of the target transaction segment; Calculate the confidence of the target transaction segment based on the time density and transaction activity of the target transaction segment; The target transaction segment with a confidence level lower than a threshold is defined as a suspicious transaction segment; In the association relationship topology structure graph, extract the nodes and edges contained in each suspicious transaction fragment to construct a suspicious transaction subgraph; The Louvain community mining algorithm is used to conduct community mining on the constructed suspicious transaction subgraph, and the community mining results are output as data analysis results.
7. The big data analysis method based on blockchain according to claim 6 is characterized in that: Calculate the time difference between every two adjacent transactions in the target transaction segment using the following formula: in, Indicates the time difference between the i-th transaction and the i+1-th transaction in the target transaction segment; represents the transaction time of the i-th transaction in the target transaction segment; Indicates the transaction time of the i+1th transaction in the target transaction segment.
8. The big data analysis method based on blockchain according to claim 7 is characterized in that: Calculate the confidence of the target transaction fragment using the following formula: Among them, Confidence represents the confidence of the target transaction fragment, and its value range is [0, 1]; Indicates the temporal closeness of the target transaction fragment; Indicates the transaction activity of the target transaction segment; in, Indicates that the time difference is less than the preset time threshold The number of transactions, N represents the total number of transactions in the target transaction segment; in, Indicates that the transaction amount is greater than the preset amount threshold Number of transactions; Indicates the transaction amount of the i-th transaction in the target transaction segment.
9. A big data analysis system based on blockchain, characterized in that: include: At least one processing unit; used to execute instructions to implement the blockchain-based big data analysis method as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Financial big data-oriented multi-way tree structure block chain integrated optimization storage method
CN111611315A
Digital archive system based on block chain
CN118350047A