Abnormal transaction behavior detection method and system based on block chain
By combining the DBSCAN algorithm and the isolated forest algorithm, blockchain transaction data is preprocessed and abnormal detection, which solves the problem of difficult detection of abnormal transaction behavior in the blockchain network, and realizes efficient and accurate abnormal transaction behavior detection, ensuring the security and reliability of blockchain transactions.
Patent Information
- Application Number
- CN202510127003.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-27
- Publication Date
- 2025-06-24
AI Technical Summary
In blockchain networks, abnormal trading behaviors exist, such as market manipulation, fraud and privacy violations, especially token airdrops, which may be maliciously exploited, making it difficult to effectively detect and prevent.
Combined with the DBSCAN algorithm and the isolated forest algorithm, the transaction data is preprocessed, clustered analysis and abnormal detection, and potential abnormal transaction behaviors are identified and confirmed. The DBSCAN algorithm is used to identify high-density areas and potential outliers, while the isolated forest algorithm further refines anomaly detection through random tree construction and path length evaluation.
It realizes efficient and accurate detection of abnormal transaction behaviors in the blockchain network, expands the coverage of abnormal detection, improves the accuracy and efficiency of detection, and ensures the security and reliability of blockchain transactions.
Smart Images

Figure CN120197084A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for detecting abnormal transaction behaviors based on blockchain, and also relates to a corresponding system for detecting abnormal transaction behaviors, belonging to the technical field of blockchain. Background Art
[0002] Blockchain is a distributed ledger system, with its core features of decentralization, transparency, and immutability. It forms a continuous chain structure by linking data in the form of blocks in chronological order. Each block contains a series of transaction records and related metadata, such as timestamps and the hash value of the previous block, ensuring the integrity and security of the data. The consensus mechanisms of blockchain, such as Proof of Work (PoW) and Proof of Stake (PoS), further guarantee the consistency of the ledger state among all participants in the network. Blockchain was initially designed for Bitcoin but has now been extended to finance and many other fields, providing a secure and transparent way of transaction and data management with its unique features.
[0003] In the context of the openness and anonymity of the blockchain network, in addition to normal transaction activities, there are also some abnormal transaction behaviors. These behaviors may involve issues such as market manipulation, fraud, and privacy infringement, posing threats to the health of the blockchain ecosystem and the asset security of users. Among them, a special promotion activity - distributing tokens to users for free to increase project awareness and user participation - may also be maliciously exploited. Its original intention is to reward users or attract new users, but it may also become a tool for lawbreakers to conduct fraud and manipulate the market.
[0004] Regarding the abnormal trading behaviors that token airdrops may trigger, researchers have developed various detection algorithms to identify and prevent these risks. These algorithms include density-based clustering algorithms (such as the DBSCAN algorithm), which can identify abnormal trading patterns based on the density characteristics of transactions; and isolation-based anomaly detection algorithms (such as the isolation forest algorithm), which can quickly identify anomaly points deviating from normal trading patterns by constructing random trees. The application of these algorithms helps to maintain the security and reliability of the blockchain network and protect users from potential damages caused by abnormal trading behaviors. Through these technical means, fraud in token airdrops can be effectively detected and prevented, ensuring the healthy application of blockchain technology and the trading security of users. In the Chinese invention patent with the patent number ZL202210248751.X, a method for detecting abnormal trading behaviors based on subgraph matching is disclosed. The method includes: processing and parsing the detailed historical transaction data of Ethereum, constructing a transaction dataset using the transaction data; extracting the behavioral characteristics of Ethereum abnormal transactions based on the transaction input address, transaction output address, transaction timestamp, and transaction amount information in the transaction dataset, and constructing an Ethereum transaction flow graph; formulating matching rules corresponding to various abnormal trading behaviors according to the Ethereum abnormal transaction behavioral characteristics; using the characteristic subgraphs of various Ethereum abnormal transactions to detect the Ethereum transaction flow graph according to the matching rules, and obtaining the Ethereum abnormal trading behaviors in the Ethereum transaction flow graph based on the detection results. Summary of the Invention
[0005] The primary technical problem to be solved by the present invention is to provide a method for detecting abnormal trading behaviors based on blockchain.
[0006] Another technical problem to be solved by the present invention is to provide a system for detecting abnormal trading behaviors based on blockchain.
[0007] To achieve the above technical objectives, the present invention adopts the following technical solutions:
[0008] According to the first aspect of the embodiments of the present invention, a method for detecting abnormal trading behaviors based on blockchain is provided, including the following steps:
[0009] Step 1: Select and preprocess the features in the transaction data, including cyclically encoding the transaction time, normalizing the features, and assigning weights according to the contribution of the features to abnormal trading behaviors; then, perform clustering analysis using the DBSCAN algorithm, divide the data points into clusters through the neighborhood radius and the minimum number of points, and the points that do not meet the density requirements are identified as noise points and regarded as potential abnormal transactions;
[0010] Step 2: Adopt a method of dynamically adjusting the neighborhood radius to optimize the clustering effect of the DBSCAN algorithm according to the quantity and density changes of the transaction data, ensuring that abnormal trading behaviors can be effectively identified in high-density areas;
[0011] Step 3: In the detection of abnormal trading behaviors, combine multiple features of the transactions and perform comprehensive matching when expanding clusters to identify abnormal trading behaviors that are similar in multiple dimensions, thereby enhancing the robustness of the clustering algorithm;
[0012] Step 4: On the basis of the DBSCAN algorithm, further use the isolation forest algorithm to identify and confirm potential abnormal points; split the data by randomly selecting features and split points, and recursively construct multiple isolation trees until the stopping condition is met; calculate the path length of the sample from the root node to the leaf node in the isolation tree, evaluate the abnormal score of the sample, and thus identify abnormal trading behaviors.
[0013] According to the second aspect of the embodiments of the present invention, there is provided an abnormal trading behavior detection system based on a blockchain, including a processor and a memory; wherein, the memory is coupled to the processor and is used to store a computer program, and when the computer program is executed by the processor, the processor implements the above-mentioned abnormal trading behavior detection method based on the blockchain.
[0014] Compared with the prior art, the present invention combines the DBSCAN algorithm and the isolation forest algorithm to achieve efficient and accurate detection of abnormal trading behaviors in the blockchain network. This method first uses the DBSCAN algorithm to identify high-density regions in the transaction data to screen abnormal trading behaviors, and then further refines the detection through the isolation forest algorithm, evaluating the abnormal score of each data point to determine its abnormality. This multi-level abnormal detection mechanism not only expands the coverage of abnormal detection, but also improves the accuracy and efficiency of detection, provides strong technical support for blockchain transaction security, effectively maintains the security and reliability of the blockchain network, and protects users from potential damages caused by abnormal trading behaviors. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 It is a schematic diagram of the detection process of the airdrop candy behavior model;
[0016] Figure 2 It is a flowchart of the working process of the abnormal trading behavior detection method provided by the first embodiment of the present invention;
[0017] Figure 3 It is a schematic structural diagram of the abnormal trading behavior detection system provided by the second embodiment of the present invention;
[0018] Figure 4 It is a working principle diagram of the abnormal trading behavior detection system provided by the second embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0019] The technical content of the present invention will be described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0020] First of all, it should be noted that abnormal trading behaviors in the blockchain network refer to activities that violate the conventional trading patterns. These activities often involve complex operation methods, such as insider trading using undisclosed information, high-frequency trading manipulation through automated programs, and abnormal transactions using the anonymity of blockchain technology. These behaviors are usually accompanied by abnormal trading volumes, frequencies, and patterns, and they may be covered up by technical means to evade supervision and monitoring. For example, abnormal trading behaviors may include a large number of high-frequency transactions in a short period of time, or abnormal fluctuations in trading patterns that are inconsistent with historical data and market trends. In addition, technology abuse, such as hacking attacks or the use of malware, is also part of abnormal trading behaviors, which may damage the security and integrity of the trading system. Therefore, advanced data analysis, machine learning models, and network security technologies need to be used to effectively address these challenges.
[0021] In the embodiments of the present invention, the abnormal trading behaviors targeted mainly refer to abnormal trading behaviors of the "airdrop candy" (also known as token airdrop) type. It refers to a series of abnormal trading activities carried out by certain individuals or organizations in the blockchain network by taking advantage of the opportunity of token airdrop (that is, distributing tokens to users for free). These activities may include, but are not limited to, fraud, market manipulation, etc., and their purpose is to illegally obtain benefits or disrupt the normal order of the blockchain network.
[0022] Analyzed from a technical perspective, such abnormal trading behaviors have some obvious characteristics, including but not limited to:
[0023] First of all, a significant characteristic of abnormal trading behaviors is the frequency of small transactions. In such behaviors, the attacker may initiate a large number of small transactions to reduce the attention and cost of a single transaction, while increasing network congestion to cover up their illegal acts. The amounts of these transactions are usually much lower than those of normal transactions, but the cumulative amount may have a significant impact on the network.
[0024] Secondly, there is often a certain correlation between the addresses involved in abnormal trading behaviors. By analyzing the interaction patterns of trading addresses, we can find some unusual fund flow paths. For example, there may be frequent fund transfers between multiple addresses, which may be a sign of fraud. In addition, these addresses may form specific patterns in the network, such as star or chain structures, which are significantly different from normal trading behaviors.
[0025] Furthermore, abnormal trading behaviors may exhibit concentration in terms of time. Such behaviors often occur intensively within a specific time period, especially at the beginning or end of token airdrop activities. Attackers may take advantage of users' attention to airdrop activities and the dynamic changes in the market to launch attacks to maximize their illegal gains.
[0026] In addition, abnormal trading behaviors also show anomalies in trading patterns. By using machine learning and data mining techniques, we can identify trading behaviors that are significantly different from normal trading patterns. For example, a sudden increase in trading frequency or an abnormal change in the trading path may be signs of abnormal trading behaviors.
[0027] Finally, abnormal trading behaviors may involve the abuse of smart contracts. In some cases, attackers may exploit vulnerabilities in smart contracts or write malicious smart contract code to manipulate transactions or steal funds. This behavior poses a direct threat to the security of the blockchain network and requires technical means to prevent it.
[0028] First Embodiment
[0029] As previously mentioned, the technical characteristics of abnormal trading behaviors of the "airdrop candy" type include the frequency of small transactions, address correlation, time concentration, anomalies in trading patterns, and the abuse of smart contracts, etc.
[0030] Based on this, the abnormal trading behavior detection method provided by the first embodiment of the present invention constructs a behavior model for airdrop candy abnormal trading based on four key features. These four features include: small transactions, a single-flow transaction structure from the attacker to the attacked party, high-frequency transactions, and trading activities within a certain time range. As Figure 1 shown, this abnormal trading behavior model takes time (Time) as the main line, details the single-flow relationship between the attacker and the attacked party, and clarifies that the initiation of the attack behavior is based on the standard of the BEI (Overall Airdrop Effect Index) index. Within a given time interval "Δt", the attacker completes all airdrop candy attack behaviors, and within this time interval, the set of all affected attacked parties is marked as {R n}.
[0031] As Figure 2 shown, the abnormal trading behavior detection method provided by the first embodiment of the present invention at least includes the following steps:
[0032] Step 1: First, reasonably select and preprocess each feature (transaction frequency, transaction amount, transaction time, and transaction address) in the transaction data. To better reflect the impact of these features on abnormal transaction behaviors, the transaction time can be cyclically encoded, and all features can be standardized. In addition, different weights can be assigned to each feature according to its contribution to abnormal transaction behaviors, so as to more accurately measure the influence of each feature in distance calculation. Then, perform clustering analysis on the transaction data. By the neighborhood radius and the minimum number of points, divide the transaction data points into different clusters, and identify the points that do not meet the density requirements as noise points, which are potential abnormal transactions;
[0033] Step 2: A key parameter in using the DBSCAN algorithm is the neighborhood radius (eps), which determines the size of the neighborhood of each data point. In practical applications, since the density of transaction behaviors may vary at different times or different addresses, a fixed neighborhood radius may not be applicable to all situations. Therefore, a method of dynamically adjusting the neighborhood radius can be adopted to optimize the clustering effect of the DBSCAN algorithm according to the quantity and density changes of the transaction data, ensuring that abnormal transaction behaviors can be effectively identified even in high-density areas;
[0034] Step 3: The clustering expansion of the DBSCAN algorithm depends on the neighborhood of core points. However, in the detection of abnormal transaction behaviors of airdropped candies, simply relying on distance metrics may not be sufficient to capture complex behavior patterns. To improve the accuracy of clustering, multiple features of the transaction (such as address, time, amount, etc.) can be combined and comprehensively matched when expanding the cluster. This multi-feature-based expansion method helps to identify abnormal transaction behaviors that are similar in multiple dimensions and enhances the robustness of the clustering algorithm;
[0035] Step 4: Finally, combine the Isolation Forest algorithm for anomaly detection. On the basis of the DBSCAN algorithm, further use the Isolation Forest algorithm to identify and confirm potential anomaly points. After the isolation tree is constructed, recursively construct multiple isolation trees; split the data by randomly selecting features and split points until the stopping condition is met, such as reaching the maximum height of the tree or the number of data points is not enough to continue splitting. Calculate the path length of the sample from the root node to the leaf node in the isolation tree to evaluate the anomaly score of the sample; use the anomaly score to identify abnormal transaction behaviors. The Isolation Forest can effectively discover transaction behaviors that deviate from the normal pattern by randomly partitioning the data points and evaluating their isolation. Combining the density clustering of the DBSCAN algorithm and the anomaly scoring of the Isolation Forest can significantly improve the accuracy of airdropped candy behavior detection, avoid abnormal patterns that may be missed by traditional clustering algorithms, and thus more comprehensively capture complex abnormal transaction behaviors.
[0036] The following will give a detailed and specific description of this.
[0037] In the process of analyzing transaction data characteristics, the key step is to extract a dataset that conforms to specific transaction behaviors, especially those "dust" attack activities related to airdropping candies. A notable feature of such activities is the dispersal of a small amount of funds to numerous different addresses. To effectively detect such behavior of dispersed funds, analyzing the output amount of transactions becomes an important means, which is crucial for identifying activities such as micropayments, donations, or distributing funds to multiple recipients.
[0038] During the transaction process, if it is found that a single address sends funds to multiple other addresses, this single-to-multiple address fund output pattern is usually marked as an abnormal transaction behavior. Step 1 in the embodiments of the present invention is based on the DBSCAN algorithm to detect and identify such abnormal transaction behaviors. The advantage of this algorithm is its ability to adapt to the dynamic changes of transaction data. By setting the output amount threshold in the abnormal transaction dataset, those abnormal distribution behaviors can be effectively identified. The output amount here refers to the total amount of all outputs in a transaction, that is, the total amount of funds obtained by all recipients. Through a detailed analysis of the output amount, the dispersed flow pattern of funds can be revealed, thereby identifying those abnormal transaction activities that may be related to "dust" attacks. Through this data feature-based analysis method, the behavior model of abnormal transactions of airdropping candies can be constructed more precisely, and a scientific basis can be provided for subsequent defense strategies.
[0039] See Figure 1 , if a transaction meets specific conditions, it can be determined that the address cluster composed of its nodes involves the "dust" injection behavior of airdropping candies. In this context, step 1 in the embodiments of the present invention is designed to process the set D of transaction feature output transaction amounts, where points represent the transaction feature points in D, and pointId refers to a specific output transaction amount. This step 1 is based on the DBSCAN algorithm for the identification and marking of abnormal transactions.
[0040] In order to optimize the detection of abnormal transaction behaviors by adjusting the key parameters of the DBSCAN algorithm, the inventor designed a total of three experiments, and each experiment focused on the distinction of neighborhood parameters and the minimum number of samples. In particular, the cosine similarity was used to determine the neighborhood parameters, which is particularly effective when dealing with ratio or frequency features, such as the proportion of the number of abnormal transaction behaviors selected in the present invention.
[0041] The three experiments were carried out for different parameter combinations, and the specific descriptions are as follows:
[0042] As shown in Table 1, Experiment 1 focused on the minimum number of points (MinPts) while keeping the neighborhood radius (ε) fixed at 0.5. The value of MinPts had three variations in this experiment, namely 5, 10, and 15, to explore the influence of different density thresholds on the clustering results.
[0043] Table 1 Parameter Configuration of Experiment 1
[0044]
[0045] As shown in Table 2, Experiment 2 fixed MinPts while varying the value of ε. In this experiment, MinPts was set as a constant, and the value of ε had three variations, namely 0.3, 0.4, and 0.6. This allowed for the evaluation of the sensitivity of the clustering algorithm to abnormal transaction detection under different neighborhood radii.
[0046] Table 2 Parameter Configuration of Experiment 2
[0047]
[0048] As shown in Table 3, Experiment 3 was the most complex, varying both ε and MinPts simultaneously. In this experiment, there were three sets of parameter settings: the first set set ε to 0.4 and MinPts to 10; the second set set ε to 0.6 and MinPts to 9; the third set set ε to 0.5 and MinPts to 8. The variations of these parameters were aimed at comprehensively evaluating how the two jointly affected the recognition effect of abnormal transaction behavior.
[0049] Table 3 Parameter Configuration of Experiment 3
[0050]
[0051] The specific experimental steps are as follows:
[0052] Feature selection: Select features suitable for processing ratio or frequency features, such as the proportion of the number of transactions of abnormal transaction behavior.
[0053] Feature value standardization: Adjust the numerical range of feature values to between 0 and 10 to eliminate the influence of different dimensions.
[0054] Similarity calculation: Use the cosine similarity formula to calculate the similarity between samples, providing a basis for clustering analysis.
[0055] Clustering or outlier detection: By identifying samples with low similarity, regarding them as potential abnormal transactions, and verifying the accuracy of these values through multiple experiments.
[0056] In an embodiment of the present invention, the cosine similarity formula is specifically as follows:
[0057]
[0058] (1) Feature selection: Suitable for processing ratio or frequency features, such as the proportion of the number of transactions of the abnormal transaction behavior selected in the present invention, the specific reference data of the amount, and the reference value of the input and output addresses.
[0059] (2) Feature value: If it is the transaction amount, the value of the feature value is between 0 and 10 to eliminate the influence of dimension. Here, two feature values, trait1 and trait2, can be selected.
[0060] (3) Similarity sim calculation: Use the cosine similarity formula to calculate the similarity between samples, and calculate the similarity from the first time to the last time of i.
[0061] (4) Clustering or outlier detection: Identify samples with low similarity as potential abnormal transactions, and verify the values after multiple experiments.
[0062] Among them, trait1 i and trait2 i are the values of the two vectors in the i-th dimension respectively.
[0063] is the sum of the products of the values in the corresponding dimensions of the two vectors, and are the Euclidean norms (i.e., the lengths of the vectors) of the two vectors respectively.
[0064] Through this series of steps, step 1 based on the DBSCAN algorithm can effectively identify abnormal transaction behaviors from blockchain transaction data, providing data support and scientific basis for defense strategies.
[0065] The following is the pseudocode of step 1, implemented in the Python language:
[0066]
[0067] The above step 1 takes the feature array {trait} as input, which contains multiple transaction data points X1 to X n , and each data point has its own transaction features, such as transaction amount and time, etc. The goal of step 1 is to output a clustering graph composed of multiple transaction clusters C1 to C n , and each cluster contains transaction data points with similar features, thereby revealing potential abnormal transaction patterns.
[0068] The specific steps of step 1 are as follows:
[0069] Input transaction array: Receive a feature array {trait} containing multiple transaction data points, and each data point X i has its own transaction features.
[0070] Initialization parameters: Set two key parameters of the DBSCAN algorithm: the neighborhood radius ε and the minimum number of points MinPts. These parameters will be used to determine whether a point is a core point and the size of its neighborhood.
[0071] Traverse all points: Traverse each point in the feature array {trait}. For each point, check whether it is a core point. A core point is a point that has at least MinPts points within the ε distance.
[0072] Construct clusters: For each core point, find all the points within its ε neighborhood and add these points to a new cluster C. If there are still unvisited points among these points, continue to expand the cluster until no more core points can be found.
[0073] Process border points and noise points: For non-core points, if they fall within the ε neighborhood of a core point, they are regarded as border points and added to the corresponding cluster. Those points that are neither core points nor border points are regarded as noise points, i.e., abnormal transactions.
[0074] Output the clustering results: Finally, output all the constructed clusters C1 to C n , as well as the identified abnormal transactions.
[0075] Generally speaking, in step 1, by adjusting the two key parameters, the neighborhood radius ε and the minimum number of points MinPts, to adapt to the fluctuations of transaction data, abnormal transactions that violate the conventional patterns can be effectively identified. In the specific implementation process, first determine these two parameters, and then traverse each data point in the transaction dataset to judge whether they are core points, that is, whether there are at least MinPts other points within the set neighborhood radius ε. For each data point identified as a core point, it will be further expanded to include all the points within its neighborhood to construct a cluster. Those non-core points but located within the neighborhood of a core point are marked as border points and classified into the corresponding cluster. Points that do not belong to any cluster are regarded as noise points, which are indications of abnormal transactions.
[0076] In the abnormal transaction behavior detection method provided in the first embodiment of the present invention, step 1 plays an important role in the data preprocessing stage. It identifies core points by calculating the number of points within the given neighborhood radius ε for each data point. Specifically, for each point in the dataset, if the number of its neighboring points within the ε range is not less than the minimum number of points MinPts, then this point is defined as a core point. Here, each point represents a transaction, and each transaction usually contains multiple features, such as transaction amount, time, user ID, and transaction address, etc.
[0077] It should be noted that the identification of core points is based on the following mathematical definition:
[0078] For each point p, the ε-neighborhood of point p here is calculated, where dist(p, q) represents the distance between point p and point q:
[0079] N ε (p) = {q ∈ D | dist(p, q) ≤ ε} (2)
[0080] Point p is marked as a core point when the number of points |N ε in its neighborhood (p)| is greater than or equal to a preset minimum number of points MinPts:
[0081] |N ε (p)| ≥ MinPts (3)
[0082] Then, each identified core point is traversed, and clustering is continuously expanded starting from the seed point as shown in formula (4). Its neighborhood is continuously explored as shown in formula (5). When a new transaction point is found and it is confirmed that the number of points in its neighborhood is greater than MinPts, it is regarded as a new clustering point and continues to expand outward until no new points can be added, as shown in formula (6):
[0083] CP = {p} (4)
[0084] for each q ∈ |N ε (p)| do (5)
[0085] CP = CP U {q} (6)
[0086] Finally, after clustering is completed, all points that do not belong to any cluster will be regarded as abnormal transactions. Usually, these points are potential "dust" injection airdrop candy transaction behaviors.
[0087] Clustered: CP i = CP (7)
[0088] Through the above steps, the dataset D of abnormal transaction behaviors can be obtained.
[0089] Next, it is necessary to separately split the normal point A and the abnormal point B. First, we randomly select a splitting value X between the maximum and minimum values of the data, and then divide the data into two groups according to this value: the data less than X and the data greater than or equal to X. On these two subsets, repeat this splitting process until the data can no longer be further divided. Usually, the abnormal point B is farther away from other data points, so it may be isolated with only a small number of splitting times; while the normal point A is usually clustered with other points and may require more splitting times to be separated. Therefore, from a statistical perspective, relatively clustered data points require more splitting times, while relatively isolated points require fewer splitting times. The isolation forest algorithm precisely utilizes this difference in the number of splitting times to measure the clustering and isolation of data points, that is, normal or abnormal. Then, the present invention selects a feature in the data set, such as the amount (not limited to this, other features are also possible), and performs steps 2, 3, and 4 based on the isolation forest algorithm, successively performing random splitting, tree construction, and then calculating the anomaly score by analyzing the average path length of each data point in all isolation trees. The shorter the path, the easier it is for the data point to be isolated, and thus it may be abnormal data.
[0090]
[0091] Step 2 aims to identify abnormal transactions from the transaction amount array D. This step receives three input parameters: the transaction amount array D, the number of trees t, and the sampling rate size k. The output of step 2 is a set of t transaction trees Trees for abnormal transactions.
[0092] The specific steps of step 2 are as follows:
[0093] Initialize the forest Trees, which will be a set for storing each isolation tree.
[0094] Calculate the height limit L for each tree, which is a logarithmic function based on the sampling rate size k, specifically L is the upper limit value of the logarithm of k to the base 2.
[0095] Traverse the index from 1 to t. For each tree, randomly select k sample points from the transaction amount array D to form a subset X'.
[0096] Use the subset X' to construct an isolation tree iTree, set the height of the root node of the tree to 0, and the maximum height limit of the tree to the previously calculated L.
[0097] Add each constructed isolation tree iTree to the forest Trees until t trees are constructed.
[0098] Through the above steps, the transaction data can be randomly split, and multiple isolation trees can be constructed. The construction of each tree is independent. By randomly sampling and splitting the sample data, Step 2 can evaluate the anomaly degree of each data point. Finally, Step 2 outputs a set of isolation trees, which are jointly used to identify and isolate abnormal transaction behaviors.
[0099]
[0100] Step 3 is the process of constructing a single Isolation Tree (iTree), which is a key component in the isolation forest algorithm for identifying anomaly points. The purpose of this step is to draw samples from the original data without replacement and construct a binary tree based on these samples to determine whether a data point is abnormal. The following are the specific steps of Step 3:
[0101] 1. Input and output definitions:
[0102] The input parameters include the transaction amount array D, which contains multiple transaction data points X1 to X n , the current height e of the tree, and the maximum height limit L of the tree.
[0103] The output is the completed transaction tree iTree.
[0104] 2. Process of constructing iTree:
[0105] At the beginning, initialize an empty transaction tree Tree and set the current height e of the tree to 0.
[0106] If the current height e of the tree reaches the maximum limit L, or the number of data points in the sample set X is less than or equal to 1 and no further splitting can be performed, stop the construction and return an external node exNode with its size attribute set to the number of data points in the current sample set X.
[0107] 3. Select features and split points:
[0108] If the current height e of the tree has not reached the limit L and the number of data points in the sample set X is greater than 1, continue to execute.
[0109] The algorithm sets P as the list of attributes in the sample set X and randomly selects an attribute q belonging to P.
[0110] Randomly select a split point point between the maximum and minimum values of the selected attribute q.
[0111] 4. Data splitting:
[0112] According to the split point point, the sample set X is split into two subsets Xl and Xr. Xl contains all data points with attribute Q values less than point, and Xr contains all data points with attribute q values greater than or equal to point.
[0113] 5. Recursively construct subtrees:
[0114] For subset X l , construct the left subtree by recursive call and update the current height to e + 1.
[0115] For subset X r , construct the right subtree by recursive call and also update the current height to e + 1.
[0116] Return an internal node inNode, which contains the left subtree, the right subtree, the split attribute SplitAtt, and the split value SplitValue.
[0117] 6. End condition:
[0118] The recursive construction process will continue until the stopping condition is met, that is, the height of the tree reaches the limit or the data cannot be divided further.
[0119] Through the above steps, step 3 can construct a binary tree that can distinguish normal and abnormal transactions. The construction of each tree is based on randomly selected features and split points. This method helps to capture abnormal patterns in the data and identify abnormal points in the isolation forest algorithm.
[0120] In the prediction process of a similar binary classification model, the probability P(y = 1) that a sample belongs to the positive class can be output. Similarly, the isolation forest algorithm can finally output the anomaly score of each sample. This anomaly score is based on the number of splits experienced by the sample when reaching the leaf node in the isolation tree, that is, the number of edges passed by the sample from the root node to the leaf node in the tree. For each tree in the isolation forest algorithm, we can calculate the average path length of the sample, and then take the average of the results of all trees as the anomaly score of the sample. This score reflects the degree of anomaly of the sample point. The higher the score, the more likely the sample is to be abnormal.
[0121]
[0122] Step 4 is used to calculate the path length of the sample in the isolation forest algorithm. The following are the detailed steps of this step:
[0123] Input and output definitions:
[0124] The input parameters include the sample x, the current tree T, and the current path length e.
[0125] The output is the path length of sample x in tree T.
[0126] Path length calculation:
[0127] At the beginning, initialize the path length e to 0.
[0128] If the current tree T is an external node (i.e., a leaf node), return e plus the correction value c multiplied by T.size. Here, T.size represents the number of samples in the same leaf node as sample x, and c(T.size) is a correction value representing the average path length of constructing a binary tree with T.size samples.
[0129] Recursively traverse the tree structure:
[0130] If the current tree T is not a leaf node, continue the recursive traversal.
[0131] Determine T.SplitAtt as the splitting attribute of the current tree T, that is, the attribute used to divide data points.
[0132] Check whether the value x_a of sample x on the splitting attribute is less than the splitting value T.SplitValue of the current tree T.
[0133] If x_a is less than T.SplitValue, recursively call the PathLength algorithm to calculate the path length of sample x in the left subtree T.left, and update the path length e to e + 1.
[0134] If x_a is greater than or equal to T.SplitValue, similarly recursively call the PathLength algorithm to calculate the path length of sample x in the right subtree T.right, and update the path length e to e + 1.
[0135] End condition: The recursive process will continue until reaching the leaf node, at which point the final path length is returned.
[0136] Through the PathLength algorithm, the path length from the root node to the leaf node of the sample in each isolation tree can be calculated. This length reflects the degree of abnormality of the sample, because in the isolation forest algorithm, outliers usually have shorter path lengths. By taking the average of the path lengths of all trees, the overall anomaly score of the sample can be obtained.
[0137] Steps 2, 3, and 4 together constitute the process of calculating the anomaly score in the isolation forest algorithm. These algorithms recursively traverse the structure of each isolation tree, starting from the root node until reaching the leaf nodes, i.e., the external nodes. During this process, the path length for each data point to reach the leaf node is recorded, and the number of data points in that leaf node is added. For each data point in the dataset, this process is repeated to calculate the path lengths of all data points.
[0138] After calculating the path lengths of all data points, the average of these lengths is obtained. This average path length not only reflects the overall structure of the dataset but also forms the basis for subsequent anomaly score calculation. Next, the anomaly score is usually calculated by comparing the path length of a data point with the average path length of the entire dataset. Specifically, formula (8) is used:
[0139]
[0140] where c is a constant, usually set to the height of the tree or other appropriate values for normalization.
[0141] In the context of blockchain transaction data, a shorter path may indicate that the transaction behavior significantly deviates from the normal pattern and may thus be marked as an anomaly. Therefore, data points with shorter path lengths receive higher anomaly scores because they are more easily isolated from the dataset, which usually means they are outliers. On the contrary, data points with longer path lengths receive lower anomaly scores, indicating that they are more likely to be normal values. This path length-based approach enables the isolation forest algorithm to effectively identify anomaly points in the dataset.
[0142] By evaluating and comparing the path lengths of all data points, the isolation forest algorithm can assign an anomaly score to each data point. The anomaly score is calculated by comparing the average path length of each data point with the average path length of all data points. This score provides a quantitative anomaly metric for each transaction. The higher the score, the more likely the transaction behavior is to be regarded as an abnormal transaction behavior. This quantitative method enables an objective distinction between normal transactions and potential abnormal transaction behaviors, providing an objective basis for further analysis and monitoring.
[0143] In the embodiments of the present invention, the DBSCAN algorithm and the isolation forest algorithm are combined to form a powerful multi-level abnormal transaction behavior detection mechanism. As the first step, the DBSCAN algorithm initially identifies possible abnormal transaction candidates by analyzing the density distribution in transaction data. These candidates, based on the density of transactions, can reveal aggregation behaviors that are not common in regular transaction patterns. Subsequently, the isolation forest algorithm further refines the anomaly detection by constructing multiple isolation trees. In each tree, data points are isolated by randomly selecting features and split points, and then the path length of each point is calculated. These path lengths are then used to evaluate the anomaly score of each transaction. The higher the anomaly score, the more likely the transaction is to be abnormal. In this way, it is possible to identify abnormal transaction behaviors that are significant in terms of transaction density and also discover transaction points that appear abnormal under random feature selection.
[0144] It should be noted that when detecting abnormal transaction behaviors of the "airdrop candy" type, in addition to the DBSCAN algorithm and the isolation forest algorithm, a subgraph matching model can also be used. The subgraph matching model accurately identifies specific types of abnormal transaction behaviors by matching predefined abnormal pattern subgraphs, and is particularly suitable for analyzing structured data such as transaction networks and conducting analysis using the topological structure of the data. However, this method requires sufficient prior knowledge of abnormal transaction behaviors to define the abnormal patterns, and subgraph matching in large-scale graph data may face high computational complexity.
[0145] The DBSCAN algorithm is applicable to the case where the distribution of clusters in the dataset is unknown. It automatically identifies the number of clusters through neighborhood parameters and identifies noise points without assigning them to any cluster, which is conducive to accurately identifying abnormal transactions. The DBSCAN algorithm can not only identify spherical clusters but also clusters of any shape, making it suitable for the analysis of data with complex distributions. However, the selection of the neighborhood size (ε) and the minimum number of points (MinPts) of this algorithm has a significant impact on the results and needs to be adjusted according to the data characteristics. In high-dimensional spaces, the performance and effectiveness of the DBSCAN algorithm may decline because the sparsity of high-dimensional data makes it difficult to define "neighborhoods".
[0146] On the other hand, the isolation forest algorithm has a low time complexity, is suitable for processing large-scale data, and does not rely on any data distribution assumptions. It detects anomalies relying on the isolation mechanism and has a wide range of applications. Compared with some traditional algorithms, the isolation forest algorithm can still maintain good anomaly detection performance on high-dimensional data. However, the random forest construction process in the algorithm introduces randomness, which may lead to different running results.
[0147] In the abnormal transaction behavior detection method provided by the first embodiment of the present invention, the DBSCAN algorithm and the isolation forest algorithm are combined to jointly detect abnormal transaction behaviors, so as to ensure the complementarity of algorithm advantages and make the detection effect of abnormal transaction behaviors better. This dual-algorithm mechanism not only utilizes the density clustering characteristics of the DBSCAN algorithm to identify abnormal clusters, but also evaluates and identifies the abnormality degree of data points through the random forest model of the isolation forest algorithm. This combined method greatly improves the recognition rate and accuracy of abnormal transaction behaviors, providing a more reliable security guarantee for the blockchain network.
[0148] Compared with the prior art, the abnormal transaction behavior detection method provided by the first embodiment of the present invention combines the results of the isolation forest algorithm with the output of the aforementioned DBSCAN algorithm, providing a multi-level abnormal detection mechanism. The DBSCAN algorithm initially screens abnormal transaction behaviors by identifying high-density regions in transaction data, while the isolation forest algorithm further refines this process by evaluating the abnormality score of each data point to determine its abnormality. The specific improvement points are described as follows:
[0149] First, based on the traditional DBSCAN algorithm, the embodiment of the present invention introduces a multi-dimensional weighted distance metric. The traditional DBSCAN algorithm mainly calculates the distance between points based on spatial positions, but transaction data usually has multiple features (such as transaction amount, time, frequency, etc.). Therefore, the embodiment of the present invention comprehensively considers these features and assigns different weights to different features through methods such as weighted Euclidean distance or Mahalanobis distance to reflect the importance of each feature in clustering. In this way, the distance between data points no longer depends only on spatial positions, but comprehensively considers multiple dimensions, thereby improving the accuracy and adaptability of clustering.
[0150] Second, the neighborhood radius (eps) in the DBSCAN algorithm is dynamically adjusted. In the traditional DBSCAN algorithm, the eps value is fixed, which may not be able to effectively capture the true clustering structure for data with uneven density distributions. Therefore, the embodiment of the present invention dynamically calculates its k-nearest neighbor distance (k-NN) based on the local density of each data point and adjusts its neighborhood radius. Specifically, a smaller eps value is used in regions with higher density, while a larger eps value is used in regions with lower density. Through this dynamic adjustment, the clustering structure of different density regions can be more flexibly identified, making the clustering effect more accurate.
[0151] Third, introduce time series analysis for time weighting. In transaction data, time is often an important feature, and transaction behaviors with time series need special attention. To better process such time series data, embodiments of the present invention calculate the time difference for data points and use a time decay function (such as exponential decay) to reduce the impact of points with large time differences on clustering. At the same time, embodiments of the present invention combine the time difference with the spatial distance to design a new weighted distance metric, making it easier for transaction data that is close in time to be clustered together, enhancing the influence of time correlation on the clustering result.
[0152] Fourth, in terms of the cluster merging mechanism, embodiments of the present invention propose an innovative solution to address the problem of overlapping or ambiguous boundaries between adjacent clusters. In some cases, the DBSCAN algorithm may generate multiple adjacent clusters, and the boundaries between these clusters may not be obvious, resulting in an unsatisfactory clustering result. To optimize this problem, embodiments of the present invention design a cluster merging algorithm. When the distance between two clusters is less than a certain threshold, they are merged into a larger clustering cluster. In addition, embodiments of the present invention also re-evaluate the boundary points to ensure that they are assigned to the most appropriate cluster, thus avoiding the emergence of boundary points that are difficult to classify.
[0153] Fifth, the handling of noise points is improved. In the traditional DBSCAN algorithm, noise points are directly labeled as the "noise" class, but these noise points may contain important abnormal information. Therefore, in embodiments of the present invention, an independent isolation forest algorithm is used to identify noise points. By labeling noise points as abnormal classes and conducting specialized analysis, abnormal transaction behaviors can be further explored, thereby enhancing the overall understanding and processing ability of the data.
[0154] Sixth, a multi-level clustering method is introduced to improve the hierarchy and interpretability of the clustering result. The traditional DBSCAN algorithm usually generates a flat clustering structure, while in complex datasets, the clustering result may exhibit a hierarchical distribution. Therefore, embodiments of the present invention apply a more fine-grained clustering algorithm, such as hierarchical clustering, to the data within each cluster after preliminary clustering to further subdivide the internal structure of each cluster. In this way, not only can a flat clustering result be obtained, but also a multi-level clustering structure with nested relationships can be generated, making the clustering result of the data richer and easier to understand.
[0155] Seventh, the isolation forest algorithm isolates points by constructing random trees and is particularly good at quickly identifying global outliers, that is, those points that are easy to distinguish from other points. The isolation forest algorithm can effectively identify single or a small number of abnormal transaction behaviors, which may not be easily detected in the DBSCAN algorithm.
[0156] In addition, the combination of the DBSCAN algorithm and the isolation forest algorithm provides complementary coverage of different types of anomalies. The DBSCAN algorithm excels at identifying density-based local anomalies, while the isolation forest algorithm complements this ability by being particularly good at detecting individual transactions that deviate from normal behavior patterns, such as unusually large transactions or transaction behaviors that are completely different from previous patterns.
[0157] Furthermore, the combined use of the DBSCAN algorithm and the isolation forest algorithm can effectively handle large-scale data sets. The DBSCAN algorithm is first used to identify dense anomaly clusters, and the isolation forest algorithm then further analyzes these clusters or is used to quickly screen out global anomalous transactions. This combined approach can dynamically adapt to different anomaly patterns and data characteristics.
[0158] Finally, this combined approach reduces the false positive rate. Using either method alone may result in a high false positive rate, but the combination of the two algorithms can mutually verify the detection results, thereby reducing false positives. For example, the DBSCAN algorithm may incorrectly identify normally distributed transactions as anomalies, while the isolation forest algorithm may incorrectly classify rare but normal transactions as anomalies. By combining these two algorithms, anomalous transaction behaviors can be more accurately identified.
[0159] In summary, the embodiments of the present invention provide an efficient, accurate, and flexible solution for detecting anomalous transaction behaviors by combining the DBSCAN algorithm and the isolation forest algorithm. This method not only expands the coverage of anomaly detection but also improves the accuracy and efficiency of detecting anomalous transaction behaviors, providing strong technical support for blockchain transaction security.
[0160] Second Embodiment
[0161] Based on the above method, the second embodiment of the present invention provides a system for detecting anomalous transaction behaviors based on a blockchain. As Figure 3 shown, the system for detecting anomalous transaction behaviors includes a processor and a memory; wherein, the memory is coupled to the processor and is used to store a computer program, and when the computer program is executed by the processor, the processor implements the above method for detecting anomalous transaction behaviors based on a blockchain.
[0162] Among them, the processor is used to control the overall operation of the system for detecting anomalous transaction behaviors to complete all or part of the steps of the above method for detecting anomalous transaction behaviors (see Figure 4)。The processor can be a central processing unit (CPU), a graphics processing unit (GPU), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a digital signal processing (DSP) chip, etc. The memory is used to store various types of data to support the operation of the abnormal transaction behavior detection system. Such data may include, for example, instructions for any application or method operating on the abnormal transaction behavior detection system, as well as application-related data. The memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, etc.
[0163] In an exemplary embodiment, the abnormal transaction behavior detection system can be specifically implemented by a computer or a microprocessor, or by a product with certain functions, for executing the above-mentioned abnormal transaction behavior detection method and achieving the same technical effects as the above method. Specifically, the computer can be, for example, a personal computer, a laptop computer, an in-vehicle human-machine interaction device, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0164] In another exemplary embodiment, the present invention also provides a computer-readable storage medium including program instructions, which, when executed by a processor, implement the steps of the abnormal transaction behavior detection method in any of the above embodiments. For example, the computer-readable storage medium can be the above-mentioned memory including program instructions, and the above program instructions can be executed by the processor to complete the above-mentioned abnormal transaction behavior detection method and achieve the same technical effects as the above abnormal transaction behavior detection method.
[0165] It should be noted that the above-mentioned multiple embodiments are only examples. The technical solutions of each embodiment can be combined, and all are within the protection scope of the present invention.
[0166] The above has described in detail the method and system for detecting abnormal transaction behavior based on blockchain provided by the present invention. For those of ordinary skill in the art, any obvious changes made without departing from the essence of the present invention will constitute an infringement of the patent right of the present invention and will bear corresponding legal responsibilities.
Claims
1. A method for detecting abnormal transaction behavior based on blockchain, characterized in that The steps include: Step 1: Select and preprocess the features in the transaction data, including periodically encoding transaction time, standardizing features, and assigning weights based on the contribution of features to abnormal transaction behavior; then, use the DBSCAN algorithm to perform cluster analysis, divide data points into clusters based on neighborhood radius and minimum number of points, and identify points that do not meet density requirements as noise points, as potential abnormal transactions; Step 2: Use the method of dynamically adjusting the neighborhood radius to optimize the clustering effect of the DBSCAN algorithm according to the number and density changes of transaction data, ensuring that abnormal transaction behaviors can be effectively identified even in high-density areas; Step 3: In the detection of abnormal trading behavior, multiple features of transactions are combined and comprehensive matching is performed when expanding clusters to identify abnormal trading behaviors with similarities in multiple dimensions and improve the robustness of the clustering algorithm; Step 4: Based on the DBSCAN algorithm, the isolation forest algorithm is further used to identify and confirm potential anomalies. The data is split by randomly selecting features and split points, and multiple isolation trees are recursively constructed until the stopping condition is met. The path length of the sample from the root node to the leaf node in the isolation tree is calculated, and the anomaly score of the sample is evaluated to identify abnormal trading behavior.
2. The abnormal transaction behavior detection method according to claim 1, characterized in that The step 1 includes the following sub-steps: Create an array containing multiple trading data points, each with its own trading characteristics; Set the neighborhood radius and minimum number of points to determine whether a point is a core point and the size of its neighborhood; Traversing each point in the array to check whether it is a core point; wherein the core point refers to a point that has at least a minimum number of points within the neighborhood radius; For each core point, find all the points in its neighborhood and add them to a new cluster; if there are unvisited points among all the points, continue to expand the cluster until no more core points are found.
3. The abnormal transaction behavior detection method according to claim 2, characterized in that: For non-core points, if they fall within the neighborhood of a core point, they are considered as boundary points and added to the corresponding cluster; Points that are neither core points nor boundary points are identified as noise points and are considered potential abnormal transactions.
4. The abnormal transaction behavior detection method according to claim 3, characterized in that: For each core point, more points are included by continuously expanding the neighborhood of the core point until no more points can be added.
5. The abnormal transaction behavior detection method according to claim 1, characterized in that The step 2 includes the following sub-steps: Create an empty collection to store each constructed isolation tree; Calculate the height limit of each tree, where the height limit is the upper limit of the logarithm of the sampling rate with base 2 as the value; Traverse and construct multiple isolation trees; randomly extract a specified number of sample points from the transaction amount array to form a subset; use the subset to construct an isolation tree, in which the root node height is set to 0 and the maximum height does not exceed the height limit of each tree; add the constructed isolation tree to the forest.
6. The abnormal transaction behavior detection method according to claim 1, characterized in that The step 3 includes the following sub-steps: Construct a single isolation tree, extract samples from the original data without replacement, and construct a binary tree based on the samples to determine whether the data point is abnormal; If the height of the current tree does not reach the height limit, and the number of data points in the sample set is greater than 1, continue to execute; Randomly select an attribute and then randomly select a split point between the maximum and minimum values of the attribute; The sample set is split into two subsets according to the split point, wherein the first subset contains all data points whose attribute values are less than the split point, and the second subset contains all data points whose attribute values are greater than or equal to the split point; For the first and second subsets, recursively call to build the left and right subtrees, and update the current height; Returns internal node information including left and right subtrees, split attributes and split values; The recursive operation is continued until a stopping condition is met; the stopping condition is that the height of the tree reaches the height limit or the data cannot be divided any further.
7. The abnormal transaction behavior detection method according to claim 6, characterized in that: In the process of building the isolation tree, a feature is randomly selected as the splitting attribute; On the selected splitting attribute, randomly select a splitting point; The data is continuously split in a recursive manner until the stopping condition is met.
8. The abnormal transaction behavior detection method according to claim 1, characterized in that The step 4 includes the following sub-steps: (1) Initially, the path length is set to 0; (2) Check whether the current tree is a leaf node; if it is a leaf node, return the path length plus the correction value multiplied by the number of samples in the leaf node; (3) If the current tree is not a leaf node, find out the splitting attribute of the current tree; (4) Compare the sample attribute value with the split point to check whether the sample value on the split attribute is less than the split point; (5) If the attribute value of the sample is less than the split point, recursively calculate the path length of the sample in the left subtree and update the path length; If the attribute value of the sample is greater than or equal to the split point, recursively calculate the path length of the sample in the right subtree and update the path length; (6) For the recursive call of the left subtree or the right subtree, repeat steps (2) to (5) until a leaf node is reached; (7) When the recursion reaches the leaf node, the final calculated path length is returned as the path length of the sample in the isolation tree.
9. The abnormal transaction behavior detection method according to claim 8, characterized in that: The path lengths of each sample in all isolated trees are calculated and averaged to obtain the average path length of the sample; the average path length of the sample is compared with the overall average path length of all samples to determine the abnormal score of the sample.
10. A blockchain-based abnormal transaction behavior detection system, characterized by It comprises a processor and a memory; wherein the memory is coupled to the processor and is used to store a computer program, and when the computer program is executed by the processor, the processor implements the abnormal transaction behavior detection method described in any one of claims 1 to 9.
Citation Information
Patent Citations
A Subgraph Matching-Based Method for Detecting Abnormal Transaction Behavior on Ethereum
CN114677217B
Cited By
Student behavior anomaly detection system based on clustering processing
CN120524399A
Mobile payment transaction behavior analysis method based on spatial data
CN120597176A
A method for analyzing mobile payment transaction behavior based on spatial data
CN120597176B
Method and system for tracing abnormal data of equipment operation behavior based on block chain
CN121071753A
A method and system for tracing abnormal device operation data based on blockchain
CN121071753B