Method for identifying Panze cheating in block chain transaction
Through the combination of multi-source data fusion and deep learning models, the spatiotemporal features in blockchain transactions are extracted, and the problems of single data and incomplete feature extraction in the existing technology are solved, achieving efficient and accurate Ponzi scam recognition.
Patent Information
- Application Number
- CN202510222783.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-05-30
AI Technical Summary
When identifying Ponzi scams in blockchain transactions, the recognition accuracy and recall rate are not ideal due to single data, incomplete feature extraction and poor algorithm adaptability.
By integrating multi-source data, refined spatiotemporal feature engineering and deep learning models (LSTM, CNN, Adaboost), features such as transaction time interval autocorrelation coefficient, address correlation index and geographic distribution concentration index are extracted to form high-dimensional feature vectors for identification.
The recognition accuracy and robustness of Ponzi schemes in blockchain transactions have been significantly improved, with the recall rate increased to 88%, the accuracy rate reached 92%, and the F1 value reached 90%.
Smart Images

Figure CN120067761A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information data processing, and provides a method for identifying Ponzi schemes in blockchain transactions. Background Art
[0002] In today's digital finance era, the anonymity and decentralization characteristics of blockchain technology have promoted an explosive growth in transaction activities. However, this technological advantage has also been exploited by criminals, and Ponzi schemes frequently occur in blockchain transactions. Such scams usually use high returns as bait, pay the returns of early investors with the funds of later investors, create a false prosperity, and ultimately lead to the breakage of the capital chain. This not only causes huge losses to investors, but also seriously threatens the stability of the financial market and the credibility of blockchain technology.
[0003] Currently, there are significant defects in the methods for identifying Ponzi schemes in blockchain transactions. Traditional identification means mainly rely on manual experience analysis and simple rule setting (such as judging by transaction amount, frequency threshold, or static statistics of the number of transaction addresses). However, such methods have the following limitations:
[0004] 1. Single data source: Most solutions only rely on the original blockchain transaction records, lacking user behavior data of the trading platform (such as trading habits, associated information) and risk marking data of regulatory agencies, resulting in insufficient data dimensions and difficulty in capturing the hidden operation patterns of scams.
[0005] 2. One-sidedness in feature extraction: The existing methods have insufficient depth and dimensions in mining spatio-temporal features, especially lacking quantitative analysis tools for the dynamic volatility features of transactions in the time dimension (such as frequent transactions during the rapid capital absorption period and abnormal stagnation during the collapse period) and the spatial features of the address association network (such as geographical distribution concentration and abnormal fund transfer patterns between associated addresses). For example, the autocorrelation coefficient of the trading time intervals of Ponzi schemes may show a significant positive correlation due to the need for fund transfer, while traditional methods do not effectively utilize such features.
[0006] 3. Poor algorithm adaptability: Solutions based on shallow machine learning models (such as random forests) have insufficient generalization ability when dealing with high-dimensional spatio-temporal sequence data and are difficult to dynamically adapt to changes in scam patterns. In addition, a single model has poor robustness to noisy data and is prone to misjudgment due to abnormal transaction nodes or data missing.
[0007] The deficiencies of the existing technologies result in unsatisfactory identification accuracy and recall rate. For example, rule-based methods are difficult to cover the changing fund transfer strategies, and the misjudgment rate is as high as 30%; the recall rate of traditional machine learning models is often lower than 80% due to incomplete feature engineering. Especially in the face of the enhanced concealment characteristics of Ponzi schemes in the later stage (such as the decentralization of the address association network and the disguise of trading frequency as normal fluctuations), the existing technologies are extremely likely to fail.
[0008] Therefore, there is an urgent need for a solution that combines multi-source data fusion, refined spatio-temporal feature engineering, and deep learning model integration to break through the bottlenecks of existing technologies and achieve efficient and accurate identification of Ponzi schemes in blockchain transactions. Summary of the Invention
[0009] By integrating technical means such as multi-source data, refined spatio-temporal feature engineering, and deep learning models (LSTM, CNN, Adaboost), the present invention solves the problem of low accuracy in the identification of Ponzi schemes in blockchain transactions by traditional methods due to single data, incomplete feature extraction, and poor algorithm adaptability, and significantly improves the accuracy and robustness of identification.
[0010] To achieve the above object, the invention adopts the following technical means:
[0011] The present invention provides a method for identifying Ponzi schemes in blockchain transactions, including the following steps:
[0012] Step 1: Multi-source data collection: Integrate blockchain transaction records, transaction platform behavior data, and regulatory agency marked data to obtain complete and diverse transaction data;
[0013] Step 2: Data preprocessing: Form structured data through noise filtering, missing value filling, and time series construction;
[0014] Step 3: Refined spatio-temporal feature engineering:
[0015] Time dimension feature extraction: Include calculating the transaction time interval sequence and its statistics, the standard deviation of transaction response time, the standard deviation of transaction confirmation time, the slope of the change trend of transaction frequency, and the autocorrelation coefficient ρ(k) of the transaction time interval;
[0016] Spatial dimension feature extraction: Include calculating the address correlation index I rel and the geographical distribution concentration index I geo ;
[0017] Step 4: Use LSTM and CNN to respectively learn the refined spatio-temporal features of transaction data, deeply mine the time series features and transaction address features, and obtain the time series feature F LSTM and the transaction address feature F CNN ;
[0018] Step 5: Adaboost integration: Align and fuse the features of F LSTM and F CNN , and form the final classifier P(y|X) by iteratively training weak classifiers and dynamically weighted combination.
[0019] In the above technical solution, Step 1 includes the following steps:
[0020] Step 1.1: Obtain transaction records from blockchain nodes, including transaction timestamps, transaction amounts, transaction initiation addresses, and receiving addresses;
[0021] Step 1.2: Collect user transaction behavior data from the trading platform, including transaction frequency, transaction habits, and associated information of both trading parties;
[0022] Step 1.3: Collaborate with regulatory agencies to obtain data on confirmed Ponzi scheme cases and label the data. Label Ponzi schemes as "1" and normal transactions as "0" to obtain complete and diverse transaction data.
[0023] In the above technical solution, Step 2 includes the following steps:
[0024] Step 2.1: Data cleaning:
[0025] Remove noise data by setting thresholds for transaction amount and frequency;
[0026] Delete duplicate data;
[0027] Eliminate invalid data, including data with a transaction amount of zero, invalid addresses, or abnormal timestamps;
[0028] Step 2.2: Missing value handling:
[0029] For continuous features, use the mean filling method. Continuous features include transaction amount and transaction frequency;
[0030] Complement missing data for transaction timestamps using interpolation;
[0031] Set missing values in discrete features to the default value "unknown". Discrete features include transaction addresses;
[0032] Step 2.3: Time series construction:
[0033] Sort the transaction data in chronological order based on the transaction timestamp to construct time series data.
[0034] Specific detailed issues that need to be overcome in refined spatio-temporal feature engineering:
[0035] Time dimension:
[0036] In blockchain transactions, the transaction time-related characteristics of Ponzi schemes are often unstable and volatile. In the early stage, due to rapid capital absorption and maintenance of the capital chain, a large amount of funds will flow in and out of a specific address in a short period of time, and the transactions will be fast and frequent. In the later stage, due to the break of the capital chain or abnormal manipulation, the capital flow is small, the transactions are slow, and the transactions are deserted. This feature is in sharp contrast to the relatively stable law of normal transactions. This feature cannot be directly reflected in the redundant data exported by the blockchain and trading platform, and it is also difficult to extract the relevant logical relationship. At present, there is a lack of comprehensive and accurate mathematical formulas to mine these features.
[0037] Therefore, mathematical methods are used to analyze and model the fine-grained spatiotemporal feature engineering, and mathematical formulas are used to calculate and refine the characteristics of the Ponzi scheme. To this end, in the above technical solution, the present invention provides the following detailed steps:
[0038] Step 3.1: Time dimension feature extraction
[0039] The time dimension feature is used to capture abnormal patterns of transaction behavior over time and reflect the time distribution law of Ponzi schemes.
[0040] Step 3.1.1: Trading time interval related indicators
[0041] Step 3.1.1.1: Record the timestamp T of each transaction i , forming a time series T 1 ,T 2 ,…,T n ;
[0042] Step 3.1.1.2: Based on the transaction time series, calculate the transaction time interval sequence I 1 ,I 2 ,…,I n , where I i =T i -T i-1 ;
[0043] Step 3.1.1.3: Calculate the average trading time interval Reflects the average time interval between transactions;
[0044] Step 3.1.1.4: Calculate the standard deviation σ of the average trading time interval T , reflecting the degree of fluctuation of trading time intervals:
[0045]
[0046] If the average transaction time interval is too short or the standard deviation is extremely small, it means that funds are flowing quickly in a specific address, which is consistent with the characteristics of frequent operations when a Ponzi scheme is trying to make money quickly; conversely, if the average transaction time interval is too long or the standard deviation is too large, it reflects that the flow of scam funds is blocked or there is abnormal manipulation.
[0047] Step 3.1.1.5: Calculate the probability density distribution p(I) of the trading time interval and show the distribution of the trading time interval:
[0048]
[0049] where n I is the number of transactions with a time interval of I, and N is the total number of transactions.
[0050] Step 3.1.1.6: Calculate the autocorrelation coefficient ρ(k) of the trading time interval to reflect the correlation between trading time intervals:
[0051]
[0052] ① Autocorrelation coefficient (ρ(k)): It is the calculation result of the formula, which measures the tightness of the linear relationship between two values separated by k time intervals in the trading time interval sequence. k is the lag order, indicating the interval number of two trading time intervals in the sequence.
[0053] ② Covariance (Cov(I i , I i+k )): Covariance is used to measure the overall error of two variables. In the context of trading time intervals, Cov(I i , I i+k ) represents the covariance between the i-th time interval I i and the i + k-th time interval I i+k in the trading time interval sequence. If the covariance is positive, it means that I i and I i+k tend to change in the same direction, that is, when I i is larger, I i+k also tends to be larger; if the covariance is negative, they tend to change in the opposite direction; when the covariance is 0, there is no linear correlation between the two. Its calculation formula is: where E represents the expectation, is the mean of the trading time interval sequence.
[0054] ③ Variance (Var(I i ), Var(I i+k )): Variance is used to measure the degree of dispersion of a single variable. Var(I i ) is the variance of the i-th trading time interval I i , and Var(I i+k ) is the variance of the i + k-th trading time interval I i+k . The larger the variance, the greater the degree to which the time interval value deviates from its mean, and the more dispersed the data. The calculation formula of variance is Here, X represents I respectively i and I i+k .
[0055] ④ Overall meaning of the formula: The autocorrelation coefficient ρ(k) is calculated by dividing the covariance by the square root of the product of the variances of two variables. This is done to standardize the covariance so that the value of the autocorrelation coefficient is always between -1 and 1. -1 indicates a perfect negative correlation between two trading time intervals, 1 indicates a perfect positive correlation, and 0 indicates no linear correlation. By calculating the autocorrelation coefficients for different values of k, we can comprehensively understand the correlation of the trading time interval sequence at different lag orders, and thus discover whether there is some potential pattern or trend between the trading time intervals. For example, if ρ(1) is large and positive, it means there is a strong positive correlation between adjacent trading time intervals, that is, when the current trading time interval is long, the next trading time interval also tends to be long.
[0056] In blockchain transactions, the time intervals of normal transactions are usually random, and the autocorrelation coefficient is close to zero. However, the time intervals of Ponzi schemes often have regularity, and the autocorrelation coefficient often significantly deviates from zero. For example:
[0057] Positive correlation: If ρ(k)>0, it means that the trading time intervals have some regularity, for example, to quickly attract funds or maintain the capital chain.
[0058] Negative correlation: If ρ(k)<0, it means that there is some alternating pattern in the trading time intervals, for example, to cover up the abnormality of capital flow.
[0059] Volatility: As the scam develops, the autocorrelation coefficient will show large fluctuations, especially when the capital chain is tight or under regulatory pressure.
[0060] By calculating and analyzing the autocorrelation coefficients of trading time intervals, these abnormal correlations and fluctuations can be captured, thus effectively identifying Ponzi schemes in blockchain transactions.
[0061] Step 3.1.2: Transaction response time-related metrics
[0062] Step 3.1.2.1: Record the transaction response time series R 1 , R 2 , …, R n , reflecting the response efficiency of each transaction;
[0063] Step 3.1.2.2: Calculate the average transaction response time based on the response time series Reflecting the average response time of transactions:
[0064]
[0065] Step 3.1.2.3: Calculate the standard deviation σ of the transaction response time based on the average response time R , reflecting the fluctuation of the transaction response time:
[0066]
[0067] Step 3.1.3: Metrics related to the transaction confirmation time
[0068] Step 3.1.3.1: Record the transaction confirmation time series C 1 , C 2 , …, C n , reflecting the completion efficiency of each transaction;
[0069] Step 3.1.3.2: Calculate the average transaction confirmation time based on the transaction confirmation time series Reflecting the average confirmation time of the transaction:
[0070]
[0071] Step 3.1.3.3: Calculate the standard deviation σ of the transaction confirmation time based on the average confirmation time C , reflecting the degree of fluctuation of the transaction confirmation time:
[0072]
[0073] A short transaction response time and transaction confirmation time mean a fast transaction speed, and vice versa. This is in line with the situation in Ponzi schemes in blockchain transactions where there are fast transactions and rebates in the early stage to attract investors, and delays in rebates or even prevention of investors from withdrawing funds in the later stage. A large standard deviation of the transaction response time and the transaction confirmation time reflects the extremely unstable characteristics of the response time and confirmation time of Ponzi scheme transactions.
[0074] Step 3.1.4: Metrics related to the transaction time deviation
[0075] Step 3.1.4.1: Calculate the transaction time deviation D = |T i - T gvg |, reflecting the difference between the transaction time and the overall market trading active time, where T avg represents the overall market trading active time;
[0076] Step 3.1.4.2: Calculate the average transaction time deviation based on the transaction time deviation D Reflecting the overall situation of the transaction time deviation:
[0077]
[0078] Step 3.1.4.3: Calculate the standard deviation σ of the trading time deviation D , reflecting the degree of fluctuation of the trading time deviation:
[0079]
[0080] Step 3.1.5: Transaction frequency related indicators
[0081] Step 3.1.5.1: Transaction number sequence N 1 , N 2 , …, N n , reflecting the number of transactions within each period of time;
[0082] Step 3.1.5.2: Calculate the total number of transactions N for a specific address total , reflecting the overall activity level of the transactions;
[0083] Step 3.1.5.3: Calculate the transaction frequency F i , reflecting the activity level of the transactions:
[0084]
[0085] Step 3.1.5.4: Calculate the average transaction frequency Reflecting the average activity level of the transactions:
[0086]
[0087] Step 3.1.5.5: Calculate the standard deviation σ of the transaction frequency F , reflecting the fluctuation of the transaction frequency:
[0088]
[0089] Step 3.1.5.6: Calculate the slope m of the transaction frequency change trend, reflecting the change trend of the transaction activity level over time:
[0090]
[0091] where the transaction frequency sequence is F 1 , F 2 , …, F n , and the corresponding time sequence is t 1 , t 2 , …, t n ;
[0092] The indicators related to the trading frequency accurately reflect the characteristics of the trading frequency in a Ponzi scheme. During the development of a Ponzi scheme, to attract investors and maintain operations, the trading frequency usually shows abnormal changes. In the initial stage, an illusion of high-frequency trading is created, and when the capital chain becomes tight in the later stage, the trading frequency drops sharply, resulting in abnormal fluctuations in the trading frequency. The slope of the change trend of the trading frequency in the early stage of a Ponzi scheme is usually positive and has a large value, indicating a rapid increase in the trading frequency; but as the scheme gradually collapses, the slope quickly becomes negative. The change in the positive and negative values of the slope more intuitively reflects the development of the Ponzi scheme. These statistics can accurately capture the trend of frequency changes and identify potential frauds.
[0093] Step 3.1.6: Periodicity indicator of trading time
[0094] Calculate the periodicity coefficient of trading time Judge whether there is periodicity in trading time, where T p represents the period, where is the average trading time interval;
[0095] Spatial dimension
[0096] To create an illusion of active capital flow, Ponzi schemes in blockchain transactions often conduct transactions by manipulating multiple related addresses. These addresses interact frequently, and the relationships between trading addresses show abnormal characteristics. For example, a few core addresses will conduct a large number of transactions with many other addresses, forming a complex trading network.
[0097] To avoid supervision and cover up the flow of funds, Ponzi schemes in blockchain transactions usually show abnormal concentration or dispersion in the geographical distribution of transactions. If it is concentrated in certain specific geographical areas, it may be to facilitate the manipulation and control of funds; if it is abnormally dispersed, it may be an attempt to confuse the line of sight through multi-location transactions and avoid supervision.
[0098] The above characteristics cannot be directly reflected in blockchain transaction data, and there is currently a lack of effective mathematical formulas to extract these characteristics.
[0099] Therefore, in the refined spatio-temporal feature engineering, mathematical methods are used for analysis and modeling, and mathematical formulas are used to calculate and extract the characteristics of Ponzi schemes. For this purpose, the present invention further provides the following detailed steps:
[0100] Step 3.2: Spatial dimension feature extraction
[0101] Step 3.2.1: Record the transaction initiation address A src , the transaction receiving address A dst ;
[0102] Step 3.2.2: Calculate the capital inflow I of each address A, reflecting the total amount of funds received at a specific address:
[0103]
[0104] where M i represents the transaction amount;
[0105] Step 3.2.3: Calculate the fund outflow O of each address A , reflecting the total amount of funds sent from a specific address:
[0106]
[0107] Step 3.2.4: According to the number of counterparty addresses n opp and the total number of transactions N total , calculate the counterparty concentration C opp , reflecting the degree of concentration of counterparties over a period of time:
[0108]
[0109] Step 3.2.5: According to the number of transactions N of new counterparty addresses new and the total number of transactions N total , calculate the transaction ratio P of new counterparties new , reflecting the frequency of transactions with new counterparts at a specific address over a period of time:
[0110]
[0111] Step 3.2.6: According to the active duration sequence T of relevant counterparty addresses act1 , T act2 , …, T actn , calculate the average active duration of counterparty addresses reflecting the active behavior of counterparties:
[0112]
[0113] Step 3.2.7: Sort out all pairs of transaction addresses of the trading entity over a period of time. The number of transactions between the address pair (A i , A j ) is N ij , the total transaction amount is V ij , and the total number of transactions and total transaction amount of all address pairs are N total and V total , calculate the address correlation index I rel , reflecting the tightness of the trading relationship between transaction addresses:
[0114]
[0115] Calculation Since both the number of transactions and the transaction amount can reflect the degree of association between address pairs, combining the two and taking the square root can comprehensively measure the association strength of each pair of addresses, and then summing them up to obtain the comprehensive value of the association strength of all address pairs.
[0116] Denominator Serves the purpose of standardization. By dividing by it, the calculated address correlation index can be made more comparable under different data sets.
[0117] Address correlation index I rel The higher the index, the closer the transaction relationship between the addresses; conversely, it indicates a weaker transaction relationship.
[0118] Using the address correlation index can effectively quantify the tightness of the transaction relationship between these address pairs. In normal transactions, the correlation between addresses is relatively dispersed and there are no obvious clustering characteristics, and the index value is within the normal range. In a Ponzi scheme, due to a large amount of funds circulating among related addresses, the total number of transactions and the total amount of transactions of related address pairs will increase, resulting in an abnormal increase in the address correlation index. If the correlation index between certain address pairs in the blockchain transaction data is too high and does not conform to the normal transaction logic, it means that there is fraudulent fund transfer. Using the address correlation index helps to identify Ponzi schemes.
[0119] Step 3.2.8: Record the scale of fund transfer Q between related addresses transfer , reflecting the scale of fund flow between addresses;
[0120] Step 3.2.9: Calculate the fund transfer frequency f between related addresses according to the number of transfers n transfer and the total number of transactions N total , reflecting the fund transfer behavior between addresses: transfer
[0121]
[0122] Step 3.2.10: Calculate the geographical distribution concentration index I according to the number of transactions N in different geographical regions in all transactions of a certain transaction address within a period of time 1 , N 2 , …, N m , and the total number of transactions is N total , measuring the degree of concentration of transactions in different geographical regions and reflecting the spatial distribution of transactions: geo
[0123]
[0124] Among them
[0125] ①N i : It represents the number of transactions in different geographical regions in all transactions of a certain trading address within a period of time. m is the total number of geographical regions, and N i varies by region, reflecting the trading activity of each region.
[0126] ② is the average value of the number of transactions in different geographical regions, obtained by averaging N i . It represents the average level of the number of transactions in each region.
[0127] ③ is the sum of the squares of the differences between the number of transactions in each geographical region and the average value. The larger this sum value is, the greater the difference between the number of transactions in each region and the average value, that is, the more uneven the distribution of transactions in geographical regions.
[0128] ④ m is the number of regions, is the square of the average transaction quantity, and the denominator is used to standardize the numerator.
[0129] ⑤Geographical distribution concentration index I geo The larger the value of, the higher the degree of concentration of transactions in geographical regions; the smaller the value, the more uniform the distribution of transactions in different geographical regions.
[0130] The geographical distribution concentration index can measure the degree of concentration of transactions in different geographical regions. Normal blockchain transactions will show a relatively balanced state in geographical distribution, and the index value is relatively stable. However, in a Ponzi scheme, due to the artificial manipulation of its trading behavior, there will be a large difference in the number of transactions in different geographical regions, causing the geographical distribution concentration index to deviate from the normal range. Calculating this index can effectively identify Ponzi schemes.
[0131] In the above technical solution, step 4 includes the following steps:
[0132] Step 4.1: Perform LSTM processing on the time dimension features to obtain time series features F LSTM :
[0133] F LSTM = LSTM(X)
[0134] X is the input time dimension feature;
[0135] Step 4.2: Perform CNN processing on the spatial dimension features to obtain trading address features F CNN :
[0136] F CNN = CNN(X)
[0137] X is the input spatial dimension feature;
[0138] Specific detailed problems that need to be overcome in the Adaboost algorithm for integrating LSTM and CNN models:
[0139] In the above technical solution, step 5 includes the following steps:
[0140] Step 5: Adaboost integration: For F LsTM and F CNN Perform feature alignment and fusion, and form the final classifier P(y|X) by iteratively training weak classifiers and dynamically weighted combination.
[0141] Step 5.1: Feature alignment:
[0142] Obtain the output features F LSTM and F CNN of the long short-term memory network LSTM model and the convolutional neural network CNN model;
[0143] If the dimension of F LSTM is less than that of F CNN , the dimension of F LSTM can be expanded through a fully connected layer to match the dimension of F CNN ;
[0144] If the dimension of F CNN is greater than that of F LSTM , dimensionality reduction can be performed on F CNN by applying global average pooling to align its dimension with that of F LSTM .
[0145] Step 5.2: Feature fusion:
[0146] Concatenate the aligned F LSTM and F CNN features to form a high-dimensional feature vector F 融合 .
[0147] When concatenating, place the features of LSTM in the first half of the vector and the features of CNN in the second half to retain their complementary information, combine the two features into a unified representation, and provide higher-quality input for the integrated model.
[0148] Step 5.3: Format regularization:
[0149] Check whether the fused feature vector F 融合 meets the input format requirements of the Adaboost integrated model,
[0150] If not, convert F 融合 to a two-dimensional array form to ensure that each sample corresponds to a row and each feature corresponds to a column.
[0151] Step 5.4: Adaboost Integration:
[0152] Use the Adaboost algorithm to iteratively train the fused feature F 融合 to generate multiple weak classifiers. Dynamically combine them with weights according to the performance of each weak classifier. For samples correctly classified by the current weak classifier, appropriately reduce their weights so that the model pays less attention to these samples in subsequent training; for samples misclassified, significantly increase their weights so that the model focuses more on these difficult-to-classify samples. At the same time, the weight allocation of the weak classifiers will also be dynamically adjusted according to the model's performance on the validation set. If a weak classifier performs well on the validation set and greatly improves the overall model performance, increase its weight proportion in the final integrated model; conversely, if a weak classifier performs poorly, reduce its weight.
[0153] Through multiple iterations, Adaboost combines multiple weak classifiers into a strong classifier P(y|X), and uses P(y|X) to identify Ponzi schemes in blockchain transaction data, and finally gives a clear classification result of whether it is a Ponzi scheme or a normal transaction.
[0154] It is expressed by the following mathematical formula:
[0155] P(y|X) = Adaboost(F CNN , F LSTM )
[0156] P(y|X) represents that Adaboost generates the final classification result P(y|X) by integrating F CNN and F LSTM , where X is the input feature after feature alignment, fusion, and regularization, and y is the target class label.
[0157] Through the above technical solutions, the present invention systematically solves the key problem of identifying Ponzi schemes in blockchain transactions and achieves remarkable technical effects:
[0158] 1. Multi-source data fusion significantly improves the comprehensiveness of features
[0159] By integrating blockchain transaction records, transaction platform behavior data, and regulatory agency marked data (Step 1), the present invention makes up for the defect of traditional methods relying on a single data source. Experiments show that the introduction of multi-source data expands the feature data to 22 items in the time dimension and 10 items in the space dimension. Among them, the mining of new features such as the autocorrelation coefficient ρ(k) of the trading time interval, the address correlation index I rel , and the geographical distribution concentration index I geo improves the feature coverage. For exampleFigure 3 As shown, compared with the method that only uses single-source data, the recall rate of this method has increased from 65% (rule-based recognition method) to 88%, solving the problem of incomplete feature extraction caused by single data in traditional methods.
[0160] 2. The accuracy rate is significantly leading, and the contributions of feature engineering and model integration optimization are prominent
[0161] Based on the multi-source data fusion (step 1) and refined spatio-temporal feature engineering (step 3) proposed by the present invention, the model successfully captures the core features of Ponzi schemes (such as the slope m of the trading frequency change trend (step 3.1.5.6), the autocorrelation coefficient ρ(k) of the trading time interval (step 3.1.1.6), the geographical distribution concentration index I geo (step 3.2.10), etc.), achieving an accuracy rate of 92% in the test set (1000 normal transactions + 100 Ponzi schemes), which is 29.6% higher than the rule-based recognition method (71%) and 10.8% higher than the random forest model (83%). The key performance gains come from:
[0162] Deep spatio-temporal feature extraction: Through features such as the slope m of the trading frequency change trend (step 3.1.5.6), the autocorrelation coefficient ρ(k) of the trading time interval (step 3.1.1.6), and the geographical distribution concentration index I geo (step 3.2.10), etc., covering the dynamic laws before and after the fraud, reducing the misjudgment rate (72.4% lower than the rule-based recognition method and 52.9% lower than the random forest).
[0163] Model integration optimization: The Adaboost algorithm fuses the features output by LSTM and CNN (step 5.4), and the recognition accuracy rate is increased to 92%, which is significantly better than a single model.
[0164] 3. The recall rate has increased significantly, and the risk coverage is more comprehensive
[0165] In terms of the recall rate index, the present invention reaches 88% (88 out of 100 Ponzi scheme samples are successfully identified), which is 35.4% higher than the rule-based method (65%) and 14.3% higher than the random forest (77%). Thanks to:
[0166] Complementary multi-source data: The fusion of trading platform behavior data (step 1.2) and regulatory marking data (step 1.3) solves the problem of missed detection of hidden Ponzi schemes in traditional methods.
[0167] Dynamic feature design: The standard deviation σ of the average trading time interval T (step 3.1.1.4), the autocorrelation coefficient ρ(k) of the trading time interval (step 3.1.1.6), the standard deviation σ of the trading response time R(Step 3.1.2.3), Standard Deviation σ of Transaction Confirmation Time C (Step 3.1.3.3), Standard Deviation σ of Transaction Time Deviation D (Step 3.1.4.3), Standard Deviation σ of Transaction Frequency F (Step 3.1.5.5), Slope m of Transaction Frequency Change Trend (Step 3.1.5.6) Captures Abnormal Fluctuations during the Break Period of the Scam Fund Chain, Improving the Recall Rate in the Middle and Late Stages of the Scam.
[0168] 4. The F1 value is more optimal comprehensively, enhancing the robustness of the model
[0169] The F1 value of the present invention reaches 90% (combining precision and recall), comprehensively superior to the rule-based method (68%) and random forest (79%), solving the problem of poor recognition effect caused by poor algorithm adaptability in traditional solutions:
[0170] Feature Alignment Technique (Step 5.1): Align the feature dimension differences output by LSTM and CNN from the initial 32 vs. 128 to a unified dimension of 64 through the fully connected layer and pooling operation, reducing feature redundancy and improving the model training convergence speed. This process not only optimizes the expression of features but also provides high-quality input data for subsequent model training, laying a foundation for the improvement of model performance.
[0171] Adaboost Dynamic Weighting:
[0172] In traditional Adaboost applications, the update of sample weights often follows fixed rules, lacking flexibility when facing complex and changing data distributions. The Adaboost dynamic weighting proposed in the present invention innovatively introduces a dynamic adjustment strategy. It monitors the classification situation of the model during training in real time, dynamically adjusts the weights of samples according to the classification difficulty of different samples in each iteration, and dynamically adjusts the weight allocation of weak classifiers according to the validation performance.
[0173] Specifically, for samples correctly classified by the current weak classifier, appropriately reduce their weights so that the model pays less attention to these samples in subsequent training; for samples misclassified, significantly increase their weights to make the model focus more on these difficult-to-classify samples. At the same time, the weight allocation of weak classifiers will also be dynamically adjusted according to the performance of the model on the validation set. If a weak classifier performs well on the validation set and contributes greatly to the overall model performance, increase its weight proportion in the final integrated model; conversely, if a weak classifier performs poorly, reduce its weight.
[0174] Through this Adaboost dynamic weighting mechanism, the problem of insufficient adaptability of the model to complex data in the traditional solution is solved. In the traditional method, when dealing with data with unbalanced distribution and complex and variable features, overfitting or underfitting easily occurs, resulting in unstable recognition effects. However, the dynamic weighting mechanism of the present invention enables the model to continuously adapt to data changes during the training process, effectively improving the robustness and classification accuracy of the model.
[0175] 5. The deep learning model of spatio-temporal features breaks through the bottleneck of algorithm adaptability, and the dynamic recognition accuracy is significantly improved
[0176] Aiming at the defect that the existing algorithms cannot adapt to the dynamic changes of Ponzi schemes, the present invention innovatively adopts an LSTM-CNN dual-channel architecture to process spatio-temporal features, significantly improving the accuracy of dynamic recognition. Specifically, the LSTM model accurately captures the abnormal trading rhythm in the life cycle of Ponzi schemes through time-dimensional features such as the autocorrelation coefficient of trading time intervals (step 3.1.1.6), the standard deviation of trading time deviation (step 3.1.4.3), and the slope of the changing trend of trading frequency (step 3.1.5.6); the CNN model effectively identifies anomalies in the topology of the fund network through spatial-dimensional features such as the address correlation index, the fund transfer frequency between associated addresses, and the geographical distribution concentration index.
[0177] Through the synergistic effect of the LSTM-CNN dual-channel architecture, the present invention not only breaks through the bottleneck of traditional algorithms in adaptability, but also gives full play to the advantages of the two models in processing spatio-temporal features, and more importantly, realizes the dynamic recognition of Ponzi schemes. This deep learning model of spatio-temporal features can comprehensively and dynamically capture the spatio-temporal features of transactions, thus more accurately identifying abnormal trading patterns.
[0178] 6. The Adaboost dynamic integration mechanism overcomes the problem of model robustness, and the recognition stability in complex scenarios is enhanced
[0179] Aiming at the problem that a single model is vulnerable to noise interference, an Adaboost integration method based on feature alignment is invented. The LSTM-CNN feature space alignment is realized by expanding the dimension through a fully connected layer and reducing the dimension through global average pooling, and the spatio-temporal feature vectors are fused by using a splicing complementary strategy. The sample weights are dynamically adjusted and the weight distribution of the classifier is dynamically adjusted during iterative training. The Adaboost dynamic integration mechanism of the present invention overcomes the problem of the robustness of the Ponzi scheme recognition model in blockchain transactions. Experimental data shows that this integration mechanism improves the F1 value from 0.68 of the rule-based recognition method to 0.90, and the recognition stability in complex scenarios is significantly enhanced.
[0180] 7. The full-process technological innovation realizes a breakthrough in comprehensive performance, and the key indicators are significantly better than the existing methods
[0181] By constructing a technical closed-loop of "data fusion - feature engineering - model integration", breakthrough results have been achieved in the test of real datasets. Compared with the rule-based method, the accuracy rate has increased from 71% to 92%, significantly improving the recognition accuracy; compared with the random forest model, the recall rate has leaped from 77% to 88%, effectively solving the problem of missed reports. The F1 value of the solution of the present invention reaches 90%, an increase of 11 percentage points compared with the existing optimal technology. BRIEF DESCRIPTION OF THE DRAWINGS
[0182] Figure 1 It is the system architecture diagram corresponding to the method of the present invention;
[0183] Figure 2 It is the schematic diagram of the Adaboost integration process;
[0184] Figure 3 It is the comparison of the accuracy rate, recall rate, and F1 value of different methods;
[0185] Figure 4 It is the bar chart of the experimental data. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0186] The following will give a detailed description of the embodiments of the present invention. Although the present invention will be described and explained in conjunction with some specific embodiments, it should be noted that the present invention is not limited to these embodiments only. On the contrary, any modifications or equivalent replacements made to the present invention should be covered within the scope of the claims of the present invention.
[0187] In addition, in order to better illustrate the present invention, numerous specific details are given in the following detailed description. Those skilled in the art will understand that the present invention can be implemented without these specific details.
[0188] The object of the present invention is to solve the following key technical problems:
[0189] 1. Traditional methods rely on manual analysis and simple rules, with low precision
[0190] In the prior art, traditional Ponzi scheme identification methods mainly rely on manual analysis and simple rule setting. This method is not only inefficient but also difficult to cope with complex and changeable scam patterns, resulting in low recognition precision.
[0191] 2. Single data source and incomplete feature extraction
[0192] Some technical solutions based on data analysis only rely on a single data source and fail to fully utilize multi-source data (such as blockchain transaction data, trading platform data, and regulatory data) for comprehensive analysis, resulting in incomplete feature extraction and difficulty in capturing the hidden features of Ponzi schemes.
[0193] 3. Poor algorithm adaptability and inability to capture complex dynamic changes
[0194] Existing algorithms have deficiencies in adaptability and are unable to effectively handle the dynamic changes of Ponzi schemes in blockchain transactions. Especially when dealing with complex spatio-temporal features, it is difficult to accurately identify abnormal patterns.
[0195] 4. Insufficient optimization of spatio-temporal feature extraction and model integration, resulting in low recognition accuracy
[0196] Existing models are not fine enough in spatio-temporal feature extraction and fail to fully exploit the time and space features in blockchain transactions. In addition, the model integration method is not optimized enough to make full use of the advantages of different models, resulting in a relatively low overall recognition accuracy.
[0197] The existence of these problems makes it difficult for existing Ponzi scheme recognition methods to achieve efficient and accurate recognition in blockchain transactions. There is an urgent need for a solution that combines spatio-temporal features and deep learning technology to improve the accuracy and robustness of recognition.
[0198] To achieve the above objectives, the present invention provides a method for identifying Ponzi schemes in blockchain transactions, including the following steps:
[0199] Step 1: Multi-source data collection: Integrate blockchain transaction records, transaction platform behavior data, and regulatory agency marked data to obtain complete and diverse transaction data;
[0200] Step 2: Data preprocessing: Form structured data through noise filtering, missing value filling, and time series construction;
[0201] Step 3: Spatio-temporal feature engineering:
[0202] Time dimension feature extraction: Include calculating the transaction time interval sequence and its statistics, the standard deviation of transaction response time, the standard deviation of transaction confirmation time, the slope of the transaction frequency change trend, and the autocorrelation coefficient ρ(k) of the transaction time interval;
[0203] Space dimension feature extraction: Include calculating the address correlation index I rel and the geographical distribution concentration index I geo ;
[0204] Step 4: Perform LSTM and CNN processing on the refined spatio-temporal features of the transaction data to extract time series features and transaction address features respectively, obtaining the time series feature F LSTM and the transaction address feature F CNN ;
[0205] Step 5: Adaboost integration: For F LsTM and F CNNPerform feature alignment and fusion, and form the final classifier P(y|X) by iteratively training weak classifiers and dynamically weighting and combining them.
[0206] For step 1, the present invention provides the following embodiments:
[0207] Step 1.1: Obtain detailed transaction records from blockchain nodes, including: ① The timestamp of each transaction, accurate to seconds or smaller time units. ② The amount of money involved in the transaction. ③ The addresses of the transaction initiator and recipient.
[0208] Step 1.2: Collect users' transaction behavior data from mainstream trading platforms, including detailed dimensions such as transaction frequency, transaction habits, and the correlation information between the two parties of the transaction.
[0209] Step 1.3: Cooperate with authoritative regulatory agencies to obtain data on confirmed Ponzi scheme cases. For blockchain transaction data that has been clearly identified as a Ponzi scheme, mark these data as "1" (indicating it is a Ponzi scheme), and mark normal transaction data as "0" (indicating it is not a Ponzi scheme).
[0210] For step 2, the present invention provides the following embodiments:
[0211] Step 2.1: Data cleaning.
[0212] Noise filtering: Remove obviously abnormal data by setting reasonable thresholds for transaction amount and transaction frequency.
[0213] Delete duplicate data: Remove duplicates from the collected blockchain transaction records to ensure data uniqueness.
[0214] Eliminate invalid data: Eliminate data entries with zero transaction amount, invalid addresses, or abnormal timestamps.
[0215] Step 2.2: Missing value handling.
[0216] Use the mean filling method to handle missing values in continuous features such as transaction amount and transaction frequency.
[0217] For missing timestamps, perform time series completion through interpolation.
[0218] For missing values in discrete features (such as transaction addresses), set them to the default value "unknown" to avoid information loss.
[0219] Step 2.3: Time series construction.
[0220] Based on the transaction timestamps, sort the transaction data in chronological order to construct time series data, ensuring that the order of the time series is increasing in chronological order for convenient subsequent data analysis and processing.
[0221] Wherein for step 3, the present invention provides the following embodiment:
[0222] 1. Time dimension:
[0223] 1. In blockchain transactions, the transaction time-related characteristics of Ponzi schemes are often unstable and volatile. In the early stage, due to the need to quickly attract funds and maintain the capital chain, a large amount of funds will flow in and out of a specific address in a short period of time, and the transactions will be fast and frequent; in the later stage, due to the break of the capital chain or abnormal manipulation, the capital flow is small, the transactions are slow, and the transactions are deserted. This feature is in sharp contrast to the relatively stable law of normal transactions. This feature cannot be directly reflected in the redundant data exported by the blockchain and trading platform, and it is also difficult to extract the relevant logical relationship. At present, there is a lack of comprehensive and accurate mathematical formulas to mine these features.
[0224] Therefore, mathematical methods are used for analysis and modeling in the refined spatiotemporal feature engineering, and mathematical formulas are used to calculate and refine the characteristics of the Ponzi scheme from six angles: transaction time interval, transaction response time, transaction confirmation time, transaction time deviation, transaction frequency, and transaction time periodicity.
[0225] Trading time interval perspective:
[0226] Record the timestamp T of each transaction i , forming a time series T 1 ,T 2 ,…,T n
[0227] Based on the transaction time series, calculate the transaction time interval series I 1 ,I 2 ,…,I n , where I i =T i -T i-1
[0228] Average transaction time interval:
[0229] Average transaction time interval standard deviation: reflects the fluctuation degree of average transaction time interval:
[0230]
[0231] If the average transaction time interval is too short or the standard deviation is extremely small, it means that funds are flowing quickly in a specific address, which is consistent with the characteristics of a Ponzi scheme that quickly makes money and operates frequently when the capital chain is tight; conversely, if the average transaction time interval is too long or the standard deviation is too large, it reflects that the flow of scam funds is blocked or there is abnormal manipulation.
[0232] Calculate the probability density distribution p(I) of the trading time interval to show the distribution of the trading time interval:
[0233]
[0234] Trading response time perspective:
[0235] The trading response time series over a period of time is denoted as R 1 , R 2 , …, R n
[0236] Average trading response time:
[0237] The standard deviation of the trading response time, reflecting the fluctuation of the trading response time:
[0238]
[0239] Trading confirmation time perspective:
[0240] The trading confirmation time series over a period of time is denoted as C 1 , C 2 , …, C n ;
[0241] Average trading confirmation time:
[0242] The standard deviation of the trading confirmation time: Reflecting the degree of fluctuation of the trading confirmation time:
[0243]
[0244] A short trading response time and trading confirmation time mean a fast trading speed, and vice versa means a slow trading speed. This is in line with the situation in blockchain transactions where Ponzi schemes conduct rapid transactions and offer rebates in the early stage to attract investors, and then delay rebates or even prevent investors from withdrawing their funds in the later stage. A large standard deviation of the trading response time and the trading confirmation time reflects the extremely unstable characteristics of the response time and confirmation time of Ponzi scheme transactions.
[0245] Trading time deviation perspective:
[0246] Calculate the trading time deviation D = |T i - T avg |, reflecting the difference between the trading time and the overall trading active time of the market, where T avg represents the overall trading active time of the market;
[0247] Calculate the average trading time deviation based on the trading time deviation D Reflecting the overall situation of the trading time deviation:
[0248]
[0249] Calculate the standard deviation σ of the trading time deviation D , reflecting the degree of fluctuation of the trading time deviation:
[0250]
[0251] From the perspective of trading frequency:
[0252] Statistical trading times sequence N 1 , N 2 , …, N n , reflecting the number of transactions within each period of time
[0253] Calculate the total number of transactions N at a specific address total , reflecting the overall activity of transactions
[0254] Trading frequency: reflecting the degree of trading activity, trading frequency The trading time series corresponding to the trading frequency series within a period of time is expressed as F 1 , F 2 , …, F n
[0255] Average trading frequency:
[0256] The standard deviation of trading frequency: reflecting the fluctuation of trading frequency:
[0257]
[0258] The slope of the trading frequency change trend: reflecting the change trend of trading activity over time. Let the trading frequency series be F 1 , F 2 , …, F n , and the corresponding time series be t 1 , t 2 , …, t n ,
[0259] Slope where
[0260] Data from the perspective of trading frequency accurately reflects the relevant characteristics of the trading frequency of Ponzi schemes. During the development of Ponzi schemes, in order to attract investors and maintain operations, the trading frequency usually shows abnormal changes. In the initial stage, an illusion of high-frequency trading is created, and when the capital chain becomes tight in the later stage, the trading frequency will drop sharply, resulting in abnormal fluctuations in the trading frequency. The slope of the change trend of the trading frequency in the early stage of the Ponzi scheme is usually positive and has a large value, indicating that the trading frequency rises rapidly; however, as the scam gradually collapses, the slope will quickly become negative. The change of the positive and negative values of the slope more intuitively reflects the development of the Ponzi scheme. These statistics can accurately capture the trend of frequency change and identify potential scams.
[0261] From the perspective of the periodicity of trading time:
[0262] Calculate the periodicity coefficient of trading time Judge whether there is periodicity in trading time, where T p represents the period, is the average trading time interval
[0263] 2. In blockchain transactions, normal transactions are usually independent of each other, and the time interval of the previous transaction has no significant impact on subsequent transactions. However, the trading behavior of Ponzi schemes shows obvious correlations. For example, in order to maintain the stability of the capital chain, scammers will conduct transactions according to a certain pattern, making the trading time intervals show regularity. If the time interval of the previous transaction is short, the time interval of the next transaction is also often short, so as to accelerate the speed of capital turnover. This regularity leads to a strong positive correlation in the trading time intervals, in sharp contrast to the randomness of normal transactions. At the same time, as the scam develops and faces situations such as financial pressure and supervision, the trading rhythm changes, and the autocorrelation coefficient will fluctuate significantly. These characteristics cannot be directly reflected in blockchain transaction data, and there is currently no accurate mathematical formula to extract this feature.
[0264] We use mathematical methods to define the autocorrelation coefficient of trading time intervals: reflecting the correlation between trading time intervals.
[0265] Let the trading time interval sequence be I 1 , I 2 , …, I n , and the autocorrelation coefficient
[0266] This formula measures the linear correlation between two time intervals separated by k intervals in the trading time interval sequence.
[0267] The autocorrelation coefficient (ρ(k)): is the calculation result of the formula, which measures the tightness of the linear relationship between two values separated by k time intervals in the trading time interval sequence. k is the lag order, indicating the number of intervals between two trading time intervals in the sequence.
[0268] Covariance (Cov(I i ,I i+k )): Covariance is used to measure the overall error between two variables. In the context of trading time intervals, Cov(I i ,I i+k ) represents the covariance between the i-th time interval I i and the (i + k)-th time interval I i+k in the trading time interval sequence. If the covariance is positive, it indicates that I i and I i+k tend to vary in the same direction, that is, when I i is large, I i+k also tends to be large; if the covariance is negative, they tend to vary in the opposite direction; when the covariance is 0, there is no linear correlation between the two. Its calculation formula is:
[0269] where E represents the expectation, is the mean of the trading time interval sequence.
[0270] Variance (Var(I i ), Var(I i+k )): Variance is used to measure the degree of dispersion of a single variable. Var(I i ) is the variance of the i-th trading time interval I i , and Var(I i+k ) is the variance of the (i + k)-th trading time interval I i+k . The larger the variance, the greater the degree to which the value of this time interval deviates from its mean, and the more dispersed the data. The calculation formula for variance is where X represents I i and I i+k respectively.
[0271] Overall meaning of the formula: The autocorrelation coefficient ρ(k) is calculated by dividing the covariance by the square root of the product of the variances of the two variables. This is done to standardize the covariance so that the value of the autocorrelation coefficient is always between -1 and 1. -1 indicates a perfect negative correlation between two trading time intervals, 1 indicates a perfect positive correlation, and 0 indicates no linear correlation. By calculating the autocorrelation coefficients for different values of k, we can comprehensively understand the correlation of the trading time interval sequence at different lag orders, and thus discover whether there is a certain potential pattern or trend between trading time intervals. For example, if ρ(1) is large and positive, it indicates a strong positive correlation between adjacent trading time intervals, that is, when the current trading time interval is long, the next trading time interval also tends to be long.
[0272] In blockchain transactions, the time intervals of normal transactions are usually random, and the autocorrelation coefficient is close to zero. However, the time intervals of Ponzi schemes often exhibit regularity, and the autocorrelation coefficient often deviates significantly from zero. For example:
[0273] Positive correlation: If ρ(k)>0, it indicates that there is a certain regularity in the transaction time interval. For example, in order to quickly attract funds or maintain the capital chain.
[0274] Negative correlation: If ρ(k)<0, it indicates that there is an alternating pattern in the transaction time interval. For example, in order to conceal the abnormality of capital flow.
[0275] Volatility: As the scam develops, the autocorrelation coefficient will show large fluctuations, especially when the capital chain is tight or under regulatory pressure.
[0276] By calculating and analyzing the autocorrelation coefficient of the transaction time interval, these abnormal correlations and fluctuations can be captured, thus effectively identifying Ponzi schemes in blockchain transactions.
[0277] II. Spatial Dimension
[0278] 1. In order to create an illusion of active capital flow, Ponzi schemes in blockchain transactions often conduct transactions by manipulating multiple associated addresses. There are frequent interactions between these addresses, and the relationships of transaction addresses show abnormal characteristics. For example, a small number of core addresses will conduct a large number of transactions with many other addresses, forming a complex transaction network. Such characteristics cannot be directly reflected in blockchain transaction data, and there is currently no effective mathematical formula to extract these characteristics.
[0279] Therefore, in the refined spatio-temporal feature engineering, mathematical methods are used to analyze and model, and mathematical formulas are used to calculate and extract the characteristics of Ponzi schemes.
[0280] Record the transaction initiation address A src , the transaction receiving address A dst
[0281] Calculate the capital inflow I of each address A , reflecting the total amount of funds received by a specific address:
[0282]
[0283] where M i represents the transaction amount
[0284] Calculate the capital outflow O of each address A , reflecting the total amount of funds sent by a specific address:
[0285]
[0286] According to the number of counterparty addresses nopp and the total number of transactions is N total , calculate the counterparty concentration index C opp , reflecting the degree of concentration of counterparties:
[0287]
[0288] According to the number of transactions N of the new counterparty address new and the total number of transactions is N total , calculate the trading ratio P of the new counterparty new , reflecting the frequency of transactions between a specific address and new counterparts:
[0289]
[0290] According to the active duration sequence T of relevant counterparty addresses act1 , T act2 , …, T actn , calculate the average active duration of the counterparty address reflecting the active behavior of counterparties:
[0291]
[0292] Record the fund transfer scale Q between associated addresses transfer , reflecting the scale of fund flow between addresses
[0293] According to the number of transfers n transfer and the total number of transactions N total calculate the fund transfer frequency f between associated addresses transfer , reflecting the fund transfer behavior between addresses:
[0294]
[0295] Address correlation index: Reflecting the tightness of the trading relationship between trading addresses.
[0296] Sort out all trading address pairs of the trading entity within a period of time. The number of transactions between address pairs (A i , A j ) is N ij , and the total sum of trading amounts is V ij , and the total number of transactions and the total trading amount of all address pairs are N total and V total , then the address correlation index can be expressed as:
[0297]
[0298] is for all address pairs (A i , Aj ) of summation. Among them, N ij represents the number of transactions between address pairs (A i , A j ). V ij represents the total transaction amount between address pairs (A i , A j ). The calculation is because both the number of transactions and the transaction amount can reflect the degree of association between address pairs. By combining the two and taking the square root, the association strength of each pair of addresses can be comprehensively measured, and then summing up to obtain the comprehensive value of the association strength of all address pairs.
[0299] N total is the total number of transactions of all address pairs, and V total is the total transaction amount of all address pairs. The denominator plays a role in standardization. By dividing by it, the calculated address correlation index can be made more comparable under different data sets.
[0300] The address correlation index I rel The higher the index, the closer the transaction relationship between the two addresses; conversely, it indicates a weaker transaction relationship.
[0301] Using the address correlation index can effectively quantify the tightness of the transaction relationship between these address pairs. In normal transactions, the correlation between addresses is relatively dispersed and there are no obvious aggregation characteristics, and the index value is within the normal range. In a Ponzi scheme, due to a large amount of funds circulating among related addresses, the number of transactions and the total amount of transactions of related address pairs will increase, resulting in an abnormal increase in the address correlation index. If the correlation index between certain address pairs in the blockchain transaction data is too high and does not conform to the normal transaction logic, it means that there is fraudulent fund transfer. Using the address correlation index helps to identify Ponzi schemes.
[0302] 2. Ponzi schemes in blockchain transactions To avoid supervision and cover up the flow of funds, they usually show abnormal concentration or dispersion in the geographical distribution of transactions. If it is concentrated in certain specific geographical areas, it may be to facilitate the manipulation and control of funds; if it is abnormally dispersed, it may be an attempt to confuse the sight through multi-location transactions and avoid supervision. These characteristics cannot be directly reflected in the blockchain transaction data, and there is currently no effective mathematical formula to extract this feature.
[0303] Therefore, in the refined spatio-temporal feature engineering, we use mathematical methods to analyze and model, define the geographical distribution concentration index of transactions, and use mathematical formulas to calculate and extract such features of Ponzi schemes.
[0304] Geographical distribution concentration index: Measures the concentration of transactions in different geographical regions, reflecting the spatial distribution of transactions.
[0305] The number of transactions in different geographical regions for all transactions of a certain transaction address within a period of time is N 1 , N 2 , …, N m , and the total number of transactions is N total , then the geographical distribution concentration index can be expressed as:
[0306]
[0307] where
[0308] N i : Represents the number of transactions in different geographical regions for all transactions of a certain transaction address within a period of time. m is the total number of geographical regions, and N i varies by region, reflecting the trading activity of each region.
[0309] is the average value of the number of transactions in different geographical regions, obtained by averaging N i . It represents the average level of the number of transactions in each region.
[0310] is the sum of the squares of the differences between the number of transactions in each geographical region and the average value. The larger this sum value, the greater the difference between the number of transactions in each region and the average value, that is, the more uneven the distribution of transactions in geographical regions.
[0311] m is the number of regions, is the square of the average transaction quantity, and the denominator is used to standardize the numerator.
[0312] The geographical distribution concentration index I geo The larger the value, the higher the concentration degree of transactions in geographical regions; the smaller the value, the more uniform the distribution of transactions in different geographical regions.
[0313] The geographical distribution concentration index can measure the concentration degree of transactions in different geographical regions. Normal blockchain transactions will show a relatively balanced state in geographical distribution, and the index value is relatively stable. However, in a Ponzi scheme, due to the artificial manipulation of its trading behavior, there will be a large difference in the number of transactions in different geographical regions, causing the geographical distribution concentration index to deviate from the normal range. Calculating this index can effectively identify Ponzi schemes.
[0314] 2. Ponzi Schemes in Blockchain Transactions To evade regulation and conceal the flow of funds, Ponzi schemes in blockchain transactions often exhibit abnormal concentration or dispersion in terms of geographical distribution. If they are concentrated in certain specific geographical areas, it may be for the convenience of manipulating and controlling funds; if they are abnormally dispersed, it may be an attempt to confuse the trail through multi-location transactions and evade regulation. These characteristics cannot be directly reflected in blockchain transaction data, and there is currently a lack of effective mathematical formulas to extract such characteristics.
[0315] Therefore, in the refined spatio-temporal feature engineering, we use mathematical methods for analysis and modeling, define the geographical distribution concentration index of transactions, and use mathematical formulas to calculate and extract such characteristics of Ponzi schemes.
[0316] Geographical Distribution Concentration Index: Measures the degree of concentration of transactions in different geographical regions and reflects the spatial distribution of transactions.
[0317] Let the number of transactions in different geographical regions for a certain transaction address in all transactions over a period of time be N 1 , N 2 , …, N m , and the total number of transactions be N total , then the geographical distribution concentration index can be expressed as:
[0318]
[0319] Where
[0320] N i : Represents the number of transactions in different geographical regions for a certain transaction address in all transactions over a period of time. m is the total number of geographical regions, and N i varies by region and reflects the trading activity in each region.
[0321] is the average value of the number of transactions in different geographical regions, obtained by averaging N i . It represents the average level of the number of transactions in each region.
[0322] is the sum of the squares of the differences between the number of transactions in each geographical region and the average value. The larger this sum, the greater the difference between the number of transactions in each region and the average value, indicating that the distribution of transactions across geographical regions is more uneven.
[0323] m is the number of regions, is the square of the average number of transactions, and the denominator is used to standardize the numerator.
[0324] Geographical Distribution Concentration Index I geoThe larger the value, the higher the concentration of transactions in the geographical area; the smaller the value, the more evenly the transactions are distributed across different geographical areas.
[0325] The geographical distribution concentration index can measure the concentration of transactions in different geographical areas. Normal blockchain transactions will show a relatively balanced state in geographical distribution, and the index value is relatively stable. However, in a Ponzi scheme, due to the artificial manipulation of its trading behavior, there will be significant differences in the number of transactions in different geographical areas, causing the geographical distribution concentration index to deviate from the normal range. Calculating this index can effectively identify Ponzi schemes.
[0326] For step 4, the present invention provides the following embodiments:
[0327] Perform LSTM processing on the time dimension features to obtain the time series features F LSTM :
[0328] F LSTM = LSTM(X)
[0329] X is the input time dimension feature.
[0330] Step 4.2: Perform CNN processing on the spatial dimension features to obtain the transaction address features F CNN :
[0331] F CNN = CNN(X)
[0332] X is the input spatial dimension feature.
[0333] For step 5, the present invention provides the following embodiments:
[0334] I. Specific detailed problems that need to be overcome in integrating the LSTM and CNN models using the Adaboost algorithm:
[0335] The reasons for performing feature alignment, feature fusion, and format regularization are to solve:
[0336] There are problems such as inconsistent dimensions, differences in numerical ranges, and format mismatches in the data output by the LSTM and CNN models, and the LSTM and CNN models cannot be directly integrated using the Adaboost algorithm. Therefore, we propose a data regularization method to process the data output by the LSTM and CNN. The following are the processing steps:
[0337] 1. Feature alignment
[0338] The outputs of the LSTM and CNN are usually inconsistent in dimension and need to be converted to the same dimension for subsequent processing.
[0339] The output features of LSTM and CNN can be aligned to the same dimension by means of dimensionality reduction or expansion. If the output dimension of LSTM is smaller, a fully connected layer can be used to expand its dimension to match the output dimension of CNN. If the output dimension of CNN is larger, global average pooling can be used to reduce its dimension to align it with the output dimension of LSTM.
[0340] 2. Feature Fusion
[0341] Combine the aligned LSTM and CNN features into a unified representation, retaining the complementary information of both to provide higher-quality input for the integrated model.
[0342] Concatenate the aligned F LSRM and F CNN features to form a high-dimensional feature vector F 融合 . When concatenating, place the features of LSTM in the first half of the vector and the features of CNN in the second half to form a high-dimensional feature vector F 融合 .
[0343] 3. Format Regularization
[0344] According to the requirements of the integrated model, check whether the fused feature vector F 融合 meets the input format requirements of the Adaboost integrated model. If not, convert F 融合 to a two-dimensional array form, ensuring that each sample corresponds to a row and each feature corresponds to a column.
[0345] Through these operations, the data features extracted by LSTM and CNN can be fully integrated, combining the advantages of LSTM and CNN, enhancing the expressive power of the model, improving the performance and robustness of the model, and better identifying Ponzi schemes in blockchain transactions.
[0346] II. Adaboost Integration:
[0347] Use the Adaboost algorithm to iteratively train the fused feature F 融合 to generate multiple weak classifiers. Dynamically weight and combine them according to the performance of each weak classifier. For those samples that are correctly classified by the current weak classifier, appropriately reduce their weights so that the model pays less attention to these samples in subsequent training; for samples that are misclassified, significantly increase their weights so that the model focuses more on these difficult-to-classify samples. At the same time, the weight allocation of the weak classifiers will also be dynamically adjusted according to the performance of the model on the validation set. If a certain weak classifier performs well on the validation set and significantly improves the overall model performance, increase its weight proportion in the final integrated model; conversely, if a certain weak classifier performs poorly, reduce its weight.
[0348] Through multiple iterations, Adaboost combines multiple weak classifiers into a strong classifier P(y|X), and uses P(y|X) to identify Ponzi schemes in blockchain transaction data, finally giving a clear classification result of whether it is a Ponzi scheme or a normal transaction.
[0349] It is expressed by the following mathematical formula:
[0350] P(y|X) = Adaboost(F CNN , F LSTM )
[0351] P(y|X) means that Adaboost generates the final classification result P(y|X) by integrating F CNN and F LSTM . Among them, X is the input feature after feature alignment, fusion, and regularization, and y is the target class label.
[0352] Verification:
[0353] To verify the accuracy of the method proposed in this patent for identifying Ponzi schemes in blockchain transactions, our experiment uses the Ponzi scheme labeled dataset in the blockchain transactions released by walletexplorer. Through the known transaction address information, we can download transaction data from the blockchain and related transaction platforms. The experiment mixes 1000 normal transaction data and 100 Ponzi scheme transaction data as the test dataset according to the ratio of non-Ponzi scheme:Ponzi scheme = 10:1.
[0354] We compare two traditional methods for identifying Ponzi schemes in blockchain transactions (rule-based identification method, random forest) with the method proposed in this patent.
[0355] 1. Rule-based identification method
[0356] Transaction amount rule: Set the thresholds for large and small transactions. For example, a single transaction amount greater than 10,000 yuan is regarded as a large transaction, and less than 100 yuan is regarded as a small transaction. Count the number and proportion of large and small transactions within a certain period of time. If large transactions appear concentratedly and small transactions are frequently swiped, it may be a Ponzi scheme.
[0357] Transaction frequency rule: Calculate the number of transactions of a specific address within a certain period of time. If the transaction frequency suddenly increases significantly and does not conform to the normal transaction pattern, such as a certain address has more than 100 transactions in a day, while the average number of transactions of a normal address is less than 10 times a day, it is regarded as abnormal.
[0358] Address association rule: Analyze the relationships between transaction addresses. If it is found that a large number of new addresses frequently transact with a small number of addresses, and the participation ratio of new addresses is too high, such as the number of transactions of new addresses accounting for more than 80% of the total number of transactions, it is suspected of being a Ponzi scheme.
[0359] 2. Traditional machine learning model (taking random forest as an example)
[0360] Based on the dataset after conventional data preprocessing, all the extracted features (such as relevant features of transaction time, amount, address, etc.) are used as inputs, and the identification label (Ponzi scheme is "1", normal transaction is "0") is used as the output to train the random forest model. By adjusting the relevant parameters of the random forest (such as the number of trees, etc.), cross-validation is performed using the validation set to select the optimal parameter combination. Through the parameter optimization process, the trained random forest model can identify Ponzi schemes in blockchain transactions.
[0361] Through experiments, the accuracy rate, recall rate, and F1 value of the rule-based identification method, traditional machine learning model (random forest), and the method of the present invention on the test set are obtained respectively. The specific experimental results are as Figure 3 、 Figure 4 shown.
Claims
1. A method for identifying a Ponzi scheme in a blockchain transaction, characterized in that: The following steps are involved: Step 1: Multi-source data collection: Integrate blockchain transaction records, transaction platform behavior data, and regulatory agency tag data to obtain complete and diverse transaction data; Step 2: Data preprocessing: structured data is formed through noise filtering, missing value filling and time series construction; Step 3: Spatiotemporal feature engineering: Time dimension feature extraction: including calculation of transaction time interval sequence and its statistics, transaction response time standard deviation, transaction confirmation time standard deviation, transaction frequency change trend slope, and transaction time interval autocorrelation coefficient ρ(k); Spatial dimension feature extraction: including calculation of address relevance index I rel , Geographical distribution concentration index I geo ; Step 4: Perform LSTM and CNN processing on the refined spatiotemporal features of the transaction data to extract the time series features and transaction address features respectively, and obtain the time series features F LSTM and transaction address feature F CNN ; Step 5: Adaboost ensemble: F LSTM and F CNN Perform feature alignment and fusion, iteratively train weak classifiers and dynamically weighted combination to form the final classifier P(y|X).
2. A method for identifying a Ponzi scheme in a blockchain transaction according to claim 1, characterized in that: Step 1 includes the following steps: Integrate transaction data from multiple sources such as blockchain, trading platforms, and regulatory authorities to obtain complete and diverse transaction data; Step 1.1: Obtain transaction records from the blockchain node, including transaction timestamp, transaction amount, transaction initiation address, and receiving address; Step 1.2: Collect user transaction behavior data from the trading platform, including transaction frequency, transaction habits, and related information between the two parties of the transaction; Step 1.3: Work with regulators to obtain confirmed Ponzi scheme case data and mark the data, with Ponzi schemes marked as "1" and normal transactions marked as "0" to obtain complete and diverse transaction data.
3. A method for identifying a Ponzi scheme in a blockchain transaction according to claim 1, characterized in that: Step 2 includes the following steps: Step 2.1: Data cleaning: Remove noise data by setting transaction amount and frequency thresholds; Remove duplicate data; Eliminate invalid data, including data with zero transaction amount, invalid address or abnormal time; Step 2.2: Missing value handling: The mean filling method is used for continuous features, including transaction amount and transaction frequency; Interpolation is used to complete data with missing transaction timestamps; Set missing values in discrete features to the default value "unknown", including transaction addresses; Step 2.3: Time series construction: According to the transaction timestamp, the transaction data is sorted in chronological order to construct time series data.
4. The method for identifying a Ponzi scheme in a blockchain transaction according to claim 1, characterized in that: Step 3 includes the following steps: Step 3.1: Time dimension feature extraction The time dimension feature is used to capture the abnormal patterns of transaction behavior over time and reflect the time distribution law of the Ponzi scheme; Step 3.1.1: Trading time interval related indicators Step 3.1.1.1: Record the timestamp T of each transaction i , forming a time series T1, T2, ..., T n ; Step 3.1.1.2: Based on the transaction time series, calculate the transaction time interval series I1, I2, …, I n , where I i =T i -T i-1 ; Step 3.1.1.3: Calculate the average trading time interval Reflects the average time interval between transactions; Step 3.1.1.4: Calculate the standard deviation σ of the average trading time interval T , reflecting the degree of fluctuation of trading time intervals: Step 3.1.1.5: Calculate the probability density distribution of the trading time interval p(I) and display the distribution of the trading time interval: Where n I is the number of transactions in time interval I, and N is the total number of transactions; Step 3.1.1.6: Calculate the autocorrelation coefficient ρ(k) of the trading time interval, which reflects the correlation between the trading time intervals: k is the lag order, which indicates the number of intervals between two trading times in the sequence. Cov(I i , I i+k ) is the covariance function, Var(I i ) and Var(I i+k ) are I i and I i+k The variance of Step 3.1.2: Transaction response time related indicators Step 3.1.2.1: Record transaction response time series R1, R2, …, R n , reflects the response efficiency of each transaction; Step 3.1.2.2: Calculate the average transaction response time based on the response time series Reflects the average response time of transactions: Step 3.1.2.3: Calculate the standard deviation σ of transaction response time based on the average response time R , reflecting the fluctuation of transaction response time: Step 3.1.3: Transaction confirmation time related indicators Step 3.1.3.1: Record transaction confirmation time series C1, C2, …, C n , reflecting the completion efficiency of each transaction; Step 3.1.3.2: Calculate the average transaction confirmation time based on the transaction confirmation time series Reflects the average confirmation time of transactions: Step 3.1.3.3: Calculate the standard deviation of transaction confirmation time σ based on the average confirmation time C , reflecting the degree of fluctuation in transaction confirmation time: Step 3.1.4 Trading time deviation related indicators Step 3.1.4.1: Calculate the transaction time deviation D = |T i -T avg |, reflects the difference between trading time and the overall market trading active time, T avg Indicates the overall trading activity time of the market; Step 3.1.4.2: Calculate the average trading time deviation based on the trading time deviation D Reflecting the overall situation of transaction time deviation: Step 3.1.4.3: Calculate the standard deviation of trading time deviation σ D , reflecting the degree of volatility of trading time deviation: Step 3.1.5: Trading frequency related indicators Step 3.1.5.1: Transaction number sequence N1, N2, ..., N n , reflecting the number of transactions in each period; Step 3.1.5.2: Calculate the total number of transactions N for a specific address total , reflecting the overall activity of transactions; Step 3.1.5.3: Calculate transaction frequency F i , reflecting the activity of transactions: Step 3.1.5.4: Calculate the average transaction frequency Reflects the average activity of transactions: Step 3.1.5.5: Calculate the standard deviation of transaction frequency σ F , reflecting the fluctuation of transaction frequency: Step 3.1.5.6: Calculate the slope m of the trading frequency change trend to reflect the trend of trading activity over time: The transaction frequency sequence is F1, F2, ..., F n , the corresponding time series is t1, t2, ..., t n ; Step 3.1.6: Periodic indicators of trading time Calculate the periodicity coefficient of trading time Determine whether there is periodicity in trading time, where T p Represents the cycle, is the average transaction time interval; Step 3.2 Spatial dimension feature extraction Step 3.2.1: Record the transaction initiation address A src , transaction receiving address A dst ; Step 3.2.2: Calculate the capital inflow I of each address A , reflecting the total amount of funds received by a specific address: Among them, M i Indicates the transaction amount; Step 3.2.3: Calculate the outflow of funds from each address O A , reflecting the total amount of funds sent to a specific address: Step 3.2.4: According to the number of counterparty addresses n opp And the total number of transactions is N total , calculate the counterparty concentration index C opp , reflecting the concentration of counterparties: Step 3.2.5: Transaction number N based on the new counterparty address new And the total number of transactions is N total , calculate the new counterparty transaction ratio P new , reflecting the frequency of transactions between a specific address and a new counterparty: Step 3.2.6: According to the active time sequence T of the relevant counterparty address act1 , T act2 ,…,T actn , calculate the average active time of the counterparty address Reflect the active behavior of counterparties: Step 3.2.7: Sort out all transaction address pairs of the transaction subject within a period of time, address pair (A i , A j ) is N ij , the total transaction amount is V ij , the total number of transactions and total transaction amount of all address pairs is N total and V total , calculate the address relevance index I rel , reflecting the closeness of the transaction relationship between transaction addresses: Step 3.2.8: Record the amount of funds transferred between the associated addresses Q transfer , reflecting the scale of fund flows between addresses; Step 3.2.9: According to the number of transfers n transfer And the total number of transactions N total Calculate the frequency f of fund transfers between related addresses transfer , reflecting the fund transfer behavior between addresses: Step 3.2.10: The number of transactions in different geographical regions of all transactions of a certain transaction address within a period of time is N1, N2, ..., N m , the total number of transactions is N total , calculate the geographical distribution concentration index I geo , which measures the concentration of transactions in different geographic areas and reflects the spatial distribution of transactions: in 5. A method for identifying a Ponzi scheme in a blockchain transaction according to claim 4, characterized in that: Step 4 includes the following steps: Step 4.1: Perform LSTM processing on the time dimension features to obtain the time series features F LSTM : F LSTM =LSTM(X) Step 4.2: Perform CNN processing on the spatial dimension features to obtain the transaction address feature F CNN : F CNN =CNN(X)。 6. A method for identifying a Ponzi scheme in a blockchain transaction according to claim 5, characterized in that: Step 5 includes the following steps: Step 5.1: Feature Alignment: Get the output feature F of the long short-term memory network LSTM model and the convolutional neural network CNN model LSTM and F CNN , If F LSTM The dimension is smaller than F CNN , through the fully connected layer to F LSTM Expand the dimension to make it equal to F CNN match; If F CNN The dimension is greater than F LSTM , through F CNN Apply global average pooling to reduce the dimension to the same size as F LSTM Alignment; Step 5.2: Feature fusion: After alignment, F LSTM and F CNN The features are concatenated to form a high-dimensional feature vector F 融合 ; When concatenating, the LSTM features are placed in the first half of the vector and the CNN features are placed in the second half to preserve the complementary information of the two. Step 5.3: Formatting: Check the fused feature vector F 融合 Whether it meets the input format requirements of the Adaboost ensemble model, If not, F 融合 Convert to a two-dimensional array, ensuring that each sample corresponds to a row and each feature corresponds to a column; Step 5.4: Adaboost Ensemble: Use the Adaboost algorithm to fused the feature F 融合 Perform iterative training to generate multiple weak classifiers, and perform dynamic weighted combination based on the performance of each weak classifier to form the final strong classifier P(y|X). Use the strong classifier P(y|X) to identify Ponzi schemes in blockchain transaction data, which can be expressed as follows: P(y|X)=Adaboost(F CNN ,F LSTM ) P(y|X) represents the sum of the values of F by Adaboost. CNN and F LSTM The final classification result P(y|X) is generated by integration, where X is the input feature after feature alignment, fusion and regularization, and y is the target category label.
Citation Information
Patent Citations
Method and device for acquiring image data of block chain node and computing device
CN108564469A
Internet financial client application fraud detection method based on AdaBoost
CN112581265A
Network encryption traffic classification method and system based on multi-feature learning
CN113037730A
Improved CNN-RF-based Ethereum Pincer cheating detection method and system
CN114511330A
Block chain-based account identification method and device, storage medium and electronic equipment
CN118333624A