Ponzi scheme detection method and system based on deep learning for an ethereum smart contract

CN117336051BActive Publication Date: 2026-09-29QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311278891.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-10-07
Publication Date
2026-09-29
Estimated Expiration
2043-10-07

AI Technical Summary

Technical Problem

然而,上述方法都存在不同的缺陷

Benefits of technology

[0093]1、获取智能合约的初始特征后,基于初始特征构建多层级特征,得到粗粒度特征和细粒度特征,粗粒度特征可以概括地描绘合约整体行为,而细粒度特征则提供了详细的操作执行信息,可以从不同层次抽取智能合约的语义特征刻画智能合约的整体行为,以多层级特征为输入、通过基于深度学习方法构建的模型进行智能合约检测,提高了庞氏骗局检测的准确度;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117336051B_ABST
    Figure CN117336051B_ABST
Patent Text Reader

Abstract

The application discloses a deep learning-based Ethereum smart contract Ponzi scheme detection method and system, and belongs to the technical field of blockchains. The technical problems to be solved are that the prior art is difficult to extract smart contract Ponzi scheme features, is difficult to detect non-open-source smart contracts, and has model class bias. The method comprises the following steps: based on a plurality of opcode features and a plurality of account features, initial features are formed; based on the initial features corresponding to each historical smart contract and the class label, a data set is constructed, and the data set is balanced by using the SMOTE-Tomek method; for the initial features in the balanced data set, multi-level features of the smart contract are obtained, and based on the multi-level features of the smart contract and the class label, a sample data set is constructed; based on the sample data set, a contract recognition model is trained; and the multi-level features of the to-be-detected smart contract are input into the trained contract recognition model, and the class label of the to-be-detected smart contract is predicted and output by using the trained recognition model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of blockchain technology, specifically to a method and system for detecting Ethereum smart contract Ponzi schemes based on deep learning. Background Technology

[0002] Ethereum, a decentralized blockchain platform with smart contract functionality, has rapidly developed in the cryptocurrency field, becoming the second-largest blockchain platform by market capitalization and holding a crucial position in the digital asset sector. While Ethereum has grown rapidly, it still faces significant challenges in security, regulation, and standardization. One of its major challenges is the lack of robust and accurate smart contract auditing solutions, leading to Ponzi schemes and other fraudulent activities deployed on Ethereum, causing severe financial losses to investors. Ponzi schemes use funds from new investors to pay off previous investors, rather than through legitimate transactions or genuine investment returns. Therefore, investors are both beneficiaries and victims. Currently, Ethereum-based Ponzi schemes are causing enormous damage to investors and the Ethereum platform, urgently requiring the design of effective methods to detect and audit Ponzi schemes using smart contracts.

[0003] To address Ponzi schemes in smart contracts, scholars both domestically and internationally have proposed various detection methods, including rule-based methods, graph analysis methods, and machine learning methods. Rule-based methods rely on predefined rules or heuristics to identify Ponzi schemes; graph analysis methods analyze the transaction relationships and data flow of smart contracts to identify potential Ponzi schemes and anomalous behavior; and machine learning methods utilize large datasets to train models for Ponzi scheme detection. However, all of these methods have different drawbacks. Rule-based methods struggle to capture the evolving characteristics of Ponzi schemes and identify complex variables within smart contracts. Graph analysis methods struggle to construct feature maps of Ponzi schemes in large-scale smart contracts. Furthermore, both rule-based and graph analysis methods rely on smart contract source code, making it impossible for auditing tools to detect non-open-source smart contracts, thus limiting their detection scope. Additionally, machine learning methods struggle to extract key features to accurately identify Ponzi contracts. Moreover, only a small number of smart contracts in public datasets are identified as Ponzi schemes, leading to a severe imbalance between legitimate and Ponzi contracts, making it impossible to train unbiased models without class bias.

[0004] Existing technologies struggle to extract features of Ponzi schemes from smart contracts, are unable to detect non-open-source smart contracts, and suffer from model category bias; these are technical problems that need to be addressed. Summary of the Invention

[0005] The technical objective of this invention is to address the above-mentioned shortcomings by providing a method and system for detecting Ponzi schemes in Ethereum smart contracts based on deep learning. This addresses the technical problems of existing technologies, such as difficulty in extracting Ponzi scheme features from smart contracts, difficulty in detecting non-open-source smart contracts, and model category bias.

[0006] In a first aspect, the present invention provides a method for detecting Ethereum smart contract Ponzi schemes based on deep learning, comprising the following steps:

[0007] Multiple historical smart contracts were obtained through the Ethereum website. Each historical smart contract was labeled with a class tag, which included two types: normal contracts and Ponzi contracts. Among them, multiple historical smart contracts included both normal contracts and Ponzi contracts.

[0008] For each historical smart contract, obtain the call frequency of multiple opcodes during the execution of the smart contract, construct opcode features, and extract multiple behavioral features as account features based on the transaction information in the smart contract. The initial features are composed of multiple opcode features and multiple account features.

[0009] A dataset is constructed based on the initial features and class labels corresponding to each historical smart contract, and the dataset is balanced using the SMOTE-Tomek method to obtain a balanced dataset.

[0010] For the initial features in the balanced dataset, fine-grained and coarse-grained features are constructed to obtain multi-level features of smart contracts, and a sample dataset is constructed based on the multi-level features and class labels of smart contracts.

[0011] A contract recognition model is constructed, and the model is trained based on a sample dataset to obtain a trained contract recognition model. The contract recognition model includes a feature extractor built based on a CNN network model and a decision classifier built based on the LightGBM algorithm. The feature extractor is used to extract features from smart contracts with multi-level features as input, and the decision classifier is used to predict the class label of smart contracts with the feature data output by the feature extractor as input.

[0012] For the smart contract under test, the call frequency of multiple opcodes during the execution of the smart contract is obtained, opcode features are constructed, and multiple behavioral features are extracted as account features based on the transaction information in the smart contract. The initial features are composed of multiple opcode features and multiple account features, and fine-grained features and coarse-grained features are constructed on the initial features to obtain the multi-level features of the smart contract under test.

[0013] The multi-level features of the smart contract under test are input into the trained contract recognition model, and the class label of the smart contract under test is predicted and output by the trained recognition model.

[0014] Preferably, the call frequency of multiple opcodes during smart contract execution and the construction of opcode characteristics include the following operations:

[0015] Obtain the bytecode information of the smart contract address from the Ethereum website;

[0016] The bytecode information is converted into opcode information using a disassembler;

[0017] Opcode features are constructed based on the frequency of opcode calls.

[0018] Preferably, the initial features are constructed by combining fine-grained and coarse-grained features, including the following steps:

[0019] Define the initial features as fine-grained features;

[0020] For each initial feature, an opcode feature set is formed based on all opcode features, and an account feature set is formed based on all account features;

[0021] The account feature set and the opcode feature set are dimensionality reduced by the t-distributed random neighborhood embedding algorithm respectively. The account feature set is mapped to a low-dimensional space to obtain coarse-grained account features, and the opcode features are mapped to a low-dimensional space to obtain coarse-grained opcode features.

[0022] By combining coarse-grained account features and coarse-grained opcode features, coarse-grained features of smart contracts can be obtained.

[0023] Preferably, the feature extractor includes a preprocessing layer, a feature extraction layer, and a classification layer.

[0024] The preprocessing layer is used to perform the following: fill operations on the input fine-grained features and coarse-grained features respectively, and concatenate the filled fine-grained features and coarse-grained features into a feature matrix;

[0025] The feature extraction layer is used to perform convolution operations on the feature matrix output by the preprocessing layer as input to obtain the feature data of the smart contract.

[0026] The classification layer is used to optimize the model by taking feature data as input, predicting the output class label through the softmax function, and calculating the value of the loss function.

[0027] The decision classifier is used to predict and output the corresponding class label by taking the feature data output by the feature extraction layer as input.

[0028] Correspondingly, the contract recognition model is trained based on the sample dataset, including the following operations:

[0029] The feature extractor is trained based on the sample dataset: the multi-level features in the sample dataset are used as the input of the feature extractor. A loss function is constructed based on the class label predicted by the feature extractor and the actual class label. The feature extractor is trained by minimizing the loss function to obtain the trained feature extractor.

[0030] The decision classifier is trained based on the sample dataset: the multi-level features in the sample dataset are used as the input of the trained feature extractor, the feature data output by the feature extraction layer of the trained feature extractor is used as the input of the decision classifier, the loss function is constructed based on the class label predicted by the decision classifier and the actual class label, and the decision classifier is trained by minimizing the loss function to obtain the trained decision classifier.

[0031] Correspondingly, the trained recognition model predicts and outputs the class label of the smart contract under test, including the following steps:

[0032] The multi-level features of the smart contract under test are used as input to the trained feature classifier, and features are extracted through the trained feature classifier.

[0033] The feature data output from the feature extraction layer in the post-trained feature classifier is used as the input to the post-trained decision classifier, which then predicts the class label of the smart contract to be tested.

[0034] Preferably, account features include:

[0035] The number of investments, Num_inv, represents the amount of investment in each smart contract. If N transactions transfer funds from other accounts to a smart contract, then the formula for calculating Num_inv is as follows:

[0036]

[0037] Where Σ represents the summation symbol, and i represents the sequence number or index of the investment;

[0038] The number of payment transactions, Num_pay, represents the number of payment transactions in the smart contract. Each contract contains L payment transactions, and each payment transaction transfers funds from the smart contract to a designated transaction account. The formula for calculating Num_pay is as follows:

[0039]

[0040] Where j represents the transaction number or index;

[0041] Max_pay represents the maximum number of transactions a smart contract can initiate to another account.

[0042] The payment rate (Paid_rate) represents the percentage of investors who have received at least one payment. It is calculated by dividing the number of investors who have received payments by the total number of participating investors.

[0043] Bal balance represents the balance of a smart contract, which is the remaining or available assets after a transaction is completed.

[0044] The difference index D_ind represents the difference between the investment amount and the payment amount of all investors in the smart contract. Assuming there are p investors in the smart contract, M[i] and N[i] represent the investment amount and payment amount of the i-th investor, respectively, and V is a vector of length p used to store the difference for each investor.

[0045] The skewness s of vector V is obtained using the following formula:

[0046] V[i] = N[i] - M[i]

[0047]

[0048] After calculating the skewness s, if the number of participating investors p≥3, then D_ind=s, to ensure that there are enough participating investors in the smart contract before calculating D_ind;

[0049] The percentage of investors who invest first and then profit (Kr) represents the proportion of investors who invest before their accounts generate profits.

[0050] In a second aspect, the present invention provides a deep learning-based Ethereum smart contract Ponzi scheme detection system for identifying Ethereum smart contract Ponzi schemes by means of the method described in any one of the first aspects, the system comprising:

[0051] The data acquisition module is used to acquire multiple smart contracts through the Ethereum website. Each smart contract is labeled with a class tag, which includes two types: normal contracts and Ponzi contracts. The multiple smart contracts include both normal contracts and Ponzi contracts. The module is also used to acquire the smart contract to be tested from the Ethereum website.

[0052] The initial feature construction module is used to obtain the call frequency of multiple opcodes during the execution of the smart contract and construct opcode features for both historical smart contracts and smart contracts to be tested. It also extracts multiple behavioral features as account features based on transaction information in the smart contract and forms the initial features based on the multiple opcode features and multiple account features.

[0053] The balancing module is used to construct a dataset based on the initial features and class labels corresponding to each historical smart contract, and to balance the dataset using the SMOTE-Tomek method to obtain a balanced dataset.

[0054] The multi-level feature construction module is used to construct fine-grained and coarse-grained features from the initial features in the balanced dataset and the initial features of the smart contract to be tested, thereby obtaining the multi-level features of the smart contract, and constructing a sample dataset based on the multi-level features and class labels of the smart contract.

[0055] The model building module is used to construct a contract recognition model and train the contract recognition model based on a sample dataset to obtain a trained contract recognition model. The contract recognition model includes a feature extractor based on a CNN network model and a decision classifier based on the LightGBM algorithm. The feature extractor is used to extract features from smart contracts with multi-level features as input, and the decision classifier is used to predict the class label of smart contracts with the feature data output by the feature extractor as input.

[0056] The contract recognition module is used to input the multi-level features of the smart contract to be tested into the trained contract recognition model, and predict and output the class label of the smart contract to be tested through the trained recognition model.

[0057] Preferably, the initial feature construction module is used to perform the following construction of opcode features:

[0058] Obtain the bytecode information of the smart contract address from the Ethereum website;

[0059] The bytecode information is converted into opcode information using a disassembler;

[0060] Opcode features are constructed based on the frequency of opcode calls.

[0061] Preferably, the multi-level feature construction module is used to perform the following fine-grained and coarse-grained feature construction on the initial features:

[0062] Define the initial features as fine-grained features;

[0063] For each initial feature, an opcode feature set is formed based on all opcode features, and an account feature set is formed based on all account features;

[0064] The account feature set and the opcode feature set are dimensionality reduced by the t-distributed random neighborhood embedding algorithm respectively. The account feature set is mapped to a low-dimensional space to obtain coarse-grained account features, and the opcode features are mapped to a low-dimensional space to obtain coarse-grained opcode features.

[0065] By combining coarse-grained account features and coarse-grained opcode features, coarse-grained features of smart contracts can be obtained.

[0066] Preferably, the feature extractor includes a preprocessing layer, a feature extraction layer, and a classification layer.

[0067] The preprocessing layer is used to perform the following: fill operations on the input fine-grained features and coarse-grained features respectively, and concatenate the filled fine-grained features and coarse-grained features into a feature matrix;

[0068] The feature extraction layer is used to perform convolution operations on the feature matrix output by the preprocessing layer as input to obtain the feature data of the smart contract.

[0069] The classification layer is used to optimize the model by taking feature data as input, predicting the output class label through the softmax function, and calculating the value of the loss function.

[0070] The decision classifier is used to predict and output the corresponding class label by taking the feature data output by the feature extraction layer as input.

[0071] Correspondingly, the model building module is used to perform model training as follows:

[0072] The feature extractor is trained based on the sample dataset: the multi-level features in the sample dataset are used as the input of the feature extractor. A loss function is constructed based on the class label predicted by the feature extractor and the actual class label. The feature extractor is trained by minimizing the loss function to obtain the trained feature extractor.

[0073] The decision classifier is trained based on the sample dataset: the multi-level features in the sample dataset are used as the input of the trained feature extractor, the feature data output by the feature extraction layer of the trained feature extractor is used as the input of the decision classifier, the loss function is constructed based on the class label predicted by the decision classifier and the actual class label, and the decision classifier is trained by minimizing the loss function to obtain the trained decision classifier.

[0074] Correspondingly, the contract identification module is used to perform the following operation: predicting and outputting the class label of the smart contract to be tested using the trained identification model:

[0075] The multi-level features of the smart contract under test are used as input to the trained feature classifier, and features are extracted through the trained feature classifier.

[0076] The feature data output from the feature extraction layer in the post-trained feature classifier is used as the input to the post-trained decision classifier, which then predicts the class label of the smart contract to be tested.

[0077] Preferably, account features include:

[0078] The number of investments, Num_inv, represents the amount of investment in each smart contract. If N transactions transfer funds from other accounts to a smart contract, then the formula for calculating Num_inv is as follows:

[0079]

[0080] Where Σ represents the summation symbol, and i represents the sequence number or index of the investment;

[0081] The number of payment transactions, Num_pay, represents the number of payment transactions in the smart contract. Each contract contains L payment transactions, and each payment transaction transfers funds from the smart contract to a designated transaction account. The formula for calculating Num_pay is as follows:

[0082]

[0083] Where j represents the transaction number or index;

[0084] Max_pay represents the maximum number of transactions a smart contract can initiate to another account.

[0085] The payment rate (Paid_rate) represents the percentage of investors who have received at least one payment. It is calculated by dividing the number of investors who have received payments by the total number of participating investors.

[0086] Bal balance represents the balance of a smart contract, which is the remaining or available assets after a transaction is completed.

[0087] The difference index D_ind represents the difference between the investment amount and the payment amount of all investors in the smart contract. Assuming there are p investors in the smart contract, M[i] and N[i] represent the investment amount and payment amount of the i-th investor, respectively. V is a vector of length p used to store the difference for each investor. The skewness s of the V vector is obtained according to the following formula:

[0088] V[i] = N[i] - M[i]

[0089]

[0090] After calculating the skewness s, if the number of participating investors p≥3, then D_ind=s, to ensure that there are enough participating investors in the smart contract before calculating D_ind;

[0091] The percentage of investors who invest first and then profit (Kr) represents the proportion of investors who invest before their accounts generate profits.

[0092] The deep learning-based Ethereum smart contract Ponzi scheme detection method and system of the present invention have the following advantages:

[0093] 1. After obtaining the initial features of the smart contract, multi-level features are constructed based on the initial features to obtain coarse-grained features and fine-grained features. Coarse-grained features can summarize the overall behavior of the contract, while fine-grained features provide detailed operation execution information. Semantic features of the smart contract can be extracted from different levels to characterize the overall behavior of the smart contract. Using multi-level features as input, a model built based on deep learning methods is used to detect smart contracts, which improves the accuracy of Ponzi scheme detection.

[0094] 2. The contract identification model is a network model built on CNN and LightGBM algorithms. The model includes a feature extractor built on the CNN model and a decision classifier built on the LightGBM algorithm. The feature extractor extracts features from multiple levels of features. After obtaining the feature data, the decision classifier classifies the contracts based on the feature data. The LightGBM algorithm is a gradient boosting algorithm widely used in classification tasks. It can handle the complex relationship between features and results, making it very suitable for detecting Ponzi schemes in smart contracts.

[0095] 3. It does not rely on the source code of smart contracts, but can detect fraudulent behavior in closed-source contracts using only bytecode and transaction information, which greatly improves the adaptability of the model.

[0096] 4. This method is scalable and can be applied to other types of blockchain fraud monitoring and to the automated monitoring of Ponzi schemes by regulatory authorities. Attached Figure Description

[0097] To more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0098] The invention will be further described below with reference to the accompanying drawings.

[0099] Figure 1 This is a flowchart of the deep learning-based Ethereum smart contract Ponzi scheme method in Example 1.

[0100] Figure 2 This is the feature information gain map in Example 1, which describes a deep learning-based method for solving Ethereum smart contract Ponzi schemes.

[0101] Figure 3 This is a flowchart illustrating the contract identification model in the deep learning-based Ethereum smart contract Ponzi scheme method of Example 1. Detailed Implementation

[0102] The present invention will be further described below with reference to the accompanying drawings and specific embodiments, so that those skilled in the art can better understand and implement the present invention. However, the embodiments are not intended to limit the present invention. In the absence of conflict, the embodiments of the present invention and the technical features in the embodiments can be combined with each other.

[0103] This invention provides a method and system for detecting Ponzi schemes in Ethereum smart contracts based on deep learning, which addresses the technical problems of existing technologies being unable to extract Ponzi scheme features from smart contracts, being unable to detect non-open-source smart contracts, and model category bias.

[0104] Example 1:

[0105] This invention discloses a deep learning-based method for detecting Ponzi schemes in Ethereum smart contracts, comprising the following steps:

[0106] S100. Obtain multiple historical smart contracts through the Ethereum website. Each historical smart contract is labeled with a class tag. The class tag types include normal contracts and Ponzi contracts. Among them, multiple historical smart contracts include both normal contracts and Ponzi contracts.

[0107] S200. For each historical smart contract, obtain the call frequency of multiple opcodes during the execution of the smart contract, construct opcode features, and extract multiple behavioral features as account features based on the transaction information in the smart contract. The initial features are composed of multiple opcode features and multiple account features.

[0108] S300: Construct a dataset based on the initial features and class labels corresponding to each historical smart contract, and balance the dataset using the SMOTE-Tomek method to obtain a balanced dataset.

[0109] S400. For the initial features in the balanced dataset, fine-grained and coarse-grained features are constructed to obtain multi-level features of the smart contract, and a sample dataset is constructed based on the multi-level features and class labels of the smart contract.

[0110] S500. Construct a contract recognition model and train the model based on a sample dataset to obtain a trained contract recognition model. The contract recognition model includes a feature extractor built on a CNN network model and a decision classifier built on the LightGBM algorithm. The feature extractor is used to extract features from smart contracts with multi-level features as input, and the decision classifier is used to predict the class label of smart contracts with the feature data output by the feature extractor as input.

[0111] S600. For the smart contract under test, obtain the call frequency of multiple opcodes during the execution of the smart contract, construct opcode features, and extract multiple behavioral features as account features based on the transaction information in the smart contract. Based on the multiple opcode features and multiple account features, form the initial features, and construct fine-grained features and coarse-grained features on the initial features to obtain the multi-level features of the smart contract under test.

[0112] S700: Input the multi-level features of the smart contract to be tested into the trained contract recognition model, and predict and output the class label of the smart contract to be tested through the trained recognition model.

[0113] In step S100 of this embodiment, the bytecode information and transaction information of the smart contract address are obtained through the Ethereum website. The bytecode information is converted into opcode information using a disassembler, and opcode features are constructed based on the opcode call frequency. At the same time, key features are extracted from the transaction information as the account features of the smart contract.

[0114] As a specific implementation, this embodiment statistically analyzed the call frequency of nine key opcodes (GASLIMIT, EXP, CALLDATALOAD, SLOAD, CALLER, LT, GAS, MOD, and MSTORE) during smart contract execution, constructing an opcode feature vector for the smart contract. These features reflect the underlying execution logic of the smart contract. Simultaneously, seven behavioral statistical features were extracted from the smart contract's historical transaction records as account features, constructed as follows.

[0115] (1) Investment number Num_inv represents the investment amount in each smart contract. If there are N transactions transferring from other accounts to the smart contract, then the formula for calculating Num_inv is as follows:

[0116]

[0117] Where Σ represents the summation symbol, and i represents the sequence number or index of the investment;

[0118] (2) Number of payment transactions Num_pay represents the number of payment transactions in the smart contract. Each contract contains L payment transactions. Each payment transaction transfers funds from the smart contract to a designated transaction account. The formula for calculating Num_pay is as follows:

[0119]

[0120] Where j represents the transaction number or index;

[0121] (3) Max_pay, which represents the maximum number of transactions that a smart contract can initiate to another account;

[0122] (4) Paid_rate represents the percentage of investors who have received at least one payment. It is calculated by the ratio of the number of investors who have received payments to the total number of participating investors.

[0123] (5) Bal balance, which represents the balance of a smart contract, is the remaining or available assets after the transaction is completed;

[0124] (6) The difference index D_ind represents the difference between the investment amount and the payment amount of all investors in the smart contract. Assuming there are p investors in the smart contract, M[i] and N[i] represent the investment amount and payment amount of the i-th investor, respectively, and V is a vector of length p used to store the difference for each investor.

[0125] The skewness s of vector V is obtained using the following formula:

[0126] V[i] = N[i] - M[i]

[0127]

[0128] After calculating the skewness s, if the number of participating investors p≥3, then D_ind=s, to ensure that there are enough participating investors in the smart contract before calculating D_ind;

[0129] (7) The proportion of investors who invest first and then profit, Kr, represents the proportion of investors who invest before making a profit in their account.

[0130] In step S200, the opcode features and account features are concatenated into an initial feature vector. This initial feature vector, along with the corresponding smart contract class labels, is used to construct a dataset Q, containing M smart contracts. Specifically, Q is represented as {(X... i ,y i )|i=1,2,3...M}, where X i It is a vector connecting the opcode of the i-th smart contract and the account characteristics, y i It is a class label. The normal contract dataset is represented as Q. n The Ponzi contract dataset is represented as Q. f And Q = Q n ∪Q f The number of contracts differs significantly between the two categories, i.e., |Q n |>>|Q f Therefore, dataset Q is class-imbalanced. In this embodiment, the SMOTE-Tomek method is used to balance the dataset, resulting in a balanced dataset Q1.

[0131] In this embodiment, nine opcode features and seven account features are summarized and concatenated to form a 16-dimensional initial feature vector, which comprehensively reflects the overall behavior of the smart contract. This lays the foundation for subsequent multi-level feature construction.

[0132] Step S300 constructs fine-grained and coarse-grained features from the initial features, including the following steps:

[0133] (1) Define the initial features as fine-grained features;

[0134] (2) For each initial feature, an account feature set is formed based on all account features, and an opcode feature set is formed based on all opcode features;

[0135] (3) The account feature set and the opcode feature set are dimensionality reduced by the t-distributed random neighborhood embedding algorithm respectively. The account feature set is mapped to the low-dimensional space to obtain coarse-grained account features, and the opcode features are mapped to the low-dimensional space to obtain coarse-grained opcode features.

[0136] (4) Combine the coarse-grained account features and coarse-grained opcode features to obtain the coarse-grained features of the smart contract.

[0137] This step constructs multi-level features on dataset Q1 to extract a comprehensive representation of smart contract Ponzi schemes. The multi-level features extract fine-grained features to reflect the detailed attributes of the smart contract, and coarse-grained features to reflect its overall properties. Finally, the multi-level feature dataset Q2 is obtained from this multi-level feature construction.

[0138] In this embodiment, during the multi-level feature construction process, the initial features are defined as fine-grained features. Then, higher-level coarse-grained features are constructed to represent the overall behavioral pattern of the smart contract. Specifically, the account feature set A (containing 7 features) and the opcode feature set B (containing 9 features) extracted from the contract are treated as two wholes to represent the higher-level semantic features of the smart contract. To reduce the dimensionality of these high-dimensional feature sets, the t-distributed random neighborhood embedding (t-SNE) algorithm is applied to map the account feature set A to a low-dimensional space to obtain coarse-grained account features (d1), and similarly, the opcode feature set B is mapped to a low-dimensional space to obtain coarse-grained opcode features (d2). Finally, these two coarse-grained feature representations are combined to form the coarse-grained feature set G = (d1, d2) representing the entire smart contract. This multi-level feature construction method can extract semantic features of the contract from different levels to characterize the overall behavior of the smart contract.

[0139] This embodiment utilizes the information gain method to evaluate the importance of the constructed multi-level features, thereby verifying their role in the smart contract Ponzi scheme detection task. Information gain can effectively measure the contribution of each feature to the classification result. The results show... Figure 2 In this embodiment, the coarse-grained features d1 and d2 exhibit high information gain, indicating that the abstracted high-level semantic features can effectively characterize the overall behavioral pattern of a Ponzi scheme. Furthermore, some fine-grained features, such as SLOAD and LT, also show significant information gain, suggesting that these fine-grained operational features also play an auxiliary role in detecting Ponzi schemes. In summary, multi-level features can construct semantic features with discriminative significance for Ponzi scheme detection tasks from different levels of abstraction. Coarse-grained features can provide a general description of the overall contract behavior, while fine-grained features offer detailed operational execution information. This multi-level feature construction method can improve the effectiveness of Ponzi scheme detection.

[0140] In step S400, the feature extractor built based on the CNN network model includes a preprocessing layer, a feature extraction layer, and a classification layer. The preprocessing layer performs the following operations: padding operations on the input fine-grained and coarse-grained features, and concatenating the padded fine-grained and coarse-grained features into a feature matrix. The feature extraction layer takes the feature matrix output by the preprocessing layer as input and performs a convolution operation on the feature matrix to obtain the feature data of the smart contract. The classification layer takes the feature data as input and predicts the output class label using the softmax function. The decision classifier built based on the LightGBM algorithm takes the feature data output by the feature extraction layer as input and predicts the corresponding class label of the output smart contract.

[0141] Correspondingly, the contract recognition model is trained based on the sample dataset, including the following operations:

[0142] (1) Training the feature extractor based on the sample dataset: The multi-level features in the sample dataset are used as the input of the feature extractor. A loss function is constructed based on the class label predicted by the feature extractor and the actual class label. The feature extractor is trained by minimizing the loss function to obtain the trained feature extractor.

[0143] (2) Training the decision classifier based on the sample dataset: The multi-level features in the sample dataset are used as the input of the trained feature extractor. The feature data output by the feature extraction layer in the trained feature extractor is used as the input of the decision classifier. A loss function is constructed based on the class label predicted by the decision classifier and the actual class label. The decision classifier is trained by minimizing the loss function to obtain the trained decision classifier.

[0144] Correspondingly, the trained recognition model predicts and outputs the class label of the smart contract under test, including the following steps:

[0145] (1) Use the multi-level features of the smart contract under test as input to the trained feature classifier, and extract features through the trained feature classifier;

[0146] (2) Use the feature data output by the feature extraction layer in the post-trained feature classifier as the input of the post-trained decision classifier, and use the post-trained decision classifier to predict the class label of the smart contract to be tested.

[0147] The working principle of the contract recognition model in this embodiment is as follows: Figure 3 As shown. First, before input, the features in dataset Q2 are divided into opcode F and account features G, and padding is performed for subsequent convolutions. After the padding step, the padded features are concatenated to form a new matrix M. In the core part of the feature extractor, since M is a single-channel feature map of size n×1, each filter in the convolutional layer performs a single-channel convolution using a convolutional kernel of size k×1. To extract key features, a set of J filters of the same size are convolved with matrix M. The following formula represents the calculation of convolutional feature extraction:

[0148]

[0149] Where C[i] represents the i-th pixel in the output feature map, M[i+j] represents the k×k region in the input feature map that is aligned with the convolution kernel, f(·) represents the activation function, b represents the bias term, and W[j] represents the weights in the convolution kernel. The calculation process is as follows: the k×k regions in the input feature map M that overlap with the convolution kernel are multiplied element-wise with the convolution kernel W, then all multiplication results are summed, the bias b is added, and finally the activation function f(·) is used to obtain a feature value in the output feature map C.

[0150] The decision classifier algorithm is LightGBM, a gradient boosting algorithm widely used for classification tasks. LightGBM can handle complex relationships between features and results, making it well-suited for detecting Ponzi schemes in smart contracts. The process is as follows:

[0151] (1) Feature extraction: Input the multi-level features of the smart contract into the trained feature extractor and extract the penultimate layer output (i.e., the key data prototype);

[0152] (2) Data preparation: Combine the extracted data prototypes with the corresponding labels to construct a new dataset, which will be used as input for the LightGBM model;

[0153] (3) Training LightGBM: Train LightGBM using the newly prepared dataset.

[0154] This embodiment integrates Convolutional Neural Networks (CNNs) and the LightGBM algorithm to detect Ponzi schemes in smart contracts. Specifically, it first uses a CNN as a feature extractor to learn and extract key data prototypes of Ponzi schemes from the constructed multi-granularity feature representation Q2. Then, these key data prototypes learned by the CNN are input into the LightGBM classifier to detect whether a smart contract is a Ponzi scheme. Furthermore, the entire contract recognition model is trained in a step-by-step manner, allowing for flexible adjustment of each module. Simultaneously, the LightGBM classifier can be directly trained based on the features extracted by the CNN, avoiding repetitive and redundant feature extraction calculations. Figure 1 The overall framework and process of the contract recognition model are demonstrated. This detection method, which integrates CNN feature learning and LightGBM classification, can not only learn the abstract features of Ponzi scheme behavior patterns, but also has efficient and powerful classification capabilities.

[0155] Example 2:

[0156] This invention discloses a deep learning-based Ethereum smart contract Ponzi scheme detection system, comprising a data acquisition module, an initial feature construction module, a balancing module, a multi-layer feature construction module, a model construction module, and a contract identification module. The system can execute the method disclosed in Example 1 to identify Ethereum smart contract Ponzi schemes.

[0157] The data acquisition module is used to obtain multiple smart contracts through the Ethereum website. Each smart contract is labeled with a class tag, which includes two types: normal contracts and Ponzi contracts. The multiple smart contracts include both normal contracts and Ponzi contracts. It is also used to obtain the smart contracts to be tested from the Ethereum website.

[0158] For historical smart contracts and smart contracts under test, the initial feature construction module is used to obtain the call frequency of multiple opcodes during the execution of the smart contract, construct opcode features, and extract multiple behavioral features as account features based on transaction information in the smart contract. The initial features are composed of multiple opcode features and multiple account features.

[0159] In this embodiment, the initial feature construction module is used to obtain the bytecode information and transaction information of the smart contract address through the Ethereum website, use a disassembler to convert the bytecode information into opcode information, and construct opcode features based on the opcode call frequency; at the same time, key features are extracted from the transaction information as the account features of the smart contract.

[0160] In a specific implementation, the initial feature construction module of this embodiment statistically analyzed the call frequency of nine key opcodes (GASLIMIT, EXP, CALLDATALOAD, SLOAD, CALLER, LT, GAS, MOD, and MSTORE) during smart contract execution, and constructed an opcode feature vector for the smart contract. These feature vectors reflect the underlying execution logic of the smart contract. Simultaneously, seven behavioral statistical features were extracted from the smart contract's historical transaction records as account features, constructed as follows.

[0161] (1) Investment number Num_inv represents the investment amount in each smart contract. If there are N transactions transferring from other accounts to the smart contract, then the formula for calculating Num_inv is as follows:

[0162]

[0163] Where Σ represents the summation symbol, and i represents the sequence number or index of the investment;

[0164] (2) Number of payment transactions Num_pay represents the number of payment transactions in the smart contract. Each contract contains L payment transactions. Each payment transaction transfers funds from the smart contract to a designated transaction account. The formula for calculating Num_pay is as follows:

[0165]

[0166] Where j represents the transaction number or index;

[0167] (3) Max_pay, which represents the maximum number of transactions that a smart contract can initiate to another account;

[0168] (4) Paid_rate represents the percentage of investors who have received at least one payment. It is calculated by the ratio of the number of investors who have received payments to the total number of participating investors.

[0169] (5) Bal balance, which represents the balance of a smart contract, is the remaining or available assets after the transaction is completed;

[0170] (6) The difference index D_ind represents the difference between the investment amount and the payment amount of all investors in the smart contract. Assuming there are p investors in the smart contract, M[i] and N[i] represent the investment amount and payment amount of the i-th investor, respectively, and V is a vector of length p used to store the difference for each investor.

[0171] The skewness s of vector V is obtained using the following formula:

[0172] V[i] = N[i] - M[i]

[0173]

[0174] After calculating the skewness s, if the number of participating investors p≥3, then D_ind=s, to ensure that there are enough participating investors in the smart contract before calculating D_ind;

[0175] (7) The proportion of investors who invest first and then profit, Kr, represents the proportion of investors who invest before making a profit in their account.

[0176] The balancing module is used to construct a dataset based on the initial features and class labels corresponding to each historical smart contract, and to balance the dataset using the SMOTE-Tomek method to obtain the balanced dataset.

[0177] In this embodiment, the balancing module concatenates the opcode features and account features into an initial feature vector. It then constructs a dataset Q using the initial feature vector and the corresponding smart contract class labels. The dataset contains M smart contracts. Specifically, Q is represented as {(X... i ,y i )|i=1,2,3...M}, where X i It is a vector connecting the opcode of the i-th smart contract and the account characteristics, y i It is a class label. The normal contract dataset is represented as Q. n The Ponzi contract dataset is represented as Q. f And Q = Q n ∪Q f The number of contracts differs significantly between the two categories, i.e., |Q n |>>|Q f Therefore, dataset Q is class-imbalanced. In this embodiment, the SMOTE-Tomek method is used to balance the dataset, resulting in a balanced dataset Q1.

[0178] In this embodiment, nine opcode features and seven account features are summarized and concatenated to form a 16-dimensional initial feature vector, which comprehensively reflects the overall behavior of the smart contract. This lays the foundation for subsequent multi-level feature construction.

[0179] For the initial features in the balanced dataset and the initial features of the smart contract to be tested, the multi-level feature construction module is used to construct fine-grained and coarse-grained features from the initial features to obtain the multi-level features of the smart contract, and to construct the sample dataset based on the multi-level features and class labels of the smart contract.

[0180] In this embodiment, the multi-level feature construction module is used to perform the following fine-grained and coarse-grained feature construction on the initial features:

[0181] (1) Define the initial features as fine-grained features;

[0182] (2) For each initial feature, an account feature set is formed based on all account features, and an opcode feature set is formed based on all opcode features;

[0183] (3) The account feature set and the opcode feature set are dimensionality reduced by the t-distributed random neighborhood embedding algorithm respectively. The account feature set is mapped to the low-dimensional space to obtain coarse-grained account features, and the opcode features are mapped to the low-dimensional space to obtain coarse-grained opcode features.

[0184] (4) Combine the coarse-grained account features and coarse-grained opcode features to obtain the coarse-grained features of the smart contract.

[0185] This multi-level feature construction module is used to construct multi-level features on dataset Q1 to extract a comprehensive representation of smart contract Ponzi schemes. The multi-level features extract fine-grained features to reflect the detailed attributes of the smart contract, and coarse-grained features to reflect the overall properties of the smart contract. Finally, the multi-level feature dataset Q2 of the smart contract is obtained through multi-level feature construction.

[0186] In this embodiment, during the multi-level feature construction process, the initial features are defined as fine-grained features. Then, higher-level coarse-grained features are constructed to represent the overall behavioral pattern of the smart contract. Specifically, the account feature set A (containing 7 features) and the opcode feature set B (containing 9 features) extracted from the contract are treated as two wholes to represent the higher-level semantic features of the smart contract. To reduce the dimensionality of these high-dimensional feature sets, the t-distributed random neighborhood embedding (t-SNE) algorithm is applied to map the account feature set A to a low-dimensional space to obtain coarse-grained account features (d1), and similarly, the opcode feature set B is mapped to a low-dimensional space to obtain coarse-grained opcode features (d2). Finally, these two coarse-grained feature representations are combined to form the coarse-grained feature set G = (d1, d2) representing the entire smart contract. This multi-level feature construction method can extract semantic features of the contract from different levels to characterize the overall behavior of the smart contract.

[0187] The model building module is used to construct a contract recognition model and train the contract recognition model based on a sample dataset to obtain a trained contract recognition model. The contract recognition model includes a feature extractor built based on a CNN network model and a decision classifier built based on the LightGBM algorithm. The feature extractor is used to extract features from smart contracts with multi-level features as input, and the decision classifier is used to predict the class label of smart contracts with the feature data output by the feature extractor as input.

[0188] In this embodiment, the feature extractor built based on a CNN network model includes a preprocessing layer, a feature extraction layer, and a classification layer. The preprocessing layer performs the following operations: padding operations on the input fine-grained and coarse-grained features, and concatenating the padded fine-grained and coarse-grained features into a feature matrix. The feature extraction layer takes the feature matrix output from the preprocessing layer as input and performs a convolution operation on the feature matrix to obtain the feature data of the smart contract. The classification layer takes the feature data as input and predicts the output class label using the softmax function. A decision classifier built based on the LightGBM algorithm takes the feature data output from the feature extraction layer as input and predicts the corresponding class label of the smart contract.

[0189] Correspondingly, the contract recognition model is trained based on the sample dataset, including the following operations:

[0190] (1) Training the feature extractor based on the sample dataset: The multi-level features in the sample dataset are used as the input of the feature extractor. A loss function is constructed based on the class label predicted by the feature extractor and the actual class label. The feature extractor is trained by minimizing the loss function to obtain the trained feature extractor.

[0191] (2) Training the decision classifier based on the sample dataset: The multi-level features in the sample dataset are used as the input of the trained feature extractor. The feature data output by the feature extraction layer in the trained feature extractor is used as the input of the decision classifier. A loss function is constructed based on the class label predicted by the decision classifier and the actual class label. The decision classifier is trained by minimizing the loss function to obtain the trained decision classifier.

[0192] Correspondingly, the trained recognition model predicts and outputs the class label of the smart contract under test, including the following steps:

[0193] (1) Use the multi-level features of the smart contract under test as input to the trained feature classifier, and extract features through the trained feature classifier;

[0194] (2) Use the feature data output by the feature extraction layer in the post-trained feature classifier as the input of the post-trained decision classifier, and use the post-trained decision classifier to predict the class label of the smart contract to be tested.

[0195] The working principle of the contract recognition model in this embodiment is as follows:

[0196] First, before input, the features in dataset Q2 are divided into opcode F and account features G, and padding is performed for subsequent convolutions. After the padding step, the padded features are concatenated to form a new matrix M. In the core part of the feature extractor, since M is a single-channel feature map of size n×1, each filter in the convolutional layer performs a single-channel convolution using a convolutional kernel of size k×1. To extract key features, a set of J filters of the same size are convolved with matrix M. The following formula represents the calculation of convolutional feature extraction:

[0197]

[0198] Where C[i] represents the i-th pixel in the output feature map, M[i+j] represents the k×k region in the input feature map that is aligned with the convolution kernel, f(·) represents the activation function, b represents the bias term, and W[j] represents the weights in the convolution kernel. The calculation process is as follows: the k×k regions in the input feature map M that overlap with the convolution kernel are multiplied element-wise with the convolution kernel W, then all multiplication results are summed, the bias b is added, and finally the activation function f(·) is used to obtain a feature value in the output feature map C.

[0199] The decision classifier algorithm is LightGBM, a gradient boosting algorithm widely used for classification tasks. LightGBM can handle complex relationships between features and results, making it well-suited for detecting Ponzi schemes in smart contracts. The process is as follows:

[0200] (1) Feature extraction: Input the multi-level features of the smart contract into the trained feature extractor and extract the penultimate layer output (i.e., the key data prototype);

[0201] (2) Data preparation: Combine the extracted data prototypes with the corresponding labels to construct a new dataset, which will be used as input for the LightGBM model;

[0202] (3) Training LightGBM: Train LightGBM using the newly prepared dataset.

[0203] The contract recognition module is used to input the multi-level features of the smart contract under test into the trained contract recognition model, and then predict and output the class label of the smart contract under test through the trained recognition model.

[0204] The present invention has been shown and described in detail above with reference to the accompanying drawings and preferred embodiments. However, the present invention is not limited to these disclosed embodiments. Based on the above embodiments, those skilled in the art will know that more embodiments of the present invention can be obtained by combining the means in the different embodiments described above, and these embodiments are also within the protection scope of the present invention.

Claims

1. A method for detecting Ethereum smart contract Ponzi schemes based on deep learning, characterized in that, Includes the following steps: Multiple historical smart contracts were obtained through the Ethereum website. Each historical smart contract was labeled with a class tag, which included two types: normal contracts and Ponzi contracts. Among them, multiple historical smart contracts included both normal contracts and Ponzi contracts. For each historical smart contract, obtain the call frequency of multiple opcodes during the execution of the smart contract, construct opcode features, and extract multiple behavioral features as account features based on the transaction information in the smart contract. The initial features are composed of multiple opcode features and multiple account features. A dataset is constructed based on the initial features and class labels corresponding to each historical smart contract, and the dataset is balanced using the SMOTE-Tomek method to obtain a balanced dataset. For the initial features in the balanced dataset, fine-grained and coarse-grained features are constructed to obtain multi-level features of smart contracts, and a sample dataset is constructed based on the multi-level features and class labels of smart contracts. A contract recognition model is constructed, and the model is trained based on a sample dataset to obtain a trained contract recognition model. The contract recognition model includes a feature extractor built based on a CNN network model and a decision classifier built based on the LightGBM algorithm. The feature extractor is used to extract features from smart contracts with multi-level features as input, and the decision classifier is used to predict the class label of smart contracts with the feature data output by the feature extractor as input. For the smart contract under test, the call frequency of multiple opcodes during the execution of the smart contract is obtained, opcode features are constructed, and multiple behavioral features are extracted as account features based on the transaction information in the smart contract. The initial features are composed of multiple opcode features and multiple account features, and fine-grained features and coarse-grained features are constructed on the initial features to obtain the multi-level features of the smart contract under test. The multi-level features of the smart contract to be tested are input into the trained contract recognition model, and the class label of the smart contract to be tested is predicted and output by the trained recognition model. This includes obtaining the call frequency of multiple opcodes during smart contract execution and constructing opcode characteristics, including the following operations: Obtain the bytecode information of the smart contract address from the Ethereum website; The bytecode information is converted into opcode information using a disassembler; Construct opcode features based on opcode invocation frequency; The process of constructing fine-grained and coarse-grained features from the initial features includes the following steps: Define the initial features as fine-grained features; For each initial feature, an opcode feature set is formed based on all opcode features, and an account feature set is formed based on all account features; The account feature set and the opcode feature set are dimensionality reduced by the t-distributed random neighborhood embedding algorithm respectively. The account feature set is mapped to a low-dimensional space to obtain coarse-grained account features, and the opcode features are mapped to a low-dimensional space to obtain coarse-grained opcode features. By combining coarse-grained account features and coarse-grained opcode features, coarse-grained features of smart contracts can be obtained.

2. The method for detecting Ethereum smart contract Ponzi schemes based on deep learning according to claim 1, characterized in that, The feature extractor includes a preprocessing layer, a feature extraction layer, and a classification layer. The preprocessing layer is used to perform the following: fill operations on the input fine-grained features and coarse-grained features respectively, and concatenate the filled fine-grained features and coarse-grained features into a feature matrix; The feature extraction layer is used to perform convolution operations on the feature matrix output by the preprocessing layer as input to obtain the feature data of the smart contract. The classification layer is used to optimize the model by taking feature data as input, predicting the output class label through the softmax function, and calculating the value of the loss function. The decision classifier is used to predict and output the corresponding class label by taking the feature data output by the feature extraction layer as input. Correspondingly, the contract recognition model is trained based on the sample dataset, including the following operations: The feature extractor is trained based on the sample dataset: the multi-level features in the sample dataset are used as the input of the feature extractor. A loss function is constructed based on the class label predicted by the feature extractor and the actual class label. The feature extractor is trained by minimizing the loss function to obtain the trained feature extractor. The decision classifier is trained based on the sample dataset: the multi-level features in the sample dataset are used as the input of the trained feature extractor, the feature data output by the feature extraction layer of the trained feature extractor is used as the input of the decision classifier, the loss function is constructed based on the class label predicted by the decision classifier and the actual class label, and the decision classifier is trained by minimizing the loss function to obtain the trained decision classifier. Correspondingly, the trained recognition model predicts and outputs the class label of the smart contract under test, including the following steps: The multi-level features of the smart contract under test are used as input to the trained feature classifier, and features are extracted through the trained feature classifier. The feature data output from the feature extraction layer in the post-trained feature classifier is used as the input to the post-trained decision classifier, which then predicts the class label of the smart contract to be tested.

3. The method for detecting Ethereum smart contract Ponzi schemes based on deep learning according to claim 1, characterized in that, Account characteristics include: Investment Num _ inv , representing the amount of investment in each smart contract, has N If a transaction is transferred from another account to a smart contract, then Num _ inv The calculation formula is as follows: , Where Σ represents the summation symbol, i Represents the sequence number or index of the investment; Number of payment transactions Num _ pay This indicates the number of payment transactions in a smart contract; each contract contains... L For each payment transaction, funds are transferred from the smart contract to a designated transaction account. The formula for calculating Num_pay is as follows: , in, j Represents the sequence number or index of the transaction; Large payment transaction volume Max _ pay This represents the maximum number of transactions that a smart contract can initiate to another account. Payment ratio Paid _ rate This represents the percentage of investors who have received at least one payment, calculated by the ratio of the number of investors who have received payments to the total number of participating investors. Balance Bal This represents the balance of a smart contract, which is the remaining or available assets after the transaction is completed; Difference Index D _ ind This represents the difference between the investment amount and the payment amount of all investors in the smart contract. Assume the smart contract has... p Individual investor, M [ i ]and N [ i ] respectively represent the first i The investment amount and payment amount of each investor, V It is a length of p A vector is used to store the difference for each investor. Obtained from the following formula V skewness of a vector s : V [ i ]= N [ i ]- M [ i ], , Calculate the skewness s Then, if the number of participating investors p ≥3, then D _ ind = s To ensure in calculation D _ ind Previously, there were a sufficient number of investors participating in the smart contract; Proportion of investors who invest first and then profit Kr This indicates the percentage of investors who made their investments before their accounts turned a profit.

4. A deep learning-based Ethereum smart contract Ponzi scheme detection system, characterized in that, The system includes: The data acquisition module is used to acquire multiple smart contracts through the Ethereum website. Each smart contract is labeled with a class tag, which includes two types: normal contracts and Ponzi contracts. The multiple smart contracts include both normal contracts and Ponzi contracts. The module is also used to acquire the smart contract to be tested from the Ethereum website. The initial feature construction module is used to obtain the call frequency of multiple opcodes during the execution of the smart contract and construct opcode features for both historical smart contracts and smart contracts to be tested. It also extracts multiple behavioral features as account features based on transaction information in the smart contract and forms the initial features based on the multiple opcode features and multiple account features. The balancing module is used to construct a dataset based on the initial features and class labels corresponding to each historical smart contract, and to balance the dataset using the SMOTE-Tomek method to obtain a balanced dataset. The multi-level feature construction module is used to construct fine-grained and coarse-grained features from the initial features in the balanced dataset and the initial features of the smart contract to be tested, thereby obtaining the multi-level features of the smart contract, and constructing a sample dataset based on the multi-level features and class labels of the smart contract. The model building module is used to construct a contract recognition model and train the contract recognition model based on a sample dataset to obtain a trained contract recognition model. The contract recognition model includes a feature extractor based on a CNN network model and a decision classifier based on the LightGBM algorithm. The feature extractor is used to extract features from smart contracts with multi-level features as input, and the decision classifier is used to predict the class label of smart contracts with the feature data output by the feature extractor as input. The contract recognition module is used to input the multi-level features of the smart contract to be tested into the trained contract recognition model, and predict and output the class label of the smart contract to be tested through the trained recognition model. The initial feature construction module is used to construct opcode features as follows: Obtain the bytecode information of the smart contract address from the Ethereum website; The bytecode information is converted into opcode information using a disassembler; Construct opcode features based on opcode invocation frequency; The multi-level feature construction module is used to perform the following fine-grained and coarse-grained feature construction on the initial features: Define the initial features as fine-grained features; For each initial feature, an opcode feature set is formed based on all opcode features, and an account feature set is formed based on all account features; The account feature set and the opcode feature set are dimensionality reduced by the t-distributed random neighborhood embedding algorithm respectively. The account feature set is mapped to a low-dimensional space to obtain coarse-grained account features, and the opcode features are mapped to a low-dimensional space to obtain coarse-grained opcode features. By combining coarse-grained account features and coarse-grained opcode features, coarse-grained features of smart contracts can be obtained.

5. The Ethereum smart contract Ponzi scheme detection system based on deep learning according to claim 4, characterized in that, The feature extractor includes a preprocessing layer, a feature extraction layer, and a classification layer. The preprocessing layer is used to perform the following: fill operations on the input fine-grained features and coarse-grained features respectively, and concatenate the filled fine-grained features and coarse-grained features into a feature matrix; The feature extraction layer is used to perform convolution operations on the feature matrix output by the preprocessing layer as input to obtain the feature data of the smart contract. The classification layer is used to optimize the model by taking feature data as input, predicting the output class label through the softmax function, and calculating the value of the loss function. The decision classifier is used to predict and output the corresponding class label by taking the feature data output by the feature extraction layer as input. Correspondingly, the model building module is used to perform model training as follows: The feature extractor is trained based on the sample dataset: the multi-level features in the sample dataset are used as the input of the feature extractor. A loss function is constructed based on the class label predicted by the feature extractor and the actual class label. The feature extractor is trained by minimizing the loss function to obtain the trained feature extractor. The decision classifier is trained based on the sample dataset: the multi-level features in the sample dataset are used as the input of the trained feature extractor, the feature data output by the feature extraction layer of the trained feature extractor is used as the input of the decision classifier, the loss function is constructed based on the class label predicted by the decision classifier and the actual class label, and the decision classifier is trained by minimizing the loss function to obtain the trained decision classifier. Correspondingly, the contract identification module is used to perform the following operation: predicting and outputting the class label of the smart contract to be tested using the trained identification model: The multi-level features of the smart contract under test are used as input to the trained feature classifier, and features are extracted through the trained feature classifier. The feature data output from the feature extraction layer in the post-trained feature classifier is used as the input to the post-trained decision classifier, which then predicts the class label of the smart contract to be tested.

6. The Ethereum smart contract Ponzi scheme detection system based on deep learning according to claim 4, characterized in that, Account characteristics include: Investment Num _ inv , representing the amount of investment in each smart contract, has N If a transaction is transferred from another account to a smart contract, then Num _ inv The calculation formula is as follows: , Where Σ represents the summation symbol, i Represents the sequence number or index of the investment; Number of payment transactions Num _ pay This indicates the number of payment transactions in a smart contract; each contract contains... L For each payment transaction, funds are transferred from the smart contract to a designated transaction account. The formula for calculating Num_pay is as follows: , in, j Represents the sequence number or index of the transaction; Large payment transaction volume Max _ pay This represents the maximum number of transactions that a smart contract can initiate to another account. Payment ratio Paid _ rate This represents the percentage of investors who have received at least one payment, calculated by the ratio of the number of investors who have received payments to the total number of participating investors. Balance Bal This represents the balance of a smart contract, which is the remaining or available assets after the transaction is completed; Difference Index D _ ind This represents the difference between the investment amount and the payment amount of all investors in the smart contract. Assume the smart contract has... p Individual investor, M [ i ]and N [ i ] respectively represent the first i The investment amount and payment amount of each investor, V It is a length of p A vector is used to store the difference for each investor. Obtained from the following formula V skewness of a vector s : V [ i ]= N [ i ]- M [ i ], , Calculate the skewness s Then, if the number of participating investors p ≥3, then D _ ind = s To ensure in calculation D _ ind Previously, there were a sufficient number of investors participating in the smart contract; Proportion of investors who invest first and then profit Kr This indicates the percentage of investors who made their investments before their accounts turned a profit.

Citation Information

Patent Citations

  • Account identification method and system based on auto-encoder and generative adversarial network

    CN114818999A

  • Intelligent contract byte code similarity detection method

    CN116627490A