Intelligent contract vulnerability detection method based on pre-training technology
By adopting pre-training technology-based methods in smart contract vulnerability detection, multimodal information is extracted and feature extraction and decision-making fusion is used to use pre-trained language models and GAT models to perform feature extraction and decision-making fusion, the problem of insufficient multimodal data utilization and feature extraction of smart contract vulnerability detection in the existing technology is solved, and higher detection accuracy and reliability are achieved.
Patent Information
- Application Number
- CN202510315271.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-06-20
AI Technical Summary
The existing smart contract vulnerability detection technology has shortcomings in multimodal data utilization, feature extraction methods, and detection model performance, which is difficult to meet the growing demand for smart contract security.
Using a smart contract vulnerability detection method based on pre-training technology, we extract multimodal information from the smart contract source code, including code structure, control flow graph structure and intermediate representation, semantic features and graph structure features are extracted using pre-trained language model and graph attention network (GAT) model, and input these features into the classifier for decision-making fusion to improve the accuracy of vulnerability detection.
This method can comprehensively capture the characteristics of smart contract nodes from a multimodal perspective, improve detection accuracy and reliability, significantly reduce security risks, and avoid the one-sidedness of single-modal decisions.
Smart Images

Figure CN120180447A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of information security technology, and in particular relates to a smart contract vulnerability detection method based on pre-training technology. Background Art
[0002] With the widespread application of blockchain technology, smart contracts play a key role in many fields such as finance, supply chain management, and the Internet of Things. However, once a smart contract is deployed, its operating mechanism is difficult to modify. If there are loopholes, it may lead to serious economic losses and security risks, such as hacker attacks, stolen funds, and data leaks. Therefore, accurately detecting loopholes in smart contracts has become a core task to ensure the safe and stable operation of blockchain systems.
[0003] At present, smart contract vulnerability detection faces many challenges. From the perspective of data characteristics, smart contract data has multimodal characteristics, including information at different levels such as code structure, control flow graph structure, and domain knowledge. Traditional single-modal analysis methods only start from a specific perspective and cannot fully capture the complex characteristics of smart contracts, resulting in limited detection accuracy.
[0004] In terms of vulnerability detection models, commonly used machine learning and deep learning models have shortcomings when processing smart contract data. For example, traditional machine learning models are highly dependent on feature engineering and require careful manual design and feature extraction, which not only consumes a lot of manpower and time, but also makes it difficult to cover all the key features of smart contract data. Deep learning models, such as some simple neural networks, lack effective feature extraction and processing capabilities when faced with complex graph structure data of smart contracts, and cannot fully explore the potential patterns in the data.
[0005] Specifically for graph structure data processing, the control flow graph of smart contracts contains rich structural information, but existing methods are not effective in extracting features from control flow graphs. For example, traditional node feature aggregation methods, such as average pooling, treat all nodes equally and fail to highlight the importance of key nodes; the maximum pooling method is easily disturbed by noisy nodes and loses important information. In addition, when using the graph attention network (GAT) to extract features, how to effectively combine multimodal information and optimize the network structure to improve the accuracy and efficiency of feature extraction is also an urgent problem to be solved.
[0006] In summary, the existing smart contract vulnerability detection technology has shortcomings in terms of multimodal data utilization, feature extraction methods, and detection model performance, and it is difficult to meet the growing security needs of smart contracts. An innovative technical solution is urgently needed to improve the accuracy and reliability of smart contract vulnerability detection. Summary of the invention
[0007] To solve the problems existing in the background art, one aspect of the present invention provides an intelligent contract vulnerability detection method based on pre-training technology, including:
[0008] S1: Extract multi-modal information of the intelligent contract source code according to the intelligent contract source code, where the multi-modal information includes: the intelligent contract source code, the control flow chart of the intelligent contract source code, and the intermediate representation IR of the intelligent contract source code;
[0009] S2: Input the intelligent contract source code into the first pre-trained language model to extract the semantic features of the intelligent contract source code;
[0010] S3: Input the control flow chart of the intelligent contract source code into the GAT model based on multi-head attention to extract the graph structure features of the intelligent contract source code;
[0011] S4: Input the intermediate representation IR of the intelligent contract source code into the second pre-trained language model to extract the execution behavior features of the intelligent contract source code;
[0012] S5: Input the semantic features of the intelligent contract source code, the execution behavior features of the intelligent contract source code, and the graph structure features of the intelligent contract source code into three classifiers respectively to obtain corresponding prediction results;
[0013] S6: Perform decision fusion on the prediction results obtained by the three classifiers to obtain the final decision.
[0014] Another aspect of the present invention provides an electronic device, including a processor, a memory, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, the intelligent contract vulnerability detection method based on pre-training technology is implemented.
[0015] On the one hand, the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the intelligent contract vulnerability detection method based on pre-training technology is implemented.
[0016] The present invention has at least the following beneficial effects
[0017] The present invention generates initial node features from multimodal perspectives such as code structure, graph structure, and domain knowledge. It comprehensively captures the characteristics of smart contract nodes, providing a rich and accurate data foundation for subsequent analysis. Compared with single-modal feature extraction, it can more completely reflect the internal information of smart contracts and improve the detection accuracy. The present invention uses a graph attention network (GAT) combined with a multi-head attention mechanism. In the multi-layer hidden layer calculation, each node can effectively fuse the information of neighboring nodes, dynamically adjust the importance of neighboring nodes, and learn more representative node features. This process fully exploits the control flow graph structure information, overcomes the limitations of traditional methods in processing graph structure data, and improves the efficiency and quality of feature learning. Sub-decisions are made based on different modal features respectively, and then the decisions are fused through a stacking method. This method synthesizes multi-modal decision-making information, avoids the one-sidedness of single-modal decision-making, and can ensure that the detection effect is not lower than the best single-modal detection effect even in the worst case, significantly improving the accuracy and reliability of smart contract vulnerability judgment and effectively reducing security risks. Description of the Drawings
[0018] Figure 1 It is a schematic flowchart of the method of the present invention. Detailed Embodiment
[0019] The following uses specific specific examples to illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. The details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the diagrams provided in the following embodiments only schematically illustrate the basic concept of the present invention. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0020] Please refer to Figure 1 , the present invention provides a smart contract vulnerability detection method based on pre-training technology, including:
[0021] S1: Extract multimodal information of the smart contract source code according to the smart contract source code, where the multimodal information includes: the smart contract source code, the control flow graph of the smart contract source code, and the intermediate representation IR of the smart contract source code;
[0022] Preferably, the extraction of the multimodal information of the smart contract source code includes: preprocessing the smart contract source code; generating the control flow graph of the smart contract source code through the Slither static analysis tool; generating the intermediate representation IR of the smart contract source code through the Slither tool.
[0023] In this embodiment, since the control flow graph and modal data of the intermediate representation of the smart contract cannot be directly obtained, it is necessary to generate the control flow graph and intermediate representation based on the source code. The control flow graph (CFG) is a graph structure used to represent the program execution flow, recording the variable calls and function call relationships of each part in the system when the EVM executes the smart contract. The control flow graph consists of basic blocks and control flow edges. Among them, the basic block is a set of sequentially executed instructions without jump or branch instructions until the last statement, and the control flow edge represents the directed edge of the possible execution paths of the code, including sequential execution, conditional jump, function call, etc. This method uses the Slither static analysis tool to generate the control flow graph. The intermediate representation (IR) is a data structure in program compilation and analysis, used as a bridge between the source code and the machine code. IR is an abstract code representation form that can help optimize, analyze, and transform the program. It is usually closer to the machine code than the source code but still maintains a certain readability for further processing by the compiler.
[0024] S2: Input the smart contract source code into the first pre-trained language model to extract the semantic features of the smart contract source code;
[0025] S3: Input the control flow graph of the smart contract source code into the GAT model based on multi-head attention to extract the graph structure features of the smart contract source code;
[0026] Preferably, the extraction of the graph structure features of the smart contract source code includes:
[0027] S31: Initialize the initial features of the nodes based on the node's code structure, graph structure, and domain knowledge;
[0028] Preferably, the initialization of the initial features of the nodes includes:
[0029] S331: Tokenize the smart contract code to generate a bag of words containing all unique words, symbols, and keywords; for each node in the control flow graph, determine its associated code fragment, and count the number of occurrences of each element in the bag of words in this code fragment to form the bag of words frequency feature of the node. For example, if the bag of words contains 1000 elements, and the word "function" appears 5 times and "uint" appears 3 times in the code fragment associated with the node, etc., then construct a 1000-dimensional vector and fill the corresponding positions with the corresponding number of occurrences.
[0030] S332: Use a dedicated AST construction tool or library to construct the abstract syntax tree of the smart contract. Extract the node type, AST depth, number of child nodes, and parent node type from the AST to form AST features. Concatenate the node bag-of-words frequency features and AST features to obtain the code structure features of the node;
[0031] In this embodiment: One-hot encode common node types (such as function definitions, variable declarations, assignment statements, conditional statements, loop statements, etc.). Assume there are 10 node types. The function definition node is represented as [1, 0, 0, …, 0], the variable declaration node is represented as [0, 1, 0, …, 0], etc.
[0032] The AST depth records the depth value of the node in the AST. The depth of the root node is 0, and the depth value increases by 1 for each level of depth.
[0033] Number of child nodes: Count the number of child nodes owned by the node as a numerical feature.
[0034] Parent node type: One-hot encode the parent node type in the same way as the node type. Combine these features into a vector. Assume the node type encoding is 10-dimensional, the depth is 1-dimensional, the number of child nodes is 1-dimensional, and the parent node type encoding is 10-dimensional. The final AST feature vector has a length of 22 dimensions.
[0035] S333: Calculate the in-degree and out-degree of each node in the control flow graph, and concatenate them to obtain the degree centrality feature; Use a graph algorithm library to calculate the betweenness centrality of each node, that is, count the number of times all the shortest paths pass through this node, and normalize the result so that its value is between 0 and 1 as the betweenness centrality feature of this node; Use a graph algorithm library to calculate the closeness centrality of each node, that is, calculate the reciprocal of the sum of the shortest paths between this node and all other nodes, and normalize it to between 0 and 1 as the closeness centrality feature of this node; Concatenate the degree centrality feature, betweenness centrality feature, and closeness centrality feature of the node to obtain the graph structure feature of the node;
[0036] In this embodiment, the degree centrality feature: Calculate the in-degree and out-degree of each node in the control flow graph, and combine them into a two-dimensional vector. For example, if the in-degree of a certain node is 8 and the out-degree is 5, the degree centrality feature vector is [8, 5].
[0037] Betweenness centrality feature: Use a graph algorithm library (such as NetworkX) to calculate the betweenness centrality of each node, that is, count the number of times all the shortest paths pass through this node, and normalize the result so that its value is between 0 and 1 as a separate numerical feature.
[0038] Closeness centrality feature: Use a graph algorithm library to calculate the closeness centrality of each node, that is, calculate the reciprocal of the sum of the shortest paths between this node and all other nodes, and normalize it to between 0 and 1 as a separate numerical feature.
[0039] S334: According to the knowledge of smart contract security, check whether the node is involved in sensitive operations, and set binary flags for each type of sensitive operation, where being involved is 1 and not being involved is 0; at the same time, check whether the node has known security vulnerability patterns, and set binary flag bits for each vulnerability pattern, where being involved is 1 and not being involved is 0; concatenate the sensitive operation flag vector and the security vulnerability pattern flag vector to obtain the domain knowledge feature of the node;
[0040] According to the knowledge of smart contract security, check whether the node is involved in sensitive operations (such as fund transfer, permission modification, external contract call, etc.), and set binary flags for each type of sensitive operation (being involved is 1 and not being involved is 0). At the same time, check whether the node has known security vulnerability patterns (such as integer overflow, re-entrancy attack, unauthenticated external call, etc.), and set binary flag bits for each vulnerability pattern. Concatenate the sensitive operation flag vector and the security vulnerability pattern flag vector. Assuming there are 3 types of sensitive operations and 5 types of security vulnerability patterns, a final 8-dimensional security feature vector is formed.
[0041] S335: Concatenate the code structure feature, graph structure feature, and domain knowledge feature of the node to obtain the initial feature of the node.
[0042] S32: Input the control flow graph data and the initial feature vector of the node into the GAT model. For each node i and its neighbor node j, use a shared linear transformation matrix to transform the node features, obtaining and where k = 1, …, K, and K is the number of attention heads; represents the weight matrix of the k-th attention head in the l-th hidden layer of GAT; represents the feature embedding of node i output from the (l - 1)-th hidden layer of GAT; represents the feature embedding of the neighbor node j of node i output from the (l - 1)-th hidden layer of GAT;
[0043] S33: Calculate the attention coefficients between the node and its neighbor nodes in the attention head through the attention function:
[0044]
[0045] where represents the transpose of the learnable parameter matrix , ∥ represents concatenation, j, m ∈ N (i), N (i) denotes the set of neighbor nodes of node i; LeakyReLU denotes the activation function; exp denotes the exponential function; denotes the attention coefficient of node i and its neighbor node j output by the k-th attention head in the (l - 1)-th hidden layer of GAT;
[0046] S34: Based on the attention coefficients of the attention heads, perform weighted summation on the neighbor node features to obtain the intermediate features of each node under each attention:
[0047]
[0048] wherein, denotes the intermediate feature output by the k-th attention head of node i in the l-th hidden layer of GAT; σ denotes the activation function;
[0049] S35: Average the intermediate features obtained by K attention heads to obtain the feature of node i after being processed by the multi-head attention mechanism in the l-th layer
[0050]
[0051] wherein, K denotes the number of attention heads in the multi-head attention;
[0052] S36: After calculating through the L-th hidden layer of GAT, obtain the features of the nodes in the last hidden layer, and generate the graph structure features of the smart contract source code by using the node attention mechanism and weighted average aggregation method.
[0053] Preferably, the generated graph structure features of the smart contract source code include:
[0054] S361: Calculate the intermediate feature c:
[0055]
[0056] wherein, n is the total number of nodes, and W is a learnable weight matrix;
[0057] S362: Calculate the attention weight between each node and the intermediate feature c wherein, σ is the Sigmoid activation function, and W A is the weight matrix;
[0058] S363: Calculate the graph structure features of the smart contract source code;
[0059]
[0060] wherein, h G denotes the graph structure features of the smart contract source code.
[0061] S4: Input the intermediate representation IR of the smart contract source code into the second pre-trained language model to extract the execution behavior features of the smart contract source code;
[0062] Preferably, both the first pre-trained language model and the second pre-trained language model adopt the BERT pre-trained language model.
[0063] In this embodiment, in order to capture high-level semantic features from the source code, this method uses the BERT pre-trained language model to extract features from the source code. First, it is necessary to perform data preprocessing on the source code of the smart contract. After data preprocessing, the average length of the code is greatly reduced. Specifically, two cleaning operations need to be performed on the source code: removing code comments and compressing white space characters. Since code comments and redundant white space characters have nothing to do with the code logic, deleting them will not affect the execution logic of the code, and they may introduce noise to interfere with feature extraction. The preprocessed code sequence is converted into three groups of embedding vectors through the embedding layer, which contain the basic information in the smart contract source code, such as code semantics, code structure, and the position information of words in the contract. The source code is converted into a vector through the input embedding layer and then sent to the encoding block to calculate the relationship between words. Since this method utilizes the overall semantic features of the smart contract, the extracted feature vector corresponds to the first character, that is, the feature vector of the [CLS] character. At the same time, in order to extract the execution behavior features and context dependencies of the smart contract from the intermediate representation, this method also uses the BERT pre-trained language model to extract features from the intermediate representation. The powerful context awareness ability of BERT is used to finely model the execution behavior and context dependencies of the contract.
[0064] S5: Input the semantic features of the smart contract source code, the execution behavior features of the smart contract source code, and the graph structure features of the smart contract source code into three classifiers respectively to obtain the corresponding prediction results;
[0065] S6: Perform decision fusion on the prediction results obtained by the three classifiers to obtain the final decision.
[0066] Preferably, the performing decision fusion on the prediction results obtained by the three classifiers to obtain the final decision includes:
[0067] S61: Use the semantic features of the smart contract source code, the execution behavior features of the smart contract source code, and the graph structure features of the smart contract source code to train the corresponding classifiers respectively;
[0068] S62: After training, for the input features, each classifier makes a judgment based on the learned modal features and gives the decision result on whether there is a vulnerability in this smart contract;
[0069] S63: The stacking method is adopted for decision fusion. The decision outputs of multiple classifiers and the true labels are used as the input information of the meta-learner. The meta-learner fuses the decision results of the classifiers to obtain the final decision.
[0070] In this embodiment, each sub-classifier makes a judgment on the input smart contract data based on the learned modal features and gives a decision result on whether there are vulnerabilities in the smart contract. Let the number of sub-classifiers be n, and their corresponding decision outputs are denoted as y1, y2, …, yn, where yi (i = 1, 2, …, n) represents the decision result of the i-th sub-classifier. The stacking method is adopted for decision fusion, and the decision outputs of multiple sub-classifiers are used as the input information. Specifically, the input of the meta-learner is composed of the decision output set {y1, y2, …, y n} and the true label Y. The true label Y is used to provide supervision information during the training of the meta-learner to ensure that the meta-learner can learn how to effectively fuse the decisions of the sub-classifiers. The present invention selects the binomial logistic regression model as the meta-learner to fuse the decision results of the sub-classifiers. The binomial logistic regression model can map the multi-modal decision information of the sub-classifiers into the final classification result to determine whether there are vulnerabilities in the smart contract. In the binomial logistic regression model, the model parameters w are determined by the method of maximizing the likelihood estimation. The specific calculation formula is as follows:
[0071]
[0072] where, x ∈ R n+1 is the input vector of the meta-learner, which includes the decision outputs of the sub-classifiers and possible bias terms, etc.; Y ∈ {0, 1} is the output of the model, representing whether there are vulnerabilities in the smart contract (Y = 1 indicates the existence of vulnerabilities, Y = 0 indicates the non-existence of vulnerabilities).
[0073] Another aspect of the present invention provides an electronic device, including a processor, a memory, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, it implements the above-mentioned smart contract vulnerability detection method based on the pre-training technology.
[0074] Another aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the above-mentioned smart contract vulnerability detection method based on the pre-training technology.
[0075] In summary, initial node features are generated from multimodal perspectives such as code structure, graph structure, and domain knowledge. Through various methods such as the bag-of-words model, AST features, and degree centrality, the characteristics of smart contract nodes are comprehensively captured, providing a rich and accurate data basis for subsequent analysis. Compared with single-modal feature extraction, it can more completely reflect the internal information of smart contracts and improve the detection accuracy. The present invention adopts a graph attention network (GAT) combined with a multi-head attention mechanism. In the calculation of multiple hidden layers, each node can effectively fuse the information of neighboring nodes, dynamically adjust the importance of neighboring nodes, and learn more representative node features. This process fully exploits the information of the control flow graph structure, overcomes the limitations of traditional methods in processing graph-structured data, and improves the efficiency and quality of feature learning. The sub-classifiers make decisions based on different modal features respectively, and then use a binomial logistic regression model to fuse the decisions through a stacking method. This method synthesizes multi-modal decision-making information, avoids the one-sidedness of single-modal decision-making, and can ensure that the detection effect is not lower than that of the best single-modal even in the worst case, significantly improving the accuracy and reliability of smart contract vulnerability judgment and effectively reducing security risks.
[0076] Any reference to a memory, storage, database, or other medium used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0077] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the present technical solution, and they should all be covered by the scope of the claims of the present invention.
Claims
1. A smart contract vulnerability detection method based on pre-training technology, characterized in that: include: S1: extracting multimodal information of the smart contract source code according to the smart contract source code, wherein the multimodal information includes: the smart contract source code, the control flow graph of the smart contract source code, and the intermediate representation IR of the smart contract source code; S2: Input the smart contract source code into the first pre-trained language model to extract the semantic features of the smart contract source code; S3: Input the control flow graph of the smart contract source code into the GAT model based on multi-head attention to extract the graph structure features of the smart contract source code; S4: Input the intermediate representation IR of the smart contract source code into the second pre-trained language model to extract the execution behavior features of the smart contract source code; S5: input the semantic features of the smart contract source code, the execution behavior features of the smart contract source code, and the graph structure features of the smart contract source code into three classifiers respectively to obtain corresponding prediction results; S6: The prediction results obtained by the three classifiers are fused to obtain the final decision.
2. According to claim 1, a smart contract vulnerability detection method based on pre-training technology is characterized in that: The extraction of multimodal information of the smart contract source code includes: preprocessing the smart contract source code; generating a control flow graph of the smart contract source code through the Slither static analysis tool; and generating an intermediate representation IR of the smart contract source code through the Slither tool.
3. According to claim 1, a smart contract vulnerability detection method based on pre-training technology is characterized in that: The first pre-trained language model and the second pre-trained language model both adopt the BERT pre-trained language model.
4. According to claim 1, a smart contract vulnerability detection method based on pre-training technology is characterized in that: The graph structure features of extracting the smart contract source code include: S31: Initialize the initial features of the node based on the node's code structure, graph structure and domain knowledge; S32: Input the control flow graph data and the initial feature vector of the node into the GAT model. For each node i and its neighbor node j, use the shared linear transformation matrix in each attention head k. Transform the node features to obtain and Where k = 1,…,K, K is the number of attention heads; represents the weight matrix of the kth attention head in the lth hidden layer of GAT; represents the feature embedding of node i output in the l-1th hidden layer of GAT; Represents the feature embedding of node i’s neighbor node j output in the l-1th hidden layer of GAT; S33: Calculate the attention coefficient of the node in the attention head and its neighboring nodes through the attention function: in, Represents the learnable parameter matrix The transpose of , ∥ represents concatenation, j,m∈N (i) , N (i) represents the set of neighbor nodes of node i; LeakyReLU represents the activation function; exp represents the exponential function; represents the attention coefficient of node i and its neighbor node j output by the kth attention head in the l-1th hidden layer of GAT; S34: Based on the attention coefficient of the attention head, the neighbor node features are weighted summed to obtain the intermediate features of each node under attention: in, represents the intermediate feature of node i output by the kth attention head in the lth hidden layer of GAT; σ represents the activation function; S35: Average the intermediate features obtained by K attention heads to obtain the features of node i after being processed by the multi-head attention mechanism at layer l Among them, K represents the number of attention heads in multi-head attention; S36: After the Lth hidden layer calculation of GAT, the features of the node in the last hidden layer are obtained, and the graph structure features of the smart contract source code are generated by using the node attention mechanism and weighted average aggregation method.
5. According to claim 4, a smart contract vulnerability detection method based on pre-training technology is characterized in that: It is characterized in that The initial features of the initialization node include: S331: Segment the smart contract code to generate a word bag containing all unique words, symbols and keywords; for each node in the control flow graph, determine its associated code snippet, count the number of occurrences of each element in the word bag in the code snippet, and form the word bag frequency feature of the node; S332: Use a dedicated AST construction tool or library to build an abstract syntax tree for the smart contract, extract the node type, AST depth, number of child nodes, and parent node type for each node from the AST, and combine the node bag-of-words frequency features and AST features to obtain the code structure features of the node; S333: Calculate the in-degree and out-degree of each node in the control flow graph, and concatenate them to obtain the degree centrality feature; calculate the betweenness centrality of each node with the help of the graph algorithm library, that is, count the number of times all shortest paths pass through the node, and normalize the result so that its value is between 0 and 1, as the betweenness centrality feature of the node; use the graph algorithm library to calculate the closeness centrality of each node, that is, calculate the inverse of the sum of the shortest paths between the node and all other nodes, and normalize it to between 0 and 1, as the closeness centrality feature of the node; concatenate the degree centrality feature, betweenness centrality feature and closeness centrality feature of the node to obtain the graph structure feature of the node; S334: Check whether the node is involved in sensitive operations based on smart contract security knowledge, and set a binary mark for each sensitive operation type, where involved is 1 and not involved is 0; at the same time, check whether the node has a known security vulnerability pattern, and set a binary mark for each vulnerability pattern, where involved is 1 and not involved is 0; concatenate the sensitive operation mark vector and the security vulnerability pattern mark vector to obtain the domain knowledge feature of the node; S335: The code structure features, graph structure features and domain knowledge features of the node are concatenated to obtain the initial features of the node.
6. According to claim 4, a smart contract vulnerability detection method based on pre-training technology is characterized in that: The graph structure features of generating smart contract source code include: S361: Calculate the intermediate feature c: Where n is the total number of nodes and W is a learnable weight matrix; S362: Calculate the attention weight between each node and the intermediate feature c Among them, σ is the Sigmoid activation function, W A is the weight matrix; S363: Calculate the graph structure characteristics of smart contract source code; Among them, h G Graph structure features representing smart contract source code.
7. According to claim 1, a smart contract vulnerability detection method based on pre-training technology is characterized in that: It is characterized in that The decision fusion of the prediction results obtained by the three classifiers to obtain the final decision includes: S61: using the semantic features of the smart contract source code, the execution behavior features of the smart contract source code, and the graph structure features of the smart contract source code to train corresponding classifiers respectively; S62: After the training is completed, each classifier makes a judgment based on the input features according to the learned modal features, and gives a decision result on whether the smart contract has a vulnerability; S63: A stacking method is used for decision fusion. The decision outputs of multiple classifiers and true labels are used as input information of the meta-learner. The decision results of the classifiers are fused by the meta-learner to obtain the final decision.
8. An electronic device, characterized in that: The method comprises a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the computer program implements the method according to any one of claims 1 to 7 when executed by the processor.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.