Intelligent contract vulnerability detection and positioning method and system based on interpretable graph neural network

Through the method based on the interpretability graph neural network, the attribute control flow graph is constructed and the graph neural network model is used to solve the problem that smart contract reentry vulnerabilities are difficult to locate code-line-level code, and efficient vulnerability repair is achieved.

CN120372629APending Publication Date: 2025-07-25HOHAI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510496491.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

In the smart contract vulnerability detection, especially reentry vulnerability detection, it is difficult to achieve accurate code-line-level positioning, resulting in low vulnerability repair efficiency.

Method used

Using an interpretability graph neural network method, by compiling the smart contract source code into bytecode, building attribute control flow graphs, using the graph neural network model for training, identifying key nodes and mapping back to the source code, and generating a sorted list of suspicious statements.

Benefits of technology

It realizes the precise positioning of smart contract reentry vulnerabilities to specific code lines, significantly improving the efficiency of vulnerability repair.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120372629A_ABST
    Figure CN120372629A_ABST
Patent Text Reader

Abstract

The invention discloses a method for detecting and positioning a smart contract re-entry vulnerability. Comprising the following steps: compiling an intelligent contract source code into a byte code, extracting an operation code sequence from the byte code, and constructing a program control flow diagram with attributes; and training by using a graph neural network model to detect whether the smart contract has a re-entry vulnerability, explaining a result by using an interpretable graph neural network model, identifying a sub-graph structure which has the greatest influence on a classification result, mapping key nodes in the sub-graph structure back to a source code, and generating a sorting list of suspicious statements. According to the method, when the intelligent contract reentry vulnerability is detected, the specific code line is accurately positioned, and the vulnerability repairing efficiency of developers can be remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of blockchain technology, and particularly to a method and system for detecting and locating vulnerabilities in smart contracts based on an interpretable graph neural network. Background Art

[0002] Smart contracts are a core component of blockchain technology and are widely used in multiple fields such as decentralized finance (DeFi), smart asset management, digital identity authentication, and supply chain management. As one of the important features of blockchain technology, smart contracts provide a solid infrastructure support for the development of decentralized applications (DApps). Through smart contracts, blockchain not only realizes the function of decentralized data storage but also can execute decentralized logic and automated operations. This has enabled blockchain to be widely used in many industries such as finance, logistics, law, and healthcare, promoting the diversified development of blockchain technology. However, smart contracts have also exposed many security problems in actual applications. Especially due to programming errors, vulnerabilities, and the lack of a complete security verification mechanism, it may lead to serious asset losses and contract attacks.

[0003] Although the vulnerability detection rate of most existing deep learning models has been improved to a certain extent, the positioning granularity is at the file level and has not been refined to the code line level. For this reason, the present invention proposes a method for detecting and locating reentrancy vulnerabilities in smart contracts based on an interpretable graph neural network.

[0004] The method proposed by the present invention compiles the smart contract source code into bytecode, extracts the opcode sequence from the bytecode, and constructs a program control flow graph with attributes; uses a graph neural network model for training to detect whether there are reentrancy vulnerabilities in the smart contract, uses an interpretable graph neural network model to interpret the results, identifies the subgraph structure that has the greatest impact on the classification result, maps the key nodes in the subgraph structure back to the source code, and generates a sorted list of suspicious statements. The present invention accurately locates specific code lines while detecting reentrancy vulnerabilities in smart contracts. Summary of the Invention

[0005] The purpose of the present invention is to provide a method and system for detecting and locating reentrancy vulnerabilities in smart contracts, and the specific steps are as follows:

[0006] Step 1: Take the smart contract source code as input and use the compiler Solc for compilation to obtain the corresponding bytecode sequence;

[0007] Step 2: Analyze the bytecode sequence to obtain the corresponding opcodes, divide the basic nodes of the control flow graph according to the corresponding opcode instructions, and assign features to the basic nodes according to different attributes of the opcodes to obtain an attribute control flow graph;

[0008] Step 3: Input the obtained representation of the smart contract control flow graph and the graph node features into a graph neural network model for iterative training and validation to construct a model classifier;

[0009] Step 4: Conduct interpretability training of the model. Input the control flow graph that the model classifier determines to have vulnerabilities, and identify the nodes that have the greatest impact on the judgment result;

[0010] Step 5: Use the source code - opcode mapping file provided by the compiler to map the opcodes in the nodes back to the source code, generate a sorted list of suspicious statements, and finally obtain the line where the vulnerable code is located.

[0011] Further, the said Step 1 includes:

[0012] Step 1-1: Collect a dataset of contract source codes through the Ethereum platform. At the same time, combine the descriptions of vulnerabilities on the open source platform to label the source codes;

[0013] Step 1-2: Bytecode conversion. Use the Solidity language compiler Solc to compile the source code of the smart contract into bytecode, and parse the bytecode into an opcode sequence.

[0014] Further, the said Step 2 includes:

[0015] Step 2-1: Based on the opcode sequence obtained in Step 1-2, identify and divide basic blocks according to jump instructions JUMP, JUMPI, JUMPDEST, RETURN, STOP, REVERT, etc. in the opcode sequence;

[0016] Step 2-2: Based on the opcode sequence obtained in Step 1-2, analyze the target addresses of the jump instructions to determine the control flow transfer relationship between basic blocks and construct a program control flow graph;

[0017] Step 2-3: Combine the identified basic blocks and the determined control flow relationship to construct a complete control flow graph, and assign attributes to the nodes in the control flow graph according to the category to which the opcode belongs.

[0018] Further, the said Step 3 includes:

[0019] Step 3-1: Input the control flow graph into a trained GNN model. The input data is G=(V, E, X), where V is the set of nodes in the control flow graph, E is the set of edges in the control flow graph, X∈R n*d is the feature matrix of the nodes, n is the number of nodes, d is the feature dimension of each node, and use A to represent the adjacency matrix of the control flow graph;

[0020] Step 3-2: Process the graph data through the convolutional layer, SortPooling layer, and final stage. For each convolution, there is:

[0021]

[0022] where $Z$ t represents the output of the $t$-th layer of the graph convolutional layer is the graph adjacency matrix with additional self-loops, and the values therein are:

[0023]

[0024] is the diagonal matrix of the graph, and each element value therein is:

[0025]

[0026] $W\in\mathbb{R}$ d*d’ is the matrix of trainable graph convolution parameters, $f$ is a non-linear activation function. In the SortPooling layer, mainly the output of the last layer of the graph convolutional layer is used to sort the nodes, and the sorted nodes are truncated or padded to generate a fixed-size feature vector;

[0027] Step 3-3: After SortPooling, enter the classification stage. The generated feature vector is input into the classifier for vulnerability identification through the following formula:

[0028] $p$ i $=\text{Softmax}(W\cdot z + b)$,

[0029] where $z$ is the generated feature vector, $W$ is the weight matrix of the classifier, $b$ is the bias term, and $p_i$ is the prediction probability.

[0030] Furthermore, the said step 4 includes:

[0031] Step 4-1: Learn the mask of node features. Define a node feature selection vector $F\in\{0,1\}^d$, and the values in the vector are either 0 or 1. Optimize the mask to interpret the classification result. If the features masked by the mask are very important, then the prediction result will change. The optimization formula is:

[0032]

[0033] where $G'$ and $F$ represent the subgraph and the feature selector respectively;

[0034] Step 4-2: Generate the subgraph structure containing key nodes and their associated features.

[0035] Furthermore, the said step 5 includes:

[0036] Step 5-1: Establish a one-to-one correspondence between the opcode sequence and the source code using the mapping file provided by the compiler Solc;

[0037] Step 5-2: Search for and return the line where the corresponding vulnerability code is located.

[0038] An intelligent contract reentrancy vulnerability detection and localization system based on an interpretable graph neural network, characterized by including:

[0039] A detection module: used to extract the opcode sequence from the bytecode of the intelligent contract, construct a control flow graph with attributes, and use a graph neural network model to detect whether the contract contains a reentrancy vulnerability;

[0040] A localization module: used to interpret the results of the detection stage using a GNN model interpreter, identify the subgraph structure that has the greatest impact on the detection results, and map the key opcode blocks back to the source code to generate a ranked list of suspicious statements.

[0041] The present invention provides a computer system, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is loaded into the processor, it implements the intelligent contract reentrancy vulnerability detection and localization method and system based on an interpretable graph neural network.

[0042] The present invention also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and is characterized in that when the computer program is executed by a processor, it implements the intelligent contract reentrancy vulnerability detection and localization method and system based on an interpretable graph neural network.

[0043] Beneficial effects: The method proposed by the present invention can accurately locate specific code lines while detecting intelligent contract reentrancy vulnerabilities, significantly improving the efficiency of developers in fixing vulnerabilities. Description of the Drawings

[0044] Figure 1 It is the overall flowchart of the embodiment of the present invention. Detailed Embodiments

[0045] The following further clarifies the present invention in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. After reading the present invention, various equivalent forms of modification of the present invention by those skilled in the art all fall within the scope defined by the appended claims of this application.

[0046] As Figure 1 shown, an intelligent contract reentrancy vulnerability detection and localization method disclosed in an embodiment of the present invention mainly includes the following steps:

[0047] Step 1: Take the source code of the smart contract as input and use the compiler Solc for compilation to obtain the corresponding bytecode sequence. Step 1 includes the following steps:

[0048] Step 1-1: Collect the contract source code dataset through the Ethereum platform, and at the same time combine the descriptions of vulnerabilities on the open source platform to tag the source code;

[0049] Step 1-2: Bytecode conversion, use the Solidity language compiler Solc to compile the source code of the smart contract into bytecode, and parse the bytecode into an opcode sequence.

[0050] Step 2: Analyze the bytecode sequence to obtain the corresponding opcodes, divide the basic nodes of the control flow graph according to the corresponding opcode instructions, and assign characteristics to the basic nodes according to the different attributes of the opcodes to obtain the attribute control flow graph. Step 2 includes the following steps:

[0051] Step 2-1: Based on the opcode sequence obtained in Step 1-2, identify and divide basic blocks according to instructions such as JUMP, JUMPI, JUMPDEST, RETURN, STOP, REVERT in the opcode sequence;

[0052] Step 2-2: Based on the opcode sequence obtained in Step 1-2, analyze the target addresses of the jump instructions to determine the control flow transfer relationship between basic blocks and construct the program control flow graph;

[0053] Step 2-3: Combine the identified basic blocks and the determined control flow relationship to construct a complete control flow graph, and assign attributes to the nodes in the control flow graph according to the category of the opcode. The opcode types are shown in Table 1 specifically.

[0054] Table 1: Classification of opcodes

[0055]

[0056] Step 3: Input the obtained smart contract control flow graph representation and graph node characteristics into the graph neural network model for iterative training and verification to construct a model classifier. Step 3 includes the following steps:

[0057] Step 3-1: Input the control flow graph into the trained GNN model. The input data is G=(V, E, X), where V is the set of nodes in the control flow graph, E is the set of edges in the control flow graph, X∈R n*d is the feature matrix of the nodes, n is the number of nodes, d is the feature dimension of each node, and use A to represent the adjacency matrix of the control flow graph;

[0058] Step 3-2: Process the graph data through the convolutional layer, SortPooling layer, and the final stage. For each convolution, there is:

[0059]

[0060] where Z t represents the output of the t-th layer of the graph convolutional layer is the graph adjacency matrix with additional self-loops, and its values are:

[0061]

[0062] is the diagonal matrix of the graph, and each element value is:

[0063]

[0064] W ∈ R d*d’ is the matrix of trainable graph convolution parameters, f is a non-linear activation function. In the SortPooling layer, mainly the output of the last layer of the graph convolutional layer is used to sort the nodes, and the sorted nodes are truncated or padded to generate a fixed-size feature vector;

[0065] Step 3-3: After SortPooling, enter the classification stage. The generated feature vector is input into the classifier for vulnerability identification through the following formula:

[0066] where z is the generated feature vector, W is the weight matrix of the classifier, b is the bias term, and pi is the prediction probability.

[0067] Step 4: Conduct interpretability training of the model. Input the control flow graph that the model classifier determines to have vulnerabilities, and identify the nodes that have the greatest impact on the determination result. Step 4 includes the following steps:

[0068] Step 4-1: Learn the mask of node features. Define a node feature selection vector F ∈ 1*d, and the values in the vector are either 0 or 1. Optimize the mask to explain the classification result. If the features masked by the mask are very important, then the prediction result will change. The optimization formula is:

[0069]

[0070] where G’ and F represent the subgraph and the feature selector respectively;

[0071] Step 4-2: Generate the subgraph structure containing the key nodes and their associated features.

[0072] Step 5: Utilize the source code - opcode mapping file provided by the compiler to map the opcodes in the nodes back to the source code, generate a sorted list of suspicious statements, and finally obtain the line where the vulnerable code is located. Step 5 includes the following steps:

[0073] Step 5-1: Use the mapping file provided by the compiler Solc to establish a one-to-one correspondence between the opcode sequence and the source code;

[0074] Step 5-2: Search for and return the line where the corresponding vulnerable code is located.

[0075] To further verify the method proposed in this application, this application verifies the effectiveness of the method by setting up the same experimental environment. The dataset used in this application contains 1000 Ethereum smart contracts. The pattern extraction tool and the graph construction tool are implemented in Python, and the graph neural network is implemented in pytorch. 70% of the functions are randomly selected as the training set, another 20% as the test set, and 10% as the validation set. The experimental results are shown in Table 2 and Table 3:

[0076] Table 2: Detection experiment results

[0077] Method Accuracy Rate Recall Rate Precision F1 Score Mythril 0.50 0.52 0.59 0.58 Oyente 0.58 0.57 0.62 0.59 Slither 0.54 0.67 0.60 0.55 CGE 0.85 0.87 0.88 0.84 TMP 0.8 0.85 0.84 0.75 DR-GCN 0.82 0.84 0.81 0.77 ReVulDL 0.87 0.84 0.89 0.88 Our Method 0.89 0.86 0.80 0.89

[0078] Table 3: Localization experiment results

[0079] Method Top-1 Top-5 Top-10 MAR MFR ReVulDL 17.7% 55.8% 68.5% 6.2 5.8 Our Method 20.0% 61.1% 74.5% 4.1 .5.1

[0080] From the experiments, it can be seen that this application achieves an 89% accuracy rate, an 86% recall rate, an 80% precision, and an 89% F1 score for detecting reentrancy vulnerabilities. For the accuracy of vulnerability detection, the method of the present invention is on average about 5% higher than the existing methods in terms of experimental metrics; this application achieves 20.0% Top-1, 61.1% Top-5, 74.5% Top-10, a MAR of 4.1, and an MFR of 5.1 for localizing reentrancy vulnerabilities, all of which are better than the existing methods. This proves the effectiveness of this application, which can accurately locate reentrancy vulnerabilities and also has satisfactory performance in terms of accuracy rate and precision.

[0081] Based on the same inventive concept, a detection and localization system for reentrancy vulnerabilities in smart contracts based on an interpretable graph neural network includes a detection module for extracting an opcode sequence from the bytecode of a smart contract, constructing a control flow graph with attributes, and using a graph neural network model to detect whether the contract contains a reentrancy vulnerability; a localization module for using a GNN model interpreter to interpret the results of the detection stage, identifying the subgraph structure that has the greatest impact on the detection results, and mapping the key opcode blocks back to the source code to generate a ranked list of suspicious statements.

[0082] For the specific working processes of the modules described above, reference may be made to the corresponding processes in the foregoing method embodiments, which will not be elaborated herein. The division of the modules is only a logical function division, and there may be other division methods in actual implementation. For example, multiple modules may be combined or integrated into another system.

[0083] Based on the same inventive concept, a computer system disclosed in an embodiment of the present invention includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is loaded into the processor, the steps of the method for detecting and locating smart contract reentry vulnerabilities based on an interpretable graph neural network are implemented.

[0084] Those skilled in the art can understand that the technical solution of the present invention, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer system (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in the embodiments of the present invention. The storage medium includes: various media that can store computer programs such as USB flash drives, mobile hard disks, read-only memory ROM, random access memory RAM, magnetic disks, or optical discs.

Claims

1. An intelligent contract vulnerability detection and localization method based on an interpretable graph neural network, characterized in that, It includes the following steps: Step 1: Take the source code of the smart contract as input and compile it using the compiler Solc to obtain the corresponding bytecode sequence; Step 2: Analyze the bytecode sequence to obtain the corresponding opcodes. According to the corresponding opcode instructions, divide the basic nodes of the control flow graph. Based on the different attributes of the opcodes, assign characteristics to the basic nodes to obtain an attribute control flow graph; Step 3: Input the obtained smart contract control flow graph representation and graph node characteristics into a graph neural network model for iterative training and verification to construct a model classifier; Step 4: Conduct interpretability training of the model. Input the control flow graph determined by the model classifier to have vulnerabilities, and identify the nodes that have the greatest impact on the determination result; Step 5: Use the source code-opcode mapping file provided by the compiler to map the opcodes in the nodes back to the source code, generate a sorted list of suspicious statements, and finally obtain the line where the vulnerable code is located.

2. The intelligent contract vulnerability detection and location method based on an interpretable graph neural network according to claim 1, characterized in that The said Step 1 includes the following steps: Step 1-1: Collect the contract source code dataset through the Ethereum platform. At the same time, combine the descriptions of vulnerabilities on the open source platform to tag the source code; Step 1-2: Bytecode conversion. Use the Solidity language compiler Solc to compile the source code of the smart contract into bytecode, and parse the bytecode into an opcode sequence.

3. The intelligent contract vulnerability detection and location method based on an interpretable graph neural network according to claim 1, characterized in that The said Step 2 includes the following steps: Step 2-1: Based on the opcode sequence obtained in Step 1-2, identify and divide basic blocks according to instructions such as the jump instructions JUMP, JUMPI, JUMPDEST, RETURN, STOP, REVERT, etc. in the opcode sequence; Step 2-2: Based on the opcode sequence obtained in Step 1-2, analyze the target addresses of the jump instructions to determine the control flow transfer relationship between basic blocks and construct a program control flow graph; Step 2-3: Combine the identified basic blocks and the determined control flow relationship to construct a complete control flow graph, and assign attributes to the nodes in the control flow graph according to the category to which the opcode belongs.

4. The intelligent contract vulnerability detection and location method based on an interpretable graph neural network according to claim 1, wherein The said Step 3 includes the following steps: Step 3-1: Input the control flow graph into the trained GNN model. The input data is G=(V, E, X), where V is the set of nodes in the control flow graph, E is the set of edges in the control flow graph, X∈R n*d is the feature matrix of the nodes, n is the number of nodes, d is the feature dimension of each node, and A is used to represent the adjacency matrix of the control flow graph; Step 3-2: Process the graph data through convolutional layers, SortPooling layers, and the final stage. For each convolution, there is: Among them, Z t represents the output of the t-th layer graph convolutional layer is the graph adjacency matrix with additional self-loops, and for its values, there are: is the diagonal matrix of the figure, and for each element value therein, there is: W ∈ R d*d’ is a matrix of trainable graph convolution parameters, and f is a non-linear activation function. In the SortPooling layer, the main operation is to sort the nodes based on the output of the last graph convolution layer, and then truncate or pad the sorted nodes to generate a feature vector of a fixed size; Step 3-3: After SortPooling, enter the classification stage. Input the generated feature vector into the classifier for vulnerability identification through the following formula: p i = Softmax(W·z + b), where z is the generated feature vector, W is the weight matrix of the classifier, b is the bias term, and pi is the prediction probability.

5. The intelligent contract vulnerability detection and location method based on an interpretable graph neural network according to claim 1, characterized in that The said Step 4 includes the following steps: Step 4-1: Learn the mask of node features. Define a node feature selection vector F ∈ 1*d, and the values in the vector are either 0 or 1. Optimize the mask to explain the classification result. If the features masked by the mask are very important, then the prediction result will change. The optimization formula is: where G’ and F represent the subgraph and the feature selector respectively; Step 4-2: Generate a subgraph structure containing key nodes and their associated features.

6. The intelligent contract vulnerability detection and localization method based on an interpretable graph neural network according to claim 1, characterized in that The said Step 5 includes the following steps: Step 5-1: Use the mapping file provided by the compiler Solc to establish a one-to-one correspondence between the opcode sequence and the source code; Step 5-2: Locate and return the line where the corresponding vulnerability code is located.

7. An intelligent contract vulnerability detection and localization system based on an interpretable graph neural network, characterized in that, Including: Detection module: used to extract the opcode sequence from the bytecode of the smart contract, construct a control flow graph with attributes, and use the graph neural network model to detect whether the contract contains a reentrancy vulnerability; Location module: used to interpret the results of the detection phase using the GNN model interpreter, identify the subgraph structure that has the greatest impact on the detection results, and map the key opcode blocks back to the source code to generate a ranked list of suspicious statements.

8. An intelligent contract vulnerability detection and location system based on an interpretable graph neural network according to claim 7, characterized in that, The detection module includes a bytecode conversion unit, a control flow graph construction unit, and a GNN model training and detection unit.

9. An intelligent contract vulnerability detection and location system based on an interpretable graph neural network according to claim 7, characterized in that The location module includes a GNN model interpretation unit and a key block mapping unit.

10. A computer system, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the computer program is loaded into the processor, it implements a method for detecting and locating smart contract vulnerabilities based on an interpretable graph neural network according to any one of claims 1-6.

11. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements a method for detecting and locating smart contract vulnerabilities based on an interpretable graph neural network according to any one of claims 1-6.