Smart contract vulnerability detection method and system based on taint analysis
By applying taint analysis and symbol execution technology in the EVM bytecode of smart contracts, identifying and distinguishing vulnerabilities related to access rights control, the shortcomings in accuracy and coverage of existing detection methods are solved, and more efficient and reliable smart contract security analysis is achieved.
Patent Information
- Application Number
- CN202510314768.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-03-18
AI Technical Summary
Existing smart contract vulnerability detection methods are difficult to accurately identify vulnerabilities related to access permission control, especially due to the lack of clear access permission control policy specifications and insufficient generalization capabilities of predefined rules, resulting in high false positive rates and incomplete vulnerability coverage.
Using a taint analysis method, the control flow analysis and taint analysis of the EVM bytecode of the smart contract is used, access control checks and status variables are identified, and the generation constraints are generated by symbol execution are distinguished from expected normal operations from actual vulnerabilities.
It significantly improves the accuracy and reliability of smart contract security analysis, reduces the false alarm rate of vulnerability detection, accurately identify access permission control vulnerabilities, and provides stronger security guarantees.
Smart Images

Figure SMS_1 
Figure SMS_2 
Figure SMS_3
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of vulnerability detection, and in particular to an intelligent contract vulnerability detection method and system based on taint analysis. Background Art
[0002] Smart contracts are computer programs that automatically execute, control, or record legal events and behaviors. They run on the Ethereum blockchain and can execute complex transactions and protocols without the need for intermediaries. However, once a smart contract is deployed on the blockchain, its code cannot be changed, and any potential security vulnerabilities can be maliciously exploited, resulting in financial losses or other serious consequences.
[0003] Programming languages such as Solidity for writing smart contracts did not consider security vulnerabilities based on access control during their design. Smart contract developers can only implement access control settings based on their own judgment and experience, which leads to a large number of vulnerabilities related to access control in smart contracts. Such vulnerabilities may allow unauthorized users to access or operate key functions in the contract, not only harming the interests of users but also having a negative impact on the stability and credibility of the Ethereum ecosystem.
[0004] Static analysis plays an important role in identifying access control vulnerabilities. However, due to the lack of clear access control policy specifications, it is difficult to accurately identify access control checks. Secondly, there may be some code patterns in smart contracts that seem to be vulnerabilities but are actually part of the contract's functions. Due to the lack of contract access control specifications, it is difficult to distinguish these patterns from real vulnerabilities during static analysis.
[0005] In addition, among the existing detection methods for access control vulnerabilities, some rely on predefined rules to identify access control checks and assume that the contract is written in a specific format, and the compiler always generates specific bytecode patterns for these checks. However, the generalization ability of these predefined rules is poor, the accuracy is insufficient, and false positives are easily caused. There are also methods that identify permission errors in smart contracts by analyzing the contract transaction history on the blockchain. In fact, not all contract functions will be called and generate transactions. For example, the function to delete a smart contract is only triggered under specific circumstances. These existing methods do not focus on access control vulnerabilities across the entire contract caused by weak or missing access control. They mainly analyze the problem of missing access control in specific code statements based on predefined access control patterns. Summary of the Invention
[0006] To solve the above problems raised in the background art, the present invention provides an intelligent contract vulnerability detection method and system based on taint analysis.
[0007] The technical solution of the present invention is as follows: An intelligent contract vulnerability detection method based on taint analysis, comprising the following steps: S1. Compile the intelligent contract into EVM bytecode, perform control flow analysis on the EVM bytecode to obtain a control flow graph, and obtain several key operation instructions in the control flow graph; Extract the key parameters of each key operation instruction, perform backward slicing on the key parameters on the control flow graph to obtain a slice containing instructions directly or indirectly controlled by the attacker, and construct a slice set; S2. Perform access permission control identification on the EVM bytecode based on the defined access permission control conditions: perform backward data flow analysis on each conditional statement in the EVM bytecode, and identify the conditional statement containing the instructions in the slice set as an access permission control check, and construct a condition set based on all access permission control checks; Identify the storage location accessed corresponding to the access permission control check as an access permission control state variable; S3. Set the user input containing the execution operation as a taint source, and set the address parameter of the key operation instruction and the SSTORE instruction writing to the access permission control state variable as taint sinks. Perform taint analysis on the control flow graph based on the taint source and taint sinks, and mark the contaminated operands as taints; S4. Perform symbolic execution on the taint flow path where the taint flowing to the access permission control state variable is located, generate the constraint conditions of the taint flow path, solve the constraint conditions, and if a value that makes the constraint conditions hold can be found, then take the negation of the constraint conditions to obtain the negative constraint conditions; Further solve the negative constraint conditions. If a value that makes the negative constraint conditions hold can be found, then mark the operand corresponding to the taint as an expected normal behavior. If not, then mark the operand corresponding to the taint as an access permission control vulnerability to obtain the vulnerability detection result.
[0008] Specifically, the access permission control conditions defined in S2 are specifically: Define access permission control check: If there is a conditional statement C in the EVM bytecode such that the caller A passes the verification of the conditional statement C, then the caller A can execute the code controlled by the conditional statement C, and this conditional statement C is an access permission control check; Define access permission control state variable: If there is an access permission control check C in the EVM bytecode, the value V in the stored data item D can be used to compare with the caller A, or the value V loaded from the stored data item D can be used to calculate the access permission control check C, then the stored data item D is an access permission control state variable; Define authorized user: If in the EVM bytecode, the access control status variable D contains the address of the caller A, then the caller A is an authorized user protected by the defined access control check C.
[0009] Furthermore, in S2, it also includes excluding non-access control checks that do not meet the access control check judgment conditions from the condition set. The access control check judgment condition is: If there is no data mapping relationship in the conditional statement, or there is a data mapping relationship in the conditional statement and the state variable it uses stores a boolean value, then this condition is an access control check.
[0010] Several key operation instructions in S1 include CALL, CALLCODE, DELEGATECALL, and SELFDESTRUCT instructions.
[0011] In S1, obtain the slice containing instructions directly or indirectly controlled by the attacker. Instructions directly controlled by the attacker include: ORIGIN, CALLER, CALLVALUE, CALLDATALOAD, CALLDATASIZE, CALLDATACOPY. Instructions indirectly controlled by the attacker include: SLOAD, MLOAD.
[0012] The key parameters in S1 include: target address, the amount of currency sent, input data.
[0013] In S3, perform taint analysis on the control flow graph based on taint sources and taint sinks, and mark the contaminated operands as tainted. Specifically: Perform backward slicing on the taint sink on the control flow graph to obtain all path segments that can reach the taint sink. Then, starting from the taint source, along the execution order of the path segments, use the TaintAnalyse function to mark the contaminated operands as tainted.
[0014] Furthermore, during the taint analysis process, it also includes starting from each taint source, traversing the control flow graph to find all paths flowing to the SSTORE instructions that write to storage slots. If the storage slot is tainted by taint, and its taint source is not the storage slot itself or the taint source has been marked as tainted, then mark the storage slot as tainted, and at the same time record the address of the SSTORE instruction that writes to this storage slot.
[0015] In S4, use the Z3 solver to solve both the constraint conditions and the negative constraint conditions.
[0016] The present invention also provides an intelligent contract vulnerability detection system based on taint analysis, including: Static analysis module: used to compile the smart contract into EVM bytecode, perform control flow analysis on the EVM bytecode to obtain a control flow graph, and acquire several key operation instructions in the control flow graph; extract the key parameters of each key operation instruction, perform backward slicing on the key parameters on the control flow graph to obtain a slice containing instructions directly or indirectly controlled by the attacker, and construct a slice set; Access control identification module: used to perform access control identification on the EVM bytecode based on the defined access control conditions: perform backward data flow analysis on each conditional statement in the EVM bytecode, identify the conditional statements containing the instructions in the slice set as access control checks, and construct a condition set based on all access control checks; identify the storage locations accessed by the access control checks as access control state variables; Taint analysis module: used to set the user input containing the execution operation as the taint source, and set the address parameter of the key operation instruction and the SSTORE instruction for writing to the access control state variable as the taint sinks. Perform taint analysis on the control flow graph based on the taint source and taint sinks, and mark the contaminated operands as tainted; Vulnerability detection module: used to perform symbolic execution on the taint flow path where the taint flowing to the access control state variable is located, generate the constraint conditions of the taint flow path, solve the constraint conditions. If a value that makes the constraint conditions hold can be found, then take the negation of the constraint conditions to obtain the negative constraint conditions; further solve the negative constraint conditions. If a value that makes the negative constraint conditions hold can be found, then mark the operand corresponding to the taint as the expected normal behavior. If not, then mark the operand corresponding to the taint as an access control vulnerability to obtain the vulnerability detection result.
[0017] The beneficial effects of the present invention are as follows: 1. A method for detecting smart contract vulnerabilities based on taint analysis provided by the present invention, by performing control flow analysis on the EVM bytecode of the smart contract, screening key information, including key operation instructions and key parameters, and combining the taint analysis method to transform the detection of access control vulnerabilities into a pollution analysis problem, without relying on predefined access control modes or existing transaction records, can significantly improve the accuracy and reliability of smart contract security analysis, and provide strong support for the security guarantee of smart contracts.
[0018] 2. The present invention generates constraint conditions for the taint flow path through symbolic execution, efficiently and accurately distinguishes the expected normal operations in the smart contract from the real security vulnerabilities, takes the negation of the constraint conditions to obtain the negative constraint conditions, and further solves the negative constraint conditions to distinguish whether an unauthorized user can obtain the access control permission of the contract in accordance with the expected functional logic under specific circumstances. Through symbolic execution analysis, the potential expected normal behaviors are distinguished from the explicit access permission control vulnerabilities, thereby reducing the false positive rate of vulnerability detection. Detailed implementation manners
[0019] The exemplary implementation manners of the present disclosure will be described in more detail below in conjunction with the embodiments.
[0020] Embodiment This embodiment provides a method for detecting smart contract vulnerabilities based on taint analysis, including the following steps: S1. Compile the smart contract into EVM bytecode, perform control flow analysis on the EVM bytecode to obtain a control flow graph, and obtain several key operation instructions in the control flow graph; Extract the key parameters of each key operation instruction, perform backward slicing on the key parameters on the control flow graph to obtain a slice containing the instructions directly or indirectly controlled by the attacker, and construct a slice set.
[0021] The present invention mainly detects access permission control vulnerabilities, and the access permission control vulnerabilities in smart contracts are mainly divided into two categories. The first type of vulnerability stems from weak access permission control checks, enabling attackers to bypass these checks. We define this type of problem as "access permission control violation"; the second type of vulnerability is due to the lack of necessary access permission control, which we name "lack of access permission control". In addition, the normal contract operations designed by developers are defined as "expected normal operations".
[0022] Regarding access control violations, as shown in the example in Table 1. The access control vulnerabilities in the smart contract (BAFCToken) in Table 1 are mainly due to the naming error of the constructor UBSecToken, which should be named the contract name BAFCToken to ensure that it is executed only once during deployment as a constructor. Due to the naming error, the UBSecToken function can still be called after the smart contract is deployed, allowing anyone to set themselves as the owner of the smart contract, thus bypassing the access control check of the onlyOwner modifier. In addition, since the owner variable can be arbitrarily changed, attackers can take advantage of this to freeze and unfreeze any account, including their own, to bypass the check of the unFrozenAccount modifier. At the same time, the variable name error in the switchLiquidity function may cause the function to not execute correctly, and the access control of the transfer function may not be strict due to the lack of definition of the onlyTransferable modifier, increasing the risk of the smart contract being maliciously exploited. These vulnerabilities together lead to the breakdown of the access control mechanism of the smart contract, allowing unauthorized users to perform critical operations such as freezing accounts and switching liquidity status.
[0023] Table 1 Access Control Violations
[0024] Regarding the lack of access control, as shown in the example in Table 2. The self-destruction instruction on line 3 in the code shown in Table 2 will delete the contract from the blockchain and send any amount in the smart contract balance to the address provided as a parameter to the instruction (msg.sender). Therefore, only authorized users (such as the owner) should be allowed to call the remContract function. In addition, the parameter of the self-destruction instruction should not be set by unauthorized users.
[0025] Table 2 Lack of Access Control
[0026] For the expected normal operation, as shown in the example of Table 3. In certain scenarios, smart contracts may be deliberately designed by developers to allow users to modify state variables related to access control under certain conditions. These situations need to be distinguished from genuine exploitable vulnerabilities to identify the true expected behavior. We define such situations as expected normal operations. Taking the code snippet in Table 3 as an example, the changeNamesymb function in it has an access control check on line 4 to ensure that only the owner of the contract (the address of the owner is stored in the owner state variable) can call this function. At the same time, the contract also includes a changeOwner function that allows users to purchase the ownership of the contract. Most existing tools will misreport the changeOwner function as a vulnerability when looking for untrusted writes to access control data because this function allows anyone to change the owner state variable. But in fact, this is a preset functional behavior: the contract designer deliberately allows users to transfer ownership by paying a certain amount. When the corresponding funds are transferred to the account of the current owner, the state variable of the owner will be updated to the address of the purchaser.
[0027]
[0028] Table 3 Expected normal operations In step S1, perform control flow analysis on the EVM bytecode to obtain the control flow graph, and obtain several key operation instructions in the control flow graph. The several key operation instructions include CALL, CALLCODE, DELEGATECALL, and SELFDESTRUCT instructions.
[0029] The main functions of the CALL instruction and the CALLCODE instruction are to be able to accurately send a specific amount of money to the target contract or the specified address. The role of the SELFDESTRUCT instruction is to be able to permanently destroy the smart contract. However, if the parameters of this instruction can be manipulated by an attacker, then a very dangerous situation will occur, and the smart contract will send its own balance to the attacker, resulting in illegal transfer of assets. The DELEGATECALL instruction has a unique operation mechanism. It will replace the code of the current smart contract with the code of the called smart contract. During this process, the called smart contract can successfully access and operate the storage space of the current smart contract, and is completely consistent with the current smart contract in terms of permissions and access control. This feature also provides an opportunity for attackers. Attackers may use the DELEGATECALL instruction to implant malicious code into the smart contract, thus posing a serious threat to the security of the smart contract.
[0030] Then, extract the key parameters of each key operation instruction. The key parameters include: target address, the amount of currency sent, and input data. Further, perform backward slicing on the key parameters on the control flow graph to obtain a slice containing the instructions directly or indirectly controlled by the attacker, and construct a slice set. Among them, the instructions directly controlled by the attacker include: ORIGIN, CALLER, CALLVALUE, CALLDATALOAD, CALLDATASIZE, CALLDATACOPY, and the instructions indirectly controlled by the attacker include: SLOAD, MLOAD.
[0031] In addition, the SSTORE instruction is responsible for storing a given key-value pair into the persistent state of the smart contract, while the corresponding SLOAD instruction can load the value of a specified key from the state storage of the smart contract. However, there are also security risks in this process. If the transfer amount of the smart contract is read through the SLOAD instruction and can be modified by the SSTORE instruction, then the attacker may interfere with or even control the transfer behavior of the smart contract by deliberately modifying the state variable, thereby disrupting the normal operation logic and fund transfer order of the smart contract.
[0032] S2. Perform access control identification on the EVM bytecode based on the defined access control conditions: Perform backward data flow analysis on each conditional statement in the EVM bytecode, and identify the conditional statements containing the instructions in the slice set as access control checks. Construct a condition set based on all access control checks; Identify the storage location accessed by the access control check as the access control state variable.
[0033] In step S2, first perform access control identification on the EVM bytecode based on the defined access control conditions. The defined access control conditions are expressed as follows: Define access control check: If there is a conditional statement C in the EVM bytecode such that the caller A passes the verification of the conditional statement C, then the caller A can execute the code controlled by the conditional statement C, and this conditional statement C is an access control check; Define access control state variable: If there is an access control check C in the EVM bytecode, and the value V in the stored data item D can be used to compare with the caller A, or the value V loaded from the stored data item D is used to calculate the access control check C, then the stored data item D is an access control state variable, and the access control state variable stores the key information for access control decision-making; Define authorized user: If in the EVM bytecode, the access control state variable D contains the address of the caller A, then the caller A is an authorized user protected by the defined access control check C.
[0034] As can be seen from the above definitions, all access control checks have a core function, that is, to verify the caller of the smart contract based on the access control status variables representing various authorized users. In a smart contract, the built-in global variable msg.sender can be used to obtain the contract caller in the source code, and in the bytecode, it is achieved through the caller-related mechanism.
[0035] Table 4 Access Control Identification Algorithm
[0036] Algorithm 1 in Table 4 shows the identification process of access control checks. This algorithm takes the conditional statements in the EVM bytecode as input and outputs the access control checks and the corresponding access control status variables. For each conditional statement in the smart contract, backward data flow analysis is performed on it. If the part of the data flow SLOAD instruction and CALLER instruction is found, the conditional statement containing this instruction is marked as an access control check, and the storage location corresponding to the access (read by the SLOAD instruction) is marked as the access control status variable.
[0037] Based on all access control checks, a set of conditions is constructed. The set of conditions may contain non-access control checks that behave similarly to access control checks, such as require(balances[msg.sender]>= amount). To solve this problem, the conditional statements returned by the backward data flow analysis need to be further analyzed to exclude non-access control checks that do not meet the judgment conditions of access control checks from the set of conditions.
[0038] The judgment condition for access control checks is: if there is no data mapping relationship in the conditional statement (lines 6 - 7), or if there is a data mapping relationship in the conditional statement and the state variable used stores a boolean value (lines 8 - 9), then this condition is an access control check. Among them, state variables include: account balance, user permissions.
[0039] S3. Set the user input containing the execution operation as the taint source, and set the address parameter of the critical operation instruction and the SSTORE instruction for writing to the access control status variable as the taint sinks. Perform taint analysis on the control flow graph based on the taint source and taint sinks, and mark the contaminated operands as tainted.
[0040] Taint analysis determines whether data can propagate from a taint source to a taint sink by studying the data dependencies between program variables without executing or modifying the code. A taint source refers to a source that may contain malicious data, while a taint sink is a critical location in the program. If tainted data can reach these taint sinks, it may lead to security risks.
[0041] When combining taint analysis to detect access control vulnerabilities, first set the pollution source and taint sinks. Set the user input containing the execution operation as the taint source, which means that the execution operation entering the program from the user input is marked as potentially tainted from the beginning, so its propagation during program execution needs to be tracked. Set the address parameter of the critical operation instruction and the SSTORE instruction that writes to the access control status variable as taint sinks. In this embodiment, in addition to the SSTORE instruction, the address parameters of the SELFDESTRUCT and DELEGATECALL instructions are also regarded as taint sinks. This setting is because if the critical locations where these instructions are located are affected by data from the pollution source, it may lead to problems in access control, such as access control violations and lack of access control, thus meeting the goal of detecting access control vulnerabilities.
[0042] Table 5 Taint Analysis Algorithm
[0043] Algorithm 2 in Table 5 shows the process of using taint analysis to detect access control vulnerabilities. Based on the taint source and taint sinks, perform taint analysis on the control flow graph. Perform backward slicing on each potential taint sink (sinks) on the control flow graph to obtain all path segments that can reach the taint sink. Traverse each path segment obtained through backward slicing, and these path segments may contain taint propagation (lines 1 - 2). Then, starting from the taint source, along the execution order of the path segment, use the TaintAnalyse function to mark the contaminated operands as tainted.
[0044] During the taint analysis process, it also includes starting from each taint source, traversing the control flow graph to find all paths that flow to the SSTORE instructions that write to the storage slot. If the storage slot is tainted and its taint source is not the storage slot itself or the taint source has been marked as tainted, then mark the storage slot as tainted, and at the same time record the address of the SSTORE instruction that writes to this storage slot (lines 9 - 11).
[0045] S4. Perform symbolic execution on the taint flow path where the taint flowing to the access control status variable is located, generate the constraint conditions of the taint flow path, solve the constraint conditions. If a value that makes the constraint conditions hold can be found, then take the negation of the constraint conditions to obtain the negative constraint conditions; Further solve the negative constraint conditions. If a value that makes the negative constraint conditions hold can be found, mark the operand corresponding to the taint as the expected normal behavior. If not, mark the operand corresponding to the taint as an access control vulnerability to obtain the vulnerability detection result.
[0046] In the analysis of access control vulnerabilities in smart contracts, a simple lack of access control does not necessarily equate to a vulnerability. This is because developers may have adopted other non-access control-based checking mechanisms. Specifically, smart contract code may intentionally permit unauthorized users to operate on access control state variables under specific non-access control condition constraints. Since there is usually no clear access control policy, it greatly increases the difficulty of identifying real vulnerabilities. For example, in a specific scenario, an unauthorized user can write to the access control state variable owner under the condition of paying no less than a specific amount, and traditional static analysis methods are very likely to misjudge this as a vulnerability.
[0047] To effectively address this issue, the present invention adopts a verification method based on symbolic execution, aiming to infer the non-access control checks implemented in these specific situations and mark the operands that meet such situations as the expected normal behavior. With the view that statements that operate on access control state variables under specific non-access control condition constraints without being restricted by traditional access protection mechanisms are likely to be the intentional design of developers.
[0048] In smart contracts, the situation where access control permission checks are not imposed but access control state variables are operated based on other non-access control constraints should generally be regarded as an intentional behavior, that is, the expected normal behavior, rather than being regarded as an access control vulnerability. A relatively simple method to filter the expected normal behavior is to mark any unprotected statement that updates access control data and is protected by non-access control checks as the expected normal behavior. However, if this method is blindly used for filtering, some actual vulnerabilities may be misjudged as the expected normal behavior. For example, when a function initializes the access control state variable operator, the developer fails to correctly set initialized to true in the subsequent code, thus triggering a vulnerability. But if only the implemented access control checks are used to determine the intentional behavior, this vulnerability will be wrongly identified as the expected normal behavior.
[0049] In view of this, the present invention uses symbolic execution technology to infer the constraint conditions followed when operating on access control state variables, and thereby accurately distinguish the expected normal behavior from the actual vulnerabilities. The reason for choosing the symbolic execution method is that the constraints implemented in the expected normal behavior are relatively few and highly tractable, and can be solved by the Z3 solver.
[0050] Specifically, when it is detected that the tainted data flows to the access control status variable, symbolic execution operations are performed on the tainted data flow path where the tainted data is located, thereby generating the constraint condition for the tainted data flowing from the tainted data source to the access control status variable. Subsequently, the Z3 solver is used to verify the feasibility of this tainted data flow path and solve the constraint condition. If the path is feasible, that is, the Z3 solver can successfully find the values that make the constraint condition hold, then the constraint condition is negated to obtain the negative constraint condition.
[0051] Furthermore, the Z3 solver is used to solve the negative constraint condition. If values that make the negative constraint condition hold can be found, the operand corresponding to the tainted data is marked as an expected normal behavior. If not, the operand corresponding to the tainted data is marked as an access control vulnerability to obtain the vulnerability detection result.
[0052] Taking a specific example, when it is found that owner is tainted at a certain line of code, symbolic execution operations are immediately performed on the tainted data flow path to generate the constraint condition for owner being tainted. After solving the constraint condition, the synthesized path constraint is!msg.value≥BecomeOwner = 2000. Since the path is feasible, the Z3 solver will find the assignment of msg.value that can satisfy the negative constraint to prove that the original constraint condition does not always hold. If values that make the negative constraint condition hold can be found, the operand corresponding to the tainted data is marked as an expected normal behavior. The marking of the expected normal behavior is because by implementing specific constraints to limit the update operations on the owner status variable, it is very likely that the developer deliberately did so to manage the write operations to the owner.
[0053] When a solution that makes the negative constraint condition hold cannot be successfully found, if the constraint condition depends on the storage variables updated in the tainted data flow path, the Z3 solver will be called again. This step aims to infer the intended behavior that is usually used to initialize the storage variables only once, and then specific storage variables are set to prevent future repeated initialization operations.
[0054] By means of the method combining symbolic execution and constraint inference, the present invention can efficiently and accurately distinguish expected normal behaviors from actual existing access control vulnerabilities, thereby significantly improving the accuracy and reliability of the security analysis of smart contracts and providing strong support for the security guarantee of smart contracts.
[0055] The present invention also provides a smart contract vulnerability detection system based on taint analysis, including: Static analysis module: It is used to compile the smart contract into EVM bytecode, perform control flow analysis on the EVM bytecode to obtain a control flow graph, and acquire several key operation instructions in the control flow graph; extract the key parameters of each key operation instruction, perform backward slicing on the key parameters on the control flow graph to obtain a slice containing instructions directly or indirectly controlled by the attacker, and construct a slice set. Access control recognition module: It is used to perform access control recognition on the EVM bytecode based on the defined access control conditions: perform backward data flow analysis on each conditional statement in the EVM bytecode, identify the conditional statements containing the instructions in the slice set as access control checks, and construct a condition set based on all access control checks; identify the storage locations accessed by the access control checks as access control state variables. Taint analysis module: It is used to set the user input containing the execution operation as the taint source, and set the address parameter of the key operation instruction and the SSTORE instruction writing to the access control state variable as the taint sinks. Perform taint analysis on the control flow graph based on the taint source and taint sinks, and mark the contaminated operands as tainted. Vulnerability detection module: It is used to perform symbolic execution on the taint flow path where the taint flowing to the access control state variable is located, generate the constraint conditions of the taint flow path, solve the constraint conditions. If a value that makes the constraint conditions hold can be found, then take the negation of the constraint conditions to obtain the negative constraint conditions; further solve the negative constraint conditions. If a value that makes the negative constraint conditions hold can be found, then mark the operand corresponding to the taint as the expected normal behavior. If not, then mark the operand corresponding to the taint as an access control vulnerability to obtain the vulnerability detection result.
[0056] Experimental content The specific experimental environment settings are shown in Table 6.
[0057] Table 6 Experimental environment
[0058] Dataset: Two datasets are used in this invention to evaluate this method. The first dataset is the smart contracts with vulnerability markings obtained from the datasets in SmartBugs, which contains 20 smart contracts with access control vulnerabilities. The second dataset is 12,000 verified real smart contracts obtained from Etherscan.
[0059] Evaluation metrics: In the experiment, three evaluation metrics were collected: Precision, Recall, and F1-score. In the experiment, we obtained three key measurement values: True Positive (TP), False Positive (FP), and False Negative (FN). Among them, TP represents the successfully detected smart contract vulnerabilities, FP represents the smart contracts mislabeled as vulnerabilities, and FN represents the smart contract vulnerabilities that failed to be detected. Based on these three measurement values, the Precision (denoted as P), Recall (denoted as R), and F1-score (denoted as F) can be calculated as follows: , , .
[0060] Experimental Results and Analysis: The method of the present invention was compared with four other different detection methods, Mythril, SPCon, Slither, and Ethainter, on two datasets respectively. Mythril is a symbolic execution detection method, Slither is a method for detecting smart contract vulnerabilities using intermediate representation, and SPCon and Ethainter are static analysis methods. Through the first dataset, according to the True Positive (TP), False Positive (FP), and False Negative (FN) values, the Precision (P), Recall (R), and F1-score (F) for detecting integer overflow vulnerabilities were further calculated. The experimental results of the four detection methods are shown in Table 7. By calculating the Precision and Recall of each detection method, it can be concluded that the Precision, Recall, and F1-score of the method of the present invention are the highest.
[0061] Table 7 Comparison of Experimental Results on the First Dataset
[0062] It can be seen that Slither and Mythril exhibit relatively high false positives and false negatives. These methods do not conduct in-depth research on access control vulnerabilities across the entire contract scope caused by weak or missing access control. They mainly rely on predefined access control patterns, focusing on analyzing the lack of access control in specific code statements, and failing to comprehensively cover all relevant vulnerabilities that may exist in the contract. The false positives and false negatives of Ethainter in detecting vulnerabilities have decreased. However, our research found that the generalization ability of these predefined rules is weak, resulting in unsatisfactory performance of Ethainter when dealing with a large number of samples. In addition, the predefined patterns for access control checks lack sufficient precision and are prone to false positives. This means that in practical applications, Ethainter may misidentify normal behaviors in some contracts as vulnerabilities, thus affecting the reliability and effectiveness of its detection.
[0063] Due to the diversity and complexity of smart contracts, it is difficult for a single predefined rule to cover all possible code structures and logics, leading to challenges in actual vulnerability detection. The way SPCon mines permission errors in smart contracts is to analyze the contract transaction history on the blockchain to obtain access control roles, and it has certain limitations. It defaults that all functions of the analyzed smart contract have transaction records, but in reality, some contract functions are only called under specific circumstances, such as the function to delete a smart contract, and such functions may not generate transaction history records. Moreover, SPCon assumes that each transaction is benign and is executed by an authorized account according to the expected security policy, but this assumption does not hold for contracts with access control vulnerabilities that can be exploited by malicious users. These limitations make it difficult for SPCon to comprehensively cover potential access control problems when identifying permission errors. Especially when facing complex contracts, its effectiveness in practical applications is greatly reduced, and it cannot accurately and completely capture all possible access control vulnerabilities and related risks.
[0064] To further verify the vulnerability detection ability of the method of the present invention in real smart contracts, the method of the present invention was applied to 12,000 smart contracts obtained from Etherscan. 3,000 smart contracts were randomly selected from the second dataset for manual inspection, and the detection results were compared with those of Mythril, SPCon, Slither, and Ethainter. By manually analyzing the vulnerabilities reported in the detection reports of each detection method and classifying them as true positives (TP), false positives (FP), or false negatives (FN), the results are shown in Table 8.
[0065] Table 8 Comparison of experimental results in the second dataset
[0066] It can be seen that the method of the present invention can correctly detect more access control vulnerabilities in real-world smart contracts, and the number of false positives and false negatives reported is relatively small.
[0067] In summary, the method and system for detecting smart contract vulnerabilities based on taint analysis used in the present invention, through control flow analysis of the EVM bytecode of smart contracts, screen key information, and combine the taint analysis method to transform the detection of access control vulnerabilities into a pollution analysis problem, without relying on predefined access control patterns or existing transaction records, can significantly improve the accuracy and reliability of smart contract security analysis, and provide strong support for the security guarantee of smart contracts. The present invention generates constraint conditions for taint flow paths through symbolic execution, efficiently and accurately distinguishes the expected normal operations in smart contracts from real security vulnerabilities, takes the negation of the constraint conditions to obtain negative constraint conditions, and further solves the negative constraint conditions to distinguish whether unauthorized users can obtain the access control permissions of the contract with the expected functional logic under specific circumstances, distinguishes potential expected normal behaviors from clear access control vulnerabilities, and reduces the false positive rate of vulnerability detection.
Claims
1. A smart contract vulnerability detection method based on taint analysis, characterized in that: The following steps are involved: S1. Compile the smart contract into EVM bytecode, perform control flow analysis on the EVM bytecode, obtain a control flow graph, and obtain several key operation instructions in the control flow graph; Extract the key parameters of each key operation instruction, slice the key parameters backward on the control flow graph, obtain the slices containing the instructions directly or indirectly controlled by the attacker, and build a slice set; S2. Perform access control identification on the EVM bytecode based on the defined access control conditions: Perform backward data flow analysis on each conditional statement in the EVM bytecode, identify the conditional statements containing instructions in the slice set as access control checks, and build a condition set based on all access control checks; Identify the storage location accessed by the access control check as an access control state variable; S3, set the user input containing the execution operation as the taint source, set the address parameter of the key operation instruction and the SSTORE instruction of writing the access permission control state variable as the taint sink, perform taint analysis on the control flow graph based on the taint source and the taint sink, and mark the contaminated operand as taint; S4. Perform symbolic execution on the taint flow path where the taint flowing to the access permission control state variable is located, generate constraints on the taint flow path, solve the constraints, and if a value that makes the constraints true can be found, negate the constraints to obtain negated constraints; The negated constraint is further solved. If a value that makes the negated constraint true can be found, the operand corresponding to the taint is marked as expected normal behavior. If not found, the operand corresponding to the taint is marked as an access permission control vulnerability to obtain the vulnerability detection result.
2. The method for detecting smart contract vulnerabilities based on taint analysis according to claim 1, characterized in that: The access permission control conditions defined in S2 are specifically: Define access control check: If there is a conditional statement C in the EVM bytecode, and caller A passes the verification of conditional statement C, then caller A can execute the code controlled by conditional statement C. This conditional statement C is an access control check; Define access control state variables: If there is an access control check C in the EVM bytecode, the access control check C can be calculated using the value V in the storage data item D and the caller A, or the value V loaded from the storage data item D, then the storage data item D is the access control state variable; Defined authorized user: If in the EVM bytecode, the access control state variable D contains the address of caller A, then caller A is an authorized user protected by the defined access control check C.
3. The method for detecting smart contract vulnerabilities based on taint analysis according to claim 1, characterized in that: The step S2 also includes excluding non-access right control checks that do not meet the access right control check judgment condition from the condition set, and the access right control check judgment condition is: If there is no data mapping relationship in the conditional statement, or if there is a data mapping relationship in the conditional statement and the state variable used by it stores a Boolean value, the condition is an access control check.
4. The method for detecting smart contract vulnerabilities based on taint analysis according to claim 1, characterized in that: The key operation instructions in S1 include CALL, CALLCODE, DELEGATECALL and SELFDESTRUCT instructions.
5. The method for detecting smart contract vulnerabilities based on taint analysis according to claim 1, characterized in that: The S1 obtains a slice containing instructions directly or indirectly controlled by the attacker. The instructions directly controlled by the attacker include: ORIGIN, CALLER, CALLVALUE, CALLDATALOAD, CALLDATASIZE, CALLDATACOPY, and the instructions indirectly controlled by the attacker include: SLOAD, MLOAD.
6. The method for detecting smart contract vulnerabilities based on taint analysis according to claim 1, characterized in that: The key parameters in S1 include: target address, amount of currency sent, and input data.
7. The method for detecting smart contract vulnerabilities based on taint analysis according to claim 1, characterized in that: In S3, taint analysis is performed on the control flow graph based on taint sources and taint sinks, and the polluted operands are marked as tainted, specifically: Backward slicing of the taint sink is performed on the control flow graph to obtain all path fragments that can reach the taint sink. Then, starting from the taint source, along the execution order of the path fragments, the tainted operands are marked as tainted using the TaintAnalyse function.
8. The method for detecting smart contract vulnerabilities based on taint analysis according to claim 7 is characterized in that: The taint analysis process also includes starting from each taint source, traversing the control flow graph to find all paths of the SSTORE instructions that flow to the storage slot. If the storage slot is contaminated by the taint and its taint source is not the storage slot itself or the taint source has been marked as tainted, then the storage slot is marked as tainted and the address of the SSTORE instruction written to the storage slot is recorded.
9. The method for detecting smart contract vulnerabilities based on taint analysis according to claim 1, characterized in that: In S4, the constraints and negated constraints are solved using the Z3 solver.
10. A smart contract vulnerability detection system based on taint analysis, characterized in that: include: Static analysis module: used to compile smart contracts into EVM bytecode, perform control flow analysis on the EVM bytecode, obtain a control flow graph, and obtain several key operation instructions in the control flow graph; extract key parameters of each key operation instruction, perform backward slicing on the key parameters on the control flow graph, obtain slices containing instructions directly or indirectly controlled by the attacker, and construct a slice set; Access control identification module: used to perform access control identification on EVM bytecode based on defined access control conditions: perform backward data flow analysis on each conditional statement in the EVM bytecode, identify the conditional statement containing instructions in the slice set as an access control check, and build a condition set based on all access control checks; Identify the storage location accessed by the access control check as an access control state variable; Taint analysis module: used to set the user input containing the execution operation as the taint source, set the address parameter of the key operation instruction and the SSTORE instruction of the write access control state variable as the taint sink, perform taint analysis on the control flow graph based on the taint source and taint sink, and mark the polluted operand as taint; Vulnerability detection module: used to perform symbolic execution on the taint flow path where the taint flows to the access permission control state variable, generate constraints on the taint flow path, solve the constraints, and if a value that makes the constraint true can be found, the constraint is negated to obtain a negated constraint; the negated constraint is further solved, and if a value that makes the negated constraint true can be found, the operand corresponding to the taint is marked as expected normal behavior, and if it cannot be found, the operand corresponding to the taint is marked as an access permission control vulnerability to obtain a vulnerability detection result.
Citation Information
Patent Citations
Vulnerability detection method for binary code of intelligent contract
CN113051574A
Intelligent contract vulnerability detection method based on symbolic execution
CN116361810A
Intelligent contract vulnerability detection method and system based on hybrid fuzzy test
CN117828616A
Intelligent contract vulnerability detection method and device, electronic equipment and storage medium
CN117909992A
Intelligent contract upgrading vulnerability detection method
CN117951711A
Cited By
Contract attack detection method and device based on symbolic execution and graph neural network
CN122197005A