A smart contract vulnerability detection method and system based on a heterogeneous graph and a large model

CN122471465BActive Publication Date: 2026-09-29HEFEI UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610943120.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-29
Publication Date
2026-09-29
Estimated Expiration
2046-06-29

AI Technical Summary

Technical Problem

[0006]为了克服现有细粒度智能合约漏洞检测技术检测精度不足、难以同时实现函数级分类与语句级精确定位的问题,本发明提出了一种基于异构图和大模型的智能合约漏洞检测方法,能够有效融合代码语义与拓扑结构、实现细粒度漏洞检测

Benefits of technology

(1)本发明提出的一种基于异构图和大模型的智能合约漏洞检测方法,将智能合约代码转换为异构图结构,针对所述异构图中的节点,构建动静结合的特征提取机制,将动态语义特征与静态结构特征进行融合,得到节点的初始特征表示。然后将初始特征输入异构图Transformer进行特征聚合,异构图Transformer基于元关系注意力机制,对不同类型节点及边进行加权建模,以实现跨节点类型及跨层级的上下文信息融合,得到节点的最终表征。多任务预测模块基于节点的最终表征执行任务,主任务对语句节点进行分类,以实现漏洞语句的精确定位;辅助任务对函数节点进行分类,以实现漏洞类型识别,通过所述辅助任务对主任务提供约束信息,以提升语句级漏洞检测精度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122471465B_ABST
    Figure CN122471465B_ABST
Patent Text Reader

Abstract

The present application relates to the cross technical field of computer security and software engineering, and particularly relates to a smart contract vulnerability detection method and system based on a heterogeneous graph and a large model. After converting the code text to be detected into a heterogeneous graph, the initial feature representation of each node is extracted; the meta relationship on the heterogeneous graph is used to fuse the attention weight and the message vector, so as to obtain the final representation of each target node; and whether the function has a vulnerability, the vulnerability category and the positioning are confirmed based on the final representation. The present application can effectively fuse the code semantics and the topological structure, realize fine-grained vulnerability detection, and overcome the problems that the existing fine-grained smart contract vulnerability detection technology has insufficient detection precision and is difficult to simultaneously realize function-level classification and statement-level accurate positioning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the interdisciplinary field of computer security and software engineering, and in particular to a method and system for detecting smart contract vulnerabilities based on heterogeneous graphs and large models. Background Technology

[0002] As a core component of decentralized blockchain applications, the security of smart contracts is paramount. However, the immutability of blockchain means that once a contract is deployed, any vulnerabilities it contains will be permanently retained, potentially leading to significant asset losses. Therefore, developing efficient vulnerability detection technologies has become a key focus of industry research.

[0003] Early smart contract vulnerability detection relied primarily on traditional methods such as static analysis and symbolic execution (e.g., tools like Slither and Oyente). These methods match vulnerabilities based on pre-defined expert rules and patterns, resulting in weak generalization capabilities, difficulty in effectively identifying novel and unknown vulnerabilities, and a tendency to encounter path explosion when analyzing complex contract logic, leading to high false positive and false negative rates.

[0004] To break free from reliance on manually generated rules, deep learning-based detection methods have gradually emerged. Based on their technical principles, these methods can be mainly divided into three categories: sequence-based methods, graph neural network (GNN)-based methods, and large language model (LLM)-based methods. Sequence-based methods (such as Bi-LSTM and Transformer) treat code as a one-dimensional sequence, but lose the inherent structural information of the code, such as control flow and data dependencies, resulting in insufficient understanding of the deep logic of the program and coarse-grained detection. Graph neural network-based methods (such as various GNN models) can better preserve code structure, but the graph topological features they rely on lack a deep understanding of the natural language semantics of the code, leading to information bottlenecks when handling multi-layered nested relationships of "contract-function-line of code." Large language model-based methods, while possessing powerful code understanding and generation capabilities, typically treat code as long text, exhibiting weak perception of the code graph topology, facing limitations in long context processing and high computational costs, and potentially producing logical "illusions."

[0005] In summary, while existing methods have made some progress, they still have significant limitations. On the one hand, most methods can only achieve coarse-grained vulnerability detection at the contract or function level, failing to pinpoint specific lines of vulnerable code, thus posing difficulties for manual auditing and remediation. On the other hand, even though some research has begun to attempt fine-grained line-level localization or type identification, they still struggle to simultaneously and efficiently consider both the deep semantic information and complex topological structure information of the code in terms of feature representation and context modeling. This leads to existing models easily losing crucial contextual dependencies when dealing with multi-level nested structures of smart contracts, failing to meet the urgent need for fine-grained, automated vulnerability detection that simultaneously achieves vulnerability type identification and precise line-of-failure localization. Summary of the Invention

[0006] To overcome the shortcomings of existing fine-grained smart contract vulnerability detection technologies, such as insufficient detection accuracy and difficulty in simultaneously achieving function-level classification and statement-level precise location, this invention proposes a smart contract vulnerability detection method based on heterogeneous graphs and large models, which can effectively integrate code semantics and topological structure to achieve fine-grained vulnerability detection.

[0007] This invention proposes a smart contract vulnerability detection method based on heterogeneous graphs and large models. First, a vulnerability detection model is trained, and then the model is used to perform vulnerability detection on the code text to be tested. The vulnerability detection model includes: a preprocessing module, a dual-stream feature extraction module, a feature aggregation module, and a multi-task prediction module; The preprocessing module transforms the smart contract source code into a heterogeneous graph containing function nodes and statement nodes; The dual-stream feature extraction module uses a pre-trained code model to encode the code corresponding to the nodes in the heterogeneous graph to obtain the dynamic semantic features of each node. Then, based on the dynamic semantic features, high-risk statement nodes are selected, and expert knowledge features are injected into the high-risk statement nodes. The dynamic semantic features and expert knowledge features are fused to obtain the initial feature representation of each node. The feature aggregation module performs attention weight modeling based on the meta-relations on the heterogeneous graph, and performs multi-layer fusion on the weighted meta-relations message transmission and the initial feature representation of each node to obtain the final representation of cross-node type and hierarchical fusion context information. The multi-task prediction module determines whether a vulnerability exists based on the final representation of the statement node, and determines the vulnerability type of the vulnerability line based on the final representation of the function node.

[0008] Preferably, the heterogeneous graph defines multiple meta-relations, including control flow edges, data flow edges, function call edges, and bidirectional containment edges; control flow edges Used to represent the execution order and conditional branching relationships between statements; data flow edges Used to indicate the definition and usage relationship of variables across different statements; call edge Used to represent calling relationships between functions; includes edges Used to represent the hierarchical relationship between function nodes and statement nodes, and includes bidirectional edges.

[0009] The preferred method for obtaining expert knowledge features for high-risk statement node injection is as follows: Extract the graph context information of high-risk statement nodes in the heterogeneous graph and construct prompt words. Then, the vulnerability reasoning text is obtained by processing the clue words through the frozen parameter model. Then, the vulnerability reasoning text Mapped to structured static expert knowledge features .

[0010] Preferably, the graph context information of high-risk statement nodes in a heterogeneous graph includes: Set of predecessor node sequences in the control flow graph This indicates that by controlling the flow edge The set of source nodes pointing to the high-risk statement node; Set of successor node sequences in the control flow graph This indicates that the high-risk statement node is used. As a control flow edge The set of target nodes of the source node; Data flow graph predecessor node sequence set , indicating through data flow edge Pointing to high-risk statement nodes The set of source nodes; Data flow graph successor node sequence set This indicates that the high-risk statement node is used. As a data flow edge The set of target nodes of the source node.

[0011] Preferably, the initial feature representation of each node is obtained in the following way: For high-risk statement nodes, their dynamic semantic features and expert knowledge features are dimensionally concatenated to obtain an initial feature representation; for the remaining nodes, their dynamic semantic features are mapped to the initial feature representation space to obtain the node's initial feature representation.

[0012] Preferably, the feature aggregation module operates as follows: First, attention weights are modeled using meta-relations to obtain the attention weights of meta-relations corresponding to each unidirectional edge in the heterogeneous graph; then, the messages of each meta-relation are calculated. Then, for the meta-relations where a node exists as a target node, the characteristics of the target node are calculated using the following formula: ; in, and These represent the updated feature representations of target node t at layers r and (r-1), respectively. This represents the set of source nodes that are neighbors of the target node t. The output mapping matrix represents the type of target node; GELU represents the nonlinear activation function. Heterogeneous attention weights for meta-relations; This represents the message vector passed from source node s to target node t; Traverse r=1, 2...L to obtain features As the final feature of the target node t; Let be the initial feature representation of the target node t.

[0013] Preferably, the feature aggregation module uses an HGT network.

[0014] Preferably, during the training process, the vulnerability detection model uses code with vulnerability annotations as the training dataset; the loss function used for training is a weighted sum of the fine-grained vulnerability location loss and the function-level vulnerability type classification loss. The loss for fine-grained vulnerability localization is: ; in, For the set of statement nodes in the heterogeneous graph of the training samples, | |for The number of nodes in; For statement nodes Does the line of code contain a genuine vulnerability label? This indicates the statement node output by the model. is the predicted probability of the vulnerability row; log represents the logarithmic function with the natural constant as the base. The loss for classifying function-level vulnerability types is: ; in, For the set of function nodes of the training samples, for The number of nodes in the table; C represents the total number of predefined vulnerability categories; function node The one-hot encoding corresponding to the true category c, Represents a multi-class probability distribution. express The probability value of category c.

[0015] Preferably, the steps are as follows: S1. Obtain the text of the code to be detected and input it into the vulnerability detection model; S2. The preprocessing model transforms the code text to be detected into a heterogeneous graph G=(V,E); S3. The dual-stream feature extraction module processes the heterogeneous graph G=(V,E) to extract the initial feature representation of each node. S4. The feature aggregation module fuses attention weights and message vectors based on meta-relations on the heterogeneous graph to obtain the final representation of each target node. ; S5. The function node classification header determines whether a function has vulnerabilities based on the final representation of the function node in each line of code. If not, then the function is considered correct and the code has no bugs. If yes, then determine the type of function vulnerability and proceed to step S6; S6. The statement node classification header is based on the final representation of the statement node in the code line. It judges the vulnerability probability of each statement node one by one and outputs the statement nodes that exceed the threshold as the vulnerability location.

[0016] The present invention proposes a smart contract vulnerability detection system, comprising a memory and a processor. The memory stores a computer program, and the processor is connected to the memory. The processor is used to execute the computer program to implement the smart contract vulnerability detection method based on heterogeneous graphs and large models.

[0017] The advantages of this invention are: (1) The present invention proposes a smart contract vulnerability detection method based on heterogeneous graphs and large models. The smart contract code is converted into a heterogeneous graph structure. For the nodes in the heterogeneous graph, a dynamic and static feature extraction mechanism is constructed to fuse dynamic semantic features with static structural features to obtain the initial feature representation of the node. Then, the initial features are input into the heterogeneous graph Transformer for feature aggregation. The heterogeneous graph Transformer is based on the meta-relation attention mechanism to perform weighted modeling of different types of nodes and edges to achieve cross-node type and cross-level context information fusion to obtain the final representation of the node. The multi-task prediction module executes tasks based on the final representation of the node. The main task classifies statement nodes to achieve accurate location of vulnerable statements; the auxiliary task classifies function nodes to achieve vulnerability type identification. The auxiliary task provides constraint information to the main task to improve the accuracy of statement-level vulnerability detection.

[0018] (2) In constructing the heterogeneous graph, this invention performs detailed classification of nodes and edges, making the heterogeneous graph information complete and rich, providing a rich information foundation for subsequent processing. Smart contract vulnerabilities are essentially manifested as abnormal interactions between control flow and data flow. Traditional homogeneous graph neural networks cannot distinguish between different types of nodes and edges during message transmission, easily leading to the smoothing or loss of structural information. The heterogeneous graph of this invention, through detailed classification of nodes and edges, ensures the accuracy of message transmission and effectively prevents the loss of structural information.

[0019] (3) The feature aggregation module enables each node in the graph to obtain a high-dimensional final representation that integrates dynamic semantics, static expert logic and global topological context, which can effectively integrate code semantics and topological structure.

[0020] (4) In the inference stage of the actual application of the model, a hierarchical detection strategy is adopted. First, based on the classification results of function nodes, it is determined whether the target function has vulnerabilities. Only for functions that are determined to have vulnerabilities, statement-level vulnerability localization is performed, thereby reducing the overall false alarm rate. During the model training process, the hyperparameters that balance the importance of the two tasks are jointly optimized, which can enable function-level vulnerability type identification to constrain and guide line-level vulnerability localization, thereby improving the overall detection effect and facilitating the simultaneous completion of vulnerability type identification and accurate localization of vulnerability code lines.

[0021] (5) As can be seen, this invention proposes a fine-grained vulnerability detection scheme that integrates heterogeneous graph representation learning and large language model (LLM) reasoning capabilities to address the security requirements of smart contracts in blockchain platforms such as Ethereum. This scheme aims to simultaneously achieve function-level vulnerability type identification and line-of-code-level vulnerability precise location, and belongs to the cutting-edge direction of smart contract security auditing, static program analysis, and artificial intelligence in software security applications. Attached Figure Description

[0022] Figure 1 A flowchart illustrating the construction process of the source code representation of a graph-based smart contract; Figure 2 This is a flowchart illustrating the workflow of the feature aggregation module in the vulnerability detection model. Figure 3 This is a flowchart of a smart contract vulnerability detection method based on heterogeneous graphs and large models proposed in this invention. Detailed Implementation

[0023] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0024] This implementation provides a smart contract vulnerability detection method based on heterogeneous graphs and large models, which is applied to the automated security analysis of blockchain smart contracts. It can identify vulnerability types at the function level and accurately locate specific vulnerable lines of code, thereby providing technical support for the security audit and vulnerability remediation of smart contracts.

[0025] This invention specifically proposes a vulnerability detection model that performs vulnerability screening based on the source code representation of smart contracts with a graph structure, and outputs function-level vulnerability type identification results and code line-level vulnerability location results.

[0026] Reference Figure 1 The process of constructing a graph-based smart contract source code representation (i.e., a heterogeneous graph) includes the following steps: Step S11: Perform abstract syntax tree parsing on the smart contract source code and extract the set of function nodes. and statement node set ; Step S12: Construct a set of control flow edges inside the function. And construct a data flow edge set based on the relationship between variable definition and usage. ; Step S13: Analyze the call relationships between functions and construct a set of call edges. ; Step S14: Establish a set of function nodes With statement node set bidirectional containment edge set ; Step S15: Integrate the above nodes and edges to obtain a heterogeneous graph G=(V,E) for subsequent feature extraction and vulnerability detection.

[0027] In a heterogeneous graph G=(V,E), V represents the set of nodes, including function nodes. and statement nodes Function nodes This represents a function entity or statement node in a smart contract. E represents the specific line of code or basic unit of execution within a function; E represents the set of edges, including control flow edges. Data flow edge Calling the edge and containing edges Control flow edge Used to represent the execution order and conditional branching relationships between statements; data flow edges Used to indicate the definition and usage relationship of variables across different statements; call edge Used to represent calling relationships between functions; includes edges Used to represent the hierarchical relationship between function nodes and statement nodes, including bidirectional edges.

[0028] Heterogeneous graphs can preserve the execution logic and hierarchical structure information of code, thus realizing the graph structure representation of code.

[0029] The vulnerability detection model includes: a preprocessing module, a dual-stream feature extraction module, a feature aggregation module, and a multi-task prediction module.

[0030] The preprocessing module is used to convert the code to be detected into a heterogeneous graph G=(V,E).

[0031] The dual-stream feature extraction module extracts dynamic semantic features and static expert knowledge features from the heterogeneous graph G=(V,E), and then fuses the two types of features to form an initial feature representation. The initial feature representation can simultaneously represent the basic semantic information of the code itself and the deep reasoning information for vulnerability analysis.

[0032] The dual-stream feature extraction module consists of two parts: a dynamic semantic fine-tuning stream and a static expert knowledge injection stream.

[0033] Dynamic semantic fine-tuning streams are used to process nodes in heterogeneous graphs and obtain the dynamic semantic features of each node. It is used to describe the syntactic information and contextual semantic information of the code.

[0034] Dynamic semantic fine-tuning flow can specifically employ the Solidity-T5 network.

[0035] For nodes The dynamic semantic fine-tuning stream first extracts its corresponding plain text code snippets. Then, the code snippet is processed using the byte-pair encoder that comes with Solidity-T5. Perform word segmentation to obtain the input sequence. .in, This indicates the maximum length of the input sequence as uniformly set.

[0036] Input sequence After inputting the encoder of Solidity-T5, the hidden layer state matrix is ​​obtained. .in, This indicates the hidden layer dimension output by the Solidity-T5 encoder. Preferably, the hidden layer dimension can be fixed at 768.

[0037] In order to transform the variable-length token sequence (i.e., the hidden layer state matrix) The feature vectors are compressed into fixed-length node-level feature vectors. This invention employs an average pooling strategy.

[0038] Thus, the dynamic semantic fine-tuning flow applies to nodes. The processing procedure is expressed by the following formula: ; in, Represents the dynamic semantic features of node i. BPE stands for byte-pair encoder, i.e. ; This refers to the encoder of Solidity-T5, namely: ; This indicates the average pooling strategy.

[0039] In dynamic semantic fine-tuning flow, by averaging across the sequence dimensions, the semantic contributions of all tokens in the code sequence can be aggregated to obtain the dynamic semantic features of node i. .

[0040] After obtaining dynamic semantic features, the static expert knowledge injection stream performs high-risk screening of statement nodes based on preset heuristic rules, and injects expert knowledge features into the screened high-risk statement nodes. The expert knowledge features are essentially hints constructed based on the topological relationships of the high-risk statement nodes in the heterogeneous graph.

[0041] The static expert knowledge injection stream pre-stores heuristic rules and uses these rules to filter the set of high-risk statement nodes. Heuristic rules include, but are not limited to: statements containing arithmetic operations, statements containing external calls, statements containing time-related variables, and statements containing calls with unchecked return values. These four heuristic rules correspond to high-risk patterns such as arithmetic vulnerabilities, reentrancy vulnerabilities, timestamp dependency vulnerabilities, and unchecked send vulnerabilities in high-risk statement nodes.

[0042] For any high-risk statement node The static expert knowledge injection stream extracts graph context information from heterogeneous graphs and constructs cue words with strong context constraints. Then, the prompt words are processed through the frozen CodeLlama-Instruct model. Obtain the vulnerability reasoning text And map it into structured static expert knowledge features ; Integrate high-risk statement nodes Dynamic semantic features and static expert knowledge characteristics Obtain high-risk statement nodes Initial feature representation Dynamic semantic features Output from the dynamic semantic fine-tuning stream.

[0043] High-risk statement nodes The graph context information in heterogeneous graphs includes the set of predecessor node sequences in the control flow graph. Set of successor node sequences in the control flow graph Set of predecessor node sequences in a data flow graph and the set of successor node sequences in the data flow graph .

[0044] Specifically, the control flow edges on the heterogeneous graph G=(V,E) Data flow edge All are unidirectional edges; Set of predecessor node sequences in the control flow graph This refers to high-risk statement nodes. By controlling the flow edge Connect and control the flow edge Pointing to high-risk statement nodes The set of nodes; Set of successor node sequences in the control flow graph This refers to high-risk statement nodes. By controlling the flow edge Connect and control the flow edge From high-risk statement nodes The set of starting nodes; Data flow graph predecessor node sequence set This refers to high-risk statement nodes. via data flow edge Connection, and data flow edge Pointing to high-risk statement nodes The set of nodes; Data flow graph successor node sequence set This refers to high-risk statement nodes. via data flow edge Connection, and data flow edge From high-risk statement nodes The set of starting nodes.

[0045] Based on high-risk statement nodes Constructing prompts with strong contextual constraints using graph context information. Time: based on the line of code to be detected With high-risk statement nodes as the core, The control flow and data flow context form corresponding prompt words. The prompt words are used to guide the large language model to perform vulnerability analysis on the node.

[0046] Vulnerability reasoning text Since the text is unstructured, it needs to be converted into a computable vector representation. This invention uses the Sentence-BERT text embedding model to process the vulnerability inference text. The mapping is performed, and its calculation formula is as follows: ; in, , This represents the static expert knowledge feature vector obtained through the text embedding model; Sentence-BERT is an existing pre-trained language analysis model; CodeLlama-Instruct is an existing large programming language model.

[0047] This mapping process transforms the natural language inference results output by the large language model into structured features, which can then be used for joint modeling with dynamic semantic features.

[0048] Dynamic semantic features and static expert knowledge feature vector The fusion can be achieved by splicing, that is: ; in, This represents the fused node features, i.e., the initial feature representation of node s; Concat represents the vector concatenation operation. Through this fusion process, the node representation can simultaneously include code semantic information and vulnerability reasoning information.

[0049] Static expert knowledge feature vector This is unique to high-risk statement nodes. For nodes that have not triggered high-risk rules, zero vectors can be used for dimension padding to ensure that the feature dimensions of different nodes are consistent.

[0050] The feature aggregation module first performs attention weight modeling and message calculation on each meta-relation containing a unidirectional edge on the heterogeneous graph, and then merges the attention weights and message vectors to update the target node features in the meta-relation.

[0051] The feature aggregation module can specifically use an HGT network (Heterogeneous Graph Transformer) to achieve graph evolution and feature aggregation for heterogeneous graphs.

[0052] Reference Figure 2 The specific working steps of the feature aggregation module are as follows: SA1. Model the attention weights using meta-relations to obtain the attention weights of the meta-relations corresponding to each unidirectional edge in the heterogeneous graph.

[0053] For any directed edge in the heterogeneous graph G=(V,E), let be... Where s is the source node and t is the target node, the meta-relation of the directed edge e is defined as follows: ,in and These represent the node types of the source node s and the target node t, respectively. Indicates the edge type.

[0054] The relationship between the order and the order is Heterogeneous attention weights are denoted as , which represents the attention weight that the source node s acts on the target node t through edge e; ; in, For activation functions; Representing edge type The corresponding learnable attention weight matrix, where d represents the scaling dimension of the attention mechanism; and These represent the Query vector of the target node and the Key vector of the source node, respectively. ; ; in, This represents the feature representation of target node t at layer r-1. This represents the feature representation of the source node s at layer r-1. This represents the Query projection matrix corresponding to the target node type. This represents the Key projection matrix corresponding to the source node type; and Let r be the learnable matrix; 1 ≤ r ≤ L, where L is the number of layers in the HGT network.

[0055] and Output from the dual-stream feature extraction module. The dual-stream feature extraction module outputs dynamic semantic features for high-risk sentence nodes. and static expert knowledge feature vector fused node features The output for function nodes and the remaining (non-high-risk) statement nodes is determined by dynamic semantic features. Fill in 0 and Initial feature representation of alignment Let the initial feature representation output by the dual-stream feature extraction module for node i be denoted as... ,but: ; In meta-relations: = And node i is the source node; = And node i is the target node.

[0056] By modeling with attention weights, control flow edges can be distinguished. Calling the edge and data stream edge The different importance of features in propagation.

[0057] SA2. Calculate the message for each meta-relation. Based on the unidirectional edges in the heterogeneous graph, the source node s passes the attention weights and messages corresponding to the meta-relations to the target node t.

[0058] Meta-relation The corresponding message vector is denoted as : ; in, This represents the message vector passed from source node s to target node t. This represents the feature representation of the source node s at layer r-1. This represents the Value projection matrix corresponding to the source node type. Representing edge type The corresponding message transformation matrix. and All of these are learnable parameter matrices.

[0059] SA3. In each meta-relation, the target node t relates to its set of neighbor nodes. The received messages are weighted and aggregated, and their features are updated. The calculation formula is as follows: ; in, This represents the feature representation of target node t after updating at layer r. Denotes the set of source nodes that are neighbors of the target node t, i.e. The meta-relationships formed by the middle node and the target node t are all unidirectional edges; The output mapping matrix represents the type of target node, and it is a learnable parameter; GELU represents the nonlinear activation function.

[0060] Since every node appears at least once as the target node, traversing all meta-relations will update all nodes.

[0061] SA4, Traverse r = 1 - L to obtain features As the final feature of the target node t.

[0062] After passing through the dual-stream feature extraction module, each node in the heterogeneous graph G=(V,E) Initial feature representations that integrate dynamic semantics and static expert knowledge were obtained for all of them. Smart contract vulnerabilities essentially manifest as abnormal interactions between control flow and data flow. Traditional homogeneous graph neural networks, when passing messages, cannot distinguish between different types of nodes and edges, easily leading to the smoothing or loss of structural information. The feature aggregation module uses an HGT network for heterogeneous message passing and feature updates, enabling cross-level feature fusion between function nodes and statement nodes, as well as between different statement nodes. After L layers of HGT propagation, all nodes in the graph can capture the structural semantic information of their L-order neighborhood. After completing HGT feature aggregation, each node in the graph obtains a high-dimensional final representation that integrates dynamic semantics, static expert logic, and global topological context.

[0063] To maximize the utilization of multi-level tag information, the multi-task prediction module constructed in this invention adopts a dual-head prediction structure: a statement node classification head and a function node classification head. The statement node classification head performs a line-level vulnerability localization task, which determines whether a vulnerability exists in the line of code corresponding to the statement node based on its final representation. The function node classification head performs a function-level vulnerability type identification task, which determines the vulnerability type of the vulnerable line based on the final representation of the function node. Line-level vulnerability localization is the primary task, while function-level vulnerability type identification is an auxiliary task.

[0064] The training process of the vulnerability detection model in this invention includes the following steps: Step 1: Construct code with vulnerability annotations as a training dataset; St2, extract code segments from the training dataset as training samples to input into the vulnerability detection model; St3, the preprocessing module of the vulnerability detection model converts the training samples into heterogeneous graphs. The heterogeneous graphs are processed by the dual-stream feature extraction module and the feature aggregation module in sequence to obtain the final features. The multi-task prediction module annotates the training samples line by line based on the final features. Step 4: Calculate the loss function for the training samples. ; ; Among them, L loc For fine-grained vulnerability localization losses, L cls Classify the losses for function-level vulnerability types. and To balance the importance of the two tasks, hyperparameters are optimized jointly. This allows function-level vulnerability type identification to constrain and guide row-level vulnerability localization, thereby improving overall detection effectiveness.

[0065] The node classifier outputs the probability that each line of code in the training samples is identified as a vulnerable line. This task uses a binary cross-entropy loss function for training, the calculation formula of which is shown in the following equation: ; in, For the set of statement nodes in the heterogeneous graph of the training samples, | |for The number of nodes in; For statement nodes Does the line of code contain a genuine vulnerability label? This indicates the statement node output by the model. This represents the predicted probability of a vulnerability. The loss function is used to calculate this probability. This can constrain the model's ability to learn line-level vulnerability location, thereby outputting the specific vulnerability location.

[0066] log refers to the logarithmic function whose base is the natural constant.

[0067] The function node classification header outputs the vulnerability type classification result based on the final representation of the function node. This task uses the cross-entropy loss function for training, and its calculation formula is shown in the following equation: ; in, For the set of function nodes of the training samples, for The number of nodes in the table; C represents the total number of predefined vulnerability categories, which includes categories without vulnerabilities; function node True category one-hot encoding. Represents a multi-class probability distribution. express The probability value of category c. This loss function constrains the model's ability to learn function-level vulnerability type identification.

[0068] Step 5: Repeat steps St2-St4 until the vulnerability detection model converges.

[0069] Reference Figure 3 This invention proposes a smart contract vulnerability detection method based on heterogeneous graphs and large models, the specific steps of which are as follows: S1. Obtain the text of the code to be detected and input it into the vulnerability detection model; S2. The preprocessing model transforms the code text to be detected into a heterogeneous graph G=(V,E); S3. The dual-stream feature extraction module processes the heterogeneous graph G=(V,E) and extracts the initial feature representation of each node. S4. The feature aggregation module fuses attention weights and message vectors based on meta-relations on the heterogeneous graph to obtain the final representation of each target node. ; S5. The function node classification header determines whether a function has vulnerabilities based on the final representation of the function node in each line of code. If not, then the function is considered correct and the code has no bugs. If yes, then determine the vulnerability type of the function and execute step S6; S6. The statement node classification header is based on the final representation of the statement node in the code line. It judges the vulnerability probability of each statement node one by one and outputs the statement nodes that exceed the threshold as the vulnerability location.

[0070] By employing the above reasoning strategy, vulnerability screening can be prioritized at the function level, followed by fine-grained localization at the statement level. This approach ensures granularity of detection while suppressing false alarms, providing a basis for subsequent manual auditing and vulnerability remediation.

[0071] The following specific embodiments illustrate and verify the smart contract vulnerability detection method based on heterogeneous graphs and large models proposed in this invention (hereinafter referred to as the method of this invention, VulnLocator4SC).

[0072] All experiments in this embodiment were run on a desktop computer with an Intel Core Ultra 7 265K 3.90 GHz CPU, an NVIDIA GeForce RTX 3090 (24GB VRAM) GPU, and 64GB of RAM. The operating system was Windows 11 23H2, and the deep learning framework used was PyTorch 2.x combined with PyTorch Geometric for constructing heterogeneous graphs and implementing HGT models. CUDA version was 12.1.

[0073] This embodiment uses a hybrid dataset of SolidiFI-Benchmark, SmartBugs-curated, and SmartBugs-wild. Vulnerable samples primarily come from SolidiFI-benchmark and SmartBugs-curated. These two datasets contain various standard vulnerable contracts that have been manually audited and injected, covering major security vulnerabilities such as arithmetic overflow / underflow and reentrancy vulnerabilities. Vulnerable samples primarily come from SmartBugs-wild. Details of the datasets are shown in Table 1.

[0074] Table 1 Dataset Information

[0075] In this embodiment, the evaluation metrics selected are accuracy (Acc), precision (Precision), recall (Recall), and F1 score; the test results are shown in Tables 2 and 3.

[0076] Table 2 Performance results of VulnLocator4SC in vulnerability line location task

[0077] Table 3 Performance results of VulnLocator4SC on vulnerability type detection tasks

[0078] To investigate the contribution of each core design mechanism in the implementation framework to the overall vulnerability detection performance of the model, six ablation experiment variants were designed by sequentially removing or replacing specific network components, and compared with the complete VulnLocator4SC model in the same test environment. The specific variants are as follows: VulnLocator4SC (w / o Fine-tuning Strategy): Freezes the parameters of the Solidity-T5 encoder, making it a static global semantic feature extractor, to verify the effectiveness of the joint fine-tuning mechanism of all parameters.

[0079] VulnLocator4SC (w / o Global Stream): Completely removes the global semantic feature extraction branch based on Solidity-T5 and uses basic Word2Vec word vectors for feature initialization to verify the advantages of large-scale code pre-trained models in basic semantic representation.

[0080] VulnLocator4SC(w / o Local Stream): Removes the local context knowledge injection stream based on CodeLlama-Instruct to verify the necessity of LLM in deep logical reasoning.

[0081] VulnLocator4SC (w / o HGT): Degenerates the constructed heterogeneous graph of smart contracts into a homogeneous graph and replaces the HGT module with a standard graph convolutional network to verify the heterogeneity of graphs and the key role of meta-relational attention mechanism in capturing abnormal code interactions.

[0082] VulnLocator4SC (w / o Guidance of auxiliary task): Removes the function-level macro-level vulnerability type classification auxiliary task, retaining only the line-level localization main task, to verify the guiding significance of multi-task joint learning in providing global constraints.

[0083] VulnLocator4SC (w / o Context-Aware Prompt): In the knowledge injection phase of the local context flow, the graph topological context, i.e., the predecessor and successor sets of CFG and DFG, is removed from the prompt words of the large model. Only isolated code text is input to verify the effect of the context-aware prompt construction strategy on the logical reasoning of large language models.

[0084] Table 4 Ablation Experiment

[0085] Of course, those skilled in the art will recognize that the present invention is not limited to the details of the exemplary embodiments described above, but also includes the same or similar structures that can be implemented in other specific forms without departing from the spirit or essential characteristics of the invention. Therefore, the embodiments should be considered illustrative and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the present invention. No reference numerals in the claims should be construed as limiting the scope of the claims.

[0086] Furthermore, it should be understood that although this specification describes embodiments, not every embodiment contains only one independent technical solution. This narrative style is merely for clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

[0087] The technologies, shapes, and structures not described in detail in this invention are all known technologies.

Claims

1. A smart contract vulnerability detection method based on heterogeneous graphs and large models, characterized in that, First, a vulnerability detection model is trained, and then the model is used to perform vulnerability detection on the code text to be tested. The vulnerability detection model includes: a preprocessing module, a dual-stream feature extraction module, a feature aggregation module, and a multi-task prediction module; The preprocessing module transforms the smart contract source code into a heterogeneous graph containing function nodes and statement nodes; The dual-stream feature extraction module uses a pre-trained code model to encode the code corresponding to the nodes in the heterogeneous graph to obtain the dynamic semantic features of each node. Then, based on the dynamic semantic features, high-risk statement nodes are selected, and expert knowledge features are injected into the high-risk statement nodes. The dynamic semantic features and expert knowledge features are fused to obtain the initial feature representation of each node. The feature aggregation module performs attention weight modeling based on the meta-relations on the heterogeneous graph, and performs multi-layer fusion on the weighted meta-relations message transmission and the initial feature representation of each node to obtain the final representation of cross-node type and hierarchical fusion context information. The multi-task prediction module determines whether a vulnerability exists based on the final representation of the statement node, and determines the vulnerability type of the vulnerability line based on the final representation of the function node. The method for obtaining expert knowledge features for high-risk statement node injection is as follows: Extract the graph context information of high-risk statement nodes in the heterogeneous graph and construct prompt words. Then, the vulnerability reasoning text is obtained by processing the clue words through the frozen parameter model. Then, the vulnerability reasoning text Mapped to structured static expert knowledge features .

2. The smart contract vulnerability detection method based on heterogeneous graphs and large models as described in claim 1, characterized in that, Heterogeneous graphs define various meta-relations, including control flow edges, data flow edges, function call edges, and bidirectional containment edges; control flow edges Used to indicate the execution order and conditional branching relationships between statements; Data Flow Edge Used to indicate the definition and usage relationship of variables across different statements; call edge Used to represent calling relationships between functions; includes edges Used to represent the hierarchical relationship between function nodes and statement nodes, and includes bidirectional edges.

3. The smart contract vulnerability detection method based on heterogeneous graphs and large models as described in claim 1, characterized in that, The graph context information of high-risk statement nodes in a heterogeneous graph includes: Set of predecessor node sequences in the control flow graph This indicates that by controlling the flow edge The set of source nodes pointing to the high-risk statement node; Set of successor node sequences in the control flow graph This indicates that the high-risk statement node is used. As a control flow edge The set of target nodes of the source node; Data flow graph predecessor node sequence set , indicating through data flow edge Pointing to high-risk statement nodes The set of source nodes; Data flow graph successor node sequence set This indicates that the high-risk statement node is used. As a data flow edge The set of target nodes of the source node.

4. The smart contract vulnerability detection method based on heterogeneous graphs and large models as described in claim 1, characterized in that, The initial feature representations of each node are obtained as follows: For high-risk statement nodes, their dynamic semantic features and expert knowledge features are dimensionally concatenated to obtain an initial feature representation; for the remaining nodes, their dynamic semantic features are mapped to the initial feature representation space to obtain the node's initial feature representation.

5. The smart contract vulnerability detection method based on heterogeneous graphs and large models as described in claim 1, characterized in that, The feature aggregation module works as follows: First, attention weights are modeled using meta-relations to obtain the attention weights of meta-relations corresponding to each unidirectional edge in the heterogeneous graph; then, the messages of each meta-relation are calculated. Then, for the meta-relations where a node exists as a target node, the characteristics of the target node are calculated using the following formula: in, and These represent the updated feature representations of target node t at layers r and (r-1), respectively. Represents the set of source nodes that are neighbors of the target node t; The output mapping matrix represents the type of target node; GELU represents the nonlinear activation function. Heterogeneous attention weights for meta-relations; This represents the message vector passed from source node s to target node t; Traverse r=1, 2...L to obtain features As the final feature of the target node t; Let be the initial feature representation of the target node t.

6. The smart contract vulnerability detection method based on heterogeneous graphs and large models as described in claim 5, characterized in that, The feature aggregation module uses an HGT network.

7. The smart contract vulnerability detection method based on heterogeneous graphs and large models as described in claim 1, characterized in that, During training, the vulnerability detection model uses code labeled with vulnerabilities as the training dataset; the loss function used for training is a weighted sum of fine-grained vulnerability location loss and function-level vulnerability type classification loss. The loss for fine-grained vulnerability localization is: in, For the set of statement nodes in the heterogeneous graph of the training samples, | |for The number of nodes in; For statement nodes Does the line of code contain a genuine vulnerability label? This indicates the statement node output by the model. is the predicted probability of the vulnerability row; log represents the logarithmic function with the natural constant as the base. The loss for classifying function-level vulnerability types is: in, For the set of function nodes of the training samples, for The number of nodes in the table; C represents the total number of predefined vulnerability categories; function node The one-hot encoding corresponding to the true category c, Represents a multi-class probability distribution. express The probability value of category c.

8. The smart contract vulnerability detection method based on heterogeneous graphs and large models as described in claim 1, characterized in that, The steps are as follows: S1. Obtain the text of the code to be detected and input it into the vulnerability detection model; S2. The preprocessing model transforms the code text to be detected into a heterogeneous graph G=(V,E); S3. The dual-stream feature extraction module processes the heterogeneous graph G=(V,E) to extract the initial feature representation of each node. S4. The feature aggregation module fuses attention weights and message vectors based on meta-relations on the heterogeneous graph to obtain the final representation of each target node. ; S5. The function node classification header determines whether a function has vulnerabilities based on the final representation of the function node in each line of code. If not, then the function is considered correct and the code has no bugs. If yes, determine the type of function vulnerability and proceed to step S6; S6. The statement node classification header is based on the final representation of the statement node in the code line. It judges the vulnerability probability of each statement node one by one and outputs the statement nodes that exceed the threshold as the vulnerability location.

9. A smart contract vulnerability detection system, characterized in that, It includes a memory and a processor, wherein the memory stores a computer program, the processor is connected to the memory, and the processor is used to execute the computer program to implement the smart contract vulnerability detection method based on heterogeneous graphs and large models as described in any one of claims 1-8.