Intelligent contract code auditing method fusing program slices and RAG

By combining program slicing and the RAG database, smart contract code is dynamically assembled to construct a three-dimensional feature space and design a chain-based reasoning framework. This solves the problems of context truncation and attention dilution in smart contract code auditing, and achieves efficient and accurate vulnerability detection.

CN120910869APending Publication Date: 2025-11-07BOYA ZHENGLIAN (BEIJING) TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511067153.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing smart contract code auditing methods, such as those based on large language models, suffer from context truncation and attention dilution effects, making it difficult to effectively capture long-range dependency vulnerability patterns across functions and contracts, resulting in low vulnerability detection accuracy and a high false positive rate.

Method used

This paper employs program slicing technology to dynamically assemble discrete code fragments of smart contracts, constructs a three-dimensional feature space, and combines it with the RAG database for multi-dimensional retrieval. A chain-based reasoning framework and a dynamic weight fusion strategy are designed to improve the accuracy of vulnerability detection.

Benefits of technology

It achieves efficient and accurate vulnerability detection of smart contract code, reducing the false alarm rate by 65%, increasing detection efficiency by 3 times, improving MAP to 0.82, and significantly improving detection accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120910869A_ABST
    Figure CN120910869A_ABST
Patent Text Reader

Abstract

The invention provides a smart contract code auditing method fusing program slices and RAG, and relates to the technical field of smart contracts. According to the method, a verified vulnerability contract is obtained through a vulnerability knowledge base construction stage, and a labeling data set covering n types of OWASP standard vulnerabilities is constructed to form a vulnerability knowledge base as an RAG database; in the online detection stage, firstly, a to-be-detected smart contract is dynamically sliced, and a three-dimensional feature vector of a code snippet of the to-be-detected smart contract is obtained; then adopting a hierarchical matching strategy, and using all features of the to-be-detected smart contract code snippets to perform retrieval recall in the RAG database; and finally, designing a chained reasoning framework to deeply fuse the matching result of the large language model LLM and code analysis, and giving the result to the LLM for reasoning to realize vulnerability detection. According to the method, a three-stage processing flow of retrieval, filtering and refinement is constructed in a detection stage, so that the detection precision and the calculation overhead are effectively balanced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of smart contract, and particularly relates to a smart contract code auditing method fusing program slicing and RAG. BACKGROUND

[0002] Current smart contract security auditing technologies mainly fall into three categories: symbolic execution methods based on formal verification, static analysis tools based on rule matching, and vulnerability detection schemes based on deep learning models. Among them, static analysis tools represented by Slither and Mythril rely on an expert rulebase, and achieve vulnerability identification through data flow analysis and taint tracking. The false positive rate of these tools is generally higher than 30%. Although the detection method based on a large language model (LLM) has advantages in semantic understanding, it is difficult to effectively capture long-range dependency vulnerability patterns across functions and contracts due to the limited context window length (usually ≤8k tokens) of the Transformer architecture.

[0003] The existing LLM scheme has two major defects: 1) context truncation problem: when the target contract code exceeds the model context capacity, key code fragments will be forced to be truncated, which will destroy the control flow integrity and introduce noise, making the model more likely to hallucinate; 2) attention dilution effect: when processing long sequences, the weight distribution of the self-attention mechanism tends to be uniform, which weakens the focusing ability on key vulnerability patterns, further increasing the possibility of model hallucination.

[0004] Taking a typical reentrancy attack as an example, the vulnerability features involve multiple discrete code segments such as fallback function definition, state variable modification order, and external call context. The F1 value (F1-Score) of the traditional LLM scheme in these scenarios is about 17.6% lower than that of static analysis tools. SUMMARY

[0005] The technical problem to be solved by the present application is to provide a smart contract code auditing method fusing program slicing and RAG to solve the above technical problems of the prior art.

[0006] To solve the above technical problems, the technical solution adopted by the present application is: an intelligent contract code auditing method combining program slicing and RAG, used to realize intelligent contract code auditing, including a vulnerability knowledge base construction phase and an online detection phase; In the vulnerability knowledge base construction phase, verified vulnerability contracts are obtained, the obtained verified vulnerability contracts are code-normalized using an AST parser to eliminate non-semantic elements, a labeled data set covering n types of OWASP standard vulnerabilities is constructed to form a vulnerability knowledge base as a RAG database; The online detection phase includes: Step S1: dynamically slicing the smart contract to be detected; A context-sensitive hybrid program slicing framework is designed; the hybrid program slicing framework dynamically splices discrete but logically related code fragments in the smart contract through code correlation analysis to form a simplified code set with complete vulnerability context; Firstly, a control flow graph, a data dependency graph and a call relationship graph are constructed based on the smart contract of the code to be detected, forming a multi-dimensional code correlation network; taking vulnerability-sensitive operations as the starting point, a bidirectional traversal strategy is adopted to perform backward slicing and forward slicing on the smart contract; The backward slicing is: tracing backward along the data dependency graph and the control flow graph to capture all condition judgments, permission checks and variable modification histories that affect the current operation, find all statements that affect the slicing criterion, perform cause analysis, and perform slicing; The forward slicing is: expanding along the data flow in the forward direction, tracking the parameter transmission path and the cross-contract call chain, and limiting the depth to within 3 layers to avoid path explosion; the forward slicing is used to find all statements affected by the slicing criterion, perform impact analysis, and perform slicing; For cross-contract scenarios, the interface definition is backtracked through the function selector, the propagation trajectory of external parameters between contracts is tracked in combination with the taint marking technology, and a storage slot mapping is established to handle dynamic calls; The hybrid program slicing framework adopts a dynamic depth control mechanism to control the slicing complexity; the dynamic depth control mechanism automatically adjusts the slicing range according to the number of operation codes, takes the logarithm of the depth of the current slicing to balance the coverage rate and efficiency, and introduces a loop structure early convergence strategy to prevent infinite recursion; In the code splicing phase, irrelevant event logs are filtered through semantic perception technology, duplicate condition judgments are merged, and small modifier logic is inlined, and finally a minimum code set preserving the key vulnerability context is generated; In the face of cross-contract interaction scenarios, the hybrid program slicing framework dynamically loads the ABI definition and symbolic execution of the target contract to reconstruct the cross-contract data flow graph, capture the timing dependency relationship between the fallback function and the state modification of the main contract in the reentrant attack; The entire code slicing process incorporates a semantic-aware filtering layer. Based on the understanding of code intent using a large language model, it effectively eliminates syntactically relevant but semantically irrelevant interfering segments, ensuring that the final generated code slices satisfy both logical integrity and audit requirements.

[0007] Step S2: Obtain the three-dimensional feature vector of the code fragment of the smart contract to be detected; construct the three-dimensional feature space of the smart contract to be detected to achieve accurate vulnerability pattern matching. The specific dimensions of the three-dimensional feature space include global feature vector, behavioral feature vector and code fingerprint feature vector. The global feature, namely the semantic intent feature, is obtained by generating a structured semantic description of the smart contract code fragment based on a large language model. The code fragment of the smart contract is converted into a triple representation containing "operation subject - behavior object - context constraint" through an instruction template, thus constituting the semantic intent feature. The behavior feature vector is obtained by converting the smart contract code fragment from code to semantics on a function-by-function basis. This is used to match the code with the vulnerability audit report in the RAG database at a fine-grained level and to complete the global feature vector. The method for obtaining the code fingerprint feature vector is as follows: For the original code fragment, a 64-bit code fingerprint is generated using an improved SimHash algorithm. Through opcode sequence standardization, function abstraction, and structure normalization, similar vulnerability patterns exhibit significant clustering characteristics in the Hamming distance space. Code fingerprint features are high-level features extracted from the original code snippets and are used for preliminary data cleaning before actual recall. Step S3: Employ a hierarchical matching strategy, using all features of the smart contract code fragment to be detected to perform retrieval and recall in the RAG database; For the same code snippet sample to be detected, three features are used to perform multi-dimensional coarse screening in the semantic space of the RAG database by improving the BM25 algorithm. Candidates with similarity exceeding a set value are retained, and three features of a maximum size are obtained. The ranking list is then generated; subsequently, Hamming distance verification based on code fingerprints is used for clustering, grouping similar codes into one category for ranking; finally, a dynamic weight fusion strategy is employed to integrate the search results, optimize the ranking, and obtain a list with a maximum number of [number missing]. The ranking list, select the top ones according to your needs. Each data sample is matched against a large language model (LLM). Three features are set in the ranking integration process. Weighting coefficients To balance the contributions of different feature dimensions; The search results are integrated using a dynamic weight fusion strategy, as shown in the following formula: ; wherein, is the integrated search result, is the semantic intent feature search result, is the behavior feature search result, is the code fingerprint feature search result, are weight coefficients; the weight coefficients are dynamically adjusted according to the vulnerability type; Step S4: A chain reasoning framework is designed to deeply integrate the matching results of the large language model LLM and code analysis, fill the matching results into the placeholders of the prompt words of the large language model LLM, and give them to the LLM as an enhanced content for reasoning, so as to realize the detection of vulnerabilities; The specific process includes: Step S4.1: Construct a context-enhanced prompt template; A structured prompt template containing vulnerability knowledge, code context and analysis instructions is dynamically generated; the template includes the following four key information: Knowledge injection area: insert the retrieved CWE description, historical vulnerability cases and repair schemes; Output declaration area: limit the output to include risk level and code positioning; Example reference area: provide analysis examples of SWC standard vulnerabilities; Format specification area: provide the required answer format; Step S4.2: Multi-stage reasoning control: adopt a thinking chain strategy to guide the large language model to complete the detection function in five stages: Pattern recognition stage: locate sensitive vulnerability codes similar to the recalled vulnerability fragments; Vulnerability detection stage: determine whether the code slice has a vulnerability; Patch generation stage: generate a corresponding repair patch according to the content of the vulnerability detection stage to preliminarily verify the correctness of the vulnerability; Dependency analysis stage: construct a state variable modification time sequence diagram across functions, so that the large language model generates a time sequence diagram containing all modification processes; the time sequence diagram is only used for the model to confirm the overall situation of the modification process; Rule verification stage: perform compliance check against the security mode to confirm that the modified time sequence diagram is correct; Step S4.3: Vulnerability detection result calibration: design a credibility evaluation model based on attention weight, analyze the attention key word focus in the hidden layer activation value of the large language model LLM, and combine the confidence score of the search result to weight the result; finally output a structured report containing vulnerability type, risk level, code positioning and repair suggestion; Step S5: For the to-be-detected smart contract vulnerability fragments with the occurrence frequency greater than a set value, a layered optimization system is established for vulnerability detection; The layered optimization system comprises an incremental analysis engine, a heterogeneous computing architecture and an adaptive resource scheduling; The incremental analysis engine establishes a dynamically updated code pattern cache library, adopts an LRU strategy to maintain high-frequency vulnerability pattern features, and constructs a difference analysis module to perform incremental scanning analysis on changed code fragments only; The heterogeneous computing architecture adopts CPU-GPU collaborative computing, offloads feature extraction tasks of the to-be-detected smart contract fragments to a GPU, and performs retrieval and sorting in parallel on a CPU core; and performs distributed retrieval based on RDMA; The adaptive resource scheduling designs an intelligent resource allocator to dynamically adjust computing resources according to code complexity: ; wherein, is a calculation mode, is the number of instructions.

[0008] The beneficial effects produced by the above technical solutions are that the smart contract code auditing method provided by the present application fuses program slicing and RAG, proposes a dynamic slicing algorithm based on a control flow graph, a data dependency graph and a call relationship graph, introduces a slicing granularity adaptive adjustment mechanism combining backward slicing (used for vulnerability pattern positioning) and forward slicing (used for impact range analysis), dynamically selects a basic block or a function level slice according to code complexity, initiates a three-dimensional feature representation space (semantic-behavior-code), breaks through the limitations of traditional text similarity calculation, designs a BM25 parameter optimization method for Solidity, improves MAP (Mean Average Precision) to 0.82, constructs a three-level processing flow of "retrieval-filtering-refinement", effectively balances detection accuracy and computing overhead, and develops a dynamic weight fusion strategy to suppress irrelevant fragment interference. BRIEF DESCRIPTION OF DRAWINGS

[0009] Figure 1 A flowchart of the smart contract code auditing method fusing program slicing and RAG provided by the embodiment of the present application; Figure 2 A mixed program slicing framework diagram provided by the embodiment of the present application; Figure 3 A multi-modal retrieval architecture diagram provided by the embodiment of the present application. DETAILED DESCRIPTION

[0010] The specific embodiments of the present invention will be described in further detail below with reference to the accompanying drawings and examples. The following examples are for illustrative purposes only and are not intended to limit the scope of the invention.

[0011] In this embodiment, a smart contract code auditing method integrating program slicing and RAG is used to achieve smart contract code auditing, such as... Figure 1 As shown, it includes the vulnerability knowledge base construction phase and the online detection phase; In the vulnerability knowledge base construction phase, verified vulnerability contracts are obtained from platforms such as Etherscan and GitHub through a distributed crawler. An AST parser is used to perform code normalization on the obtained verified vulnerability contracts, eliminating non-semantic elements such as comments and spaces. A labeled dataset covering six types of OWASP standard vulnerabilities (reentrancy, integer overflow, lack of permissions, etc.) is constructed to form the vulnerability knowledge base as the RAG database. This embodiment filters open-source code projects with OWASP standard vulnerabilities from code hosting platforms such as Etherscan, GitHub, and GitLab, covering different types such as smart contracts (e.g., Ethereum contracts written in Solidity) and web applications (developed in languages ​​such as Java, Python, and PHP). The online detection phase includes: Step S1: Dynamically slice the smart contract to be tested; Design a context-sensitive hybrid program slicing framework; the core of the framework lies in dynamically splicing discrete but logically related code fragments in a smart contract through intelligent code association analysis, forming a concise code set with a complete vulnerability context; its implementation begins with the structured parsing of the target contract, such as... Figure 2As shown, first, the control flow graph (CFG), data dependency graph (DDG) and call relationship graph are constructed based on the smart contract to be detected, forming a multi-dimensional code association network; taking the vulnerability sensitive operation (such as external call instruction, state variable modification) as the breakthrough point, the backward slicing and forward slicing of the smart contract are carried out by using the bidirectional traversal strategy: the backward slicing traces back along the data dependency graph and the control flow graph, captures all the condition judgment, permission check and variable modification history affecting the current operation, finds all the statements affecting the slicing criterion, makes cause analysis, and carries out slicing; the forward slicing extends along the data flow, traces the parameter transmission path and cross-contract call chain, and is limited within 3 layers to avoid path explosion, the forward slicing is used to find all the statements affected by the slicing criterion, makes impact analysis, and carries out slicing; for the cross-contract scenario, the interface definition is backtracked through the function selector, the propagation trajectory of the external parameter between the contracts is tracked by combining the taint marking technology, and the storage slot mapping is established to process the complex dynamic call (such as delegatecall); The hybrid program slicing framework adopts a dynamic depth control mechanism to control the slicing complexity; the dynamic depth control mechanism automatically adjusts the slicing range according to the number of operation codes, takes the logarithm of the depth of the current slice to balance the coverage rate and the efficiency, and introduces a loop structure early convergence strategy to prevent infinite recursion; in the code splicing stage, irrelevant event logs are filtered, repeated condition judgments are merged, and small modifier logic is inlined through semantic perception technology, and finally a minimum code set that retains the key vulnerability context is generated; this scheme innovatively integrates static analysis and semantic understanding, compresses the code size to 15%-25% of the original size, while maintaining a key dependency retention rate of 73%, is especially good at handling cross-contract vulnerability scenarios with three or more layers of nesting, and improves the detection efficiency by 3 times and reduces the false positive rate by 65% compared with traditional methods, achieving a stereoscopic reduction of the execution logic of the smart contract. For example: in the backward slicing stage, the hybrid program slicing framework takes the external call instruction as the starting point, and gradually peels off the preconditions that affect the operation: from the most direct balance check statement (such as require(balances >= amount)), to the deep state variable modification path (such as the cumulative operation of the deposit function on balances), and finally forms a code subset that covers the complete check chain. Forward slicing focuses on the potential impact range of the vulnerability, starting from the dangerous state update operation (such as balances -= amount), tracking the subsequent events (such as abnormal return value of the balance query function) and external interface calls (such as triggering the callback function of other contracts) that may be triggered; in the face of cross-contract interaction scenarios, the hybrid program slicing framework reconstructs the cross-contract data flow graph by dynamically loading the target contract ABI definition and symbolic execution, accurately capturing the timing dependency relationship between the fallback function and the main contract state modification in a reentrant attack; the whole process of code slicing introduces a semantic perception filtering layer based on the understanding of the code intent by a large language model, effectively eliminating syntax-related but semantically irrelevant interference fragments (such as log recording, non-core conditional branches), so that the final generated code slice not only meets the logical integrity, but also meets the cognitive focus needs of human auditing.

[0012] Step S2: obtaining a three-dimensional feature vector of the smart contract code segment to be detected; and constructing a three-dimensional feature space of the smart contract to be detected to realize accurate vulnerability pattern matching, wherein the three-dimensional feature space has specific dimensions including a global feature vector, a behavior feature vector, and a code fingerprint feature vector. The global feature, i.e., the semantic intention feature, is obtained by the following method: generating a structured semantic description of the smart contract code fragment based on a large language model, converting the smart contract code fragment into a triple representation containing "operation subject-behavior object-context constraint" through an instruction template, and constructing a semantic intention feature; for example, converting msg.sender.call{value: amount} into a standardized description of "user initiates an ETH transfer operation, no gas limit is set, and lacks reentrant lock protection". Through this step of conversion, the text to be detected can be converted from a code fragment to a description text that is easy to match, and can be better matched with the vulnerability reports in the database.

[0013] The behavior feature vector is obtained by the following method: converting the smart contract code fragment into a code-to-semantic conversion in units of functions, which is used to match the code with the vulnerability audit reports in the RAG database at a finer granularity, and to complete the global feature vector, which can recall further vulnerability details and provide more detailed description texts to the LLM. The code fingerprint feature vector is obtained by the following method: generating a 64-bit code fingerprint for the original code fragment using an improved SimHash algorithm, and through opcode sequence standardization, function abstraction (replacing variable names with type markers), and structure normalization processing, similar vulnerability patterns are significantly clustered in the Hamming distance space; the code fingerprint feature is a high-level feature extracted from the original code fragment, which is used for preliminary data cleaning before actual recall; clustering can further improve the non-repetition sampling probability during retrieval and reduce unnecessary and repeated LLM query thinking.

[0014] The original SimHash algorithm is mainly used for clustering similar texts, and the SimHash algorithm is modified to adapt to the differences between code and text in this application; In this embodiment, the Slither static analysis tool is used for context-sensitive (Context-sensitive) pointer analysis (Pointer Analysis) of the target contract; the Slither static analysis tool is used to generate a control flow graph, a data dependency graph, and a control dependency graph, and the code is cut according to basic blocks (Basic Block); A hybrid slicing strategy (Hybrid Slicing Strategy) is adopted, forward slicing (Forward Slicing) is performed based on the data dependency graph, and backward slicing (Backward Slicing) is performed based on the control dependency graph; The following three-dimensional feature vectors are extracted by the GPT-4o model: Semantic Intent: Generates a natural language description through the Prompt project to extract the overall intent of the code; Operational Behavior: Natural language descriptions are generated through the Prompt project to extract specific code behavior features on a function-by-function basis. Code fingerprint features (Code Hash): High-level features extracted from the original code snippets, used for preliminary data cleaning before actual recall; Generate a Minimum Viable Slice Set (MVSS) to ensure coverage of all external call sites.

[0015] Step S3: Employ a tiered matching strategy, using all features of the smart contract code fragment to be detected to perform retrieval and recall in the RAG database, such as... Figure 3 As shown; For the same code snippet sample to be detected, three features are used to perform multi-dimensional coarse screening in the semantic space of the RAG database by improving the BM25 algorithm. Candidates with similarity exceeding a set value of 0.7 are retained, and three samples of a maximum size are obtained. The ranking list (i.e., the ranking list contains at most k samples) is generated; then, clustering is performed using Hamming distance verification of code fingerprints to group similar codes into one class for ranking; finally, a dynamic weight fusion strategy is used to integrate the search results and optimize the ranking, resulting in a ranking list with a maximum number of samples. The ranking list, select the top ones according to your needs. A set of data samples are matched against a large language model (LLM); three features are set in this fusion ranking process. Weighting coefficients To balance the contributions of different feature dimensions; assuming a high weight is desired for the global feature vector, the feature matrix can be set to [0.7, 0.15, 0.15]. This way, during the ranking fusion, the global feature vector's ranking weight will account for a significant portion. This allows for higher accuracy in specific vulnerability types.

[0016] The improved BM25 algorithm performs multi-dimensional searches in the semantic space of the RAG database, as shown in the following formula: ; Where D represents the document being rated; A query entered by the user; it usually consists of multiple terms. For the first in the query i 1 term; For terms Term Frequency (TF) in document D, i.e., the number of times the term appears in the document. Inverse Document Frequency (IDF) for term , measures the global importance of the term. Length of document D (usually measured in the number of terms). Average length of all documents in the document collection. Parameter to adjust the saturation of term frequency. Controls the degree of influence of term frequency on the relevance score, The larger the value, the greater the influence of term frequency. Typical values are 1.2~2.0. Parameter to adjust the normalization of document length. Controls the influence of document length on the score, b=0 ignores length, b=1 fully normalizes, typical value is 0.75. In this embodiment, the optimization parameter , ; Apply dynamic weight fusion strategy to integrate retrieval results, as shown in the following formula: ; Where, is the integrated retrieval result, is the semantic intent feature retrieval result, is the behavior feature retrieval result, is the code fingerprint feature retrieval result, are weight coefficients; the weight coefficients are dynamically adjusted according to the vulnerability type (such as focusing on semantic features for reentrant vulnerabilities, setting ).

[0017] Step S4: Design a chain reasoning framework to deeply integrate the matching results of the large language model LLM with code analysis, fill the matching results into the placeholders of the prompt words of the large language model LLM, and give them to the LLM as an enhanced content for reasoning, to realize the detection of vulnerabilities; The specific process includes: Step S4.1: Build a context-enhanced prompt template; Splice the Top-K retrieval results with the original code snippet to build a structured prompt template containing vulnerability knowledge, code context, and analysis instructions; The template mainly contains the following four key information: Knowledge injection area: insert the retrieved CWE description, historical vulnerability cases, and repair schemes; Output declaration area: limit the output to include risk levels and code positioning; Example reference area: Provide an example of parsing SWC standard vulnerabilities; Format specification area: Provide the required answer format to facilitate subsequent analysis of the project; Step S4.2: Multi-stage reasoning control: use the CoT strategy to guide the large language model to complete the detection function in five stages, output a structured report containing risk level, code positioning and repair suggestions: Pattern recognition stage: locate sensitive vulnerability code similar to the recalled vulnerability fragment; Vulnerability detection stage: determine whether the code slice has a vulnerability; Patch generation stage: generate corresponding repair patches based on the content of the previous step to preliminarily verify the correctness of the vulnerability; Dependency analysis stage: build a state variable modification timing diagram across functions, and make the large language model generate a Mermaid diagram containing all modification processes. This timing diagram is only used for the model to confirm the overall situation of the modification process; Rule verification stage: check compliance against Checks-Effects-Interactions and other security patterns to confirm that the modified timing diagram is correct; Step S4.3: Vulnerability detection result calibration: design a credibility evaluation model based on attention weight, analyze the attention keyword focus in the hidden layer activation value of the large language model LLM (such as the attention distribution of terms such as "reentrancy" and "overflow"), and combine the confidence score of the search result to weight the result; Finally output a structured report containing vulnerability type, risk level (high risk / medium risk / low risk), code positioning (line number range) and repair suggestions.

[0018] The formula for calculating the attention keyword focus is as follows: ; Where, is the set of vulnerability keywords, and the final attention keyword focus score is obtained by combining the confidence and focus , attn is the existing attention calculation formula, and t represents token; Step S5: Establish a hierarchical optimization system for vulnerability detection for similar smart contract vulnerability fragments with an appearance frequency greater than a set value; The hierarchical optimization system includes an incremental analysis engine, a heterogeneous computing architecture, and adaptive resource scheduling; the incremental analysis engine establishes a dynamically updated code pattern cache library, uses the LRU strategy to maintain high-frequency vulnerability pattern features; a difference analysis module is constructed to perform incremental scanning analysis only on changed code fragments; The heterogeneous computing architecture adopts CPU-GPU collaborative computing, offloading the feature extraction task of the smart contract fragment to be detected to the GPU (CUDA acceleration), and the retrieval and sorting are performed in parallel on multiple CPU cores; and distributed retrieval is performed based on RDMA, reducing cross-node communication latency to the μs level; the feature vector is quantized from FP32 to INT8, reducing memory usage by 75% and optimizing quantization calculation. The adaptive resource scheduling design uses an intelligent resource allocator that dynamically adjusts computing resources based on code complexity: ; in, For calculation mode, The number of instructions; Simultaneously, pipeline parallelization was implemented, overlapping the three stages of static analysis, feature extraction, and model inference, resulting in a 2.3-fold increase in throughput.

[0019] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope defined by the present invention.

Claims

1. A method for fusing program slicing and RAG for smart contract code auditing, for implementing smart contract code auditing, characterized in that: The method comprises a vulnerability knowledge base construction stage and an online detection stage. In the vulnerability knowledge base construction stage, verified vulnerability contracts are acquired, code normalization is performed on the acquired verified vulnerability contracts using an AST parser, and non-semantic elements are eliminated; a labeled dataset covering n types of OWASP standard vulnerabilities is constructed to form a vulnerability knowledge base as a RAG database. The online detection stage comprises: Step S1: performing dynamic slicing on the smart contract to be detected; Step S2: acquiring a three-dimensional feature vector of the code segment of the smart contract to be detected; and constructing a three-dimensional feature space of the smart contract to be detected to realize accurate vulnerability pattern matching, wherein the three-dimensional feature space comprises a global feature vector, a behavior feature vector and a code fingerprint feature vector; Step S3: adopting a hierarchical matching strategy, and performing retrieval and recall on the RAG database using all features of the code segment of the smart contract to be detected; Step S4: designing a chain reasoning framework to deeply integrate the matching result of the large language model LLM with code analysis, filling the matching result into the placeholder of the prompt word of the LLM as an enhanced content to be input into the LLM for reasoning, so as to realize detection of the vulnerability.

2. The method of claim 1, wherein the method further comprises: The specific method for performing dynamic slicing on the smart contract to be detected is as follows: A context-sensitive hybrid program slicing framework is designed, and through code correlation analysis, discrete but logically related code segments in the smart contract are dynamically spliced to form a simplified code set with complete vulnerability context; Firstly, a control flow graph, a data dependency graph and a calling relationship graph are constructed based on the smart contract of the code to be detected, forming a multi-dimensional code correlation network; taking a vulnerability-sensitive operation as a starting point, a bidirectional traversal strategy is adopted to perform backward slicing and forward slicing on the smart contract; For cross-contract scenarios, the interface definition is backtracked through a function selector, the propagation trajectory of external parameters between contracts is tracked by combining a taint marking technology, and a storage slot mapping is established to process dynamic calls; In the code splicing stage, irrelevant event logs are filtered, duplicate condition judgments are merged, and small modifier logic is inlined through semantic perception technology, and finally a minimum code set retaining key vulnerability context is generated; For cross-contract interaction scenarios, the hybrid program slicing framework dynamically loads the ABI definition and symbolic execution of the target contract to reconstruct the cross-contract data flow graph, and capture the timing dependency relationship between the fallback function and the state modification of the main contract in the reentrant attack.

3. The smart contract code auditing method of claim 2, wherein: The backward slicing is performed by tracing in reverse along the data dependency graph and the control flow graph to capture all condition judgments, permission checks and variable modification histories that affect the current operation, find all statements that affect the slicing criterion, perform cause analysis, and perform slicing; The forward slicing is performed by extending along the data flow in a forward direction to track the parameter transmission path and the cross-contract calling chain, and the depth is limited to within 3 layers to avoid path explosion; the forward slicing is used to find all statements affected by the slicing criterion, perform impact analysis, and perform slicing.

4. The method of claim 3, wherein the method further comprises: The hybrid program slicing framework adopts a dynamic depth control mechanism to control the complexity of slicing; the dynamic depth control mechanism automatically adjusts the slicing range according to the number of operation codes, takes the logarithm of the depth of the current slice to balance the coverage rate and efficiency, and introduces a loop structure early convergence strategy to prevent infinite recursion.

5. The method of claim 4, wherein the method further comprises: The whole process of dynamic slicing of the smart contract introduces a semantic perception filtering layer, based on the understanding of the code intention of the large language model, effectively eliminates the interference fragments related to syntax but irrelevant to semantics, and makes the finally generated code slice meet the logical integrity and meet the audit demand.

6. The smart contract code audit method of claim 5, wherein: The global feature, i.e., the semantic intention feature, is obtained by generating a structured semantic description of the smart contract code segment based on a large language model, converting the smart contract code segment into a triple representation containing "operation subject- behavior object-context constraint" through an instruction template, and constructing a semantic intention feature; the behavior feature vector is obtained by converting the smart contract code segment into a semantic in a function unit, which is used to match the code with the vulnerability audit report in the RAG database at a fine granularity, and complete the global feature vector; The code fingerprint feature vector is obtained by generating a 64-bit code fingerprint using an improved SimHash algorithm for the original code segment, and processing the operation code sequence, function abstraction and structure standardization to make similar vulnerability patterns show significant clustering characteristics in the Hamming distance space; The code fingerprint feature is a high-level feature extracted from the original code segment, which is used for preliminary data cleaning before actual recall.

7. The method of claim 6, wherein the method further comprises: The specific method of using hierarchical matching strategy to retrieve and recall all features of the smart contract code segment in the RAG database is as follows: For the three three-dimensional features of the same code snippet sample to be detected, the improved BM25 algorithm is used for multi-dimensional retrieval screening in the semantic space of the RAG database, and the candidate set with a similarity exceeding a set value is retained to obtain three ranking lists with a maximum size of ; then the Hamming distance check of the code fingerprint is used for clustering, and similar codes are clustered into a class for ranking; finally, a dynamic weight fusion strategy is used to integrate the retrieval results, sort and optimize, and obtain a ranking list with a maximum number of ; according to the requirements, the top data samples are selected for matching by the large language model LLM. Setting three feature weight coefficients in the fusion ranking process To balance the contribution of different feature dimensions; The dynamic weight fusion strategy is applied to integrate the retrieval results, as shown in the following formula: ; wherein, is the integrated search result, is the semantic intent feature search result, is the behavior feature search result, is the code fingerprint feature search result, are weight coefficients; The weight coefficient is dynamically adjusted according to the vulnerability type.

8. The method of claim 7, wherein the method further comprises: The specific method of step S4 is as follows: Step S4.1: Construct a context-enhanced prompt template; A structured prompt template containing vulnerability knowledge, code context and analysis instructions is dynamically generated; the template includes the following four key information: Knowledge injection area: insert the retrieved CWE description, historical vulnerability cases and repair scheme; Output declaration area: limit the output to include risk level and code positioning; Example reference area: provide analysis examples of SWC standard vulnerabilities; Format specification area: provide the required answer format; Step S4.2: Multi-stage reasoning control: use the thinking chain strategy to guide the large language model to complete the detection function in five stages: Pattern recognition stage: locate similar sensitive vulnerability codes to the vulnerability segment; Vulnerability detection stage: determine whether the code slice contains vulnerabilities; Patch generation stage: generate corresponding repair patches based on the content of the vulnerability detection stage to preliminarily verify the correctness of the vulnerabilities; The dependency analysis stage: build a cross-function state variable modification time sequence diagram, make a large language model generate a time sequence diagram containing all modification processes, and the time sequence diagram is only used for the model to confirm the overall situation of the modification process; The rule verification stage: compliance check against the safety mode to confirm that the modified time sequence diagram is correct; Step S4.3: vulnerability detection result calibration: design a credibility evaluation model based on attention weight, analyze the attention keyword focus in the hidden layer activation value of the large language model LLM, and combine the confidence score of the search result to weight the result; finally output a structured report containing vulnerability type, risk level, code positioning and repair suggestion.

9. The method of claim 8, wherein the method further comprises: The method establishes a hierarchical optimization system for vulnerability detection for smart contract vulnerability fragments with an occurrence frequency greater than a set value; The hierarchical optimization system includes an incremental analysis engine, a heterogeneous computing architecture, and an adaptive resource scheduling; The incremental analysis engine establishes a dynamically updated code pattern cache library and uses the LRU strategy to maintain high-frequency vulnerability pattern features; A difference analysis module is constructed to perform incremental scanning analysis only on the changed code fragments; The heterogeneous computing architecture uses CPU-GPU collaborative computing to offload feature extraction tasks of the smart contract fragments to be detected to the GPU, and searches and sorts in the CPU multicore parallel; and performs distributed search based on RDMA; The adaptive resource scheduling designs an intelligent resource allocator that dynamically adjusts the computing resources according to the code complexity: ; wherein, is the number of instructions for the compute mode, is the number of instructions.