Intelligent Contract Library Misuse Detection Method and System Based on Large Language Model and Static Analysis

The method uses large language models and static analysis to detect library misuse in smart contracts, addressing precision issues by refining prompts and utilizing code similarity matching, enhancing the safety of smart contract development.

CN119397546BActive Publication Date: 2025-07-15HAINAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411428940.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-14
Publication Date
2025-07-15
Estimated Expiration
2044-10-14

AI Technical Summary

Technical Problem

The existing technology lacks effective tools to detect misuse of libraries in smart contracts, resulting in security vulnerabilities in the development of smart contracts.

Method used

Using a method based on large language model and static analysis, the library misuse pattern in smart contracts is accurately detected by constructing prompt word methods, code similarity matching and result iterative optimization.

Benefits of technology

It improves the accuracy of library misuse detection, can accurately identify potential security issues in smart contracts, and improves the security of smart contracts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119397546B_ABST
    Figure CN119397546B_ABST
Patent Text Reader

Abstract

The present invention relates to a method and system for detecting misuse of intelligent contract libraries based on large language models and static analysis. The method includes: extracting library misuse information; constructing a prompt method of generate-then-select; using the prompt method to perform intelligent contract detection of the large language model, obtaining a preliminary detection result as part of the next input for result iterative feedback optimization to obtain a target large language model; creating a static matching BCSSM structure; performing weighted matching on the detection results of the target large language model and the static matching structure, and outputting a final detection result. By constructing a prompt method to use the large language model for intelligent contract detection, and utilizing the code parsing and logical reasoning capabilities of the large language model as well as static matching, it is possible to more accurately detect the library misuse patterns carried by intelligent contracts. Applied to the development and maintenance stages of intelligent contracts, it can achieve more secure intelligent contracts.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of software engineering vulnerability detection, and in particular to an intelligent contract library misuse detection method and system based on large language models and static analysis. Background Art

[0002] Blockchain technology is renowned for its decentralized and immutable nature and excellent security, finding extensive applications in critical fields such as finance and healthcare. Ethereum, as the first public blockchain platform supporting smart contracts, has revolutionized the blockchain field by introducing a programmable blockchain that enables the creation of smart contracts and has become one of the most popular blockchain platforms. A smart contract is a self-executing contract whose protocol terms are directly written in lines of code and run on the blockchain. The Ethereum platform allows developers to write smart contracts using its main language, Solidity, and deploy them on the Ethereum Virtual Machine (EVM). The EVM is a decentralized virtual machine that executes code using an international public node network. This approach brings a higher degree of transparency and trust as transactions and contract states are immutably recorded on the blockchain. The Solidity library is a set of reusable functions that can be called by other contracts. Developers often encapsulate common operations or functions to solve specific problems. When other developers wish to implement similar functionality, they can import these libraries and use their methods to solve specific problems. Incorrect implementation or application of libraries during the smart contract development process may lead to defects in the contract, and currently, there is a lack of tools that can effectively detect and identify library misuse patterns in smart contracts.

[0003] Library misuse is a common problem in smart contract development, involving three key roles: the developer who creates the library, the contract developer who uses the library, and the community that maintains the library. Effective collaboration among these three roles contributes to the development of smart contracts. However, improper behavior by the developer who creates the library or the contract developer who uses the library is called library misuse, including but not limited to incorrect function calls, incomplete understanding of the capabilities of library functions, or failure to update to the secure version of the library in a timely manner, thus introducing security vulnerabilities, etc. Therefore, traditional library misuse detection methods have the problem of inaccurate detection. Summary of the Invention

[0004] Based on this, in order to solve the above technical problems, an intelligent contract library misuse detection method and system based on large language models and static analysis are provided, which can improve the detection accuracy of library misuse.

[0005] An intelligent contract library misuse detection method based on large language models and static analysis includes:

[0006] Extract library misuse information corresponding to various library misuse patterns of smart contracts from the database;

[0007] Construct a prompting method of pre-generated-then-selected based on Automatic Chain of Thought (Auto-CoT) and zero-shot Chain of Thought (zero-shot-CoT);

[0008] According to the library misuse information, use the prompting method to perform intelligent contract detection on the large language model to obtain a preliminary detection result;

[0009] Use the preliminary detection result as a part of the prompt in the next intelligent contract detection of the large language model, perform result iterative feedback optimization on the large language model to obtain a target large language model, and obtain a first library misuse detection result through the target large language model;

[0010] Create a static matching BCSSM structure based on code similarity, and obtain a second library misuse detection result through the static matching BCSSM structure;

[0011] Perform weighted matching on the first library misuse detection result and the second library misuse detection result, and output a target library misuse detection result.

[0012] In one embodiment, extract library misuse information corresponding to various library misuse patterns of intelligent contracts from a database, including:

[0013] Search for the pattern name, scenario description, and code features of various library misuse patterns in the data, and use the pattern name, scenario description, and code features as library misuse information;

[0014] The library misuse patterns include invalid wrappers in the library, unhandled exceptions in the library, exception library extensions, abnormal usage, incomplete function replacements, abnormal library capacity estimations, and abnormal library usages.

[0015] In one embodiment, use the prompting method to perform intelligent contract detection on the large language model to obtain a preliminary detection result, including:

[0016] Locate the role of the large language model to determine the source code of the intelligent contract to be detected;

[0017] Perform code understanding and logical reasoning matching on the source code of the intelligent contract to be detected according to the library misuse information to obtain a matching result;

[0018] Based on the matching result, determine the recognition ratio of identifying the library misuse pattern in the source code of the intelligent contract to be detected as another type of library misuse pattern;

[0019] Obtain a preliminary detection result according to the matching result and the recognition ratio;

[0020] Determine the output format type and apply the format type to the output of the preliminary detection result.

[0021] In one embodiment, use the preliminary detection result as part of the prompt in the next large language model smart contract detection, and perform iterative feedback optimization on the large language model to obtain a target large language model, including:

[0022] Extract the detection result error rate and false positive ratio from the preliminary detection result;

[0023] Fuse the extracted error rate and false positive ratio, and use the fused information as part of the prompt in the next large language model smart contract detection for cyclic optimization to obtain an evaluation index for the detection result;

[0024] If the evaluation index reaches the evaluation threshold, stop the iteration, and use the last evaluation result as the output result of the large language model to obtain the target large language model.

[0025] In one embodiment, obtain the second library misuse detection result through the static matching BCSSM structure, including:

[0026] Use the TF-IDF technology through the static matching BCSSM structure to convert the code fragments and input code of the smart contract into a TF-IDF matrix;

[0027] Calculate the cosine similarity between the input code and each code fragment according to the TF-IDF matrix to obtain a calculation result;

[0028] Obtain the second library misuse detection result according to the calculation result.

[0029] In one embodiment, perform weighted matching on the first library misuse detection result and the second library misuse detection result, including:

[0030] Compare the first library misuse detection result and the second library misuse detection result to obtain a comparison result;

[0031] When the comparison result shows that the first library misuse detection result and the second library misuse detection result are consistent, output the target library misuse detection result;

[0032] When the comparison result shows that the first library misuse detection result and the second library misuse detection result are inconsistent, weight the first library misuse detection result and the second library misuse detection result.

[0033] In one embodiment, the method further includes:

[0034] Determine the result output format, and parse and output the target library misuse detection result according to the described result output format.

[0035] An intelligent contract library misuse detection system based on large language models and static analysis, comprising:

[0036] An information extraction module, configured to extract library misuse information corresponding to various library misuse patterns of intelligent contracts from a database;

[0037] A prompt method construction module, configured to construct a pre-generated-then-selected prompt method based on Auto-CoT (Automatic Chain of Thought) and zero-shot-CoT (zero-shot Chain of Thought);

[0038] A preliminary detection module, configured to perform intelligent contract detection on a large language model by using the prompt method according to the library misuse information, and obtain a preliminary detection result;

[0039] A model optimization module, configured to use the preliminary detection result as a part of the prompt in the next intelligent contract detection of the large language model, perform result iterative feedback optimization on the large language model, obtain a target large language model, and obtain a first library misuse detection result through the target large language model;

[0040] A static analysis module, configured to create a static matching BCSSM (Based on Code Similarity Static Matching) structure, and obtain a second library misuse detection result through the static matching BCSSM structure;

[0041] A detection module, configured to perform weighted matching on the first library misuse detection result and the second library misuse detection result, and output a target library misuse detection result.

[0042] The above intelligent contract library misuse detection method and system based on large language models and static analysis use the constructed prompt method to perform intelligent contract detection on a large language model, and utilize the code parsing and logical reasoning capabilities of the large language model as well as static matching, with strong generalization ability. It can be applied to intelligent contract codes of different types or scales, can more accurately detect the library misuse patterns carried by intelligent contracts, and can be applied to the development and maintenance stages of intelligent contracts to achieve more secure intelligent contracts. Description of the Drawings

[0043] Figure 1 It is an application environment diagram of an intelligent contract library misuse detection method based on large language models and static analysis in an embodiment;

[0044] Figure 2 It is a flowchart of an intelligent contract library misuse detection method based on large language models and static analysis in an embodiment;

[0045] Figure 3 It is a diagram for parsing prompt words used in an embodiment;

[0046] Figure 4 It is a schematic diagram for iterative feedback optimization of the results of the large language model detection module in an embodiment;

[0047] Figure 5 It is a schematic diagram of static matching BCSSM based on code similarity in an embodiment;

[0048] Figure 6 It is a four - evaluation - index table for each module and the final result in the experimental results;

[0049] Figure 7 It is a structural block diagram of an intelligent contract library misuse detection system based on a large language model and static analysis in an embodiment;

[0050] Figure 8 It is an internal structure diagram of a computer device in an embodiment. Detailed implementation manners

[0051] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0052] It can be understood that the terms "first", "second", etc. used in the present application can be used herein to describe the library misuse detection results, but these library misuse detection results are not limited by these terms. These terms are only used to distinguish the first library misuse detection result from another library misuse detection result. For example, without departing from the scope of the present application, the first library misuse detection result can be called the second library misuse detection result, and similarly, the second library misuse detection result can be called the first library misuse detection result. Both the first library misuse detection result and the second library misuse detection result are library misuse detection results, but they are not the same library misuse detection result.

[0053] The intelligent contract library misuse detection method based on a large language model and static analysis provided by the embodiments of the present application can be applied to an application environment as Figure 1 shown. As Figure 1As shown, the application environment includes a computer device 110. Manually extract the library misuse information corresponding to various library misuse patterns of smart contracts from the database; the computer device 110 can construct a pre-generated-then-selected prompt method based on Auto-CoT and zero-shot-CoT; the computer device 110 can use the prompt method to detect smart contracts of the large language model according to the library misuse information to obtain a preliminary detection result; the computer device 110 can use the preliminary detection result as part of the prompt in the next large language model smart contract detection to perform result iteration feedback optimization on the large language model to obtain a target large language model, and obtain a first library misuse detection result through the target large language model; the computer device 110 can create a static matching BCSSM structure based on code similarity and obtain a second library misuse detection result through the static matching BCSSM structure; the computer device 110 can perform weighted matching on the first library misuse detection result and the second library misuse detection result and output a target library misuse detection result. Among them, the computer device 110 can be, but is not limited to, various personal computers, laptop computers, smartphones, robots, unmanned aerial vehicles and other devices.

[0054] In one embodiment, as Figure 2 shown, a smart contract library misuse detection method based on a large language model and static analysis is provided, including the following steps:

[0055] Step 202, extract the library misuse information corresponding to various library misuse patterns of smart contracts from the database.

[0056] An operator can manually extract the library misuse information corresponding to various library misuse patterns of smart contracts from the database, and the computer device can preprocess the data after collecting the smart contract dataset for subsequent operations.

[0057] In one embodiment, a smart contract library misuse detection method based on a large language model and static analysis can also include the process of extracting information. The specific process includes: finding the pattern name, scenario description, and code features of various library misuse patterns from the data, and using the pattern name, scenario description, and code features as library misuse information; the library misuse patterns include invalid wrappers in the library, unhandled exceptions in the library, abnormal library extensions, abnormal usage, incomplete function replacement, abnormal library ability estimation, and abnormal library usage.

[0058] Among them, the code features of each type of library misuse pattern can be manually extracted from the database, including the scenario description and code-level features of each type of library misuse pattern. Specifically, the operator can extract various library misuse patterns of smart contracts, including scenario descriptions and code features, that is, extract information on eight library misuse patterns from empirical research, including the name of each type of library misuse pattern, the scenario description of each type of library misuse pattern, and the code features of each type of library misuse pattern. In specific implementation, manually extract information on eight library misuse patterns from the database provided by empirical research, specifically including: the pattern name of each type of library misuse pattern, the scenario description of each type of library misuse pattern, and the code features of each type of library misuse pattern.

[0059] In this embodiment, the eight types of library misuse patterns extracted include: invalid wrapper checks in the library, unhandled exceptions in the library, inappropriate library extensions, inappropriate use of "using for", incomplete function replacement, overestimation of library capabilities, underestimation of library capabilities, and unnecessary library use.

[0060] Specifically, the information on the eight library misuse patterns is as follows:

[0061] Invalid wrapper checks in the library: The check logic implemented in the library function has defects and cannot handle all expected scenarios, resulting in security issues. For example, the SafeERC20 library is designed to protect ERC-20 token operations, but developers may fail to use the library correctly, such as failing to check the return value, thus leading to potential risks.

[0062] Unhandled exceptions in the library: The library is not designed to consider all possible exceptions, resulting in contract execution failure in some cases. For example, a library function may interact with other contracts during execution but does not handle the exception responses from these contracts.

[0063] Inappropriate library extensions: Multiple functions are mixed in the library function, making them too complex to maintain and understand, increasing the risk of errors. For example, a library function is responsible for calculating and validating different business logics at the same time, increasing the possibility of errors.

[0064] Inappropriate use of "using for": The 'using for' statement is misused in the contract, binding the library function to an incompatible or unnecessary data type. For example, the SafeMath library is designed for the uint256 type but is misapplied to the int256 type.

[0065] Incomplete function replacement: In the contract, the original insecure function is not completely replaced by the safer version provided by the library. For example, when developers use the SafeERC20 library, they fail to replace all approve calls with safeApprove, missing some replacements.

[0066] Overestimating the capabilities of libraries: Developers incorrectly assume that library functions have certain unimplemented capabilities. For example, using a library function to check if an address is a contract, but the function does not verify the deployed code of the contract.

[0067] Underestimating the capabilities of libraries: Developers fail to fully utilize the functions provided by library functions, resulting in redundant implementation of functions already provided by the library in the contract.

[0068] Unnecessary use of libraries: Unnecessary libraries are included in the contract, which are either not used or are natively supported by the new version of the compiler, resulting in waste of resources. For example, although Solidity version 0.8 has built-in overflow checks, developers still introduce the SafeMath library for overflow protection, causing unnecessary gas consumption.

[0069] Integrating information on eight library misuse patterns is crucial for large language models to deeply understand the semantics of code, because it can better understand the functions of code through semantic reasoning and improve the accuracy of detecting library misuse behaviors in target contracts.

[0070] Step 204, construct a prompt method of "Pre-generated-Then Selected" based on Auto-CoT (Automatic Chain of Thought) and zero-shot-CoT (zero-shot Chain of Thought).

[0071] In this embodiment, the computer device can combine the step-by-step thinking idea of Auto-CoT and zero-shot-CoT to construct a prompt method of "Pre-generated-Then Selected" to reduce the interference of complex contracts on the large model. Among them, the prompt can include smart contract code, descriptions of library misuse patterns, and code features related to each pattern.

[0072] Constructing a prompt method of "Pre-generated-Then Selected" based on the step-by-step thinking idea of Auto-CoT and zero-shot-CoT can greatly improve the accuracy of the large language model in detecting library misuse behaviors in target contracts. The specific steps can include:

[0073] Initial prompt generation: The initial prompt (GenerateInitialPrompt) combines smart contract code, descriptions of library misuse patterns, and code snippets related to each pattern; The formula is expressed as: GenerateInitialPrompt(C, P, G) = C + P + G; where C represents smart contract code, P represents library misuse patterns and their attribute descriptions, and G represents code features related to each pattern.

[0074] Composition of prompt words: Smart contract code: The actual contract code input to the LLM for pattern matching and logical reasoning; Library misuse pattern description: A detailed description of each pattern to help the LLM understand the characteristics of each misuse pattern; Code snippet: A code example related to each misuse pattern for the LLM to perform pattern matching in the contract code.

[0075] The "Pre-generated-ThenSelected" prompting strategy: To improve the response quality of the LLM and reduce the impact of complex contracts on its reasoning ability, the "pre-generated - then select" prompting strategy is adopted; This strategy guides the LLM to perform more in-depth logical reasoning through a series of analytical thinking processes, and combines the reasoning results with the questions to form a more guiding prompt.

[0076] Step 206, according to the library misuse information, use the prompt word method to perform intelligent contract detection on the large language model to obtain a preliminary detection result.

[0077] The computer device can perform intelligent contract detection on the large language model module using the Pre-generated-Then Selected prompt word method based on the scenario description and code-level features of the library misuse pattern.

[0078] Among them, in this embodiment, when using the Pre-generated-Then Selected prompt words for intelligent contract detection of the large language model module, two models, GPT-4o and GPT-4Turbo, are used for detection, and the GPT-4Turbo model with the best effect is selected.

[0079] In one embodiment, a method for detecting library misuse of intelligent contracts based on a large language model and static analysis may further include a process of preliminary detection. The specific process includes: positioning the role of the large language model to determine the source code of the intelligent contract to be detected; performing code understanding and logical reasoning matching on the source code of the intelligent contract to be detected according to the library misuse information to obtain a matching result; based on the matching result, determining the recognition ratio of misidentifying the library misuse pattern in the source code of the intelligent contract to be detected as another type of library misuse pattern; obtaining a preliminary detection result according to the matching result and the recognition ratio; determining the output format type and applying the format type to the output of the preliminary detection result.

[0080] The computer device constructs the Pre-generated-Then Selected prompt word method based on the step-by-step thinking idea of Auto-CoT and zero-shot-CoT. The prompt word parsing diagram used in this embodiment is as Figure 3 shown, and the prompt word structure specifically includes:

[0081] Give the large language model a role. In this embodiment, the large language model is positioned as an experienced smart contract security auditor;

[0082] Determine the original code of the smart contract to be detected;

[0083] Determine the detailed information of eight types of library misuse patterns, including the pattern name, scenario description, and code characteristics of each pattern;

[0084] Determine the functional requirements, which mainly involve code understanding and logical reasoning matching of the input smart contract according to the scenario descriptions and code characteristics of the eight types of library misuse patterns;

[0085] The result of the nth detection, which mainly shows the proportion of misidentifying a certain type of library misuse pattern as another type of library misuse pattern;

[0086] The given JSON output format specifically includes: the library misuse pattern numbers carried by the target contract, divided into P1 - P8 and NONE, where NONE indicates that the target contract does not contain any library misuse behaviors; the library misuse pattern names carried by the target contract, and the pattern name for NONE is "NONE"; the scenario descriptions corresponding to the library misuse patterns carried by the target contract, and the scenario description for NONE is "NONE"; the code snippets corresponding to the library misuse patterns carried by the target contract, and the code snippet for NONE is "NONE";

[0087] It is required to conduct five background detections on the target contract, and select the result with the highest frequency in the five detections as the output result;

[0088] Constrain the output format.

[0089] Step 208: Use the preliminary detection result as part of the prompt in the next large language model smart contract detection, perform result iteration feedback optimization on the large language model to obtain the target large language model, and obtain the first library misuse detection result through the target large language model.

[0090] The schematic diagram of the result iteration feedback optimization of the large language model detection module is as Figure 4 shown. The computer device can integrate the detection results and use them as part of the prompt for the next large model detection to perform result iteration feedback optimization. Take the result R1 of the initial detection of the large language model as part of the new prompt, conduct the next iteration detection of the large language model, and loop this step, repeating the analysis with the new prompt to refine the result. This process will continue until a satisfactory accuracy level is reached.

[0091] Specifically, in one embodiment, an intelligent contract library misuse detection method based on a large language model and static analysis may further include a process of model iteration optimization. The specific process includes: extracting the detection result error rate and false positive ratio from the preliminary detection results; performing information fusion on the extracted error rate and false positive ratio, and using the fused information as part of the prompt in the next large language model intelligent contract detection, and performing cyclic optimization to obtain the evaluation index of the detection result; if the evaluation index reaches the evaluation threshold, stop the iteration, and use the last evaluation result as the output result of the large language model to obtain the target large language model.

[0092] In this embodiment, in the large language model result feedback iteration optimization stage, it specifically includes: refining the initial detection results, and extracting the error rate and false positive ratio of the target contract detection results; fusing the extracted information into the prompt, and using this prompt as the result of the second detection; repeating the above steps, and calculating the evaluation index for each result; when the evaluation index reaches the accuracy threshold, stop the iteration; using the last evaluation result as the final output result of the large language model module; the above steps are performed using two models, GPT-4o and GPT-4Turbo, and finally the GPT-4Turbo model with the best effect is selected.

[0093] Refining the results of the initial detection of the large language model, and using the refined results as part of the prompt for the next detection, and performing result feedback iteration optimization can further reduce the interference of complex contracts on the detection of the large language model.

[0094] In each iterative analysis, based on the preliminary result R1 of the large language model LLM, a new prompt is constructed. This new prompt includes the results of the preliminary analysis and any negative feedback; the formula is expressed as: GenerateFeedbackPrompt(R1) = R1 + NegativeFeedback(R1); through this iterative optimization process, the output of the LLM is gradually adjusted to improve the detection accuracy; in each iteration, the output of the LLM is used as part of the prompt for the next iteration, forming a feedback loop to achieve iterative optimization of the LLM output; in this way, the LLM can adjust its reasoning process according to the previous analysis results, and gradually improve the detection accuracy and reliability.

[0095] Step 210, create a static matching BCSSM structure based on code similarity, and obtain the second library misuse detection result through the static matching BCSSM structure.

[0096] In this embodiment, two static analysis structures are developed, namely the character-level static matching BCLSM and the code similarity-based static matching BCSSM. The BCSSM structure is selected according to the evaluation metrics. This structure detects misuse patterns in the smart contract library by calculating the cosine similarity between the input code and each code snippet.

[0097] Among them, for the character-level static matching (Character-Level Static Matching, BCLSM), the specific steps are as follows: comprehensively scan the input smart contract code and compare it with the feature code snippets of each library misuse pattern; determine the consistency between the input contract code and the code snippets of a specific pattern through quantitative matching; if the contract code has the highest matching degree with the code snippets of a certain pattern, it is considered that the input contract exhibits this library misuse pattern.

[0098] In one embodiment, a smart contract library misuse detection method based on a large language model and static analysis may further include the process of detection through the static matching BCSSM structure. The specific process includes: using the TF-IDF technology through the static matching BCSSM structure to convert the code snippets of the smart contract and the input code into a TF-IDF matrix; calculating the cosine similarity between the input code and each code snippet according to the TF-IDF matrix to obtain the calculation result; obtaining the second library misuse detection result according to the calculation result.

[0099] The schematic diagram of the code similarity-based static matching BCSSM is as Figure 5 shown. For the code similarity-based static matching (Code Similarity Static Matching, BCSSM), the specific steps are as follows: use the TF-IDF (Term Frequency-Inverse Document Frequency) technology to convert the code snippets and the input code into a TF-IDF matrix; TF-IDF is a text representation technology that takes into account the term frequency (TF) and the inverse document frequency (IDF) to reduce the influence of common words and improve the ability to identify similarities; calculate the cosine similarity between the input code and each code snippet. The cosine similarity quantifies the angular distance between two vectors, and a value close to 1 indicates a higher similarity.

[0100] In this embodiment, a character-level static matching BCLSM structure is developed. Pattern detection is performed by quantifying the match to determine the consistency between the input contract code and the code snippets of a specific pattern. A static matching BCSSM structure based on code similarity is developed. The TF-IDF (Term Frequency-Inverse Document Frequency) technique is used to convert the code snippets and the input code into TF-IDF matrices, and pattern detection is performed by calculating the cosine similarity between the input code and each code snippet.

[0101] In this embodiment, by calculating evaluation metrics (including accuracy, recall, F1-score, precision) for the output results of the BCLSM and BCSSM frameworks, the BCSSM framework with relatively better results is selected as the method for static analysis.

[0102] Step 212, perform weighted matching on the first library misuse detection result and the second library misuse detection result, and output the target library misuse detection result.

[0103] In one embodiment, an intelligent contract library misuse detection method based on a large language model and static analysis may further include a process of weighted matching. The specific process includes: comparing the first library misuse detection result and the second library misuse detection result to obtain a comparison result; when the comparison result is that the first library misuse detection result and the second library misuse detection result are consistent, output the target library misuse detection result; when the comparison result is that the first library misuse detection result and the second library misuse detection result are inconsistent, weight the first library misuse detection result and the second library misuse detection result.

[0104] That is, in the weight ratio allocation stage of this embodiment, the results of the large language model and static analysis are combined to obtain the final detection result. When integrating the results, the weights of the patterns are considered, and these weights are proportional to the occurrence frequencies in real-world contracts. When the module results are inconsistent, the consistent results are given priority, and the decision is optimized based on the weighted evaluation of the two results. The computer device can match the results of the large language model and the results of static analysis according to the weight ratio. When performing the matching, weight allocation is performed according to the proportions of eight library misuse patterns in real-world intelligent contracts, and the results of the large language model module and the results of the static analysis module are matched according to the weight ratio.

[0105] In one embodiment, an intelligent contract library misuse detection method based on a large language model and static analysis may further include a process of returning the detection result to the user. The specific process includes: determining the result output format, and parsing and outputting the target library misuse detection result according to the result output format.

[0106] Specifically, the computer device can use a pre-defined json format to parse the finally returned result and obtain the detection result of the library misuse pattern of the target contract and the detailed information of this pattern.

[0107] The intelligent contract library misuse detection method based on large language models and static analysis provided in this application can be applied in the development and maintenance stages of intelligent contracts, helping developers understand and maintain potential security issues, so as to achieve more secure intelligent contracts. Through the code parsing and logical reasoning capabilities of large language models and the matching of static analysis, the library misuse patterns carried by intelligent contracts can be detected more accurately. This method based on large prediction models and static analysis has strong generalization capabilities and can be applied to intelligent contract codes of different types or scales.

[0108] In one embodiment, in order to verify the effectiveness of the intelligent contract library misuse detection method based on large language models and static analysis, extensive experiments were conducted using 1018 intelligent contracts in the real world as the dataset. This method comes from the intelligent contract datasets on Etherscan and Github. According to the proportion of various library misuse patterns provided in empirical research in the real world, 183 intelligent contracts were randomly selected as the test dataset.

[0109] This method calculated evaluation metrics such as accuracy, recall value, F1 index, and precision for the above test dataset to verify the detection ability of this method for library misuse behaviors in intelligent contracts in the real world. The evaluation metrics of large language model detection, static analysis detection, and the final result in this method are as Figure 6 shown, Figure 3 indicating the four evaluation metrics of each module and the final result.

[0110] Weighted matching is performed on the results of large language model detection and static analysis detection. According to the parsing of the source codes of 1018 intelligent contracts in the real world, the proportion of contracts with eight library misuse patterns in these contracts is used as the weight for each library misuse pattern. When the results are inconsistent, the consistent results are given priority, and the decision is optimized based on the weighted evaluation of the two results.

[0111] As Figure 6As shown, accuracy is an indicator for evaluating the correctness of model predictions, calculated as the number of all correct predictions (true positives TP and true negatives TN) divided by the total number of all predictions, reflecting the model's ability to correctly identify positive and negative examples; recall is an indicator for evaluating the model's ability to identify all actual positive examples, calculated as the number of true positives TP divided by the sum of true positives TP and false negatives FN. The higher the recall, the fewer positive examples the model misses; the F1 score is the harmonic mean of accuracy and recall, which balances between the two and is an indicator that comprehensively considers accuracy and recall; precision is an indicator for evaluating the proportion of actual positive examples among the examples predicted as positive by the model, calculated as the number of true positives TP divided by the sum of true positives TP and false positives FP. The higher the precision, the more reliable the model's prediction results.

[0112] Despite the inherent challenges in detecting complex patterns, this method still achieved promising evaluation metrics. Specifically, the accuracy reached 68%, indicating the overall effectiveness of the model. The precision of 86% highlights the model's ability to accurately identify true positives, while the F1 score of 68% indicates the balance between precision and recall. Although the recall was 57%, it is crucial to recognize the trade - offs involved in detecting nuanced patterns in different datasets. These results confirm the potential of this method, providing a solid foundation for future improvements and practical applications.

[0113] In summary, the method for detecting library misuse patterns in smart contracts based on large - language models and static analysis provided by the present invention can accurately detect library misuse behaviors in smart contracts; it can be applied to the development and maintenance phases of smart contracts, thus enabling more secure smart contracts.

[0114] It should be understood that although the steps in the above - mentioned flowchart are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless clearly stated in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least some of the steps in the above - mentioned flowchart may include multiple sub - steps or multiple stages. These sub - steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub - steps or stages is not necessarily sequential either, but can be executed alternately or in rotation with at least a part of other steps or sub - steps or stages of other steps.

[0115] In one embodiment, as Figure 7 shown, a system for detecting library misuse in smart contracts based on large - language models and static analysis is provided, including: an information extraction module 710, a prompt - word method construction module 720, a preliminary detection module 730, a model optimization module 740, a static analysis module 750, and a detection module 760, where:

[0116] An information extraction module 710 for extracting library misuse information corresponding to various library misuse patterns of smart contracts from a database;

[0117] A prompt method construction module 720 for constructing a pre-generated-then-selected prompt method based on the Auto-CoT (Automatic Chain of Thought) and zero-shot-CoT (zero-shot Chain of Thought);

[0118] A preliminary detection module 730 for performing smart contract detection of a large language model using the prompt method according to the library misuse information to obtain a preliminary detection result;

[0119] A model optimization module 740 for using the preliminary detection result as part of the prompt in the next smart contract detection of the large language model, performing iterative feedback optimization on the large language model to obtain a target large language model, and obtaining a first library misuse detection result through the target large language model;

[0120] A static analysis module 750 for creating a static matching BCSSM structure based on code similarity and obtaining a second library misuse detection result through the static matching BCSSM structure;

[0121] A detection module 760 for performing weighted matching on the first library misuse detection result and the second library misuse detection result and outputting a target library misuse detection result.

[0122] In one embodiment, the information extraction module 710 is further configured to find the pattern name, scenario description, and code features of various library misuse patterns from the data, and use the pattern name, scenario description, and code features as library misuse information; the library misuse patterns include invalid wrappers in the library, unhandled exceptions in the library, abnormal library extensions, abnormal usage, incomplete function replacement, abnormal library capacity estimation, and abnormal library usage.

[0123] In one embodiment, the preliminary detection module 730 is further configured to perform role positioning on the large language model to determine the source code of the smart contract to be detected; perform code understanding and logical reasoning matching on the source code of the smart contract to be detected according to the library misuse information to obtain a matching result; based on the matching result, determine the recognition ratio of identifying the library misuse pattern in the source code of the smart contract to be detected as another type of library misuse pattern; obtain a preliminary detection result according to the matching result and the recognition ratio; determine the output format type and apply the format type to the output of the preliminary detection result.

[0124] In one embodiment, the model optimization module 740 is further configured to extract the detection result error rate and false alarm ratio from the preliminary detection results; perform information fusion on the extracted error rate and false alarm ratio, and use the fused information as part of the prompt words in the next large language model intelligent contract detection for cyclic optimization to obtain the evaluation index of the detection results; if the evaluation index reaches the evaluation threshold, stop the iteration, and use the last evaluation result as the output result of the large language model to obtain the target large language model.

[0125] In one embodiment, the static analysis module 750 is further configured to convert the code fragments and input code of the intelligent contract into a TF-IDF matrix by using the TF-IDF technology through static matching of the BCSSM structure; calculate the cosine similarity between the input code and each code fragment according to the TF-IDF matrix to obtain the calculation result; obtain the second library misuse detection result according to the calculation result.

[0126] In one embodiment, the detection module 760 is further configured to compare the first library misuse detection result and the second library misuse detection result to obtain a comparison result; when the comparison result is that the first library misuse detection result and the second library misuse detection result are consistent, output the target library misuse detection result; when the comparison result is that the first library misuse detection result and the second library misuse detection result are inconsistent, weight the first library misuse detection result and the second library misuse detection result.

[0127] In one embodiment, the detection module 760 is further configured to determine the result output format, and parse and output the target library misuse detection result according to the result output format.

[0128] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as Figure 8 shown. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements an intelligent contract library misuse detection method based on a large language model and static analysis. The display screen of the computer device may be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device may be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the outer shell of the computer device, or an external keyboard, a touchpad, or a mouse, etc.

[0129] Those skilled in the art can understand,Figure 8 The structure shown is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0130] In one embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the steps of the misuse detection method for the intelligent contract library based on the large language model and static analysis are implemented.

[0131] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, the steps of the misuse detection method for the intelligent contract library based on the large language model and static analysis are implemented.

[0132] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to the memory, storage, database or other media used in the various embodiments provided in this application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0133] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0134] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several variations and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.

Claims

1. An intelligent contract library misuse detection method based on large language models and static analysis, characterized in that, The method includes: Extracting library misuse information corresponding to various library misuse patterns of smart contracts from a database; Constructing a pre-generated-then-selected prompting method based on Auto-CoT (Automatic Chain of Thought) and zero-shot-CoT (zero-shot Chain of Thought); According to the library misuse information, using the prompting method to perform smart contract detection on a large language model to obtain a preliminary detection result; Using the preliminary detection result as part of the prompt in the next smart contract detection of the large language model, performing iterative feedback optimization on the large language model to obtain a target large language model, and obtaining a first library misuse detection result through the target large language model; Creating a static matching BCSSM (Based on Code Similarity Static Matching) structure and obtaining a second library misuse detection result through the static matching BCSSM structure; Performing weighted matching on the first library misuse detection result and the second library misuse detection result, and outputting a target library misuse detection result.

2. The method for detecting misuse of an intelligent contract library based on a large language model and static analysis according to claim 1, wherein Extracting library misuse information corresponding to various library misuse patterns of smart contracts from a database, including: Searching in the data for the pattern name, scenario description, and code features of various library misuse patterns, and using the pattern name, scenario description, and code features as library misuse information; The library misuse patterns include invalid wrappers in the library, unhandled exceptions in the library, abnormal library extensions, abnormal usage, incomplete function replacement, abnormal library capability estimation, and abnormal library usage.

3. The method for detecting misuse of an intelligent contract library based on a large language model and static analysis according to claim 1, characterized in that Using the prompting method to perform smart contract detection on a large language model to obtain a preliminary detection result, including: Performing role positioning on the large language model to determine the source code of the smart contract to be detected; According to the library misuse information, performing code understanding and logical reasoning matching on the source code of the smart contract to be detected to obtain a matching result; Based on the matching result, determining the recognition ratio of misidentifying the library misuse pattern in the source code of the smart contract to be detected as another type of library misuse pattern; Obtaining a preliminary detection result according to the matching result and the recognition ratio; Determining the output format type and applying the format type to the output of the preliminary detection result.

4. The method for detecting misuse of an intelligent contract library based on a large language model and static analysis according to claim 1, characterized in that Using the preliminary detection result as part of the prompt in the next smart contract detection of the large language model, performing iterative feedback optimization on the large language model to obtain a target large language model, including: Extracting the detection result error rate and false alarm ratio from the preliminary detection result; Performing information fusion on the extracted error rate and false alarm ratio, and using the fused information as part of the prompt in the next smart contract detection of the large language model for cyclic optimization to obtain an evaluation index of the detection result; If the evaluation index reaches the evaluation threshold, stop the iteration, and use the last evaluation result as the output result of the large language model to obtain a target large language model.

5. The misuse detection method for intelligent contract libraries based on large language models and static analysis according to claim 1, characterized in that, Obtaining a second library misuse detection result through the static matching BCSSM structure, including: Using the TF-IDF (Term Frequency-Inverse Document Frequency) technology through the static matching BCSSM structure to convert the code fragments of the smart contract and the input code into a TF-IDF matrix; Calculate the cosine similarity between the input code and each of the code snippets according to the TF-IDF matrix to obtain a calculation result; Obtain a second library misuse detection result according to the calculation result.

6. The method for detecting misuse of an intelligent contract library based on a large language model and static analysis according to claim 1, wherein Perform weighted matching on the first library misuse detection result and the second library misuse detection result, including: Compare the first library misuse detection result and the second library misuse detection result to obtain a comparison result; When the comparison result is that the first library misuse detection result and the second library misuse detection result are consistent, output the target library misuse detection result; When the comparison result is that the first library misuse detection result and the second library misuse detection result are inconsistent, weight the first library misuse detection result and the second library misuse detection result.

7. The method for detecting misuse of an intelligent contract library based on a large language model and static analysis according to claim 1, characterized in that The method further includes: Determine the result output format, and parse and output the target library misuse detection result according to the result output format.

8. An intelligent contract library misuse detection system based on large language models and static analysis, characterized in that, The system includes: An information extraction module, configured to extract library misuse information corresponding to various library misuse patterns of smart contracts from a database; A prompt method construction module, configured to construct a pre-generated-then-selected prompt method based on Auto-CoT (Automatic Chain of Thought) and zero-shot-CoT (zero-shot Chain of Thought); A preliminary detection module, configured to perform smart contract detection on a large language model according to the library misuse information by using the prompt method to obtain a preliminary detection result; A model optimization module, configured to use the preliminary detection result as a part of the prompt in the next smart contract detection of the large language model, perform result iterative feedback optimization on the large language model to obtain a target large language model, and obtain a first library misuse detection result through the target large language model; A static analysis module, configured to create a static matching BCSSM (Block-based Code Similarity Matching) structure based on code similarity, and obtain a second library misuse detection result through the static matching BCSSM structure; A detection module, configured to perform weighted matching on the first library misuse detection result and the second library misuse detection result, and output a target library misuse detection result.

9. The misuse detection system for intelligent contract libraries based on large language models and static analysis according to claim 8, wherein The information extraction module is further configured to find the pattern name, scenario description, and code features of various library misuse patterns from the data, and use the pattern name, scenario description, and code features as library misuse information; the library misuse patterns include invalid wrappers in the library, unhandled exceptions in the library, exception library extensions, abnormal usage, incomplete function replacement, abnormal library capability estimation, and abnormal library usage.

10. The misuse detection system for intelligent contract libraries based on large language models and static analysis according to claim 8, wherein The preliminary detection module is further configured to perform role positioning on the large language model to determine the source code of the smart contract to be detected; Perform code understanding and logical reasoning matching on the source code of the smart contract to be detected according to the library misuse information to obtain a matching result; based on the matching result, determine the recognition ratio of misidentifying the library misuse pattern in the source code of the smart contract to be detected as another type of library misuse pattern; obtain a preliminary detection result according to the matching result and the recognition ratio; determine the output format type, and apply the format type to the output of the preliminary detection result.

Citation Information

Patent Citations

  • Intelligent contract vulnerability detection system construction method based on GPT model

    CN116663010A

  • Source code defect static auditing detection system

    CN116796334A