Enhanced thinking chain prompt software vulnerability detection method based on ChatGPT
By enhancing the thinking chain prompt strategy, combining static program analysis and multi-layer self-verification, the accuracy and false positive rate problems of LLM in software vulnerability detection in the existing technology are solved, and a more efficient vulnerability detection effect is achieved.
Patent Information
- Application Number
- CN202510467894.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-07-11
AI Technical Summary
The existing software vulnerability detection methods based on large language models (LLM) have shortcomings in accuracy and false positive rates, and lack effective thinking chain prompt processes and attention to model randomness and hallucination problems.
The enhanced thinking chain prompt strategy is adopted, abstract syntax tree information is extracted through static program analysis, code features are generated by combining preset prompt words, and vulnerability detection is used using the Few-shot learning strategy, and accuracy is improved through multi-layer self-verification strategies, including the combination of data flow graphs, control flow graphs, program dependency graphs and API call sequences.
It improves the accuracy of software vulnerability detection, reduces the false positive rate, and enhances the concentration and reliability of the model in vulnerability detection tasks.
Smart Images

Figure CN120296747A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and particularly relates to a method for detecting software vulnerabilities based on enhanced chain-of-thought prompting of ChatGPT. Background Art
[0002] Software vulnerabilities are weaknesses or defects in software code that may cause unexpected behaviors such as system crashes, unauthorized access, information leakage, etc., resulting in significant losses. To defend against security vulnerabilities, on the basis of the development of traditional vulnerability detection technologies, deep learning-based vulnerability detection methods have also been continuously proposed. However, such methods also face challenges such as the lack of large-scale and high-quality labeled task-specific datasets. In recent years, large language models (LLMs) such as ChatGPT have received increasing attention due to their excellent capabilities, and they have effectively demonstrated their capabilities in code-related tasks. At the same time, LLMs have shown good results in performing vulnerability detection tasks. Prompt engineering has proven to be promising in exploring the potential of LLMs, but it has not been fully considered in software vulnerability detection. In particular, the chain-of-thought (CoT) prompt has demonstrated impressive potential in various fields. Currently, some research works have proposed methods for designing and improving prompt strategies for the vulnerability detection tasks of large models. For example, the literature "Prompt-Enhanced Software Vulnerability Detection Using ChatGPT" has obtained better results than basic prompts by using more complex prompt strategies involving API call sequences and data flow descriptions. The literature "Large Language Model for Vulnerability Detection: Emerging Results and Future Directions" has a unique advantage over the fine-tuned CodeBERT by designing a variety of enhanced prompts. However, the effectiveness of LLMs in detecting software vulnerabilities has not been fully explored to a large extent. On the one hand, some methods still do not fully consider the characteristics of LLMs and lack attention to the randomness and hallucination problems of the models, resulting in reduced accuracy and increased false positives. On the other hand, although some works have given their respective specific prompt designs, they have not fully combined the advantages of different prompt strategies and formed a complete software vulnerability detection chain-of-thought prompt process. Summary of the Invention
[0003] Object of the Invention: The object of the present invention is to provide a method for detecting software vulnerabilities based on enhanced chain-of-thought prompting of ChatGPT, which enables the model to learn specific reasoning steps through an enhanced chain-of-thought prompting strategy and solves the problems of reducing randomness and hallucination in the vulnerability detection task.
[0004] Technical solution: A method for detecting software vulnerabilities based on enhanced chain-of-thought prompting of ChatGPT of the present invention includes the following steps:
[0005] (1) Perform static program analysis on the function to be tested, extract its Abstract Syntax Tree (AST) information, and input it into the ChatGPT model in combination with preset prompt words to generate corresponding code feature information;
[0006] (2) Based on the enhanced chain-of-thought prompting process, perform vulnerability detection on the function to be tested; including the following steps:
[0007] (21) Use the Few-shot learning strategy to input vulnerability detection examples containing complete reasoning steps into the model. The format of the prompt template for the examples is the same as that of the formal test samples;
[0008] (22) Input the function to be tested into the model, and guide the model to analyze the source code function through the first prompt template;
[0009] (23) Combine the code feature information generated in step (1), and require the model to judge whether there are vulnerabilities in the function to be tested through the second prompt template, and output a binary classification result;
[0010] (24) Execute self-verification strategy 1: Require the model to re-verify the result of step (23) through the third prompt template. If the two responses are inconsistent, repeat step (23) until the results are consistent;
[0011] (25) If step (24) determines that there are vulnerabilities, guide the model to identify potential dangerous statements and their data flows through the fourth prompt template;
[0012] (26) Require the model to locate the vulnerability location and identify the vulnerability type CWE-ID through the fifth prompt template;
[0013] (27) Execute self-verification strategy 2: Require the model to re-verify the vulnerability type in step (26) through the sixth prompt template. If the two responses are inconsistent, require the model to explain the judgment basis and re-analyze the vulnerability semantics and then return to step (26) until the results are consistent.
[0014] Further, in step (1), the code feature information includes Data Flow Graph (DFG), Control Flow Graph (CFG), Program Dependence Graph (PDG), and API call sequence;
[0015] Further, step (1) includes the following steps:
[0016] (11) Use the tree-sitter library of Python to parse the source code of the function to be tested and generate the corresponding abstract syntax tree;
[0017] (12) Generate prompt words for design code features;
[0018] (13) Input the source code, AST, and prompt words into the ChatGPT model to generate DFG, CFG, PDG, and API call sequences.
[0019] Furthermore, the design of each prompt template is as follows: The first prompt template: "Please act as a program analysis assistant and analyze what the function of this source code is?"; The second prompt template: "Please combine the function of the function, data flow diagram, control flow diagram, program dependency graph, and API call sequence obtained from the above analysis to determine whether there are vulnerabilities in the source function, and answer YES or NO"; The third prompt template: "Please review and analyze the above reasoning process and determine again whether there are vulnerabilities in the source code, and answer YES or NO"; The fourth prompt template: "Please identify the potential dangerous statements in the source code and their related data flows"; The fifth prompt template: "Please locate the specific location that causes the vulnerability and identify its CWE-ID"; The sixth prompt template: "Please give the CWE-ID of the vulnerability type of the source code again."
[0020] An enhanced thought chain prompt software vulnerability detection system based on ChatGPT according to the present invention includes:
[0021] Analysis module: used to perform static program analysis on the function to be tested, extract its Abstract Syntax Tree (AST) information, and input it into the ChatGPT model in combination with preset prompt words to generate corresponding code feature information;
[0022] Vulnerability detection module: used to detect vulnerabilities in the function to be tested based on the enhanced thought chain prompt process; including:
[0023] Strategy module: used to input vulnerability detection examples containing complete reasoning steps into the model using the Few-shot learning strategy, and the format of the examples is the same as that of the formal test samples;
[0024] The first prompt module: used to input the function to be tested into the model and guide the model to analyze the function of the source code through the first prompt template;
[0025] The second prompt module: used to combine the code feature information generated by the analysis module and require the model to determine whether there are vulnerabilities in the function to be tested through the second prompt template, and output a binary classification result;
[0026] The first verification strategy module: used to execute self-verification strategy 1: require the model to re-verify the result of the second prompt module through the third prompt template. If the two responses are inconsistent, repeat the second prompt module until the results are consistent;
[0027] Fourth Prompt Module: If the First Verification Policy Module determines that there is a vulnerability, it is used to guide the model to identify potential dangerous statements and their data flows through the Fourth Prompt Template;
[0028] Fifth Prompt Module: It is used to require the model to locate the vulnerability location and identify the vulnerability type CWE-ID through the Fifth Prompt Template;
[0029] Second Verification Policy Module: It is used to execute Self-Verification Policy 2: Through the Sixth Prompt Template, require the model to re-verify the vulnerability type of the Fifth Prompt Module. If the two responses are inconsistent, require the model to explain the basis for judgment and re-analyze the vulnerability semantics and then return to the Fifth Prompt Module until the results are consistent.
[0030] Further, in the analysis module, the code feature information includes Data Flow Graph (DFG), Control Flow Graph (CFG), Program Dependence Graph (PDG), and API call sequence;
[0031] Further, the analysis module includes the following steps:
[0032] (11) Use the tree-sitter library of Python to parse the source code of the function to be tested and generate the corresponding Abstract Syntax Tree;
[0033] (12) Design code feature generation prompt words;
[0034] (13) Input the source code, AST, and prompt words into the ChatGPT model to generate DFG, CFG, PDG, and API call sequence.
[0035] Further, in the vulnerability detection module, each prompt template is designed as follows: First Prompt Template: "Please act as a program analysis assistant and analyze what the function of this source code is?"; Second Prompt Template: "Please combine the function of the function, data flow graph, control flow graph, program dependence graph, and API call sequence obtained from the above analysis to determine whether the source function has a vulnerability, and answer YES or NO"; Third Prompt Template: "Please review and analyze the above reasoning process and determine again whether the source code has a vulnerability, and answer YES or NO"; Fourth Prompt Template: "Please identify the potential dangerous statements in the source code and their related data flows"; Fifth Prompt Template: "Please locate the specific location that causes the vulnerability and identify its CWE-ID"; Sixth Prompt Template: "Please give the CWE-ID of the vulnerability type of the source code again."
[0036] An electronic device according to the present invention includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is loaded into the processor, it implements any one of the enhanced thinking chain prompt software vulnerability detection methods based on ChatGPT.
[0037] A storage medium according to the present invention, wherein the storage medium stores a computer program, and when the computer program is executed by a processor, it implements any one of the method for detecting software vulnerabilities based on enhanced chain-of-thought prompting of ChatGPT.
[0038] Advantages: Compared with the prior art, the present invention has the following remarkable advantages: The present invention uses the most advanced model and combines effective prompting strategies in existing work to improve the prompting template, while taking into account the characteristics of large language models, designs an improved chain-of-thought prompting process for vulnerability detection tasks, enabling the model to focus on the task itself, reducing randomness, improving accuracy in binary classification tasks and vulnerability type identification, and reducing false positive rates. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 is a flowchart of the present invention;
[0040] Figure 2 is the sample to be tested in the implementation example of the present invention;
[0041] Figure 3 is the AST of the sample to be tested in the implementation example of the present invention;
[0042] Figure 4 is based on the present invention Figure 2 and Figure 3 is a vulnerability detection implementation example. DETAILED DESCRIPTION OF THE INVENTION
[0043] The technical solution of the present invention will be further described below with reference to the accompanying drawings.
[0044] As Figure 1 shown, the embodiment of the present invention provides a method for detecting software vulnerabilities based on enhanced chain-of-thought prompting of ChatGPT, comprising the following steps:
[0045] 1) Extract abstract syntax tree information for the function to be tested, and input the prompt words in combination with the model to generate corresponding code feature information, including the following steps:
[0046] 1.1) Use the tree-sitter library in python to perform static analysis on the function to be tested, parse the source code and generate the corresponding abstract syntax tree (AST).
[0047] 1.2) The design of the prompt words for obtaining code features is as follows:
[0048] "The following is a C++ / C source code and its generated abstract syntax tree. Please act as a software security expert and, in combination with the abstract syntax tree information, seriously and rigorously generate data flow diagrams, control flow diagrams, program dependency diagrams, and API call sequence information for this source code:"
[0049] Source code: [code]
[0050] Abstract Syntax Tree: [AST]
[0051] 1.3) Combine the features of the function under test and its abstract syntax tree information to obtain the prompt words and input them into the model together, and require it to respond to generate corresponding code feature information such as control flow graph CFG, data flow graph DFG, program dependence graph PDG, API call sequence, etc.
[0052] 2) In the model, according to the specific chain-of-thought prompting steps, use the prompt template to perform vulnerability detection tasks on the function under test, including the following steps:
[0053] 2.1) Use the Few-shot learning strategy to perform few-shot learning for the model, and input the complete examples of the reasoning steps in the vulnerability detection task into the model. The examples can be referred to Figure 4 。
[0054] 2.2) Input the function under test and the prompt words into the model, and require the model to analyze its function. The prompt template is designed as follows:
[0055] "Please act as a program analysis assistant and analyze what the function of this source code is?"
[0056] 2.3) Enter the binary classification task. The prompt template is designed as follows:
[0057] "Please act as an expert in the field of software vulnerability detection. Combining the function of the function, data flow graph, control flow graph, program dependence graph, and API call sequence obtained from the above analysis, please judge whether the source function has vulnerabilities. Please answer YES or NO."
[0058] 2.4) Perform self-verification strategy 1. The prompt template is designed as follows:
[0059] "Please review and analyze the above reasoning process and judge again whether the source code has vulnerabilities. Please answer YES or NO."
[0060] 2.5) If the responses under self-verification strategy 1 are inconsistent, re-enter the binary classification task until the responses are consistent, and then use its response as the result of whether the source code has vulnerabilities.
[0061] 2.6) For the tested samples classified as having vulnerabilities in 2.5), the model needs to identify the potential dangerous statements and data flows in the function under test. The prompt template is designed as follows:
[0062] "Please act as an expert in the field of software vulnerability detection. Combining the code feature information in the above analysis, please identify the potential dangerous statements and related data flows in the source code."
[0063] 2.7) Locate the specific vulnerability location in the source code and identify the vulnerability type. The prompt template is designed as follows:
[0064] "Please act as a software vulnerability analysis expert. Based on the above analysis, first locate the specific location in the source code that causes the vulnerability. And to better fix the vulnerability in the future, please further identify the vulnerability type it belongs to."
[0065] 2.8) Perform self-verification strategy 2. The prompt template is designed as follows:
[0066] "Please review and analyze the above reasoning process and give the vulnerability type (CWE-ID) of the source code again."
[0067] 2.9) If the responses are consistent under self-verification strategy 2, use this CWE-ID as the vulnerability type. If the responses are inconsistent, the model needs to explain its specific judgment basis (guide it to analyze the problematic reasoning steps) and re-analyze the vulnerability semantics (vulnerability-related statements and their data flows). The prompt template is designed as follows:
[0068] "Compare the different vulnerability type classification results from the last time. Please explain the specific judgment basis and re-analyze the statements related to the vulnerability and their data flows."
[0069] 2.10) After the link of re-analyzing the vulnerability semantics, go back to step 2.7). The model needs to locate the specific vulnerability location of the sample to be tested, and then further identify the vulnerability type. Go back to self-verification strategy 2 until the responses are consistent, and then use the final CWE-ID as its vulnerability type.
Claims
1. An enhanced chain-of-thought prompting software vulnerability detection method based on ChatGPT, characterized in that, It includes the following steps: (1) Perform static program analysis on the function under test, extract its Abstract Syntax Tree (AST) information, and input it into the ChatGPT model together with preset prompt words to generate corresponding code feature information; (2) Based on an enhanced chain-of-thought prompting process, perform vulnerability detection on the function under test; it includes the following steps: (21) Use the Few-shot learning strategy to input vulnerability detection examples containing complete reasoning steps into the model. The format of the prompt template for the examples is the same as that of the formal test samples; (22) Input the function under test into the model and guide the model to analyze the function of the source code through the first prompt template; (23) Combine the code feature information generated in step (1) and require the model to judge whether the function under test has vulnerabilities through the second prompt template, and output a binary classification result; (24) Execute self-verification strategy 1: Require the model to re-verify the result of step (23) through the third prompt template. If the two responses are inconsistent, repeat step (23) until the results are consistent; (25) If step (24) determines that there are vulnerabilities, guide the model to identify potential dangerous statements and their data flows through the fourth prompt template; (26) Require the model to locate the vulnerability location and identify the vulnerability type CWE-ID through the fifth prompt template; (27) Execute self-verification strategy 2: Require the model to re-verify the vulnerability type in step (26) through the sixth prompt template. If the two responses are inconsistent, require the model to explain the basis for judgment and re-analyze the vulnerability semantics, and then return to step (26) until the results are consistent.
2. The method for detecting software vulnerabilities by enhancing the chain of thought prompt based on ChatGPT according to claim 1, wherein In step (1), the code feature information includes Data Flow Graph (DFG), Control Flow Graph (CFG), Program Dependence Graph (PDG), and API call sequence.
3. The method for detecting software vulnerabilities by enhancing the chain of thought prompt based on ChatGPT according to claim 1, wherein In step (1), it includes the following steps: (11) Use the tree-sitter library in Python to parse the source code of the function under test to generate the corresponding abstract syntax tree; (12) Design prompt words for generating code features; (13) Input the source code, AST, and prompt words into the ChatGPT model to generate DFG, CFG, PDG, and API call sequence.
4. The method for detecting software vulnerabilities by enhancing the thought chain prompt based on ChatGPT according to claim 1, wherein, The design of each prompt template is as follows: The first prompt template: "Please act as a program analysis assistant and analyze what the function of this source code is?”; The second prompt template: "Please combine the function of the function obtained from the above analysis, Data Flow Graph, Control Flow Graph, Program Dependence Graph, and API call sequence to judge whether the source function has vulnerabilities, and answer YES or NO”; The third prompt template: "Please review and analyze the above reasoning process and judge again whether the source code has vulnerabilities, and answer YES or NO”; The fourth prompt template: "Please identify the potential dangerous statements in the source code and their related data flows”; The fifth prompt template: "Please locate the specific location that causes the vulnerability and identify its CWE-ID”; The sixth prompt template: "Please give the CWE-ID of the vulnerability of the source code again.” 5. An enhanced chain-of-thought prompting software vulnerability detection system based on ChatGPT, characterized in that, It includes: Analysis module: Used to perform static program analysis on the function under test, extract its Abstract Syntax Tree (AST) information, and input it into the ChatGPT model together with preset prompt words to generate corresponding code feature information; Vulnerability Detection Module: Used to detect vulnerabilities in the function under test based on an enhanced chain-of-thought prompting process; including: Policy Module: Used to input vulnerability detection examples containing complete reasoning steps into the model using the Few-shot learning strategy, and the format of the examples is the same as the prompting template of the formal test samples; First Prompting Module: Used to input the function under test into the model and guide the model to analyze the source code function through the first prompting template; Second Prompting Module: Used to combine the code feature information generated by the analysis module and require the model to judge whether there are vulnerabilities in the function under test through the second prompting template, and output a binary classification result; First Verification Policy Module: Used to execute Self-Verification Policy 1: Require the model to re-verify the result of the second prompting module through the third prompting template. If the two responses are inconsistent, repeat the second prompting module until the results are consistent; Fourth Prompting Module: Used to, if the first verification policy module determines that there is a vulnerability, guide the model to identify potential dangerous statements and their data flows through the fourth prompting template; Fifth Prompting Module: Used to require the model to locate the vulnerability location and identify the vulnerability type CWE-ID through the fifth prompting template; Second Verification Policy Module: Used to execute Self-Verification Policy 2: Require the model to re-verify the vulnerability type of the fifth prompting module through the sixth prompting template. If the two responses are inconsistent, require the model to explain the basis for judgment and re-analyze the vulnerability semantics and then return to the fifth prompting module until the results are consistent.
6. The enhanced thought chain prompting software vulnerability detection system based on ChatGPT according to claim 5, characterized in that, In the analysis module, the code feature information includes Data Flow Graph (DFG), Control Flow Graph (CFG), Program Dependence Graph (PDG), and API call sequence.
7. The method for detecting software vulnerabilities by enhancing the chain of thought prompt based on ChatGPT according to claim 6, wherein, In the analysis module, the following steps are included: (11) Use the tree-sitter library of Python to parse the source code of the function under test and generate the corresponding Abstract Syntax Tree; (12) Design code feature generation prompting words; (13) Input the source code, AST, and prompting words into the ChatGPT model to generate DFG, CFG, PDG, and API call sequence.
8. The software vulnerability detection system based on Chain of Thought prompting of ChatGPT according to claim 5, characterized in that, In the vulnerability detection module, the design of each prompting template is as follows: First Prompting Template: "Please act as a program analysis assistant and analyze what the function of this source code is?”; Second Prompting Template: "Please combine the function of the source function, data flow graph, control flow graph, program dependence graph, and API call sequence obtained from the above analysis to judge whether there are vulnerabilities in the source function, and answer YES or NO”; Third Prompting Template: "Please review and analyze the above reasoning process and judge again whether there are vulnerabilities in the source code, and answer YES or NO”. Fourth Prompting Template: "Please identify the potential dangerous statements in the source code and their relevant data flows”; Fifth Prompting Template: "Please locate the specific location that causes the vulnerability and identify its CWE-ID”; Sixth Prompting Template: "Please give the CWE-ID of the vulnerability type of the source code again.” 9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the computer program is loaded into the processor, it implements a ChatGPT-based enhanced chain-of-thought prompting software vulnerability detection method according to any one of claims 1-4.
10. A storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements a method for detecting software vulnerabilities based on enhanced thought chain prompting of ChatGPT according to any one of claims 1-4.
Citation Information
Cited By
Code defect reasoning instruction template generation method and system for large language model
CN120631737A
Code scanning analysis method and system based on large model
CN120893046A