Web application program bug PoC generation method based on large language model
By applying a large language model in the generation of vulnerability PoCs for web application, the problem of inefficient vulnerability identification and PoC generation in traditional methods is solved, and automated vulnerability understanding and effective PoC generation are achieved.
Patent Information
- Application Number
- CN202510184996.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-02-19
AI Technical Summary
The prior art has limitations in the generation of web application vulnerability PoCs, including the inability to handle situations outside the rules, the lack of ability to automatically update vulnerability knowledge, and the insufficient semantic constraint solution capabilities, resulting in inefficient PoC generation.
Using a method based on the large language model (LLM), we build a vulnerability benchmark set, extract vulnerability information and navigation information at the function level, and split the PoC generation process into multiple subtasks. We use the thinking chain (CoT) to guide LLM to gradually reason, and generate effective attack payloads and PoCs.
It realizes automatic identification and understanding of vulnerabilities, generates effective attack payloads and PoCs, significantly improving the efficiency and accuracy of vulnerable PoC generation, and overcoming the limitations of traditional methods.
Smart Images

Figure CN120030553A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of software security, and in particular relates to a Web application vulnerability PoC generation method based on a large language model. Background Art
[0002] In software security, vulnerabilities can be exploited to undermine the confidentiality, integrity, and availability of information. As the main carrier for providing information and services on the modern Internet, web applications are often the first choice of attackers and are often threatened by cross-site scripting (XSS) and SQL injection (SQLi) vulnerabilities. Proof of concept (PoC) provides a way to prove the feasibility of a specific vulnerability. However, according to statistics, of the approximately 250,000 vulnerabilities disclosed by the National Vulnerability Database (NVD) of the United States, only 34% have links that may point to PoC. In the absence of PoC, security experts usually need to spend a lot of manpower and time resources to manually infer and build PoC again, which seriously hinders the process of mitigating Web vulnerabilities.
[0003] Traditional XSS and SQL injection PoC generation usually relies on techniques such as static analysis, dynamic analysis, or symbolic execution. For example, static analysis is combined with dynamic analysis, and attack templates are defined at the same time to automatically generate PoC for identified vulnerabilities. Application execution paths are systematically inferred through static analysis, and the application execution flow is directed to vulnerable sinks to generate PoCs across multiple HTTP requests. Symbolic sockets are used for taint propagation, while concrete symbolic execution is used to solve constraints. Symbolic execution and dynamic taint tracking are combined to generate PoCs for vulnerabilities using predefined attack payloads. A goal-oriented model checking system is designed to automatically generate PoCs for XSS and SQL injection vulnerabilities using model checking techniques by accepting vulnerability specifications.
[0004] However, existing PoC generation works have many limitations. First, they are designed to identify and verify new vulnerabilities on a large scale through rules, and cannot handle situations outside the rules, and false negatives are acceptable. For PoC generation tasks targeting specific n-day vulnerabilities, more flexible vulnerability understanding and constraint solving capabilities are required. Second, traditional PoC generation works mainly rely on predefined rules and heuristic methods to model vulnerabilities, including source, sink, sanitization, and attack payload. Therefore, they lack the ability to automatically update vulnerability knowledge, and it is difficult to obtain vulnerability information from disclosed vulnerability descriptions and patches. Finally, modern web applications usually have complex page dependencies and complex data transfer. However, traditional PoC generation works are insufficient in solving constraints related to code semantics. When dealing with complex constraints related to program logic and semantics, the performance is poor and effective attack payloads cannot be generated.
[0005] In order to overcome the above limitations, the present invention proposes a Web application vulnerability PoC generation method based on a large language model. Summary of the invention
[0006] In view of the problems that traditional PoC generation is limited to specific rule vulnerabilities, cannot automatically update vulnerability knowledge, and has insufficient semantic constraint solving capabilities, the present invention proposes a Web application vulnerability PoC generation method based on a large language model. By relying on the security domain knowledge embedded in LLM, the present invention can automatically identify vulnerabilities from vulnerability descriptions and codes, overcoming the limitations of manually predefined rules; at the same time, relying on the code understanding ability of LLM, the present invention can process semantic and grammatical constraints, generate effective attack payloads, and finally construct a usable PoC.
[0007] The present invention is achieved through the following technical solutions:
[0008] A method for generating PoC of a Web application vulnerability based on a large language model comprises the following steps:
[0009] Step 1: Build a vulnerability benchmark set. The vulnerability information includes the vulnerability CVE number, vulnerability description, CVSS3 score, patch link, patch modification information corresponding to the patch link, CWE number, and affected software configuration.
[0010] Step 2: For each vulnerability in the vulnerability benchmark set constructed in step 1, extract the function-level vulnerability information and navigation information as the input for subsequent PoC generation;
[0011] Step 3: Split the PoC generation process into multiple subtasks;
[0012] Step 4: Based on the subtasks divided in Step 3, use the Chain of Thought (CoT) to guide the LLM for step-by-step reasoning and output the content of each subtask.
[0013] In the above technical solution, Step 1 includes the following steps:
[0014] Step 1.1: Obtain all publicly available vulnerability information by accessing NVD, including vulnerability CVE numbers, vulnerability descriptions, CVSS3 scores, patch links, CWE numbers, and affected software configurations, and obtain the corresponding patch modification information through the patch links.
[0015] Step 1.2: Filter vulnerabilities with a release date after 2018, a CVSS3 score greater than 7.0, and CWE numbers CWE-79 (XSS) and CWE-89 (SQLi) based on the CVE number; then, randomly select one vulnerability from them, build a local vulnerability environment according to the affected software configuration, and reproduce it. If the reproduction is successful, add the vulnerability to the vulnerability benchmark set; otherwise, discard the vulnerability and perform random sampling of vulnerabilities again; repeat the above process until the number of the vulnerability benchmark set reaches the set number.
[0016] In the above technical solution, Step 2 includes the following steps:
[0017] Step 2.1, Extract vulnerability information
[0018] First, design prompt words to prompt the LLM to extract key information from the vulnerability description corresponding to each CVE number for each vulnerability, including variables and the file names of the files where they are located. Consider the variables extracted from the vulnerability description as potential vulnerability variables, and also consider all variables involved in the patch information corresponding to the patch link of the CVE number as potential vulnerability variables;
[0019] Then, parse the vulnerability application source code into a code property graph, track the assignment statements and data dependency statements of the above potential vulnerability variables in the propagation path until reaching the sink; consider the variables in the sink as the final vulnerability variables, and perform backward taint analysis until reaching the attacker-controllable source point source; extract all relevant functions in the path. If the code does not exist in a specific function, extract the content of the entire file;
[0020] Step 2.2, Extract navigation information
[0021] Start from the vulnerability information module, perform backward analysis, locate the first publicly accessible page, and extract the navigation information in the path.
[0022] In the above technical solution, in step 3, the generation process of PoC is divided into 14 subtasks, including: ① identifying the sink, ② identifying the vulnerable variable, ③ identifying the source, ④ identifying the data flow constraints encountered in the process of propagating the vulnerable variable from the source to the sink, ⑤ identifying the control flow constraints encountered in the process of propagating the vulnerable variable from the source to the sink, ⑥ identifying the grammatical constraints at the sink, ⑦ solving the attack payload for generating the vulnerable variable, ⑧ identifying the file navigation chain, ⑨ identifying the file navigation code, ⑩ identifying the path constraint code, Solution path constraint variables and values, Determine the request parameters required to build PoC, Determine the request method required to build the PoC and Determine the request URL required to build the PoC. All these subtask contents need to be identified or solved by LLM.
[0023] In the above technical solution, the specific process of step 4 is as follows:
[0024] First, identify the sink, vulnerable variables, and source;
[0025] Then, the data flow constraints, control flow constraints, and grammatical constraints at the sink encountered during the propagation of the vulnerability variable from the source to the sink are analyzed;
[0026] Then, the attack payload of the vulnerability variable is generated according to the above three constraint information;
[0027] Then, the global path of the application execution flow is analyzed to identify the file navigation chain and file navigation code;
[0028] Then, the path constraint code that each file needs to satisfy in order to reach the navigation code is analyzed, and based on the path constraint code, the path constraint variables and values are solved;
[0029] Finally, determine the request parameters, request method, and request URL required to build the PoC.
[0030] In the above technical solution, prompt words are designed for each subtask, and the content includes four parts: ① input information, ② subtask explanation, ③ few sample prompt cases and ④ subtask output template.
[0031] In the above technical solution, in step 4, an attack payload verifier and an execution trajectory verifier are designed to collect feedback information and combine it with CoT to optimize PoC generation;
[0032] For the attack payload verifier, the identified sink, vulnerability variable, source, data flow constraint, control flow constraint, syntax constraint and vulnerability information are collected, and prompt words are designed to let LLM synthesize a local vulnerability execution environment; at the same time, in order to debug the attack payload, LLM inserts the code that collects the three constraint information in the local execution environment; if the attack payload generation fails, the actual manifestation of the three constraints is fed back to LLM, allowing LLM to regenerate the attack payload;
[0033] For the execution trajectory verifier, run the generated PoC and obtain the log file of the program execution trajectory; if the PoC fails to execute, extract the function call and file call information in the log file, compare it with the identified file navigation chain nodes, find out the navigation file nodes that have not been triggered, and feed it back to LLM to let it re-solve the path constraint variables and values and rebuild the PoC.
[0034] The advantages and beneficial effects of the present invention are:
[0035] Firstly, the present invention splits the PoC generation subtasks of XSS vulnerabilities and SQL injection vulnerabilities to form a standardized PoC generation method; secondly, the present invention constructs a PoC generation prompt word system based on LLM to significantly improve the PoC generation effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 It is a flow chart of constructing a vulnerability benchmark set according to the present invention.
[0037] Figure 2 It is the overall flow chart of the present invention.
[0038] For ordinary technicians in this field, other relevant drawings can be obtained based on the above drawings without any creative work. DETAILED DESCRIPTION
[0039] In order to enable those skilled in the art to better understand the technical solution of the present invention, the technical solution of the present invention is further described below in conjunction with specific embodiments.
[0040] A method for generating PoC of a Web application vulnerability based on a large language model comprises the following steps:
[0041] Step 1: Build a vulnerability benchmark set.
[0042] Step 1.1: Access the National Vulnerability Database (NVD) to obtain all public vulnerability information, including the CVE number, vulnerability description, CVSS3 score, patch link, CWE number, and affected software configuration, and obtain the corresponding patch modification information through the patch link.
[0043] CVE number: refers to a unique identifier assigned to each vulnerability (such as CVE-2021-44228).
[0044] Vulnerability description: including the impact of the vulnerability, attack vector, affected software / hardware, etc.
[0045] CVSS3 score: It is the severity score of the vulnerability, which uses CVSS (Common Vulnerability Scoring System) to quantify the risk level of the vulnerability.
[0046] CWE number: Common Weakness Enumeration. It is a list maintained by the community that is used to systematically identify, describe, and classify security weaknesses (i.e. potential vulnerabilities) in software, hardware, and other systems. The CWE number is a unique identifier for each security weakness, helping developers, security researchers, and tools to uniformly describe security issues.
[0047] Step 1.2: Based on the CVE number, filter out vulnerabilities with a release date after 2018, a CVSS3 score greater than 7.0, and CWE numbers CWE-79 (XSS) and CWE-89 (SQLi). Then, randomly select a vulnerability from them, and build a local vulnerability environment based on the affected software configuration to reproduce it. If the reproduction is successful, add the vulnerability to the vulnerability benchmark set. Otherwise, discard the vulnerability and re-sample the vulnerabilities randomly. Repeat the above process until the number of vulnerability benchmark sets reaches the set number (the number set in this embodiment is 100).
[0048] Step 2: For each vulnerability in the vulnerability benchmark set constructed in step 1, extract the function-level vulnerability information and navigation information as input for subsequent PoC generation.
[0049] Step 2.1, extract vulnerability information.
[0050] First, a prompt word is designed. For each CVE numbered vulnerability prompt, the Large Language Model (LLM) extracts key information from the vulnerability description corresponding to the CVE number, including variables and the file names of the files they are located in. The variables extracted from the vulnerability description are regarded as potential vulnerability variables, and all variables involved in the patch information corresponding to the patch link corresponding to the CVE number are also regarded as potential vulnerability variables.
[0051] Then, with the help of tools such as phpjoern, the vulnerable application source code is parsed into a code property graph (CPG), and combined with the Neo4j database for query, the assignment statements and data dependency statements of the above-mentioned potential vulnerable variables in the propagation path are tracked until the sink is reached; the variables in the sink are regarded as the final vulnerable variables, and backward taint analysis is performed until the source controllable by the attacker is reached; all related functions in the path are extracted, and if the code does not exist in a specific function, the contents of the entire file are extracted.
[0052] Step 2.2, extract navigation information.
[0053] Starting from the vulnerability information module, we perform a backward analysis to locate the first publicly accessible page and extract the navigation information in the path.
[0054] Specifically, starting from the vulnerability information module, add it to the navigation information and check whether it is publicly accessible. If the conditions are met, directly output the navigation information; otherwise, combine the HTML parser and CPG to locate the code responsible for navigation, such as redirection, form, include statement, and function call, etc., identify the files that can reach the vulnerability information module, add them to the navigation information, and continue to check their accessibility; repeat the above process until the first publicly accessible page is identified. During the path tracing process, record all navigation-related functions.
[0055] Step 3: Split the PoC generation process into multiple subtasks.
[0056] PoC generation is a complex task. In order to improve the accuracy of LLM generation, the complete PoC generation process is divided into multiple subtasks, which are solved one by one by LLM and finally spliced into a complete PoC.
[0057] The present invention divides the generation process of PoC into 14 subtasks, including ① identifying sink, ② identifying vulnerability variables, ③ identifying source, ④ identifying data flow constraints encountered in the process of vulnerability variables propagating from source to sink, ⑤ identifying control flow constraints encountered in the process of vulnerability variables propagating from source to sink, ⑥ identifying grammatical constraints at sink, ⑦ solving attack payloads for generating vulnerability variables, ⑧ identifying file navigation chains, ⑨ identifying file navigation codes, ⑩ identifying path constraint codes, Solution path constraint variables and values, Determine the request parameters required to build PoC, Determine the request method required to build the PoC and Determine the request URL required to build the PoC. All these subtask contents need to be identified or solved by LLM.
[0058] Step 4: Based on the subtasks divided in step 3, the Chain of Thought (CoT) is used to guide LLM to perform step-by-step reasoning and output the content of each subtask.
[0059] The specific process is as follows:
[0060] First, identify the sink, vulnerable variables, and source.
[0061] Then, the data flow constraints, control flow constraints, and grammatical constraints at the sink encountered during the propagation of the vulnerable variable from the source to the sink are analyzed.
[0062] Then, the attack payload of the vulnerability variable is generated according to the above three constraint information.
[0063] Then, the global path of the application execution flow is analyzed, i.e., the file navigation chains and file navigation codes are identified.
[0064] Then, analyze the path constraint code that each file needs to satisfy to reach the navigation code. Based on the path constraint code, solve the path constraint variables and values.
[0065] Finally, determine the request parameters, request method, and request URL required to build the PoC.
[0066] Furthermore, in order to improve the information recognition and extraction capabilities of LLM, samples of each subtask are collected from related works, github, Google and other sources to perform few-shot prompts. Prompt words are designed for each subtask, and the content includes four parts: ① based on what input information, ② subtask explanation, ③ few-shot prompt examples, and ④ subtask output template.
[0067] Furthermore, in order to reduce the hallucination phenomenon of LLM, two verifiers are designed: attack payload verifier and execution trace verifier, which are used to collect feedback information and combine it with CoT to optimize PoC generation.
[0068] For the attack payload verifier, the identified sink, vulnerability variable, source, data flow constraint, control flow constraint, syntax constraint and vulnerability information are collected, and prompt words are designed to let LLM synthesize a local vulnerability execution environment. At the same time, in order to debug the attack payload, LLM is required to insert the code that collects the three constraint information in the local execution environment. If the attack payload generation fails, the actual manifestation of the three constraints is fed back to LLM, allowing LLM to regenerate the attack payload.
[0069] For the execution trace verifier, run the generated PoC and obtain the log file of the program execution trace. If the PoC fails to execute, extract the function call and file call information in the log file, compare it with the identified file navigation chain nodes, and find the navigation file nodes that have not been triggered. Feed it back to LLM to let it re-solve the path constraint variables and values and rebuild the PoC. The maximum number of feedbacks for both verifiers is set to 5.
[0070] The present invention is described above by way of example. It should be noted that, without departing from the core of the present invention, any simple deformation, modification or other equivalent replacement that can be made by those skilled in the art without inventive effort falls within the protection scope of the present invention.
Claims
1. A method for generating PoC of Web application vulnerability based on large language model, characterized in that: The following steps are involved: Step 1: Build a vulnerability benchmark set. The vulnerability information includes the vulnerability CVE number, vulnerability description, CVSS3 score, patch link, patch modification information corresponding to the patch link, CWE number, and affected software configuration. Step 2: For each vulnerability in the vulnerability benchmark set constructed in step 1, extract the function-level vulnerability information and navigation information as the input for subsequent PoC generation; Step 3: Split the PoC generation process into multiple subtasks; Step 4: Based on the subtasks divided in step 3, CoT is used to guide LLM to perform step-by-step reasoning and output the content of each subtask.
2. The method for generating PoC of a Web application vulnerability based on a large language model according to claim 1, characterized in that: Step 1 includes the following steps: Step 1.1: Access NVD to obtain all public vulnerability information, including vulnerability CVE number, vulnerability description, CVSS3 score, patch link, CWE number, and affected software configuration, and obtain the corresponding patch modification information through the patch link. Step 1.2: Based on the CVE number, filter out vulnerabilities with a release date after 2018, a CVSS3 score greater than 7.0, and CWE numbers CWE-79 and CWE-89; then, randomly select a vulnerability from them, and build a local vulnerability environment based on the affected software configuration to reproduce it. If the reproduction is successful, add the vulnerability to the vulnerability benchmark set; otherwise, discard the vulnerability and re-sample vulnerabilities randomly; repeat the above process until the number of vulnerability benchmark sets reaches the set number.
3. The method for generating PoC of a Web application vulnerability based on a large language model according to claim 1, characterized in that: Step 2 includes the following steps: Step 2.1, extract vulnerability information First, the prompt words are designed. For each CVE number, the LLM extracts key information from the vulnerability description corresponding to the CVE number, including variables and the file names of the files they are located in. The variables extracted from the vulnerability description are considered as potential vulnerability variables, and all variables involved in the patch information corresponding to the patch link corresponding to the CVE number are also considered as potential vulnerability variables. Then, the vulnerable application source code is parsed into a code property graph, and the assignment statements and data dependency statements of the above potential vulnerable variables in the propagation path are tracked until the sink is reached. The variables in the sink are regarded as the final vulnerable variables, and backward taint analysis is performed until the source point controlled by the attacker is reached. All related functions in the path are extracted, and if the code does not exist in a specific function, the contents of the entire file are extracted. Step 2.2, extract navigation information Starting from the vulnerability information module, we perform a backward analysis to locate the first publicly accessible page and extract the navigation information in the path.
4. The method for generating PoC of a Web application vulnerability based on a large language model according to claim 1, characterized in that: In step 3, the generation process of PoC is divided into 14 subtasks, including: ① identifying the sink, ② identifying the vulnerable variable, ③ identifying the source, ④ identifying the data flow constraints encountered in the process of propagating the vulnerable variable from the source to the sink, ⑤ identifying the control flow constraints encountered in the process of propagating the vulnerable variable from the source to the sink, ⑥ identifying the grammatical constraints at the sink, ⑦ solving the attack payload for generating the vulnerable variable, ⑧ identifying the file navigation chain, ⑨ identifying the file navigation code, ⑩ identifying the path constraint code, Solve path constraint variables and values, Determine the request parameters needed to build the PoC, Determine the request method needed to build the PoC and Determine the request URL required to build the PoC. All of these subtasks require LLM to identify or solve.
5. The method for generating PoC of a Web application vulnerability based on a large language model according to claim 1, characterized in that: Step 4: The specific process is as follows: First, identify the sink, vulnerable variables, and source; Then, the data flow constraints, control flow constraints, and grammatical constraints at the sink encountered during the propagation of the vulnerability variable from the source to the sink are analyzed; Then, the attack payload of the vulnerability variable is generated according to the above three constraint information; Then, the global path of the application execution flow is analyzed to identify the file navigation chain and file navigation code; Then, the path constraint code that each file needs to satisfy in order to reach the navigation code is analyzed, and based on the path constraint code, the path constraint variables and values are solved; Finally, determine the request parameters, request method, and request URL required to build the PoC.
6. The method for generating PoC of a Web application vulnerability based on a large language model according to claim 1, characterized in that: Prompt words are designed for each subtask, and the content includes four parts: ① input information, ② subtask explanation, ③ few-sample prompt cases, and ④ subtask output template.
7. The method for generating PoC of a Web application vulnerability based on a large language model according to claim 1, characterized in that: In step 4, the attack payload verifier and execution trajectory verifier are designed to collect feedback information and combine it with CoT to optimize PoC generation; For the attack payload verifier, the identified sink, vulnerability variable, source, data flow constraint, control flow constraint, syntax constraint and vulnerability information are collected, and prompt words are designed to let LLM synthesize a local vulnerability execution environment; at the same time, in order to debug the attack payload, LLM inserts the code that collects the three constraint information in the local execution environment; if the attack payload generation fails, the actual manifestation of the three constraints is fed back to LLM, allowing LLM to regenerate the attack payload; For the execution trace verifier, run the generated PoC to obtain the log file of the program execution trace; If the PoC fails to execute, the function call and file call information in the log file is extracted and compared with the identified file navigation chain nodes to find out the navigation file nodes that have not been triggered. The nodes are fed back to the LLM to let it re-solve the path constraint variables and values and rebuild the PoC.
Citation Information
Patent Citations
Static taint analysis and symbolic execution-based Android application vulnerability discovery method
CN106709356A
Sequence decision and probability sampling guided vulnerability detection statement level interpretable method
CN118690346A
LLM Agent-based Web application vulnerability dynamic detection method and system
CN118761060A
Automated Identification Of Vulnerable Software Components
US20240427902A1
Cited By
Supply chain cross-packet vulnerability detection method and device, equipment and storage medium
CN121051762A