Automatic cross-project vulnerability verification method and device based on large language model

By adopting an automated cross-project vulnerability verification method based on a large language model, the problem of triggerability judgment in cross-project vulnerability verification is solved, and efficient and automated vulnerability verification is achieved, which is applicable to various types of software projects.

CN120909906AActive Publication Date: 2025-11-07XIAMEN UNIV OF TECH

Patent Information

Application Number
CN202511431439.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-09
Publication Date
2025-11-07
Estimated Expiration
2045-10-09

AI Technical Summary

Technical Problem

Existing technologies are insufficient to effectively determine whether vulnerable code is actually triggerable in cross-project vulnerability verification, resulting in high false positive rates and low remediation efficiency. Furthermore, existing methods have significant limitations in cross-project migration and format adaptation.

Method used

An automated cross-project vulnerability verification method based on a large language model is adopted. By constructing prompt words, generating candidate parameter combinations and cross-format seeds, and combining RAG technology and ReAct agent, cross-project vulnerability verification is achieved.

Benefits of technology

It improves the success rate and efficiency of cross-project vulnerability verification, reduces manual intervention, and is suitable for various types of software projects, especially performing well in cross-library and cross-framework vulnerability propagation verification scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120909906A_ABST
    Figure CN120909906A_ABST
Patent Text Reader

Abstract

The invention provides an automatic cross-project vulnerability verification method and device based on a large language model, relates to the technical field of vulnerability verification, and aims to solve the problem that whether an upstream vulnerability can still be triggered or not is difficult to efficiently confirm in the prior art in a scene that significant differences exist between parameter semantics and input formats of a source project and a target project. According to the method, a multi-agent parallel target sensing parameter migration framework is constructed, and a legal command line parameter combination related to a vulnerability function is automatically generated from a target software manual in combination with an RAG retrieval enhancement technology; a ReAct intelligent agent is further introduced, a source PoC file is embedded into an acceptable input shell of target software in a cross-format mode under the condition that manual intervention is not needed, and an initial cross-format seed is formed. And iteratively optimizing the parameter-seed pair by taking the function-level trajectory similarity as a feedback index, and driving the parameter sensitive grey box fuzzy test to quickly converge to accurate input capable of triggering vulnerabilities.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of vulnerability verification, and particularly relates to an automatic cross-project vulnerability verification method and device based on a large language model. BACKGROUND

[0002] In recent years, with the explosive growth of the open source ecosystem, code reuse has become the norm of software development. In order to improve development efficiency, software developers tend to reuse verified open source project components. Although this approach speeds up the software development process, it also poses potential security risks. Since a specific code fragment may be widely reused by multiple projects, the propagation range of a single vulnerability expands exponentially, posing a serious challenge to the security of the entire open source project ecosystem.

[0003] To address this problem, existing technologies propose a vulnerability propagation analysis technology based on code clone detection. This technology can effectively identify potential vulnerability clone instances in software projects by building a code feature database and combining a similarity comparison algorithm. Its application enables developers to quickly locate known vulnerabilities that have been propagated to projects and develop appropriate repair strategies accordingly. However, existing detection technologies still have significant limitations in triggerable path verification. Specifically, although current mainstream methods can accurately identify code cloning phenomena, they cannot effectively determine whether the detected vulnerability code is actually triggerable. This problem is mainly influenced by three factors: 1. The code region identified as a vulnerability clone may not be actually executed by the program; 2. The compilation optimization process may exclude some vulnerability code segments, resulting in the program not being executed; 3. The developer may insert a patch to locally repair the vulnerability, and such static detection tools cannot detect this repair measure.

[0004] Third-party libraries, frameworks, and even complete modules are indiscriminately introduced into various projects, significantly shortening the delivery cycle, but also leading to the rapid amplification of systemic risks: a vulnerability disclosed in an upstream library often spreads to hundreds of downstream products through copy-paste, dependency reference, or binary packaging. Traditional static clone detection can quickly locate similar code fragments, but can only give a vague conclusion that there may be potential risks, and cannot answer whether the fragment will be actually executed in the target program or whether the developer has implanted a targeted patch during the transplantation process. As a result, security teams face a heavy workload of verifying each false positive, while development teams are forced to accept online with defects due to "fixing is not possible", ultimately forming a deadlock of "vulnerability database continues to expand, repair rate stagnates".

[0005] Although it is theoretically reasonable to fix all identified vulnerabilities comprehensively, significant resource constraints will be faced in actual operation. Excessive repair not only consumes a large amount of development resources, but also may cause code compatibility problems, and even abnormal functions. In view of the above difficulties, the establishment of a trigger verification mechanism becomes the key to optimizing the priority of vulnerability repair. Existing solutions mainly develop along two technical routes: one is to use the historical knowledge of third-party library PoC (Proof of Concept, concept verification) to guide the dynamic test engine to migrate the existing PoC to the vulnerability PoC suitable for the target software. However, this method can only migrate the file part of the PoC, and cannot migrate the whole PoC. And for the case where the PoC file changes too much, it is also difficult to handle. The second is to use parameter-sensitive fuzzing to automatically infer parameters and generate PoC, but the existing work generates parameters mainly to increase code coverage, and performs poorly in verifying vulnerability propagation tasks.

[0006] In short, to break through the bottleneck of "easy to detect, difficult to confirm", the academic circle has proposed two technical routes. Please refer to Figure 1 One is PoC migration: the concept verification sample (Proof-of-Concept, PoC) that has triggered the vulnerability in the source project is modified into an input suitable for the target program through symbolic execution, path alignment or gray box fuzzing test, so as to directly observe whether the vulnerability can still be triggered. The second is parameter-sensitive fuzzing: the parameter space of program command line options, configuration files, environment variables, etc. is combined with file input to increase code coverage and improve crash probability. Both routes have achieved certain results in their respective scenarios, but have exposed shortcomings in real "cross-project" verification tasks. The PoC migration scheme generally assumes that the source and target programs are "of the same origin", and once the input formats of the two are different (such as the upstream is an original image encoding library, and the downstream is a PDF reader integrating the library), the original PoC is directly rejected due to format verification failure; the existing symbolic execution engine lacks scalability in front of large-scale binary programs, and gray box mutation lacks precise positioning of "key bytes", resulting in a sharp drop in migration success rate. Parameter-sensitive fuzzing can automatically generate a large number of parameter combinations, but only takes improving overall coverage as the optimization target, without the awareness of "directed" convergence to upstream vulnerability code, and the algorithmic overhead brought by blind search conflicts sharply with business delivery cycle. More seriously, modern large software has hundreds of parameters, and the combination space explodes exponentially, making it impossible for humans to write sufficient and legal configurations, and existing algorithms also lack prior knowledge of whether the parameter is related to the vulnerability to be tested, ultimately falling into the predicament of "running more and triggering less".

[0007] In view of the above, the present application is proposed. SUMMARY

[0008] The application provides an automatic cross-project vulnerability verification method and device based on a large language model, which can at least partially improve the above problems.

[0009] To achieve the above object, the application adopts the following technical solutions: An automatic cross-project vulnerability verification method based on a large language model, comprising: S1, obtaining a plurality of work nodes to be processed, calling a plurality of agents to perform parallel processing, constructing a prompt word, and querying a target software manual based on RAG technology to generate a candidate parameter combination; S2, constructing a cross-format seed prompt word based on the candidate parameter combination, calling a ReAct agent to perform code writing processing on the cross-format seed prompt word, and generating a cross-format seed; S3, performing corpus optimization processing on the candidate parameter combination and the cross-format seed to obtain corpus information of downstream software of a related function that is most likely to trigger an upstream vulnerability, and performing parameter sensitive fuzz testing based on the corpus information to generate a PoC of a target vulnerability.

[0010] The application also provides an automatic cross-project vulnerability verification device based on a large language model, comprising: A command line parameter combination generation unit is configured to obtain a plurality of work nodes to be processed, call a plurality of agents to perform parallel processing, construct a prompt word, and query a target software manual based on RAG technology to generate a candidate parameter combination; A cross-format seed generation unit is configured to construct a cross-format seed prompt word based on the candidate parameter combination, call a ReAct agent to perform code writing processing on the cross-format seed prompt word, and generate a cross-format seed; A corpus optimization unit is configured to perform corpus optimization processing on the candidate parameter combination and the cross-format seed to obtain corpus information of downstream software of a related function that is most likely to trigger an upstream vulnerability, and perform parameter sensitive fuzz testing based on the corpus information to generate a PoC of a target vulnerability.

[0011] In summary, the automatic cross-project vulnerability verification method based on the large language model is aimed at the governance pain point of "one vulnerability, multiple project diffusion" in the open source era, and proposes an automatic cross-project vulnerability verification scheme driven by a large language model as the core. Through the three-stage architecture of "automatic parameter generation-cross-format seed generation-runtime feedback iteration", the system can automatically complete legal parameter combination mining, PoC format migration and vulnerability trigger confirmation under the condition that the source program and the target program parameter semantics and input format are completely different, so as to compress the traditional manual checking of several weeks to hours or even minutes. The PoC migration success rate of the method in the real open source software combination scene is significantly better than that of the existing symbolic execution and gray box fuzzing scheme, and without any manual debugging or field expert experience, the vulnerability trigger input that can be directly used for repair decision can be stably output. With the above technical effects, the present application provides an extensible, low-cost and high-precision vulnerability propagation verification method for large open source ecology, which can be directly embedded into the continuous integration / continuous delivery pipeline to realize the active defense target of "upstream vulnerability is public, and downstream risk is zero".

[0012] Compared with the prior art, the present application has the following advantages: 1. Strong cross-project adaptability: can automatically migrate vulnerability verification conditions in the case of differences in parameter semantics and input format between source programs and target programs. 2. High automation: relying on intelligent agents and dynamic feedback mechanisms, manual parameter debugging or seed file writing is not required, significantly reducing verification costs. 3. Improved vulnerability verification efficiency: through corpus optimization and parameter sensitive fuzzing driven by function trajectory similarity, PoC that can trigger vulnerabilities in target software can be quickly generated. 4. Strong universality: suitable for various types of software projects, especially in cross-library and cross-framework vulnerability propagation verification scenarios. Effectively solves the problems of parameter migration difficulty and format adaptation difficulty in cross-project vulnerability verification, and provides a feasible technical solution for large-scale automated vulnerability propagation verification. BRIEF DESCRIPTION OF DRAWINGS

[0013] Figure 1 is a cross-project vulnerability verification example in the prior art provided by the present application.

[0014] Figure 2 is a flowchart of the automatic cross-project vulnerability verification method based on the large language model provided by the first embodiment of the present application.

[0015] Figure 3 is a general diagram of the automatic cross-project vulnerability verification method based on the large language model provided by the embodiment of the present application.

[0016] Figure 4 is a candidate parameter combination generation flowchart provided by the embodiment of the present application.

[0017] Figure 5 is a prompt word diagram for initial parameter generation provided by an embodiment of the application.

[0018] Figure 6 is a domain knowledge query architecture based on RAG provided by an embodiment of the application.

[0019] Figure 7 is an event-driven parallel agent cooperation model diagram provided by an embodiment of the application.

[0020] Figure 8 is a cross-format seed generation diagram provided by an embodiment of the application.

[0021] Figure 9 is a prompt word for seed generation provided by an embodiment of the application.

[0022] Figure 10 is a module diagram of an automatic cross-project vulnerability verification device based on a large language model provided by a second embodiment of the application. DETAILED DESCRIPTION

[0023] The core challenge of verifying whether a disclosed vulnerability exists in other projects is how to locate suspected vulnerabilities from massive codes and prove their triggerability. The technical process can be divided into two stages: (1) candidate vulnerability location based on vulnerability code clone detection; (2) triggerability confirmation combined with dynamic verification methods. First, introduce the related technologies for cross-project verification, including PoC migration technology and parameter-sensitive fuzzing technology.

[0024] The first one is PoC migration techniques, which aims to adapt the existing PoC to the target software system. Currently, this technique is mainly applied to two scenarios: cross-version migration and cross-project migration, and their technical implementation paths are significantly different. In the cross-version vulnerability verification scenario, when a specific version of the software is known to have a vulnerability, it is necessary to verify the existence of the same vulnerability in its historical or updated version. Empirical research shows that although the original PoC can directly trigger the vulnerability in most cases, there are still 21.17% of the cases that need to be adjusted at the byte level to adapt to the target version. To meet this demand, VulScope[3] proposes a path alignment algorithm, which establishes a mapping relationship between the execution trajectories of cross-version vulnerabilities, and combines directed gray-box fuzzing technology to realize the automatic migration of PoC. However, this method has significant limitations in the cross-project migration scenario. The root cause of its failure is that when the version difference between programs is too large, the basic assumption of the path alignment algorithm fails, and the target system in the cross-project scenario and the source PoC belong to completely different code systems. To break through the technical bottleneck of cross-project migration, Kwon et al. proposed the OCTOPOCS framework, which first realized cross-project PoC migration. The innovation of this technology lies in: 1. Crash primitive extraction: extract key constraints (such as 0xffff1111==AABBCCDD) from the source PoC; 2. Constraint-guided execution: realize path exploration of the target system through a symbolic execution engine; 3. Dynamic constraint solving: generate an adaptive PoC by solving constraints in real time at the target execution point.

[0025] As Figure 1In the typical case, the source vulnerability exists in the opj_dump component of the OpenJPEG library (CVE-2020-27823), and the PoC is a specially crafted j2k file. The target project mutool reuses the OpenJPEG library code, but because of the difference in file format verification mechanism (only accepts PDF format input), the original PoC cannot trigger the vulnerability. OCTOPOCS achieves migration through the following technical path: parsing the crash primitive constraints in the j2k file; building a PDF format shell to wrap the core constraint conditions; triggering a null pointer exception at the decoding function. This case verifies the effectiveness of the framework in cross-project scenarios, successfully generating a vulnerability triggering PoC that meets the input specifications of mutool. However, the OCTOPOCS framework faces significant constraints in engineering practice: its symbolic execution engine is limited by computational complexity and faces scalability challenges in large binary program applications. To address this bottleneck, TransferFuzz proposes a method based on directed gray-box fuzzing. Its core idea is consistent with OCTOPoCs, that is, fully utilizing the information of the source PoC to guide the directed testing engine to efficiently verify the triggerability of the source project vulnerability in the target software. But unlike OCTOPoCs, TransferFuzz extracts the key bytes of the PoC through taint analysis and uses them as a mutation dictionary for fuzz testing to verify the target vulnerability.

[0026] The second, software parameters define the program behavior and functionality of the software, providing users with the ability to customize settings. This configuration design enhances the adaptability of software in different environments, with configuration ranges covering network parameter settings, user interface customization, and more. Although flexible configuration mechanisms improve software usability and user experience, they also increase system complexity and the difficulty of software testing. Since each configuration parameter can change the program path, multiple parameter combinations can cause unexpected behavior, so in-depth analysis of the relationships and combination effects between options can effectively ensure software quality and security. To meet this need, parameter-sensitive fuzz testing has emerged, with the core idea being to dynamically adjust parameter combinations to explore code coverage paths under different parameters in real time, providing support for system security verification.

[0027] There are two main types of work, mutation-based testing techniques and filter-based testing techniques. The core idea of the former is to generate option combinations through mutation strategies to find unexpected vulnerabilities by constructing illegal options. The earliest work, AFL-argv, directly randomizes the parameter as a test case byte, but may introduce a large number of invalid parameters, making it impossible to test the program deeply. To this end, ToFo adopts a structured mutation strategy to try to generate more legal option combinations. ConfigFuzz encodes program options into input files, reusing the mutation strategies of existing fuzzers, to achieve coordinated testing of configurations and inputs. Power significantly improves the crash detection capability in 30 real programs by actively selecting the option configuration with the maximum difference and designing different mutation strategies for option domains and file domains. The core idea is to expand the code execution range through configuration diversity. CarpetFuzz first proposed automatically extracting option constraint relationships from program documents through natural language processing to filter invalid combinations. Its core contribution is to reduce 67.91% of invalid option combinations using NLP technology, increasing the path coverage of AFL by 45.97%, and discovering 57 vulnerabilities in 20 open source projects. This work provides an automated constraint extraction framework for subsequent research, but its dependence on document quality also leads to more intelligent prediction methods such as ProphetFuzz. ProphetFuzz first introduced a large language model to automatically predict high-risk option combinations through prompts and directed fuzz testing. In experiments on 52 programs, 12.30% of the high-risk combinations predicted actually had vulnerabilities, improving detection efficiency by 32.85% compared to traditional methods, and the average cost per program was only $8.69. ZigZagFuzz further proposed an alternating mutation strategy to decouple the mutation process of command line options and file inputs, and dynamically refined the test set based on function-level coverage information to improve testing efficiency and effectiveness. Experimental results show that ZigZagFuzz significantly outperforms existing tools such as AFL++, CarpetFuzz, and POWER in terms of vulnerability detection capability, detecting 1.9 to 10.6 times more vulnerabilities. The core idea of the latter is to filter invalid combinations through constraints. For example, CrFuzz addresses the input validation problem of multi-purpose programs such as FFmpeg by predicting input validity through clustering analysis, enhancing the path coverage of fuzzers such as AFL and QSYM by 19.3%.

[0028] In general, the current technology mainly faces the following two challenges, making it difficult to apply to large programs and real-world cross-project vulnerability verification tasks. 1. The parameter combination search space is extremely large. Given a source software vulnerability, to verify whether it can be triggered on the target software. First, find the command line parameters related to the function of the vulnerability in the target software. However, the current PoC migration technology needs to rely on manual configuration of command line parameters. The parameters of modern software are complex, and multiple parameters need to be combined to trigger a function of the program. It is difficult for humans to complete this task. Although parameter-sensitive fuzzing can automatically generate parameter combinations for testing, this work does not have target-oriented capabilities at present, and can only improve the coverage of software testing, but cannot efficiently and quickly verify whether the source vulnerability can be triggered in the target software, resulting in poor performance in cross-project vulnerability verification tasks. 2. Cross-format file migration. Given a source software vulnerability, to verify whether it can be triggered on the target software, there are cases where the input formats accepted by the source software and the target software are different. For example, the source software is a picture library, and the target software is a PDF reader that calls the source software as a third-party library. Existing work has not considered this case, and it is difficult to vary the PoC of the source vulnerability into the PoC of the target software vulnerability of another format by relying solely on the mutation function of fuzzing.

[0029] Based on this, the method mentioned in the present application is designed. In order to make the purpose, technical scheme and advantages of the present application more clear and apparent, the present application will be further described in detail below in combination with embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and not to limit the present application.

[0030] Reference Figure 2 , Figure 3 The first embodiment of the present application discloses an automatic cross-project vulnerability verification method based on a large language model, which can be executed by an automatic cross-project vulnerability verification device based on a large language model (hereinafter referred to as a verification device). In particular, it is executed by one or more processors in the verification device to implement the following method: S1, obtaining a plurality of work nodes to be processed, calling a plurality of agents to perform parallel processing, constructing prompt words, and querying the target software manual based on RAG technology to generate candidate parameter combinations; Specifically, step S1 further comprises: obtaining a plurality of work nodes to be processed, initializing them, creating a corresponding ProcessEvent event for each node, and setting a corresponding maximum number of retries for it; Based on the event loop mechanism, multiple agents are used for parallel processing, and different event types will trigger different processing functions. When processing the ProcessEvent event, the prompt word is constructed, and the large language model is called to generate the parameter combination. When the GenerationDone event is received, the format verification is performed, and the error information is transmitted back to the large language model for re-generation when the format verification fails. The successful result is collected through the ResultEvent distributed collection. The successful results of the collected multiple events are integrated to generate the final output candidate parameter combination.

[0031] Preferably, when processing the ProcessEvent event, the prompt word is constructed, and the large language model is called to generate the parameter combination, specifically: the large model is driven to query the description in the target software manual to understand the core function of the target program, and the large model is driven to find the options of the possible calling source project in the target software manual; Ensure that no conflicting options are used when building the command, and the input file uses a preset placeholder. When it is judged that the target software manual does not specify the output file requirement, the preset requirement is used. When it is judged that there is a specified suffix in the target software manual, the requirement of the target software manual is followed; Determine the fmt_convert attribute according to the compatibility of the PoC format. The fmt_convert attribute is used to confirm whether the command directly supports the PoC format. If it supports, use the PoC format. If it does not support, specify the format to be converted to confirm the format of the generated file, and the prompt word is constructed.

[0032] The prompt word is used as the input of the large model, and the agent is driven to query the target software manual through the RAG technology. The large model infers the core query content that needs to be retrieved through semantic understanding. The query engine filters out the text segment closest to the semantic query in the pre-generated vector database based on the semantic query, and uses the vector distance index as the enhanced context information returned to the large model; The vector distance index includes Euclidean distance and cosine similarity. The database is constructed by vector embedding technology (i.e. converting the text in the original document library into a numerical semantic representation in a high-dimensional space); The large model outputs a text reply based on the enhanced context information and the given prompt word input to obtain the parameter combination.

[0033] Please refer to Figure 4In this embodiment, the command line parameter combination generation aims to generate legal parameters associated with the target (source vulnerability) for the program to be verified. This step is based on the user manual of the target software to be tested, and generates legal and possible command line parameter combinations that may trigger the function related to the upstream vulnerability through the RAG (Retrieval-Augmented Generation) technology combined with a large model.

[0034] Specifically, given the upstream vulnerability information, it is structured as a prompt word for parameter combination generation, and a plurality of intelligent agents with tool calling capabilities are constructed to generate candidate parameter combinations in parallel. Among them, the target software manual is vectorized into a knowledge base, and the intelligent agent calls the tool (queries the document) to ensure that the generated parameter combination is legal. That is, the parameter combination is generated by the large model for understanding the target software manual.

[0035] First, the structured prompt word is designed, and the structured prompt word is the input of the parameter generation agent. The parameter combination needs to meet the constraint condition, so the prompt word needs to be designed to improve the quality of the generated parameter combination. The parameter combination is obtained by the intelligent agent querying the user manual of the target software through the RAG technology. Since the initial parameter combination needs to meet the double constraint conditions, it needs to ensure that the parameter combination can pass the legality check of the target project and can trigger the related function of the source project. To achieve this goal, the method designs the parameter combination generation prompt word framework as shown in Figure 4 The prompt word first sets the two constraint conditions to guide the intelligent agent to generate appropriate parameter combinations, and emphasizes the use of RAG-based domain knowledge query tools by the intelligent agent to achieve this. Secondly, in order to facilitate subsequent processing, the output format is strictly set. Finally, the large model is prompted to think step by step and build a thinking chain to answer the question.

[0036] This is because it is relatively complex to construct a parameter combination that meets the constraint condition. Using the thinking chain (Chain-of-thought, CoT) is a relatively feasible method, and the thinking chain contains a series of continuous reasoning processes, which has been proven to improve the ability of large models to handle complex problems. As Figure 5 The prompt word shown in the display, this study uses the thinking chain to divide it into five steps to guide the large language model to complete the task step by step. At the beginning, the method lets the large model understand the core function of the target program by querying the description in the manual ( Figure 5 Step 1 in Figure 5Step 2) in the above. The subsequent steps 3 and 4 are to remind the large model of the format of the generated parameter combination to facilitate the subsequent dynamic execution verification. Finally, the generated "fmt_convert" attribute is designed for subsequent file generation. When generating a file for a given parameter combination, this method needs to refer to this attribute to confirm the format of the generated file. At the same time, if the large model judges that the generated parameter combination directly supports the PoC format, this study will directly execute the parameter combination and the source PoC on the target program in the subsequent verification link. Otherwise, it will skip the dynamic execution verification phase. After the end of these five steps, the study uses the instruction "Let's take a deep breath and think step by step. Please show your ideas at each step." to encourage the large model to reason carefully and meticulously.

[0037] Secondly, the design of the domain knowledge query tool based on RAG technology is carried out. RAG is a callable tool for agents. As the input of the agent, the agent can query the manual of the target software through RAG technology to obtain parameter text descriptions related to the function of the vulnerability. The agent replies based on this text and generates parameter combinations for output in a structured format (such as json). Specifically, RAG is a lightweight method that can add its own data to large language models without using large model fine-tuning and other time and hardware overhead-intensive techniques to make large models understand the parameter information of the target software. In addition, another reason for using RAG instead of directly feeding the manual of the target software to the large model is that the context window length of the large model is limited, and the length of the manual of some software exceeds the limit of the current context window of the large model. For example, the character number of the manual of GraphicsMagick exceeds 400K.

[0038] Please refer to Figure 6 In the present invention, given the prompt word as input, the agent will construct a query string according to the content of the prompt word and input it into the RAG engine to obtain parameter text descriptions related to the function of the vulnerability. The agent replies based on this text and outputs the parameter combination in a structured form. The parameter combination here is also similar to text, such as (mutool draw-f 100 @@, where @@ represents an input placeholder, which is also the cross-format seed that needs to be generated in this step). To illustrate how this method generates parameter combinations based on RAG, the following example is used for explanation: Before starting to generate parameters, a vector database is constructed based on the parameter manual of the target software and the help information output by the binary program. The reason for choosing these two sources is twofold: first, the parameter manual of the target software provides more detailed option descriptions, but some options are missing. In the process of this study, it was found that the run option was not mentioned in the manual of mupdf 1.9 version. However, this option exists in the help information of the compiled binary. If only the manual is used as the source of constructing the vector database, it is difficult to infer the source vulnerability related to the run option. Second, although the help information provides more complete options, the explanation of each option is relatively rough. Therefore, this method combines these two information sources to construct the vector database of Vulcan. In this example, the prompt word input gives the source project name openjpeg, the target software mupdf name, and the format of the source vulnerability PoC. Then the large language model infers the query text "mutool" to the query engine to find similar text information in the vector database. The returned context contains the explanations of multiple options of mutool, such as "draw", "convert", and "run". The large model determines whether further queries are needed based on the context information. If further queries are needed, the query text will be more refined, such as "mutool draw". Finally, based on the given context and prompt words, parameters such as "mutool draw @@ -f pdf" are generated.

[0039] Finally, the design of parallel agents, Figure 7 The parameter generation framework of parallel agents is demonstrated. Multiple worker nodes are initialized, and a ProcessEvent event is created for each node with a maximum number of retries set. Each worker node uses a different large language model, as the content generated by large language models has randomness, and different large language models have performance differences when processing different software. Therefore, to improve the quality of content generated each time, multiple large language models are used to generate content concurrently. An event loop mechanism is used to trigger different processing logic based on event types to achieve parallel processing of agents. When all the results are collected, they are integrated to generate the final output. Subsequent dynamic execution will be used to further optimize.

[0040] It should be noted that only when ProcessEvent is processed, the above-mentioned structured prompt word design and RAG manual query steps will be executed.

[0041] S2, based on the candidate parameter combination, construct cross-format seed prompt words, call ReAct agent to generate cross-format seeds by processing cross-format seed prompt words; Specifically, the step S2 further includes: constructing a cross-format seed prompt word according to the candidate parameter combination and a preset prompt word template, the cross-format seed prompt word being an input constraint; Taking the cross-format seed prompt word as an input of the ReAct agent, analyzing the task demand therein, determining the two formats to be generated, and planning a code implementation idea, based on the code implementation idea, writing and executing a code program, and generating a candidate cross-format seed; Monitoring the output result of the code execution, when judging that the code execution fails, adjusting the code according to error feedback information, when judging that the code execution succeeds, generating a cross-format seed, wherein the error feedback information includes a syntax error occurring when the code generated by the execution agent is executed, the code can be corrected according to the syntax error information, and information of detecting whether the generated file meets the format requirement.

[0042] In the embodiment, according to the format of the vulnerability PoC and the input sample format of the target software to be tested, a ReAct agent is constructed based on a large model, and a cross-format seed is generated by writing a Python code, that is, a file containing both the source vulnerability PoC format and the input format of the target software to be tested. For example, for the cross-project vulnerability verification example of Figure 1 , this step aims to generate a PDF containing a j2k picture.

[0043] Please refer to Figure 8 , Figure 8 shows the workflow of the cross-format seed generation agent. The core of the process is the prompt word driven seed generation: first, give the prompt word for seed generation, and then the large language model tries to write code under the guidance of the prompt word; then the code is executed by calling external tools and the result is observed, if the execution is successful, the initial seed is output, otherwise the code is modified according to the feedback information until the generation is successful or the maximum number of iterations is reached. The prompt word for generating the seed is shown in Figure 9 . The prompt word adopts a thinking chain design, which decomposes the complex cross-format seed generation problem into several steps, and the goal is to guide the large language model to write a piece of code for each parameter combination, so as to generate a cross-format hybrid seed (for example, a PDF file embedded with a PNG picture format). Therefore, the prompt word is not the final product, but an input constraint that drives the agent to generate a cross-format seed.

[0044] Among them, the reason why the method adopts ReAct (Reasoning and Acting) agent is that the large language model cannot guarantee that the code is completely correct in a single call, and through the cycle mechanism of "writing code - execution - feedback correction", the success rate and robustness of seed generation can be significantly improved.

[0045] S3, corpus optimization processing is performed on the candidate parameter combination and the cross-format seed to obtain corpus information of downstream software of a related function of the upstream vulnerability most likely triggered, and parameter sensitive fuzzy testing is performed according to the corpus information to generate a PoC of the target vulnerability.

[0046] Specifically, step S3 further includes: based on the candidate parameter combination and the cross-format seed, dynamically executing a program to obtain execution feedback information, and performing corpus optimization processing on the candidate parameter combination and the cross-format seed according to the execution feedback information until an optimization end standard is reached to obtain corpus information of downstream software of a related function of the upstream vulnerability most likely triggered, wherein the execution feedback information includes source project function coverage (used to determine which source project functions are covered by the corpus, and if any source project is not covered, it means that the related function of the source project is not triggered, guiding the agent to generate parameters associated with the source project or seeds embedded with the related format of the source project), program execution output results (in the program execution process, the terminal returns useful prompt information to provide optimization direction for the agent. For example, if the returned result indicates that the given format is not supported, the agent is guided to try to modify the seed to the correct file format. If the returned result prompts a parameter contradiction, the agent is guided to modify the parameters), and historical attempt records (guiding the agent to reflect on the historical error records to infer the direction of optimizing the corpus). The optimization end standard includes exceeding a maximum iteration number and the function set executed containing a vulnerability function set exceeding a preset size, and the vulnerability function set is obtained by dynamically executing the PoC of the vulnerability to record the functions executed.

[0047] Preferably, the quality of the candidate parameter combination and the cross-format seed is evaluated based on function trajectory similarity, and the formula is: s is a corpus, which is composed of a candidate parameter combination and a cross-format seed, is a function set of the corpus s on the execution trajectory in the target program T, and P is a call stack function set of the source PoC in the source program. For a corpus, the higher the function trajectory similarity, the closer the corpus is to the core function of the source PoC vulnerability.

[0048] In this embodiment, the seeds generated by the first two steps are only based on the existing vulnerability information and the target software manual, and the quality of the generated corpus (parameter combination + file) is not necessarily high. Therefore, in this step, the initial corpus generated in the previous step is input into the target software to be tested to obtain the feedback information during execution. The feedback information is used to update the prompt words of the large model in the first two steps to generate a corpus with better quality, which is ultimately used for parameter sensitive fuzz testing to generate a PoC of the target vulnerability. It should be noted that the corpus here refers to candidate parameter combinations and cross-format seeds. The purpose of corpus optimization is to use the corpus to dynamically execute the program and continuously optimize the corpus based on the feedback information of the program.

[0049] After collecting the execution feedback information of the program, it is integrated into the prompt words as the input of the agent to guide the agent to generate the next batch of corpus until the function trajectory similarity threshold is met or the maximum number of iterations is exceeded. The function trajectory similarity threshold in this method is set to |P| / 3, because the call stack function of some vulnerabilities is short, and if the threshold is set too high, high-quality corpus may be missed. The maximum number of iterations is set to 10 to avoid high overhead of LLM generating corpus.

[0050] In summary, the main purpose of this method is to verify whether a known vulnerability existing in the source software (or upstream software such as third-party libraries, etc.) can be triggered in the target software (for example, software that uses the source software as a component), and if it can be triggered, an input example (i.e. PoC) that triggers the vulnerability is generated. The existing technology cannot generate parameters associated with the source vulnerability, which requires a lot of time to search for suitable parameter combinations, resulting in low verification efficiency. This method uses a target-aware parameter migration algorithm to use a large language model to construct multiple agents to reason about the initial parameter combination associated with the vulnerability in parallel, and to obtain the vulnerability call stack function coverage information through instrumentation execution to drive the agent to optimize the parameter combination, thereby reducing the time overhead caused by traditional work of traversing the parameter search space. In addition, the existing technology has difficulty in generating high-quality cross-format seeds, which makes it difficult to verify the efficiency of the upstream and downstream software input formats. This method uses the ReAct (Reasoning and Acting) agent to generate an initial usable cross-format seed, and executes it on the instrumented program, and combines the execution of the vulnerability call stack function coverage, the program execution output result and the historical attempt record to further optimize the cross-format seed, and finally improves the efficiency of PoC migration.

[0051] Referring to Figure 10 The second embodiment of the present application provides an automatic cross-project vulnerability verification device based on a large language model, which comprises: The command line parameter combination generation unit 101 is configured to acquire a plurality of work nodes to be processed, call a plurality of agents to perform parallel processing on the plurality of work nodes, construct a prompt word, query a target software manual based on a RAG technology, and generate a candidate parameter combination; The cross-format seed generation unit 102 is configured to construct a cross-format seed prompt word based on the candidate parameter combination, call a ReAct agent to perform code writing processing on the cross-format seed prompt word, and generate a cross-format seed. The corpus optimization unit 103 is configured to perform corpus optimization processing on the candidate parameter combination and the cross-format seed, obtain corpus information of a downstream software related to a function most likely to trigger an upstream vulnerability, and perform parameter sensitive fuzz testing according to the corpus information to generate a PoC of a target vulnerability.

[0052] The above describes the preferred embodiments of the present application. It should be noted that, for those skilled in the art, without departing from the principles of the present application, a number of improvements and refinements can be made, which are also considered within the scope of protection of the present application.

Claims

1. A large language model-based automated cross-project vulnerability verification method, characterized in that, Comprise: S1, obtain a plurality of work nodes to be processed, call a plurality of agents for parallel processing, construct a prompt word, and query the target software manual based on the RAG technology to generate a candidate parameter combination; S2, based on the candidate parameter combination, construct a cross-format seed prompt word, call the ReAct agent to process the cross-format seed prompt word, and generate a cross-format seed; S3, perform corpus optimization processing on the candidate parameter combination and the cross-format seed to obtain corpus information of the downstream software related to the function that is most likely to trigger the upstream vulnerability, and perform parameter sensitive fuzz testing based on the corpus information to generate a PoC of the target vulnerability.

2. The automatic cross-project vulnerability verification method based on a large language model according to claim 1, characterized in that, The step S1 is specifically: Obtain a plurality of work nodes to be processed, initialize them, create a corresponding ProcessEvent event for each node, and set the corresponding maximum number of retries; Based on the event loop mechanism, a plurality of agents are used for parallel processing, different event types trigger different processing functions, wherein when processing the ProcessEvent event, the prompt word is constructed, the large language model is called to generate the parameter combination, when the GenerationDone event is received, the format verification is performed, if failed, the error information is transmitted back to the large language model for re-generation, and the successful result is collected through the ResultEvent for distributed collection; Collect the successful results of a plurality of events for integration to generate the final output candidate parameter combination. 3.The large language model based automated cross-project vulnerability verification method of claim 2, wherein, When processing the ProcessEvent event, the prompt word is constructed, and the large language model is called to generate the parameter combination, which is specifically: Drive the large model to query the description in the target software manual to understand the core function of the target program, and drive the large model to find the options of the possible calling source project in the target software manual; Ensure that no conflicting options are used when building the command, use a preset placeholder for the input file, and when it is judged that the target software manual does not specify the output file requirement, use the preset requirement, and when it is judged that there is a specified suffix in the target software manual, follow the requirement of the target software manual; Determine the fmt_convert attribute according to the compatibility of the PoC format, the fmt_convert attribute is used to confirm whether the command directly supports the PoC format, if it supports, use the PoC format, if it does not support, specify the format to be converted, to confirm the format of the generated file, and the prompt word is constructed.

4. The automated cross-project vulnerability verification method based on a large language model according to claim 3, characterized in that, Further comprising: The prompt word is used as the input of the large model, and the agent is driven to query the target software manual through the RAG technology, wherein the large model infers the core query content to be retrieved through semantic understanding, the query engine filters out the text segment closest to the query semantics as enhanced context information returned to the large model based on the semantic query in the pre-generated vector database using the vector distance index; Wherein, the vector distance index includes Euclidean distance and cosine similarity, and the database is constructed by vector embedding technology based on the target software manual; The large model outputs a text reply based on the enhanced context information and the given prompt word input to obtain the parameter combination.

5. The automated cross-project vulnerability verification method based on a large language model according to claim 1, characterized in that, The step S2 is specifically: According to the candidate parameter combination and the preset prompt word template, a cross-format seed prompt word is constructed, the cross-format seed prompt word being an input constraint; The cross-format seed prompt word is taken as an input of the ReAct intelligent agent, task requirements in the cross-format seed prompt word are analyzed, two formats that need to be generated are determined, a code implementation idea is planned, code programs are written and executed based on the code implementation idea, and a candidate cross-format seed is generated; Output results of code execution are monitored, when it is judged that code execution fails, the code is adjusted according to error feedback information, when it is judged that code execution succeeds, a cross-format seed is generated, wherein the error feedback information includes a syntax error that occurs when the code generated by the intelligent agent is executed, the code can be corrected according to the syntax error information, and information about whether the generated file meets format requirements is detected.

6. The automated cross-project vulnerability verification method based on a large language model according to claim 1, characterized in that, The step S3 is specifically: A dynamic execution program is executed based on the candidate parameter combination and the cross-format seed, execution feedback information is obtained, and corpus optimization processing is performed on the candidate parameter combination and the cross-format seed according to the execution feedback information, until an optimization end standard is reached, corpus information of downstream software of a related function that is most likely to trigger an upstream vulnerability is obtained, wherein the execution feedback information includes source project function coverage, program execution output results, and historical attempt records; The optimization end standard includes exceeding a maximum iteration number and a function set that is executed containing a vulnerability function set that exceeds a preset size, the vulnerability function set being obtained by dynamically executing a PoC of a vulnerability, recording functions executed by the PoC, and obtaining a function name set.

7. The automatic cross-project vulnerability verification method based on a large language model according to claim 6, characterized in that, The quality of the candidate parameter combination and cross-format seed is evaluated based on a function trace similarity, which is formulated as: s is a corpus, which is composed of the candidate parameter combination and the cross-format seed, is a set of functions of the corpus s on the execution trace of the target program T, and P is a set of call stack functions of the source PoC in the source program.

8. A large language model-based automated cross-project vulnerability verification device, characterized by, It includes: A command line parameter combination generation unit configured to obtain a plurality of work nodes to be processed, call a plurality of intelligent agents to perform parallel processing on the plurality of work nodes, construct a prompt word, and query a target software manual based on RAG technology to generate a candidate parameter combination; A cross-format seed generation unit configured to construct a cross-format seed prompt word based on the candidate parameter combination, call a ReAct intelligent agent to perform code writing processing on the cross-format seed prompt word, and generate a cross-format seed; A corpus optimization unit configured to perform corpus optimization processing on the candidate parameter combination and the cross-format seed to obtain corpus information of downstream software of a related function that is most likely to trigger an upstream vulnerability, and generate a PoC of a target vulnerability based on the corpus information.

Citation Information

Patent Citations

  • Intelligent contract vulnerability detection method and device based on large language model and retrieval enhancement generation, equipment and medium

    CN120541845A

  • Intelligent contract vulnerability detection method and system based on type perception and graph enhancement

    CN120671153A

  • Automatic bug patch migration method for large language model based on grammar and semantic enhancement

    CN120744926A

  • Vulnerability detection method and related device

    WO2025001089A1

Cited By

  • GUI (Graphical User Interface) program directional fuzzy testing method and system based on large model agent

    CN121302377A

  • Vulnerability POC generation method based on AI agent

    CN121841858A