Vulnerability tracing method and system based on large model deduction and symbolic execution cooperation

CN122437730BActive Publication Date: 2026-09-18QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610902699.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-23
Publication Date
2026-09-18
Estimated Expiration
2046-06-23

AI Technical Summary

Technical Problem

当后续的路径敏感验证发现该净化点可被绕过时,验证结果将此点排除,但该分支之后的代码序列中可能并不存在其他有效的状态改变事件,导致溯源过程陷入停滞性回溯,始终无法收敛到位于分支条件之前的真正安全配置缺失点或输入验证缺失点

Benefits of technology

本发明将大型语言模型线性推演与分支感知符号执行深度融合进行闭环溯源,确立了由状态推演到符号验证及分歧捕获再到重定位推演及迭代再验证的自愈式循环核心;通过预设自适应迭代阈值控制确保状态收敛,结合大型语言模型的模拟,并运用符号计算的精确分支拓扑探索能力对模型的前序推理进行持续纠偏和补充,从根本上跨越了代码非线性控制流分支与模型单向线性序列推理之间因数据表征断裂而引发的幻象根源停滞循环陷阱,最终在面向深层跨函数变量调用链路的工业级复杂代码资产审计场景中,实现了依赖漏洞根源定位体系的精确寻址、稳定运行与完整技术闭环。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122437730B_ABST
    Figure CN122437730B_ABST
Patent Text Reader

Abstract

The present application belongs to the field of computer and network space security, and provides a vulnerability tracing method and system based on large model deduction and symbol execution cooperation, which comprises constructing a cross-function program dependency graph and extracting a variable call dependency path therefrom; combining a pre-trained large language model for deduction to determine a first quasi-root source introduction point; performing path-sensitive symbol execution on the cross-function program dependency graph starting from the first quasi-root source introduction point to check whether there is contaminated data that has not been effectively purified; when there is contaminated data that has not been effectively purified, performing secondary deduction until the first quasi-root source introduction point is confirmed as a real vulnerability root source introduction point; and generating a tracing positioning result starting from the confirmed real vulnerability root source introduction point. The present application can actively perceive conditional divergence structures and realize stable convergence of vulnerability tracing through secondary reasoning and iterative verification guided by counterexamples on the basis of large model hop-by-hop pollution state deduction along the variable call link.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of computer and cyberspace security, and in particular relates to a vulnerability tracing method and system based on large model deduction and symbolic execution collaboration. Background Technology

[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.

[0003] As software systems continue to grow in size and complexity, external contamination data for security vulnerabilities that rely on variable call chains propagation in the code propagates across functions through multiple variable assignments, function parameter passing, and return value reception, eventually reaching a security-sensitive convergence point and triggering dangerous operations. Security analysts relying on manually tracing the contamination propagation path line by line is not only inefficient but also prone to overlooking the true root cause of the vulnerability due to missed branch conditions or cross-function semantics.

[0004] In real-world vulnerability attribution scenarios, there is a semantic break between the control flow branches of conditional cleansing and the linear inference of the large model. Cleansing functions are often located within conditional branches; for example, a variable might only be cleansed when the user's role is administrator. When the large model performs linear deduction along the sequence of events propagating from the variable, because the input sequence is presented chronologically, the model easily overlooks the existence of control flow conditions, judging cleansing operations within branches as unconditionally executed. This creates a fictitious cleansing point that exists in the sequence but can be bypassed in the actual program. When subsequent path-sensitive verification discovers that this cleansing point can be bypassed, the verification result excludes it. However, there may not be other valid state change events in the code sequence after this branch, causing the attribution process to stagnate and backtrack, never converging to the true security configuration or input verification gap before the branch condition. This conflict between the sequence illusion created by the linear inference of the large model and the omniscient path exploration of symbolic execution makes existing solutions unable to reliably locate the root cause of vulnerabilities when faced with conditional cleansing. In actual code auditing and security assessment work, this results in repeated backtracking of root cause location, inaccurate reporting, and misallocation of remediation resources. Summary of the Invention

[0005] To address the aforementioned technical issues, this invention provides a vulnerability tracing method and system based on large model deduction and symbolic execution collaboration. It can proactively perceive the conditional divergence structure and achieve stable convergence of root cause location through counterexample-guided secondary reasoning and iterative verification, based on large model hop-by-hop contamination state deduction along the variable call chain.

[0006] To achieve the above objectives, the present invention adopts the following technical solution: The first aspect of this invention provides a vulnerability tracing method based on the collaboration of large model deduction and symbolic execution.

[0007] In one or more embodiments, a vulnerability attribution method based on large model deduction and symbolic execution collaboration is provided, including: Based on the source code of the software to be analyzed, a cross-function program dependency graph is constructed, from which a variable call dependency path from the external data introductory node to the security-sensitive convergence node is extracted and converted into a variable propagation event sequence. The pre-built hop-by-hop contamination state inference prompt template is concatenated with the variable propagation event sequence to form the inference input text. Combined with a pre-trained large-scale language model, the flow of contaminated data is simulated and the contamination state is judged on an event-by-event basis to obtain the contamination state label sequence and natural language inference log corresponding to each event. Parse the natural language inference log, extract the event number that first marks the change from pollution to non-pollution in the pollution state label sequence, backtrack the event corresponding to the non-pollution event number in the variable propagation event sequence, and determine the code line position corresponding to the event as the first quasi-root source introduction point; Using the first quasi-root source initiation point as the anchor point, backtrack upwards along the control dependency edge and the control flow predecessor edge in the cross-function program dependency graph to determine the nearest condition judgment node or the nearest common predecessor node of the control flow that directly controls whether the code corresponding to the first quasi-root source initiation point is executed. This node is then used as the exploration starting point for path-sensitive symbol execution. Path-sensitive symbol execution is then performed from this exploration starting point to check whether there is a satisfiable path where contaminated data bypasses the code corresponding to the first quasi-root source initiation point and reaches the safety-sensitive convergence node without effective purification. This determines whether to generate source tracing results or backtrack upwards from the first quasi-root source initiation point to find information related to the dominating variables. When there is contaminated data that has not been effectively cleaned, the dominant variable is extracted from its definition node to the data dependency sub-path used in the branch where the path condition is located in the cross-function program dependency graph. Then, a second deduction is performed in combination with a large language model until the first quasi-root source introduction point is confirmed as the real vulnerability root source introduction point. Starting from the confirmed root cause of the vulnerability, the system combines the cross-function program dependency graph to generate a source tracing and localization result that includes the root source code line, the pollution propagation sequence, and the final state determination.

[0008] As one implementation method, the data dependency sub-paths from the definition node of the dominant variable to the branch used for the path condition are extracted from the cross-functional program dependency graph and converted into dominant variable propagation event segments. The dominant variable propagation event segments are then concatenated with the variable propagation event sequence at the position preceding the event corresponding to the first quasi-root point to form an expanded event sequence. A secondary inference hint template for the source of dominant variable contamination is constructed, concatenated with the expanded event sequence, and then fed into a large language model to drive the model to determine whether the dominant variable can be controlled by an external attacker or affected by other contamination variables, thus obtaining a secondary inference log.

[0009] As one implementation method, a secondary inference log is obtained through secondary deduction. If the secondary inference log indicates that the dominant variable is contaminated by external factors, the location of the contamination-introducing code of the dominant variable pointed out in the secondary inference log is extracted and set as a new quasi-root source introduction point. If it indicates that the dominant variable is not contaminated by external factors, the first quasi-root source introduction point is confirmed as the real vulnerability root source introduction point. When the number of consecutive secondary deductions reaches a preset threshold, the loop is forcibly terminated and the most recently confirmed quasi-root source introduction point is taken as the real vulnerability root source introduction point.

[0010] As one implementation method, the preset threshold is determined as follows: Preset threshold based on depth parameter express, The value of is determined by the function call depth on the variable call dependency path. According to the formula The calculation yielded, where The number of call points along the path from the external data import node to the security-sensitive aggregation node; This indicates the floor function.

[0011] As one implementation method, the process of extracting the dominant variable from its definition node to the data dependency subpath used in the branch containing the path condition is as follows: Starting with the definition node of the dominant variable and ending with the branch judgment node of the event corresponding to the first quasi-root introduction point, the data dependency transitive closure computation is performed on the cross-function program dependency graph. Only the path from the starting node to the ending node with all edges being data dependency edges is retained. If there are multiple paths, the one with the fewest nodes is selected as the data dependency propagation sub-path of the dominant variable.

[0012] As one implementation method, the process of extracting a variable call dependency path from an external data import node to a security-sensitive convergence node from a cross-function program dependency graph is as follows: The return value nodes, output parameter nodes, or nodes written to the buffer in the predefined source function set are marked as external data import nodes. The nodes corresponding to function call parameters in the predefined aggregation function set are marked as security-sensitive aggregation nodes. A graph search algorithm is run to obtain a path connecting the external data import nodes and the security-sensitive aggregation nodes. This path consists of a node sequence and an edge sequence. The node sequence contains variable definition nodes, variable assignment nodes, function call nodes, and parameter passing nodes. The edge sequence indicates data dependency or control dependency relationships.

[0013] As one implementation method, the hop-by-hop contamination state deduction prompt template includes a static prefix part and a dynamic insertion part. The static prefix part describes the code security auditing role that the large language model needs to play, and clarifies that the task is to deduce the propagation state of contaminated data event by event along a given event sequence. It requires the model to maintain an internal contamination flag, and the initial value is set to the contamination state when an external data introduction event is encountered. For each event in the sequence, the model needs to update the pollution flag based on the operation type and variable semantics. If the operation type is variable definition and the right side contains a polluted variable, the newly defined variable is marked as polluted. If the operation type is argument-to-parameter binding, the corresponding parameter is marked as polluted in the same way as the argument. If the operation type is function call and the called function is a known cleansing function, the return value is marked as unpolluted. If the operation type is variable assignment and the right side is a constant or an unpolluted variable, the left side is marked as unpolluted. The prompt template further requires the model to output a brief description of the state before the transition, the state after the transition, the event number, the line number, and the judgment criteria at each state transition.

[0014] As one implementation method, the process of performing path-sensitive symbolic execution across a function program dependency graph includes: Starting from the code line where the first quasi-root point is introduced, the cross-function program dependency graph node corresponding to this starting point is set as the initial position of symbolic execution; the symbolic execution engine divides all program variables that reach this starting point into a polluted variable set and a non-polluted variable set, and creates symbolic values ​​for them. The variables in the polluted variable set carry a pollution marker attribute. Starting from the initial position, the engine explores a path along the control flow graph towards the safety-sensitive convergence node. When encountering a conditional branch, it adds the path condition constraint to the path constraint set and explores both true and false branches simultaneously. During each state update, if an assignment statement is executed, the pollution mark of the variable on the left in the symbolic state is updated based on whether the expression on the right contains a pollution variable. If a cleanup function call is executed, and the function is in the pre-configured cleanup function list, the pollution mark of the variable corresponding to the function's return value is cleared. When an execution path reaches a security-sensitive convergence node, the pollution flag of the variables used by the convergence node is checked. If the pollution flag is still in a polluted state and the current path constraint set is solvable, it is determined that there is a satisfiable path where polluted data reaches the convergence node without effective purification. This path is recorded as a bypass path, and the path condition that causes the code corresponding to the first quasi-root point to be skipped is extracted from the path constraint set of this path. This path condition is the logical condition expression required to make the execution flow not pass through the code corresponding to the first quasi-root point.

[0015] As one implementation method, the specific process of finding the dominating variables and their defining nodes for the dominating path conditions along the control dependency edges is as follows: In the cross-function program dependency graph, starting from the node corresponding to the first quasi-root point, we traverse backward along the control dependency edge to visit all control dependency source nodes that directly control whether the node is executed or not. For each control dependency source node, we obtain its corresponding condition expression, use the symbolic execution engine to identify the atomic constraints related to the condition expression in the path constraint set of the bypass path, and extract the variable names that appear in the atomic constraints as candidate dominating variables. If there are multiple candidate dominant variables, the one whose definition node is closest to the first quasi-root point in the cross-functional program dependency graph is selected as the dominant variable. After determining the dominant variable, its definition node is located in the cross-functional program dependency graph. The definition node contains the variable name and the line number of the code where the definition is located. The dominant variable name, the line number of the definition, and the path conditions are recorded together as a conditional branch snapshot.

[0016] A second aspect of the present invention provides a vulnerability tracing system based on large model deduction and symbolic execution collaboration.

[0017] In one or more embodiments, a vulnerability tracing system based on large model deduction and symbolic execution collaboration includes: The cross-function program dependency graph construction module is used to construct a cross-function program dependency graph based on the source code of the software to be analyzed, and extract a variable call dependency path from the external data introductory node to the security-sensitive convergence node, and convert it into a variable propagation event sequence. The event simulation and inference module is used to concatenate the pre-built hop-by-hop contamination state inference prompt template with the variable propagation event sequence to form the inference input text. Combined with the pre-trained large-scale language model, it simulates the flow of contaminated data and judges the contamination state for each event, and obtains the contamination state label sequence and natural language inference log corresponding to each event. The first quasi-root source introduction point determination module is used to parse the natural language inference log, extract the event number that first marks the change from pollution to non-pollution in the pollution state label sequence, backtrack the event corresponding to the non-pollution event number in the variable propagation event sequence, and determine the code line position corresponding to the event as the first quasi-root source introduction point. The contamination data inspection module is used to perform path-sensitive symbolic execution on the cross-function program dependency graph, starting from the first quasi-root source introduction point, to check whether there is contamination data that has not been effectively cleaned, so as to determine whether to generate source tracing results or to trace back from the first quasi-root source introduction point to find relevant information of the dominant variable. The secondary deduction module is used to extract the dominant variable from its definition node to the data dependency sub-path used in the branch where the path condition is located when there is contaminated data that has not been effectively cleaned. Then, it combines the large language model to perform secondary deduction until the first quasi-root source introduction point is confirmed as the real vulnerability root source introduction point. The source tracing and localization module uses the confirmed root cause of the vulnerability as a starting point and combines it with the cross-function program dependency graph to generate source tracing and localization results that include the root source code line, the pollution propagation sequence, and the final state determination.

[0018] Compared with the prior art, the beneficial effects of the present invention are: This invention deeply integrates linear deduction of large-scale language models with branch-aware symbolic execution for closed-loop tracing, establishing a self-healing loop core from state deduction to symbolic verification and divergence capture, then to relocation deduction and iterative re-verification. By controlling the preset adaptive iteration threshold to ensure state convergence, combined with the simulation of large-scale language models, and using the precise branch topology exploration capability of symbolic computation to continuously correct and supplement the model's pre-order inference, it fundamentally overcomes the stagnation loop trap of the illusion root cause caused by the break in data representation between the nonlinear control flow branches of the code and the unidirectional linear sequence inference of the model. Finally, in the industrial-grade complex code asset auditing scenario facing deep cross-function variable call links, it achieves precise addressing, stable operation and complete technical closed loop of the vulnerability root cause localization system. Attached Figure Description

[0019] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.

[0020] Figure 1 This is a schematic diagram illustrating the vulnerability tracing principle based on large model deduction and symbolic execution collaboration in an embodiment of the present invention. Figure 2 This is the main flowchart of vulnerability tracing based on large model deduction and symbolic execution collaboration in an embodiment of the present invention; Figure 3This is an internal flowchart of the graphic-text joint deduction and counterexample guidance closed loop of an embodiment of the present invention; Figure 4 This is a schematic diagram illustrating the vulnerability root cause relocation principle in an embodiment of the present invention. Detailed Implementation

[0021] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0022] It should be noted that the following detailed description is illustrative and intended to provide further explanation of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0023] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of exemplary embodiments according to the invention. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0024] In this embodiment of the invention, the large language model (LLM) can be implemented using existing artificial intelligence models, such as GPT-4.5 or OpenAI, which are trained on large amounts of text and can perform various natural language tasks.

[0025] In this embodiment of the invention, "hop-by-hop" refers to sequentially deducing the pollution state transition between adjacent variable propagation nodes according to the variable propagation event sequence. Each variable assignment, parameter passing, function call, or return value passing constitutes a propagation hop.

[0026] Contamination status: refers to whether a variable is affected by untrusted input.

[0027] Inference: Based on the current code statement, the previous hop state, and the pollution propagation rules, the large language model determines how the variable state will change after this hop.

[0028] Prompt template: refers to a framework of prompt words that predefines the input content, analysis rules, and output format.

[0029] according to Figure 1 and Figure 2 The vulnerability tracing method based on large model deduction and symbolic execution collaboration in this embodiment may include the following steps S1 to S6.

[0030] The specific implementation process of steps S1 to S6 is as follows: Step S1: Based on the source code of the software to be analyzed, construct a cross-function program dependency graph and extract a variable call dependency path from the external data introductory node to the security-sensitive convergence node, and convert it into a variable propagation event sequence.

[0031] In this embodiment, each event in the event sequence records the variable name, the function name, the line number, the operation type, and the control condition information associated with the event. The operation type includes variable definition, variable assignment, actual parameter to formal parameter binding, return value reception, function call, and condition judgment. Among them, the condition judgment event is used to record the branch expression, the line number of the branch, and the variables used by the branch to control whether the pollution propagation path or the purification operation is executed.

[0032] Specifically, the process of constructing a cross-function program dependency graph includes: Lexical and syntactic analysis is performed on the source code to generate an abstract syntax tree for each function, and a control flow graph is constructed for each function based on the abstract syntax tree. Data flow analysis is performed on the basic blocks in each control flow graph to reach the setpoint, definition-use chains are established inside the function, and cross-function program dependency graphs are generated within the process. Nodes represent statements or expressions, data dependency edges connect variable definitions and uses, and control dependency edges connect conditional judgments and the statements they control. Extract the call graph of all functions, bind the actual parameters of the call point to the formal parameters of the called function, and bind the return value of the called function to the variable received by the call point, thereby connecting the cross-function program dependency graphs within each procedure to obtain the cross-function program dependency graph.

[0033] Specifically, the process of extracting a variable call dependency path from an external data import node to a security-sensitive convergence node from a cross-function program dependency graph is as follows: The return value nodes, output parameter nodes, or nodes written to the buffer in the predefined source function set are marked as external data import nodes. The nodes corresponding to function call parameters in the predefined aggregation function set are marked as security-sensitive aggregation nodes. A graph search algorithm is run to obtain a path connecting the external data import nodes and the security-sensitive aggregation nodes. This path consists of a node sequence and an edge sequence. The node sequence contains variable definition nodes, variable assignment nodes, function call nodes, parameter passing nodes, and condition judgment nodes. The edge sequence indicates data dependency or control dependency relationships.

[0034] Specifically, the method for converting the variable call dependency path from external data introduced to the security-sensitive aggregation node into a variable propagation event sequence is as follows: Traverse the path node sequence and generate an event record for each node based on its type; For the variable definition node, the variable name, the name of the function it belongs to, the line number, and the operation type are recorded for the variable definition. For the variable assignment node, the variable name to be assigned, the name of the function it belongs to, the line number, and the operation type are recorded for the variable assignment. For a function call node, if the node corresponds to the actual parameter passing at the call point, then the actual parameter variable name, the call point function name, the line number, and the operation type are recorded as the actual parameter to formal parameter binding. If the node corresponds to the return value receiving, then the receiving variable name, the called function name, the line number, and the operation type are recorded as the return value receiving. For the function call event itself, the called function name, the line number of the call point, and the operation type are recorded as a function call. The event sequence is arranged in the order of path traversal, forming a complete variable propagation event sequence. The mathematical expression is as follows: ; in This represents a sequence of propagated variable events arranged in the order of dependency propagation. Indicates time step The corresponding discrete variable propagation event, Represents the last event in the sequence. This indicates the total length of the sequence.

[0035] Step S2: Concatenate the pre-built hop-by-hop contamination state inference prompt template with the variable propagation event sequence to form the inference input text. Combine the pre-trained large-scale language model to simulate the flow of contaminated data and judge the contamination state for each event to obtain the contamination state label sequence and natural language inference log corresponding to each event.

[0036] The hop-by-hop contamination state deduction prompt template requires the model to specify the event number and corresponding line of code when determining that the contamination state changes from contaminated to non-contaminated.

[0037] In the specific implementation process, the method for constructing the jump-by-jump pollution state simulation prompt template is as follows: The prompt template includes a static prefix part and a dynamic insertion part. The static prefix part describes the code security auditing role that the large language model needs to play. The task is to deduce the propagation state of polluted data event by event along a given event sequence. The large language model is required to maintain an internal pollution flag, with the initial value set to the polluted state when an external data introduction event is encountered. For each event in the sequence, the large language model needs to update the pollution flag based on the operation type and variable semantics. If the operation type is variable definition and the right side contains a polluted variable, the newly defined variable is marked as polluted. If the operation type is argument-to-parameter binding, the corresponding parameter is marked with the same pollution state as the argument. If the operation type is function call and the called function is a known cleansing function, the return value is marked as unpolluted. If the operation type is variable assignment and the right side is a constant or an unpolluted variable, the left side is marked as unpolluted. The hint template further requires the large language model to output a brief description of the state before the transition, the state after the transition, the event number, the line number, and the judgment criteria at each state transition.

[0038] The dynamically inserted part is the variable propagation event sequence text generated in step S1, which is placed after the static prefix part to form the inference input text. After being fed into the large language model, the large language model performs inference according to the above constraints and outputs a structured inference log. This log contains a sequence of contaminated state labels and the basis text corresponding to each state transition. The contaminated state label is either contaminated or uncontaminated. In this process, the implicit computational relationship of the state transition in the large language model at time step t is as follows: ; in, Indicates time step Updated pollution status label, time step =1 Before starting the deduction, initialize the empty state sequence. , This represents the autoregressive decoding and semantic mapping function of a large language model. This represents the accumulated state context from the preceding time steps. This indicates the standardized step-by-step deduction prompt template text.

[0039] Step S3: Parse the natural language inference log, extract the event number that first marks the change from pollution to non-pollution in the pollution state label sequence, backtrack the event corresponding to the non-pollution event number in the variable propagation event sequence, and determine the code line position corresponding to the event as the first quasi-root source introduction point.

[0040] Step S4: Starting from the first quasi-root point, perform path-sensitive symbolic execution on the cross-function program dependency graph to check whether there is contaminated data that has not been effectively cleaned, in order to determine whether to generate source tracing results or to trace back upward from the first quasi-root point to find information related to the dominant variables.

[0041] Combination Figure 3 and Figure 4In step S4, starting from the first quasi-root source introductory point, path-sensitive symbolic execution is performed on the cross-functional program dependency graph. The path constraints of the symbolic path from the starting point to the security-sensitive convergence node are calculated, and the propagation of contaminated markers along each path is tracked to check whether there is a satisfyable path where contaminated data reaches the convergence node without effective purification. If no such path exists, the first quasi-root source introductory point is confirmed as the true vulnerability root source introductory point, and the process proceeds to step S6. If such a path exists, the first satisfyable path is selected as the bypass path. The path conditions on the bypass path that cause the code corresponding to the first quasi-root source introductory point to be skipped directly are extracted. The process is then backtracked upwards from the first quasi-root source introductory point along the control dependency edge in the cross-functional program dependency graph to find the dominating variable and its definition node that govern the path conditions. A conditional branch snapshot containing the dominating variable name, the location of the definition code, and the path conditions is generated.

[0042] In the specific implementation process, the execution of path-sensitive symbols and the verification of whether a bypass path exists are performed as follows: Starting from the code line where the first quasi-root point is introduced, the cross-function program dependency graph node corresponding to this starting point is set as the initial position of symbolic execution; the symbolic execution engine divides all program variables that reach this starting point into a polluted variable set and a non-polluted variable set, and creates symbolic values ​​for them. The variables in the polluted variable set carry a pollution marker attribute. Starting from the initial position, the engine explores a path along the control flow graph towards the safety-sensitive convergence node. When encountering a conditional branch, it adds the path condition constraint to the path constraint set and explores both true and false branches simultaneously. During each state update, if an assignment statement is executed, the pollution mark of the variable on the left in the symbolic state is updated based on whether the expression on the right contains a pollution variable. If a cleanup function call is executed, and the function is in the pre-configured cleanup function list, the pollution mark of the variable corresponding to the function's return value is cleared. When an execution path reaches a security-sensitive convergence node, the pollution flag of the variables used by the convergence node is checked. If the pollution flag is still in a polluted state and the current path constraint set is solvable, it is determined that there is a satisfiable path where polluted data reaches the convergence node without effective purification. This path is recorded as a bypass path, and the path condition that causes the code corresponding to the first quasi-root point to be skipped is extracted from the path constraint set of this path. This path condition is the logical condition expression required to make the execution flow not pass through the code corresponding to the first quasi-root point.

[0043] Specifically, the process of finding the dominating variables and their defining nodes that govern the path conditions along the control dependency edges is as follows: In the cross-function program dependency graph, starting from the node corresponding to the first quasi-root point, we traverse backward along the control dependency edge to visit all control dependency source nodes that directly control whether the node is executed or not. For each control dependency source node, we obtain its corresponding condition expression, use the symbolic execution engine to identify the atomic constraints related to the condition expression in the path constraint set of the bypass path, and extract the variable names that appear in the atomic constraints as candidate dominating variables. If there are multiple candidate dominant variables, the one whose definition node is closest to the first quasi-root point in the cross-functional program dependency graph is selected as the dominant variable. After determining the dominant variable, its definition node is located in the cross-functional program dependency graph. The definition node contains the variable name and the line number of the code where the definition is located. The dominant variable name, the line number of the definition, and the path conditions are recorded together as a conditional branch snapshot.

[0044] Step S5: When there is contaminated data that has not been effectively cleaned up, extract the dominant variable from its definition node to the data dependency sub-path used in the branch where the path condition is located in the cross-function program dependency graph, and then perform secondary inference in combination with the large language model until the first quasi-root source introduction point is confirmed as the real vulnerability root source introduction point.

[0045] In step S5, the data dependency sub-path from the definition node of the dominant variable to the branch used for the path condition is extracted from the cross-function program dependency graph, and the data dependency sub-path is converted into a dominant variable propagation event segment. The dominant variable propagation event segment and the variable propagation event sequence are concatenated at the position before the event corresponding to the first quasi-root source introduction point to form an expanded event sequence. A secondary inference hint template for the source of pollution of the dominant variable is constructed, concatenated with the expanded event sequence, and fed into the large language model again to drive the model to determine whether the dominant variable can be controlled by an external attacker or affected by other pollution variables, and outputs a secondary inference log. If the secondary inference log indicates that the dominant variable is externally polluted, the pollution introduction code position of the dominant variable indicated in the log is extracted and set as a new quasi-root source introduction point, and the process returns to step S4. If it indicates that the dominant variable is not externally polluted, the first quasi-root source introduction point is confirmed as the real vulnerability root source introduction point, and the process proceeds to step S6. When the number of consecutive executions of step S5 reaches a preset threshold, the loop is forcibly terminated and the most recently confirmed quasi-root source introduction point is taken as the real vulnerability root source introduction point.

[0046] The specific method for extracting the dominant variable from its definition node to the data dependency subpath used in the branch containing the path condition is as follows: Specifically, taking the definition node of the dominant variable as the starting node and the branch judgment node of the event corresponding to the first quasi-root introduction point as the ending node, data dependency transitive closure computation is performed on the cross-function program dependency graph. Only paths from the starting node to the ending node with all edges being data dependency edges are retained. If there are multiple paths, the one with the fewest nodes is selected as the dominant variable propagation data dependency sub-path. The rule for converting the data dependency sub-path into a dominant variable propagation event segment is the same as the rule for converting the path into an event sequence in step S1. The difference is that the variable names in the event records are uniformly the dominant variable name or its alias.

[0047] Specifically, the method for concatenating the propagation event segment of the dominant variable with the propagation event sequence at the position preceding the event corresponding to the first quasi-root point is as follows: Determine the index position k of the event corresponding to the first quasi-root source introduction point in the variable propagation event sequence. Insert all events of the dominant variable propagation event segment sequentially between index positions k-1 and k to form an expanded event sequence. Construct a secondary inference prompt template for the source of pollution of the dominant variable. This template first describes that there is a branch condition controlling the execution of the purification operation, and then gives the expanded event sequence. The large language model is required to additionally determine whether the dominant variable originates from external data introduction or whether it is assigned a value by other polluting variables during the inference process, and output a secondary inference log. The secondary inference log clearly indicates the event number and corresponding code line of the first time the dominant variable is polluted.

[0048] The preset threshold is determined as follows: Preset threshold based on depth parameter express, The value of is determined by the function call depth on the variable call dependency path. According to the formula The calculation yielded, where The number of call points along the path from the external data import node to the security-sensitive aggregation node; This indicates the floor function.

[0049] The formula in this embodiment This means that the maximum backtracking depth increases with the square root of the call chain length, plus a constant offset, to balance probing depth and termination; the count of consecutive executions of step S5 increments each time step S4 is returned, and when this count reaches a certain value... When the loop is terminated, the most recently confirmed quasi-root source entry point in step S5 is taken as the actual vulnerability root source entry point, and the process proceeds to step S6.

[0050] Step S6: Starting from the confirmed true root cause of the vulnerability, generate a source tracing and location result that includes the root source code line, the pollution propagation sequence, and the final state determination by combining the cross-function program dependency graph.

[0051] In step S6, the contamination propagation path from the cross-function program dependency graph to the security-sensitive convergence node is extracted, the root source introduction point, key propagation nodes and convergence nodes on the path are marked, the corresponding source code fragments are extracted, and the source location result containing the root source code line, the contamination propagation sequence and the final state determination is generated.

[0052] The following describes the vulnerability tracing method based on large model inference and symbolic execution collaboration according to the present invention, with examples. The system is deployed on a server cluster equipped with high-performance GPU computing cards. The specific hardware configuration includes multiple NVIDIA A100 GPUs with 40GB of video memory, ensuring sufficient video memory support when processing large-scale code assets containing deep cross-function calls. On the software side, a Linux-based distributed architecture is adopted, and a vLLM inference acceleration engine is integrated for inference task scheduling. This configuration effectively supports the high throughput requirements of large language models when processing long context variable propagation event sequences and significantly reduces the latency of complex dependency graph path search.

[0053] The method of this invention will be described in detail below with reference to a specific code example.

[0054] The source code snippet to be analyzed is located in the file example.c: 1: void process() { 2: char input

[256] ; 3: read(0, input, 256); / / External data import point 4: int role = atoi(getenv("ROLE")); / / Environment variable read, externally controllable 5: char *p = input; 6:if (role == 1) { 7: p = sanitize(p); / / Conditional sanitization 8:} 9:execl(p, NULL); / / Security-sensitive convergence point 10:} The variable input is introduced from the outside via read, and the variable role is introduced from the environment variable via getenv, both of which can be controlled by external attackers; the execl call in line 9 is the command execution convergence point; the sanitize in line 7 is a known sanitization function, but its execution is controlled by the conditional branch role == 1 in line 6. The branch condition depends on the externally controllable role, forming a conditional sanitization vulnerability.

[0055] The specific implementation process of each step of the method of the present invention is as follows: Step S1: Perform lexical and syntactic analysis on the source code to generate an abstract syntax tree and control flow graph for the function `process`. Based on the control flow graph, perform reach-constant analysis to establish a cross-function program dependency graph within the process: data dependency edges connect the `input` definition to the assignment at line 5p, and the `execl` parameter from line 5p to line 9; the `role` definition from line 4 to the conditional expression at line 6; the control dependency edge for the conditional judgment at line 6 connects to the `sanitize` call at line 7 and the control flow merging point after line 8. The predefined source function set includes `read` and `getenv`, and the convergence function set includes `execl`. Mark the `input` definition node and `role` definition node as external data import nodes, and mark the node using the parameter `p` of the `execl` call as a security-sensitive convergence node. Running the graph search algorithm yields a path starting from the external data ingress node, connecting along data dependency edges and control dependency edges to the security-sensitive convergence node. This path's node sequence includes: line 3 (input definition), line 4 (getenv call), line 4 (atoi call), line 4 (role definition), line 6 (conditional judgment using role), line 5 (p assignment), line 7 (argument binding and return reception of the sanitize call), and line 9 (argument binding of the execl call). Converting this path into a variable propagation event sequence according to the propagation order yields: Event 1: Variable = input, Function = process, Line = 3, Operation = variable definition (externally introduced); Event 2: Function call = getenv, Function = process, Line = 4, Operation = function call; Event 3: Function call = atoi, Function = process, Line = 4, Operation = function call; Event 4: Variable = role, Function = process, Line = 4, Operation = variable definition (externally introduced); Event 5: Variable = p, Function = process, Line = 5, Operation = Variable assignment; Event 6: Actual argument = p, function = process, line = 7, operation = actual argument to formal parameter binding; Event 7: Function call = sanitize, Function = process, Line = 7, Operation = function call; Event 8: Variable = p, Function = process, Line = 7, Operation = Return value reception; Event 9: Argument = p, Function = process, Line = 9, Operation = Argument to parameter binding; Event 10: Function call = execl, Function = process, Line = 9, Operation = function call.

[0056] Step S2: Construct a hop-by-hop pollution state inference prompt template, static prefixes describe the analysis task and pollution propagation rules, dynamic parts are inserted into the above event sequence to form inference input text, and then fed into a pre-trained large-scale language model.

[0057] Furthermore, the base model selected in this embodiment is a large-scale language model with reinforcement training using code-specific corpora, and its context processing window is set to 32K words. To ensure the determinism and consistency of the vulnerability tracing and deduction logic, the model hyperparameters are configured as follows during the inference phase: the sampling temperature is set to 0 to suppress the randomness of generation; the kernel sampling threshold is set to 1.0 and the duplication penalty factor is turned off. Through such strict parameter constraints, the model is forced to perform deterministic inference along the code logic, thereby eliminating semantic shifts and logical illusions in the autoregressive decoding process. To achieve automated closed-loop processing of the inference log, the prompt template contains explicit format constraint instructions, forcing the model to return the inference results in a structured text format. Each record output by the model must contain four fixed fields: event sequence number, line number, pollution flag update, and judgment logic description. Through this structured protocol constraint, the backend parsing module can use regular expressions or a structured parser to accurately extract the key physical coordinates of the first change from pollution to non-pollution in the long text inference log, providing standardized input assertions for the subsequent bifurcation capture of the symbolic execution engine. The model infers in the rule order and outputs a sequence of pollution state labels: [Pollution, pollution, pollution, pollution, pollution, pollution, pollution, non-pollution, non-pollution, non-pollution]; The state transition record indicates that at the point where the return value of line 7 is received in event 8, the contaminated state changes from contaminated to uncontaminated because sanitize is a known cleanup function.

[0058] Step S3: Parse the inference log and extract the sequence number of the first event in the contaminated state label sequence that transitions to non-contaminated, obtaining sequence number 8. Backtrack the event sequence; sequence number 8 corresponds to event 8, i.e., the return value of line 7p = sanitize(p). Identify line 7 as the first quasi-root cause introduction point.

[0059] Step S4: Perform local path-sensitive symbolic execution starting from the first quasi-root point, line 7. The symbolic execution engine does not perform a global blind search from the function entry point, but directly inherits the symbolic state of the propagation event sequence before reaching line 7, treating variables `input` and `role` as initial symbolic inputs carrying pollution labels. The engine focuses on probing the logical pathway from line 7 to the convergence point, line 9, along the control flow graph. When processing the control flow decision corresponding to line 6, the engine simultaneously explores both true and false branches. The true branch passes through line 7; since `sanitize` is a purification function, the return value is marked as unpolluted, and the pollution label of variable `p` is updated to false. The false branch skips line 7, with the path constraint `role != 1`. At this time, variable `p` remains a reference to `input`, and the pollution label remains true. When the false branch reaches the convergence point, line 9, the pollution label of parameter `p` is true, and the current path constraint is satisfied. The system determines that there is a bypass path where polluted data reaches the convergence point without effective purification, and records the path condition of this bypass path as `role != 1`. Subsequently, tracing back upwards along the control dependency edge of the cross-functional program dependency graph from node 7, the source of the control dependency in line 7 is the condition judgment node in line 6. The condition expression of this node is role == 1. Since the atomic constraint variable involved in bypassing the path constraint is role, role is confirmed as the dominant variable. Further, the definition node of role in the cross-functional program dependency graph is located at line 4, and a conditional branch snapshot is generated: the dominant variable name is set to role, the definition code location is set to line 4, and the path condition is set to role != 1.

[0060] Step S5: On the cross-functional program dependency graph, calculate the transitive closure of data dependencies from the definition node of role (line 4) to the node used for branch judgment (line 6), obtaining the shortest data dependency sub-path that only contains the definition from line 4 to line 6. Convert this sub-path into a dominator propagation event segment containing one event: variable = role, line = 4, operation = variable definition (external introduction). In the original event sequence, determine the position before the first quasi-root source introduction point event 8, i.e., insert the dominator propagation event segment between events 7 and 8, obtaining the expanded event sequence. Construct a secondary inference hint template for the source of dominator pollution, concatenate it with the expanded event sequence, and feed it back into the large language model. The model's secondary inference log indicates: the dominator variable role is introduced from the environment variable in line 4 through getenv, which is unverified and can be externally controlled, and should be judged as pollution, with line 4 as the pollution introduction point. Set line 4 as the new quasi-root source introduction point.

[0061] Calculate function call depth The initial variable calls along the dependency path pass through five function call nodes in sequence: read, getenv, atoi, sanitize, and execl. =5. Maximum backtracking depth = The current number of S5 executions is 1, which is insufficient. Return to step S4.

[0062] Step S4 is executed again: Starting from the new root cause introduction point line 4, inheriting the previous context state, local path-sensitive symbolic execution is re-executed. The symbolic engine marks the role as polluted and explores all paths. Under the false branch role != 1, the polluted data input bypasses line 7 via p to reach line 9. The pollution is marked as true and the path constraint is solvable, so a bypass path still exists. At this time, line 4 does not perform any verification on the role, which is consistent with the fact that a bypass path exists, indicating that line 4 is reasonable as the vulnerability root cause introduction point. S4 will extract the conditional branch snapshot of the bypass path again, but further backtracking of the dominator variables may still point to the role itself or the environment variable introduction point. Continue executing S5 until the count reaches Assuming that subsequent deductions fail to provide a valid root cause at a higher level, then upon reaching... The loop is forcibly terminated, and the most recently confirmed quasi-root source entry point, i.e., line 4, is confirmed as the actual vulnerability root source entry point.

[0063] Step S6: Starting from the confirmed root cause of the vulnerability, line 4, re-extract the contamination propagation path from line 4 to the security-sensitive convergence node, line 9, on the cross-function program dependency graph. Key nodes in this path include: line 4 (role introduction, root cause), line 6 (conditional branch), line 5 (p assignment), and line 9 (execl call). Mark line 4 as the root cause, lines 5 and 6 as the key propagation nodes, and line 9 as the convergence node. Extract the source code fragments corresponding to each node to generate the source tracing results: The root cause of the vulnerability is the lack of security verification for the environment variable ROLE in line 4, allowing attackers to control the role value; when role is not equal to 1, the cleanup in line 7 is bypassed, and the contaminated data reaches the command execution point along the input→p→execl path, leading to command injection risk. The results include the root source code line, the contamination propagation sequence, and a description of the security impact, providing developers with clear remediation guidance.

[0064] This embodiment converts the extracted variable call dependency paths into an event sequence arranged in the propagation order, and constructs a structured hop-by-hop inference prompt template with embedded contamination propagation rules and state transition annotation requirements. This drives a large language model to sequentially determine the hazard status of contaminated data for each variable definition, assignment, parameter binding, and function call event in the sequence, outputting a contamination status label sequence corresponding to each event and an inference log containing the basis for state transitions. Compared to existing methods that only perform overall macro-level semantic inference on candidate paths, this invention can precisely release the code security semantic prior knowledge of a large language model to each propagation hop in the link, accurately pinpointing the absolute location of the code where the first change in contamination status occurs at the granularity of a single event. This overcomes the defect of excessively large model inference granularity that makes it impossible to locate specific root causes, providing direct and extremely fine physical coordinate basis for vulnerability root cause localization.

[0065] This embodiment constructs a path-sensitive symbolic execution verification and divergence capture mechanism for conditional cleansing. After obtaining the first valid cleansing point through large-scale language model inference and using it as a quasi-root point, path-sensitive symbolic execution is performed starting from this physical point. The mechanism explores along the cross-function program dependency graph to see if there is a satisfyable path that bypasses this point. When a bypass path is found, the branch condition constraint that caused the point to be skipped is precisely extracted. Furthermore, the mechanism backtracks along the control dependency edge to locate the dominating variable governing the branch condition and its definition position, generating a condition divergence snapshot. Compared to existing techniques that fail to detect branch conditions, leading to repeated stagnation in valid cleansing point verification, this invention seamlessly connects the omniscient path exploration precision of symbolic execution at the mathematical level with the point determination depth of large-scale language models at the semantic level. It can proactively discover and capture phantom cleansing illusions caused by hidden control flow branches, thus providing highly targeted and precise contextual input for subsequent correction of linear inference biases.

[0066] Based on a counterexample-guided secondary inference and root cause relocation method, this invention extracts the data dependency sub-paths of the dominant variable from its definition node to the branch decision point from the cross-functional program dependency graph, converts them into event segments, and splices them into the original event sequence to form an expanded sequence. Simultaneously, it constructs directional inference hints targeting the source of pollution of the dominant variable to trigger secondary logical reasoning in a large language model. If the analysis log indicates that the dominant variable is controlled by external pollution, the root cause introduction point is relocated from the previous illusory cleansing point to the actual pollution introduction code location of the dominant variable. Symbolic execution verification is then restarted using these new root cause coordinates to form an iterative evolution closed loop. Compared to existing single-linear inference schemes that are prone to root cause defocusing, this invention cleverly utilizes the semantic source tracing potential of large language models. By using counterexample-guided directional inference to accurately migrate the location target from the false cleansing point to the true safe logic missing point, it achieves proactive closed-loop correction and accurate spatial migration of root cause location coordinates under complex nested conditional branch structures.

[0067] This embodiment deeply integrates linear deduction of large-scale language models with branch-aware symbolic execution for closed-loop tracing, establishing a self-healing loop core from state deduction to symbolic verification and divergence capture, then to relocation deduction and iterative re-verification. The entire method ensures state convergence through preset adaptive iteration threshold control, deeply injects the semantic understanding of neural computing into every micro-step of contamination propagation, and uses the precise branch topology exploration capability of symbolic computing to continuously correct and supplement the model's preceding inference. It fundamentally overcomes the stagnation loop trap of the phantom root cause caused by the break in data representation between the nonlinear control flow branches of the code and the unidirectional linear sequence inference of the model. Finally, in the industrial-grade complex code asset auditing scenario facing deep cross-function variable call chains, it achieves precise addressing, stable operation, and a complete technical closed loop relying on the vulnerability root cause localization system.

[0068] In one or more embodiments, the vulnerability tracing system based on large model deduction and symbolic execution collaboration provided by the present invention can be implemented in software. The vulnerability tracing system based on large model deduction and symbolic execution collaboration includes the following software modules: The cross-function program dependency graph construction module is used to construct a cross-function program dependency graph based on the source code of the software to be analyzed, and extract a variable call dependency path from the external data introductory node to the security-sensitive convergence node, and convert it into a variable propagation event sequence. The event simulation and inference module is used to concatenate the pre-built hop-by-hop contamination state inference prompt template with the variable propagation event sequence to form the inference input text. Combined with the pre-trained large-scale language model, it simulates the flow of contaminated data and judges the contamination state for each event, and obtains the contamination state label sequence and natural language inference log corresponding to each event. The first quasi-root source introduction point determination module is used to parse the natural language inference log, extract the event number that first marks the change from pollution to non-pollution in the pollution state label sequence, backtrack the event corresponding to the non-pollution event number in the variable propagation event sequence, and determine the code line position corresponding to the event as the first quasi-root source introduction point. The contamination data inspection module is used to perform path-sensitive symbolic execution on the cross-function program dependency graph, starting from the first quasi-root source introduction point, to check whether there is contamination data that has not been effectively cleaned, so as to determine whether to generate source tracing results or to trace back from the first quasi-root source introduction point to find relevant information of the dominant variable. The secondary deduction module is used to extract the dominant variable from its definition node to the data dependency sub-path used in the branch where the path condition is located when there is contaminated data that has not been effectively cleaned. Then, it combines the large language model to perform secondary deduction until the first quasi-root source introduction point is confirmed as the real vulnerability root source introduction point. The source tracing and localization module is used to generate source tracing and localization results, including the root source code line, the pollution propagation sequence, and the final state determination, starting from the confirmed true vulnerability root cause introduction point and combining the cross-function program dependency graph.

[0069] It should be noted that each module in the vulnerability tracing system based on large model deduction and symbolic execution collaboration in this embodiment corresponds one-to-one with each step in the vulnerability tracing method based on large model deduction and symbolic execution collaboration in the above embodiments, and their specific implementation processes are the same, so they will not be repeated here.

[0070] The structure of the electronic device according to embodiments of the present invention is described in detail below. The electronic device provided in embodiments of the present invention includes: at least one processor, a memory, a user interface, and at least one network interface. The various components in the vulnerability tracing system based on large model deduction and symbolic execution coordination are coupled together through a bus system. It can be understood that the bus system is used to realize the connection and communication between these components. In addition to a data bus, the bus system also includes a power bus, a control bus, and a status signal bus. The user interface may include a display, keyboard, mouse, trackball, click wheel, buttons, a touchpad, or a touch screen, etc.

[0071] It is understood that the memory can be volatile memory or non-volatile memory, or both. The memory in this embodiment of the invention is capable of storing data to support the operation of the terminal. Examples of this data include any computer programs used to operate on the terminal, such as operating systems and applications. The operating system includes various system programs, such as the framework layer, core library layer, driver layer, etc., used to implement various basic services and handle hardware-based tasks. Applications can include various applications.

[0072] In some embodiments, the vulnerability tracing system based on large model deduction and symbolic execution collaboration provided in this invention can be implemented using a combination of hardware and software. For example, the vulnerability tracing system based on large model deduction and symbolic execution collaboration provided in this invention can be a processor in the form of a hardware decoding processor, programmed to execute the vulnerability tracing method based on large model deduction and symbolic execution collaboration provided in this invention. For instance, the processor in the form of a hardware decoding processor can employ one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.

[0073] As an example, a processor can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where a general-purpose processor can be a microprocessor or any conventional processor, etc.

[0074] As an example of the hardware implementation of the vulnerability tracing system based on large model deduction and symbolic execution collaboration provided in this embodiment of the invention, the device provided in this embodiment of the invention can be directly executed by a processor in the form of a hardware decoding processor. For example, it can be executed by one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components to implement the vulnerability tracing method based on large model deduction and symbolic execution collaboration provided in this embodiment of the invention.

[0075] The memory in this embodiment of the invention is used to store various types of data to support the operation of a vulnerability tracing system based on large model deduction and symbolic execution collaboration, or to store data for execution. Figure 1The program code for the method shown. Examples of this data include: any executable instructions for operating on a vulnerability tracing system based on large model deduction and symbolic execution collaboration, such as executable instructions that can be included in the executable instructions to implement the vulnerability tracing method based on large model deduction and symbolic execution collaboration of the embodiments of the present invention.

[0076] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program including functions for executing... Figure 1 The program code for the method shown. In such an embodiment, the computer program can be downloaded and installed from a network via a communication component, and / or installed from a removable medium. When the computer program is executed by the central processing unit, it performs the various functions defined in the apparatus of this application.

[0077] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, as well as combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0078] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A vulnerability tracing method based on large model deduction and symbolic execution cooperation, characterized in that, include: Based on the source code of the software to be analyzed, a cross-function program dependency graph is constructed, from which a variable call dependency path from the external data introductory node to the security-sensitive convergence node is extracted and converted into a variable propagation event sequence. The pre-built hop-by-hop contamination state inference prompt template is concatenated with the variable propagation event sequence to form the inference input text. Combined with a pre-trained large-scale language model, the flow of contaminated data is simulated and the contamination state is judged on an event-by-event basis to obtain the contamination state label sequence and natural language inference log corresponding to each event. Parse the natural language inference log, extract the event number that first marks the change from pollution to non-pollution in the pollution state label sequence, backtrack the event corresponding to the non-pollution event number in the variable propagation event sequence, and determine the code line position corresponding to the event as the first quasi-root source introduction point; Starting from the first quasi-root point, path-sensitive symbolic execution is performed on the cross-function program dependency graph to check whether there is contaminated data that has not been effectively cleaned, in order to determine whether to generate source tracing results or to trace back upward from the first quasi-root point to find information related to the dominant variables. When there is contaminated data that has not been effectively cleaned, the dominant variable is extracted from its definition node to the data dependency sub-path used in the branch where the path condition is located in the cross-function program dependency graph. Then, a second deduction is performed in combination with a large language model until the first quasi-root source introduction point is confirmed as the real vulnerability root source introduction point. Starting from the confirmed root cause of the vulnerability, the system generates a source tracing and localization result that includes the root source code line, the pollution propagation sequence, and the final state determination, based on the cross-function program dependency graph.

2. The vulnerability tracing method based on large model deduction and symbolic execution cooperation according to claim 1, wherein, The dominant variable is extracted from the cross-functional program dependency graph, from its definition node to the data dependency sub-path used in the branch where the path condition is located, and converted into a dominant variable propagation event segment. The dominant variable propagation event segment is concatenated with the variable propagation event sequence at the position before the event corresponding to the first quasi-root point to form an expanded event sequence. A secondary inference hint template for the source of pollution of the dominant variable is constructed, concatenated with the expanded event sequence, and fed into the large language model again to drive the model to determine whether the dominant variable can be controlled by an external attacker or affected by other pollution variables, thus obtaining the secondary inference log.

3. The vulnerability tracing method based on large model deduction and symbolic execution collaboration as described in claim 1 or 2, characterized in that, A secondary inference log is obtained through a second inference. If the secondary inference log indicates that the dominant variable is contaminated by external factors, the location of the contamination-introducing code for the dominant variable pointed out in the secondary inference log is extracted and set as a new quasi-root source introduction point. If it indicates that the dominant variable is not contaminated by external factors, the first quasi-root source introduction point is confirmed as the true vulnerability root source introduction point. When the number of consecutive secondary inferences reaches a preset threshold, the loop is forcibly terminated and the most recently confirmed quasi-root source introduction point is taken as the true vulnerability root source introduction point.

4. The vulnerability tracing method based on large model deduction and symbolic execution collaboration as described in claim 3, characterized in that, The preset threshold is determined as follows: Preset threshold based on depth parameter express, The value of is determined by the function call depth on the variable call dependency path. According to the formula The calculation yielded, where The number of call points along the path from the external data import node to the security-sensitive aggregation node; This indicates the floor function.

5. The vulnerability tracing method based on large model deduction and symbolic execution collaboration as described in claim 1, characterized in that, The process of extracting the dominant variable from its definition node to the data dependency subpath used in the branch containing the path condition is as follows: Starting with the definition node of the dominant variable and ending with the branch judgment node of the event corresponding to the first quasi-root introduction point, the data dependency transitive closure computation is performed on the cross-function program dependency graph. Only the path from the starting node to the ending node with all edges being data dependency edges is retained. If there are multiple paths, the one with the fewest nodes is selected as the data dependency propagation sub-path of the dominant variable.

6. The vulnerability tracing method based on large model deduction and symbolic execution collaboration as described in claim 1, characterized in that, The process of extracting a variable call dependency path from an external data import node to a security-sensitive convergence node from a cross-function program dependency graph is as follows: The return value nodes, output parameter nodes, or nodes written to the buffer in the predefined source function set are marked as external data import nodes. The nodes corresponding to function call parameters in the predefined aggregation function set are marked as security-sensitive aggregation nodes. A graph search algorithm is run to obtain a path connecting the external data import nodes and the security-sensitive aggregation nodes. This path consists of a node sequence and an edge sequence. The node sequence contains variable definition nodes, variable assignment nodes, function call nodes, and parameter passing nodes. The edge sequence indicates data dependency or control dependency relationships.

7. The vulnerability tracing method based on large model deduction and symbolic execution collaboration as described in claim 1, characterized in that, The hop-by-hop contamination state deduction prompt template includes a static prefix part and a dynamic insertion part. The static prefix part describes the code security auditing role that the large language model needs to play. The task is to deduce the propagation state of contaminated data one event at a time along a given event sequence. The model is required to maintain an internal contamination flag, and the initial value is set to the contamination state when an external data introduction event is encountered. For each event in the sequence, the model needs to update the pollution flag based on the operation type and variable semantics. If the operation type is variable definition and the right side contains a polluted variable, the newly defined variable is marked as polluted. If the operation type is argument-to-parameter binding, the corresponding parameter is marked as polluted in the same way as the argument. If the operation type is function call and the called function is a known cleansing function, the return value is marked as unpolluted. If the operation type is variable assignment and the right side is a constant or an unpolluted variable, the left side is marked as unpolluted. The prompt template further requires the model to output a brief description of the state before the transition, the state after the transition, the event number, the line number, and the judgment criteria at each state transition.

8. The vulnerability tracing method based on large model deduction and symbolic execution collaboration as described in claim 1, characterized in that, The process of performing path-sensitive symbolic execution across a cross-functional program dependency graph includes: Starting from the code line where the first quasi-root point is introduced, the cross-function program dependency graph node corresponding to this starting point is set as the initial position of symbolic execution; the symbolic execution engine divides all program variables that reach this starting point into a polluted variable set and a non-polluted variable set, and creates symbolic values ​​for them. The variables in the polluted variable set carry a pollution marker attribute. Starting from the initial position, the engine explores a path along the control flow graph towards the safety-sensitive convergence node. When encountering a conditional branch, it adds the path condition constraint to the path constraint set and explores both true and false branches simultaneously. During each state update, if an assignment statement is executed, the pollution mark of the variable on the left in the symbolic state is updated based on whether the expression on the right contains a pollution variable. If a cleanup function call is executed, and the function is in the pre-configured cleanup function list, the pollution mark of the variable corresponding to the function's return value is cleared. When an execution path reaches a security-sensitive convergence node, the pollution flag of the variables used by the convergence node is checked. If the pollution flag is still in a polluted state and the current path constraint set is solvable, it is determined that there is a satisfiable path where polluted data reaches the convergence node without effective purification. This path is recorded as a bypass path, and the path condition that causes the code corresponding to the first quasi-root point to be skipped is extracted from the path constraint set of this path. This path condition is the logical condition expression required to make the execution flow not pass through the code corresponding to the first quasi-root point.

9. The vulnerability tracing method based on large model deduction and symbolic execution collaboration as described in claim 1, characterized in that, The specific process of finding the dominating variables and their defining nodes along the control dependency edges for the dominating path conditions is as follows: In the cross-function program dependency graph, starting from the node corresponding to the first quasi-root point, we traverse backward along the control dependency edge to visit all control dependency source nodes that directly control whether the node is executed or not. For each control dependency source node, we obtain its corresponding condition expression, use the symbolic execution engine to identify the atomic constraints related to the condition expression in the path constraint set of the bypass path, and extract the variable names that appear in the atomic constraints as candidate dominating variables. If there are multiple candidate dominant variables, the one that is closest to the first quasi-root point in the cross-functional program dependency graph is selected as the dominant variable. After determining the dominant variable, locate its definition node in the cross-function program dependency graph. This definition node contains the variable name and the line number of the code where the definition is located. Record the dominant variable name, the line number of the definition, and the path conditions together as a conditional branch snapshot.

10. A vulnerability tracing system based on large-scale model deduction and symbolic execution collaboration, characterized in that, The vulnerability tracing method based on large model deduction and symbolic execution collaboration as described in any one of claims 1-9 includes: The cross-function program dependency graph construction module is used to construct a cross-function program dependency graph based on the source code of the software to be analyzed, and extract a variable call dependency path from the external data introductory node to the security-sensitive convergence node, and convert it into a variable propagation event sequence. The event simulation and inference module is used to concatenate the pre-built hop-by-hop contamination state inference prompt template with the variable propagation event sequence to form the inference input text. Combined with the pre-trained large-scale language model, it simulates the flow of contaminated data and judges the contamination state for each event, and obtains the contamination state label sequence and natural language inference log corresponding to each event. The first quasi-root source introduction point determination module is used to parse the natural language inference log, extract the event number that first marks the change from pollution to non-pollution in the pollution state label sequence, backtrack the event corresponding to the non-pollution event number in the variable propagation event sequence, and determine the code line position corresponding to the event as the first quasi-root source introduction point. The contamination data inspection module is used to perform path-sensitive symbolic execution on the cross-function program dependency graph, starting from the first quasi-root source introduction point, to check whether there is contamination data that has not been effectively cleaned, so as to determine whether to generate source tracing results or to trace back from the first quasi-root source introduction point to find relevant information of the dominant variable. The secondary deduction module is used to extract the dominant variable from its definition node to the data dependency sub-path used in the branch where the path condition is located when there is contaminated data that has not been effectively cleaned. Then, it combines the large language model to perform secondary deduction until the first quasi-root source introduction point is confirmed as the real vulnerability root source introduction point. The source tracing and localization module is used to generate source tracing and localization results, including the root source code line, the pollution propagation sequence, and the final state determination, starting from the confirmed true vulnerability root cause introduction point and combining the cross-function program dependency graph.

Citation Information

Patent Citations

  • Security code automatic generation method and system based on full information theory and mechanism ambiguity

    CN120447881A

  • Automatic mining method for firmware vulnerabilities of Internet of Things equipment based on deep learning

    CN120849239A