A cross-function vulnerability detection method and system based on security obligation propagation

CN122365525BActive Publication Date: 2026-09-29HUAZHONG UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610831485.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-10
Publication Date
2026-09-29
Estimated Expiration
2046-06-10

AI Technical Summary

Benefits of technology

[0021]总体而言,通过本申请所构思的以上技术方案与现有技术相比,具有以下有益效果。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122365525B_ABST
    Figure CN122365525B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of software security detection, and specifically discloses a cross-function vulnerability detection method and system based on security obligation propagation, which comprises the following steps: based on a target function containing a sensitive sink, a test driver program driven by dependency is generated by constructing a call graph and extracting a code slice compatible with verification; based on the test driver program, a security obligation centered on the sensitive sink is extracted; wherein the security obligation is a predicate condition acting on a function entry visible variable, and when the security obligation is established, it can prevent unsafe execution at the sensitive sink; the security obligation is translated to the caller context, the caller is verified and checked, and is classified into one of the three states of having been fulfilled, exposed or propagated, tracking is performed, and a vulnerability judgment result and responsibility attribution explanation are obtained. Through the application, the problems of information accumulation and correlation decay under a deep call chain are effectively solved, and the accuracy and stability of deep cross-function vulnerability detection are greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of software security testing technology, and more specifically, relates to a cross-function vulnerability detection method and system based on security obligation propagation. Background Technology

[0002] With the continuous growth in the scale and complexity of software systems, automated vulnerability detection has become a key technology for ensuring software security. In recent years, researchers have widely adopted deep learning and large language model techniques to automatically learn vulnerability patterns in code. Representative works include VulDeePecker, Devign, and LineVul. The basic principle of such methods is: first, the source code is converted into a structural representation suitable for model learning, and then a neural network model is used to predict whether the code contains vulnerabilities.

[0003] However, the methods described above typically model program logic only within a single function. In real-world software, many vulnerabilities are cross-functional: unsafe operations execute within a single function, while the checks, assumptions, or resource constraints that determine its security are established or omitted elsewhere in the call chain. Existing research indicates that a significant proportion of vulnerabilities in real-world vulnerability corpora are cross-functional, with a relatively deep average call chain. This means that deep cross-functional reasoning is a common requirement in practical vulnerability detection.

[0004] For cross-function vulnerability detection, existing methods primarily address the issue by propagating richer inter-procedural behavioral representations. One class of methods extends the program representation across function boundaries, such as inter-procedural graphs, program slices, or value flow structures, and then performs learning and inference on top of these representations. Representative works include IVDetect and SnapVuln, which use code attribute graphs combined with graph neural networks for vulnerability detection. Another class of methods utilizes program analysis or large language models to generate semantic summaries of called functions for the detector. For example, VulnSC uses a large language model to generate natural language behavioral summaries of called functions and provides them to downstream classifiers.

[0005] While the aforementioned methods have achieved some success in shallow cross-function scenarios, they share the same goal: to propagate behavioral representations of "what the called functions did," enabling the final decision to obtain a broader cross-function context. This goal has a fundamental limitation in deep call chains: as the call depth increases, the detector must carry more and more partially relevant states across function boundaries, while simultaneously handling multiple tasks such as summary generation, relevance filtering, and responsibility localization, leading to a gradual decrease in the effectiveness of the behavioral representations.

[0006] It is evident that existing cross-function vulnerability detection technologies have the following obvious shortcomings.

[0007] (1) The performance of behavior-oriented methods degrades significantly under deep call chains because the information density of behavior representation decreases and noise increases as the depth accumulates.

[0008] (2) Existing methods lack the ability to pinpoint the responsibility for vulnerabilities and can only provide binary vulnerability assessment results, making it difficult to indicate the specific call location where the root cause of the vulnerability lies.

[0009] (3) The generation and propagation of behavioral summaries lack formal verification, and their accuracy depends on the model's capabilities rather than reliable semantic guarantees.

[0010] Therefore, developing an accurate, stable, and interpretable detection method and system for detecting deep cross-function vulnerabilities has become an urgent technical problem to be solved. Summary of the Invention

[0011] In view of the shortcomings of the existing technology, the purpose of this application is to develop an accurate, stable and interpretable detection method and system for the problem of deep cross-function vulnerability detection.

[0012] To achieve the above objectives, in a first aspect, this application provides a cross-function vulnerability detection method based on security obligation propagation, the method comprising: Based on the objective function containing sensitive convergence points, a dependency-driven test driver is generated by constructing a call graph and extracting verification-compatible code slices. Based on the test driver, a security obligation centered on the sensitive convergence point is extracted; the security obligation is a predicate condition that acts on the visible variables at the function entry point, and when the security obligation is true, it can prevent unsafe execution at the sensitive convergence point. By translating security obligations into the caller context, verifying and checking the caller, and classifying them into one of three states—fulfilled, exposed, or propagated—the vulnerability assessment results and liability attribution explanations can be obtained.

[0013] In one possible implementation, the above-mentioned target function containing sensitive convergence points generates a dependency-driven test driver by constructing a call graph and extracting verification-compatible code slices, including: Based on the objective function, with the objective function as the root node, a call graph is constructed through static code analysis. Based on the call graph, the target function and the called functions with transitive dependencies, related type definitions, macro definitions and global declarations are extracted to form verification-compatible code slices. The code slices are used to form verification units so that the verifier can analyze independently without the entire codebase. Based on sensitive convergence points, perform data flow and control flow analysis to obtain a subset of visible states of function entry points that can affect unsafe execution, denoted as convergence point-related states; Nondeterministic modeling is performed on the state related to the convergence point to generate a dependency-driven test driver program.

[0014] One possible implementation also includes achieving compatibility between the test driver and the verifier through the following steps: Based on the dependency summary of code slices and convergence point related states, test driver code is generated through a large language model; Based on the generated test driver code, a verifier is used to perform a compatibility check and obtain the compatibility check results. Based on the error information contained in the compatibility check results, the error information is fed back to the large language model for repair, forming a generation-checking-repair loop until a test driver compatible with the verifier is obtained.

[0015] In one possible implementation, the above-mentioned test driver-based approach extracts security obligations centered on sensitive convergence points, including: The initial verification process, based on the test driver, involves running the verifier to obtain specific counterexamples of unsafe execution when sensitive convergence points are reachable. The candidate obligation proposal step, based on specific counterexamples and code semantics, proposes candidate obligations that only apply to variables visible at the function entry point through a large language model; The candidate obligation verification step involves injecting the candidate obligation as a hypothesis into the test driver program, and checking whether the sensitive convergence point is still reachable by running the validator. If the validator finds a new counterexample, the new counterexample is added to the history and the candidate obligation proposal step is repeated to strengthen the obligation until the validator can no longer reproduce the unsafe execution, and the synthesis converges. By simplifying the process and based on the obligations after synthesis convergence, redundant clauses are removed by trying to delete each clause one by one and having the validator verify the validity of the remaining obligations, resulting in a compact safety obligation.

[0016] It should be noted that, in this application, the compact security obligation may be referred to simply as the compact obligation.

[0017] In one possible implementation, the above translates the security obligations into the caller context, including: For each call point, the formal parameters in the security obligations of the called function are replaced with the corresponding actual parameter expressions, and field accesses are rewritten based on the caller's variables to obtain the translated obligations in the caller's context. The post-translation obligation expresses the conditions that must be met in the caller's variable space before the call.

[0018] In one possible implementation, the above-mentioned verification and classification of the caller includes: Based on the translated obligations, a dependency-driven test driver is built on the caller side, the translated obligations are injected into the relevant call points, and the verification results are obtained by running the verifier. Based on the verification results, the caller will be classified into one of the following three states: Fulfilled: The caller established post-translation obligations on all feasible execution paths prior to the call, the rendezvous point-related responsibilities were resolved locally, and upward tracing terminated on the current path; Exposure: The translated obligation is violated at the call point, and the witness condition of each violated clause is entirely determined by the caller's local computation and state, with responsibility belonging to the current call location; Propagation: If the obligation is not fulfilled after translation, and the witness condition of at least one violated clause still depends on the input or visible state inherited by the caller from a higher level, the obligation must continue to be traced upwards.

[0019] In one possible implementation, the termination condition and responsibility attribution rules for upward tracking are as follows: If all callers of feasible execution paths are classified as fulfilled, then the sensitive convergence point is determined to be protected on the corresponding path. If the caller is traced back to an exposed state, the vulnerability is reported and the responsibility is attributed to the exposed caller. If the obligation is not resolved by the time it reaches the program entry point, the vulnerability will be reported, and the responsibility will be attributed to the highest level that forwarded the unresolved obligation to the program boundary. The obtained vulnerability assessment results also include call chain interpretation, which records the location of sensitive convergence points, extracted security obligations, caller translation sequences, classification status of each caller, and the final location where obligations are fulfilled or exposed.

[0020] Secondly, this application provides a cross-function vulnerability detection system based on security obligation propagation, the system comprising: The code context building module is used to generate dependency-driven test drivers based on target functions containing sensitive convergence points by constructing call graphs and extracting verification-compatible code slices. The security obligation extraction module is used to extract security obligations centered on sensitive convergence points based on the test driver. The security obligation is a predicate condition that acts on visible variables at the function entry point. When the security obligation is true, it can prevent unsafe execution at the sensitive convergence point. The upward tracing and vulnerability assessment module is used to trace vulnerabilities by translating security obligations into the caller context, verifying and checking the caller and classifying them into one of three states: fulfilled, exposed, or propagated, and obtaining vulnerability assessment results and explanations of responsibility attribution.

[0021] Overall, the technical solutions conceived in this application have the following beneficial effects compared with the prior art.

[0022] (1) This application proposes an obligation-oriented cross-function reasoning paradigm. Traditional methods propagate a behavioral representation of "what the called function did," while this application propagates a security obligation of "what must be guaranteed." The security obligation is a compact predicate for a specific sensitive convergence point, containing only the caller-visible conditions required to prevent unsafe execution at that point, rather than the complete semantics of the called function. This paradigm shift transforms the object of cross-function propagation from a broad behavioral context to a compact unresolved core, effectively solving the problems of information accumulation and relevance decay under deep call chains, and significantly improving the accuracy and stability of deep cross-function vulnerability detection.

[0023] (2) This application designs a counterexample-guided obligation synthesis mechanism that combines large language model proposals with formal verifier checks. The large language model proposes candidate obligation clauses based on counterexamples and code semantics. The verifier checks whether the candidate obligations can prevent the recurrence of unsafe execution within the extracted verification context. Clauses are only retained after verification-guided refinement, ensuring that the propagated obligations have formal semantic guarantees rather than relying on the uncertain output of model inference, thus guaranteeing the formal reliability of information propagated across functions.

[0024] (3) This application proposes a three-category caller determination mechanism based on fulfillment, exposure, and propagation, distinguishing between three different situations: "the current caller has established an obligation," "the current caller has violated the obligation locally," and "the obligation still depends on upper-level input." This mechanism not only supports accurate vulnerability detection but also locates the specific call position where the vulnerability responsibility lies, outputs an explanation at the call chain level, and effectively enhances the interpretability and practical value of the detection results.

[0025] (4) The security obligations propagated in this application remain compact and do not expand with depth. The security obligations only isolate the caller-visible conditions related to specific unsafe execution, rather than all the conditions required for the correct execution of the function as a whole. During multi-hop tracing, most obligations are locally fulfilled on the caller side rather than accumulating into a large-scale unresolved state, enabling the method of this application to maintain stable performance in deep scenarios. Attached Figure Description

[0026] Figure 1 This is a flowchart illustrating the cross-function vulnerability detection method based on security obligation propagation provided in an embodiment of this application; Figure 2 This is a schematic diagram of the code context construction process provided in the embodiments of this application; Figure 3 This is a schematic diagram of the security obligation extraction process provided in the embodiments of this application; Figure 4This is a schematic diagram of the upward tracing and vulnerability determination process provided in the embodiments of this application; Figure 5 This is a schematic diagram of the cross-function vulnerability detection system based on security obligation propagation provided in an embodiment of this application; Figure 6 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0027] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0028] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.

[0029] In the description of the embodiments of this application, unless otherwise stated, "multiple" means two or more, for example, multiple processing units means two or more processing units, multiple elements means two or more elements, etc.

[0030] First, the technical terms involved in the embodiments of this application will be introduced.

[0031] Security Obligation is a core concept of this application. Formally, a security obligation is a predicate acting on variables visible at the function entry point. When a security obligation is injected as a hypothesis into the test driver, the verifier cannot reproduce unsafe execution at the sensitive convergence point under the current verification settings. Therefore, a security obligation is a relatively sufficient condition for preventing violations at the analyzed convergence point.

[0032] Safety obligations differ from traditional function preconditions or integrity contracts. Integrity preconditions describe all the conditions required for the correct execution of the function as a whole, while safety obligations only isolate caller-visible conditions related to unsafe execution at specific sensitive convergence points. This narrower objective makes safety obligations more compact and suitable for propagation across functions.

[0033] Entry-level visible variables refer to the complete set of variables visible to the caller at the function entry point, encompassing formal parameters, global variables, and other variables or fields that can be read and written by the caller at the function call boundary. A safety obligation is defined as a predicate acting on entry-level visible variables; that is, all variables in the conditional expression included in the safety obligation belong to this set, thus ensuring that the safety obligation can be directly interpreted and verified in the caller's context.

[0034] Sensitive convergence points refer to program locations that may lead to unsafe execution, including but not limited to: array index access, pointer dereferencing, memory allocation size calculation, divisor in division operations, and resource release operations.

[0035] Counterexample-Guided Synthesis is the method used in this application to extract safety obligations. The basic idea is as follows: run a validator to obtain specific counterexamples that lead to unsafe execution, the large language model proposes candidate obligations based on the counterexamples, the candidate obligations are injected into the validator for inspection, if new counterexamples are still found, the obligations are strengthened and the iteration is repeated until the validator can no longer reproduce unsafe execution.

[0036] A call graph is a directed graph that depicts the function call relationships in a program, where nodes represent functions and directed edges represent call relationships. This application constructs a call graph with the target function as the root node, capturing all functions directly and indirectly called by the target function and their call relationships.

[0037] Verification-compatible code slices refer to self-contained subsets extracted from the target function and its call graph, including the source code of the target function, the source code of the called functions with transitive dependencies, related type definitions (structs, enumerations, type aliases, etc.), macro definitions, and global variable declarations, enabling the verifier to analyze independently without the need for the entire codebase.

[0038] A verification unit is a self-contained collection of code generated during the code context building phase that can be independently analyzed by the verifier. A verification unit includes the source code of the target function, the source code of the called functions with transitive dependencies, related type definitions, macro definitions, global variable declarations, and dependency-driven test and driver programs. It can fully support the verifier's analysis process without relying on the entire codebase.

[0039] The verifier refers to the program verification tool used in this application to perform formal analysis on the verification unit. The verifier takes the test driver as its entry point, performs reachability analysis on the objective function, determines whether unsafe execution at sensitive convergence points is still reproducible under given assumptions, and outputs specific counterexamples (i.e., variable assignments that lead to unsafe execution) when reachable. The verifier's analysis conclusions have formal semantic guarantees and form a reliable basis for the obligation composition and caller classification in this application.

[0040] The test driver is the wrapper code that calls the target function and is responsible for initializing the program state. The dependency-driven strategy only performs nondeterministic modeling on states relevant to the rendezvous point (using symbolic variables from the validator), while irrelevant fields and objects are omitted or given default values ​​to avoid the state explosion problem caused by full-state nondeterministic modeling.

[0041] The wrapper code refers to the main structure of the test driver, that is, the code framework that encapsulates and calls the target function. The wrapper code is responsible for declaring and initializing the entry state required by the target function (modeling the state related to the convergence point nondeterministically and assigning default values ​​to irrelevant states), and then calling the target function, thereby providing the validator with a verification entry point that can be executed and analyzed independently.

[0042] Dependency-driven is the modeling strategy adopted in this application when generating test drivers. The idea is to perform nondeterministic modeling only on a subset of the visible states of function entry points that could affect unsafe execution of sensitive sink points (i.e., sink-related states), while assigning default values ​​or omitting irrelevant fields and objects. This strategy precisely limits the modeling scope to states related to safety obligations, thereby effectively suppressing the state space explosion problem caused by full-state modeling of large structures and nested pointers.

[0043] The subset of visible state at function entry points refers to specific parts of the program state that are visible to the caller at the start of a function call (i.e., at the function entry point) and that, according to data flow and control flow analysis, can affect the unsafe execution of sensitive convergence points. This includes function parameters, global variables, and other variables or fields that are observable (and can be observed) at function boundaries. Unlike the entire entry state of a function, this subset only retains variables directly related to the security of a specific convergence point to support precise dependency-driven modeling.

[0044] Data flow and control flow analysis are static program analysis techniques employed in this application to identify the relevant states of the convergence point. Data flow analysis traces the definition, use, and propagation relationships of variables in the program, determining which entry-level visible variables' values ​​can affect the computational results at the sensitive convergence point. Control flow analysis characterizes the branching and jump structures of the program execution path, determining which paths reach the sensitive convergence point and the predicate constraints along the way. Using both together, the scope of the relevant states of the convergence point can be precisely defined.

[0045] A dependency summary is a structured description of the relevant states of a convergence point and their dependencies on the parameters and fields of the objective function. Generated by the dependency analysis submodule, the dependency summary serves as reference input when the large language model generates test driver wrapper code. This enables the large language model to accurately identify which states require nondeterministic modeling and which fields can be omitted, thereby generating a compact test driver that is compatible with the validator.

[0046] The current verification setup refers to the complete configuration combination of code slices, test drivers, and injected hypotheses used by the verifier in a specific verification run. A security obligation is a sufficient condition under the current verification setup; that is, after injecting a security obligation under this setup, the verifier cannot reproduce unsafe execution. If the verification setup changes (such as code slice expansion or hypothesis adjustment), the validity of the security obligation needs to be re-evaluated. Clearly defining this scope of the current verification setup is a crucial prerequisite for understanding the formal semantics of security obligations.

[0047] Caller-visible conditions refer to program state constraints that are observable and controllable from the caller's perspective; that is, predicate clauses whose objects are visible variables at the function entry point. Safety obligations are designed to include only caller-visible conditions, excluding local states within the called function. This allows safety obligations to be directly inspected and propagated from the caller's side without exposing the internal implementation details of the called function.

[0048] Caller classification is a key mechanism in the upward tracing phase of this application. Based on the verification results, the caller is classified into one of three states: (1) Discharged: The caller establishes post-translation obligations on all feasible execution paths before the call, and the responsibilities related to the convergence point are resolved locally; (2) Exposed: The translated obligation is violated at the call point, and the violation condition is entirely determined by the caller's local logic, and the responsibility belongs to the call location; (3) Propagated: The obligation is not fulfilled after translation, and at least one witness condition of the violation clause still depends on the input or state inherited by the caller from a higher level. The obligation needs to continue to be traced upwards.

[0049] A feasible execution path refers to the actual execution path that the program can reach from the function entry point under the current verification settings. In other words, it's the set of paths that will not produce contradictions (such as division by zero, null pointer dereferencing, or other early termination conditions) under given code logic and input constraints. The caller is categorized as fulfilled, requiring that the translated obligation be fulfilled before the call on all feasible execution paths, not just on some paths, thus ensuring full path coverage of the safety obligation.

[0050] Witness conditions refer to the specific constraint descriptions provided by the validator when it discovers that a certain obligation clause has been violated, explaining why the clause could be violated at the point of call. Witness conditions characterize the minimum prerequisites for a violation to occur and are the core basis for this application to distinguish between the two caller states of "exposure" and "propagation": if the witness condition is entirely determined by the caller's local computation and local state, it is classified as exposure; if the witness condition still depends on parameters or states inherited by the caller from the upper level, it is classified as propagation.

[0051] A caller translation sequence is an ordered record formed by the sequential translation of security obligations through each caller context during upward tracing. Each item in the sequence corresponds to a call level, recording the translated obligation of that caller, its classification status (fulfilled, exposed, or propagated), and the parameter-to-actual substitution relationships involved in the obligation translation. The caller translation sequence is a crucial component of call chain interpretation, enabling developers to trace the propagation path of security obligations layer by layer along the call chain and clearly identify the attribution of vulnerability responsibility.

[0052] The embodiments of this application are described below with reference to the accompanying drawings.

[0053] like Figure 1 As shown, this application provides a cross-function vulnerability detection method based on security obligation propagation, including the following steps S101 to S103.

[0054] S101, Code Context Construction: For target functions containing sensitive convergence points, construct a call graph and extract verification-compatible code slices to generate dependency-driven test drivers; S102, Security Obligation Extraction: Based on the verification unit obtained by building the code context, through an iterative synthesis process guided by counterexamples, security obligations centered on sensitive convergence points are extracted, and weakening and simplification are performed to obtain compact obligations; S103, Upward Tracing and Vulnerability Assessment: Translate the extracted security obligations into the caller context, verify and classify the caller, trace upwards layer by layer, and output the vulnerability assessment results and explanation of responsibility attribution.

[0055] like Figure 2 As shown, by constructing the S101 code context, a verification-compatible code context is built for the target function containing sensitive convergence points. The specific process is as follows.

[0056] First, using the target function as the root node, a call graph is constructed using static analysis tools. The call graph captures all functions directly and indirectly called by the target function and their call relationships. Based on the call graph, verification-compatible code slices are extracted, including: the target function's source code, the source code of transitively dependent called functions, related type definitions (structs, enumerations, type aliases, etc.), macro definitions, and global variable declarations. The goal of these code slices is to form self-contained verification units, enabling the verifier to analyze independently without requiring the entire codebase.

[0057] Subsequently, dependency analysis is performed starting from the sensitive convergence point. Through data flow and control flow analysis, a subset of visible states at function entry points that could affect unsafe execution at the sensitive convergence point is identified; these are called convergence point dependent states.

[0058] Based on the relevance of the relevance state, a dependency-driven test driver is generated. The dependency-driven strategy only performs nondeterministic modeling on the relevance state (using the validator's symbolic variables), while irrelevant fields and objects are omitted or given default values, thus avoiding the state explosion problem caused by full-state nondeterministic modeling of large structures and nested pointers.

[0059] To handle diverse initialization patterns in actual code (e.g., actual C code), this application uses a large language model to generate specific test driver code based on code slices and dependency summaries. The generated test driver undergoes a validator compatibility check, and any errors are fed back to the large language model for repair, forming a generation-checking-repair loop to ensure that the test driver is both compact and compatible with the validator.

[0060] like Figure 3 As shown, through S102 security obligation extraction, based on the verification unit obtained by constructing the code context, a counterexample-guided iterative synthesis process is used to extract security obligations centered on sensitive convergence points. The extraction of security obligations adopts a counterexample-guided synthesis method, and the specific process is as follows.

[0061] Step 1, Initial Validation Run. Run the validator on the test driver without adding any additional assumptions. If the sensitive convergence point is reachable, the validator generates a specific counterexample, witnessing an unsafe execution. The counterexample contains the specific variable assignment that led to the unsafe execution.

[0062] Step two, candidate obligation proposal. The test driver, target function, function entry visible variables, current counterexamples, and historical counterexamples are provided to the large language model. Based on counterexample patterns and code semantics, the large language model proposes candidate obligations that only apply to entry visible variables. Candidate obligations are expressed in the form of hypothetical statements using symbolic variable names.

[0063] Step 3, Candidate Obligation Verification. The candidate obligations are injected as hypotheses into the test driver program, and the validator is run to check if the sensitive convergence point is still reachable. If the validator finds a new counterexample (i.e., the candidate obligation fails to prevent all unsafe executions), the new counterexample is added to the history, and the process returns to Step 2 to strengthen the obligation; if the validator can no longer reproduce the unsafe executions, then the synthesis converges.

[0064] Step four, weakening and simplification. The synthesized and converged obligations may contain redundant clauses. This application performs a weakening process: it attempts to delete each clause one by one, running a validator to check whether the remaining obligations can still prevent unsafe execution from recurring. If the obligation remains valid after deleting a clause, then that clause is redundant and is removed. The weakening process produces a more compact final safety obligation, facilitating cross-function tracing.

[0065] In one possible implementation, a multi-model ensemble strategy can be employed during the candidate proposal stage of the synthesis guided by counterexamples. Multiple large language models are invoked in parallel to propose candidate obligations, and the outputs of each model are semantically deduplicated and merged to form a pool of candidate obligations. The validator sequentially checks the candidates in the pool, adopting the first valid obligation, or combining and strengthening multiple partially valid candidates. This approach leverages the complementary capabilities of different models, improving the coverage and quality of candidate proposals while reducing the number of synthesis iterations.

[0066] In one possible implementation, the extraction of safety obligations can be derived directly using symbolic execution techniques, rather than relying on large language model proposals guided by counterexamples. Specifically, targeting sensitive convergence points, backward symbolic execution is performed on the objective function to collect path conditions leading to the convergence points. All path conditions that could lead to unsafe execution are negated and conjuncted to obtain sufficient conditions preventing unsafe execution. These conditions are then simplified and projected, retaining only clauses acting on variables visible at the function entry point, thus forming safety obligations. This approach offers stronger formal completeness compared to the original scheme but may face the path explosion problem, making it suitable for functions with relatively simple control flows.

[0067] In one possible implementation, an abstract interpretation technique can be used to approximate the safety obligation. First, an abstract domain representation of the unsafe state is defined for the sensitive convergence point. Then, the objective function is abstractly interpreted to compute the abstract set of inputs that can reach the unsafe state. The complement of this set is taken and materialized into predicates on the entry-visible variables to obtain the approximate safety obligation. Abstract interpretation naturally handles loops and recursion, avoiding the path explosion problem of symbolic execution, but the obligation may be conservative. This approach is suitable for functions containing complex control flows.

[0068] In one possible implementation, a specialized security obligation extraction mechanism can be designed to address resource management vulnerabilities (such as memory leaks, double-free, and use-and-release). First, a type-state automaton is defined for each security-critical resource (memory, file handles, locks, etc.) to describe the legal state transitions from resource creation to release. Sensitive convergence points correspond to specific type-state transitions. Security obligations are expressed as constraints that require the resource to be in a specific type state at the point of invocation. When tracing upwards, the type state of the resource is tracked, rather than general predicates. This approach provides more accurate semantic modeling capabilities for resource management vulnerabilities.

[0069] like Figure 4 As shown, after the security obligation is extracted through S103 upward tracing and vulnerability assessment, this obligation becomes the responsibility of the caller. The goal of the upward tracing phase is to determine where this responsibility is handled in the call chain, and the specific process is as follows.

[0070] Step one, obligation translation. For each call point, the security obligations from the called function's side are translated into the caller context. This translation is achieved through the following substitutions: replacing formal parameters with their corresponding actual parameter expressions; and rewriting field accesses based on the caller's variables. The translated obligations express the conditions that must be true in the caller's variable space before the called function is invoked.

[0071] Step two, caller-side inspection and classification. After translation, a dependency-driven test driver is built on the caller side, the translated obligations are injected into the relevant call points, and a validator is run for inspection. Based on the verification results, the caller is classified into one of three states.

[0072] Discharged: The caller established post-translation obligations prior to the call on all feasible execution paths under the current validation settings. At this point, responsibilities related to the rendezvous point are resolved locally, and upward tracing on that path terminates.

[0073] Exposed: The translated obligation is violated at the call point, and the witness condition for each violated clause is entirely determined by the caller's local computation and state. In this case, the current caller itself determines the violation value through its own logic, and the responsibility belongs to that location.

[0074] Propagated: The translated obligation has not been fulfilled, and the witness condition of at least one violated clause still depends on the input or visible state inherited by the caller from a higher level. At this point, the current caller cannot determine the safety of the convergence point and instead forwards the unresolved obligation upwards. This application elevates the obligation one level and continues tracking it.

[0075] Distinguishing between fulfilled obligations and those exposed or disseminated: If the translated obligation is not violated at the point of invocation, it is classified as fulfilled; if the obligation is violated, further analysis of the evidence of violation is required.

[0076] Distinguishing between exposure and propagation: For each violated obligation clause, the validator provides witness conditions explaining why the clause can be violated at the call point. This application analyzes whether these witness conditions are entirely determined by the caller's locally computed and locally established state, or whether they still depend on the caller's parameters or visible state inherited from higher levels. If each violated clause is witnessed only by local conditions, it is classified as exposure; if at least one violated clause is witnessed by inherited inputs, it is classified as propagation. When the evidence is not precise enough to clearly distinguish between exposure and propagation, this application conservatively classifies it as propagation rather than exposure.

[0077] Step 3, Termination and Attribution. Upward tracing terminates in one of three ways.

[0078] Scenario 1: All callers on the relevant paths are classified as fulfilled. In this case, the convergence point is protected on these paths and is determined to be non-vulnerable (or the path is secure).

[0079] Scenario 2: The vulnerability is traced back to the caller classified as exposed. In this case, the vulnerability is reported, and the responsibility is attributed to that caller.

[0080] Scenario 3: The obligation remains unresolved even after propagating to the program entry point. In this case, external input can drive the convergence point into insecure execution, report the vulnerability, and the responsibility falls on the highest level, which will forward the unresolved obligation to the program boundary.

[0081] Step four, output the call chain explanation. The output of this application includes not only the vulnerability assessment results but also a call chain explanation. The explanation records: the location of the sensitive convergence point, the extracted security obligations, the caller translation sequence, the classification status of each caller, and the final location where the obligation is fulfilled or exposed. This explanation enables developers to understand not only that the convergence point can be unsafely reached, but also the specific location where the lack of protection responsibility lies.

[0082] One possible implementation optimizes the overhead of repeated verification during the upward tracing process. The core idea is to reuse the validator state and proven local properties across multiple verification calls. Specifically, the function-level verification summary is cached during the first verification; when the same function is encountered again during subsequent tracing, the cache is queried first rather than re-verified. Furthermore, when changes to the translated obligation involve only variable renaming rather than structural changes, existing verification results can be directly mapped. This approach can significantly reduce verification overhead in deep call chains and large-scale codebases.

[0083] The cross-function vulnerability detection system based on security obligation propagation provided in this application is described below. The cross-function vulnerability detection system based on security obligation propagation described below can be referred to in correspondence with the cross-function vulnerability detection method based on security obligation propagation described above.

[0084] like Figure 5 As shown, this application provides a cross-function vulnerability detection system based on security obligation propagation, including: Code context building module 10 is used to build a call graph and extract verification-compatible code slices for target functions containing sensitive convergence points, and generate dependency-driven test drivers; The security obligation extraction module 20 is used to extract security obligations centered on sensitive convergence points based on the verification units obtained by the code context building module through an iterative synthesis process guided by counterexamples, and to perform weakening and simplification to obtain compact obligations. The upward tracing and vulnerability assessment module 30 is used to translate the extracted security obligations into the caller context, perform verification checks on the caller and classify it into one of three states: fulfilled, exposed, or propagated, trace upwards layer by layer, and output vulnerability assessment results and explanations of responsibility attribution.

[0085] The code context building module includes: The Call Graph Construction Submodule is used to construct a call graph with the target function as the root node through static code analysis, capturing all functions directly and indirectly called by the target function and their call relationships. The code slice extraction submodule is used to extract verification-compatible code slices based on the call graph, including the source code of the target function, the source code of the called function with transitive dependencies, related type definitions, macro definitions and global variable declarations, forming a self-contained verification unit; The dependency analysis submodule is used to identify, starting from sensitive convergence points, a subset of visible states of function entry points that can affect unsafe execution through data flow and control flow analysis. The test driver generation submodule is used to generate dependency-driven test drivers based on the sink-related states, performing nondeterministic modeling only on the sink-related states.

[0086] The safety obligation extraction module includes: The initial verification submodule is used to run the verifier on the test driver to obtain specific counterexamples when sensitive convergence points are reachable. The candidate obligation proposal submodule is used to provide the test driver, target function, function entry visible variables and counterexamples to the large language model, which then proposes candidate obligations that only apply to the entry visible variables. The candidate obligation verification submodule is used to inject candidate obligations as hypotheses into the test driver, run the verifier to check whether sensitive convergence points are still reachable, and if new counterexamples are found, return to the candidate obligation proposal submodule to strengthen the obligation. The weakened and simplified submodule is used to attempt to delete each clause one by one after the synthesis converges, removing redundant clauses to obtain a compact final safety obligation.

[0087] The upward tracing and vulnerability assessment module includes: The Obligation Translation submodule is used to translate the security obligations of the called function into translated obligations in the caller context for each call point by substituting formal parameters into actual parameters and rewriting field accesses. The caller classification submodule is used to build a test driver on the caller side and inject the translated obligation for verification checks. Based on the verification results, the caller is classified into one of three states: fulfilled, exposed, or propagated. The Tracing Termination and Attribution submodule is used to determine the tracing termination conditions based on the caller classification results, and output the vulnerability assessment results and the location of responsibility attribution. The call chain interpretation generation submodule is used to record the location of sensitive convergence points, extracted security obligations, caller translation sequences, classification status of each caller, and the final location where obligations are fulfilled or exposed, and to generate call chain interpretations.

[0088] It is understood that the detailed functional implementation of each of the above units / modules can be found in the description in the aforementioned method embodiments, and will not be repeated here.

[0089] It should be understood that the above system is used to execute the methods in the above embodiments. The corresponding program modules in the system are similar in implementation principle and technical effect to those described in the above methods. The working process of the system can be referred to the corresponding process in the above methods, and will not be repeated here.

[0090] Based on the methods in the above embodiments, this application provides an electronic device that may include: a processor 810, a communications interface 820, a memory 830, and a communication bus 840, wherein the processor 810, the communications interface 820, and the memory 830 communicate with each other through the communication bus 840. The processor 810 may call logical instructions in the memory 830 to execute the methods in the above embodiments.

[0091] Furthermore, the logical instructions in the aforementioned memory 830 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application.

[0092] Based on the methods in the above embodiments, this application provides a computer-readable storage medium storing a computer program that, when run on a processor, causes the processor to execute the methods in the above embodiments.

[0093] Based on the methods in the above embodiments, this application provides a computer program product that, when run on a processor, causes the processor to execute the methods in the above embodiments.

[0094] It is understood that the processor in the embodiments of this application can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. A general-purpose processor can be a microprocessor or any conventional processor.

[0095] The method steps in this application embodiment can be implemented in hardware or by a processor executing software instructions. The software instructions can consist of corresponding software modules, which can be stored in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, portable hard disks, CD-ROMs, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor, enabling the processor to read information from and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can reside in an ASIC.

[0096] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially as a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted through the computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).

[0097] It is understood that the various numerical designations used in the embodiments of this application are merely for the convenience of description and are not intended to limit the scope of the embodiments of this application.

[0098] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. A cross-function vulnerability detection method based on security obligation propagation, characterized in that, include: Based on the objective function containing sensitive convergence points, a dependency-driven test driver is generated by constructing a call graph and extracting verification-compatible code slices. Sensitive convergence points are program locations where there is a possibility of unsafe execution; Based on the test driver, a security obligation centered on the sensitive convergence point is extracted; the security obligation is a predicate condition that acts on the visible variables at the function entry point, and when the security obligation is true, it can prevent unsafe execution at the sensitive convergence point. By translating security obligations into the caller context, verifying and checking the caller and classifying them into one of three states—fulfilled, exposed, or propagated—we can track and obtain vulnerability assessment results and explanations of responsibility attribution. The method, based on a target function containing sensitive convergence points, generates a dependency-driven test driver by constructing a call graph and extracting verification-compatible code slices, including: Based on the objective function, with the objective function as the root node, a call graph is constructed through static code analysis. Based on the call graph, the target function and its transitively dependent called functions, related type definitions, macro definitions, and global declarations are extracted to form verification-compatible code slices. These code slices are used to form verification units, enabling the verifier to analyze independently without the entire codebase. Related type definitions include structs, enumerations, and type aliases. A verification unit is a self-contained set of code generated during the code context construction phase and is available for independent analysis by the verifier. Based on sensitive convergence points, perform data flow and control flow analysis to obtain a subset of visible states of function entry points that can affect unsafe execution, denoted as convergence point-related states; Nondeterministic modeling is performed on the state related to the convergence point to generate a dependency-driven test driver program; The extraction of security obligations centered on sensitive convergence points based on the test driver includes: The initial verification process, based on the test driver, involves running the verifier to obtain specific counterexamples of unsafe execution when sensitive convergence points are reachable. The candidate obligation proposal step, based on specific counterexamples and code semantics, proposes candidate obligations that only apply to variables visible at the function entry point through a large language model; The candidate obligation verification step involves injecting the candidate obligation as a hypothesis into the test driver program, and checking whether the sensitive convergence point is still reachable by running the validator. If the validator finds a new counterexample, the new counterexample is added to the history and the candidate obligation proposal step is repeated to strengthen the obligation until the validator can no longer reproduce the unsafe execution, and the synthesis converges. By simplifying the steps and based on the obligations after synthesis convergence, redundant clauses are removed by trying to delete each clause one by one and having the validator verify the validity of the remaining obligations, resulting in a compact safety obligation. The verification and classification of callers includes: Based on the translated obligations, a dependency-driven test driver is built on the caller side, the translated obligations are injected into the relevant call points, and the verification results are obtained by running the validator. The translated obligations express the conditions that need to be true in the caller's variable space before the call. Based on the verification results, the caller will be classified into one of the following three states: Fulfilled: The caller established post-translation obligations on all feasible execution paths prior to the call, the rendezvous point-related responsibilities were resolved locally, and upward tracing terminated on the current path; Exposure: The translated obligation is violated at the call point, and the witness condition of each violated clause is entirely determined by the caller's local computation and state, with responsibility belonging to the current call location; Propagation: If the obligation is not fulfilled after translation, and the witness condition of at least one violated clause still depends on the input or visible state inherited by the caller from a higher level, the obligation must continue to be traced upwards.

2. The cross-function vulnerability detection method based on security obligation propagation according to claim 1, characterized in that, This also includes achieving compatibility between the test driver and the verifier through the following steps: Based on code slicing and dependency summaries of regress point-related states, test driver code is generated through a large language model; dependency summary refers to a structured description of regress point-related states and their dependencies on parameters and fields of the objective function. Based on the generated test driver code, a verifier is used to perform a compatibility check and obtain the compatibility check results. Based on the error information contained in the compatibility check results, the error information is fed back to the large language model for repair, forming a generation-checking-repair loop until a test driver compatible with the verifier is obtained.

3. The cross-function vulnerability detection method based on security obligation propagation according to claim 1, characterized in that, The translation of security obligations into the caller context includes: For each call point, the formal parameters in the security obligations of the called function are replaced with the corresponding actual parameter expressions, and field accesses are rewritten based on the caller's variables to obtain the translated obligations in the caller's context.

4. The cross-function vulnerability detection method based on security obligation propagation according to claim 1, characterized in that, The termination conditions and liability attribution rules for upward tracking are as follows: If all callers of feasible execution paths are classified as fulfilled, then the sensitive convergence point is determined to be protected on the corresponding path. If the caller is traced back to an exposed state, the vulnerability is reported and the responsibility is attributed to the exposed caller. If the obligation is not resolved by the time it reaches the program entry point, the vulnerability will be reported, and the responsibility will be attributed to the highest level that forwarded the unresolved obligation to the program boundary. The obtained vulnerability assessment results also include call chain interpretation, which records the location of sensitive convergence points, extracted security obligations, caller translation sequences, classification status of each caller, and the final location where obligations are fulfilled or exposed.

5. A cross-function vulnerability detection system based on security obligation propagation, characterized in that, The application of the cross-function vulnerability detection method based on security obligation propagation as described in claim 1 includes: The code context building module is used to generate dependency-driven test drivers based on target functions containing sensitive convergence points by constructing call graphs and extracting verification-compatible code slices. The security obligation extraction module is used to extract security obligations centered on sensitive convergence points based on the test driver. The security obligation is a predicate condition that acts on visible variables at the function entry point. When the security obligation is true, it can prevent unsafe execution at the sensitive convergence point. The upward tracing and vulnerability assessment module is used to trace vulnerabilities by translating security obligations into the caller context, verifying and checking the caller and classifying them into one of three states: fulfilled, exposed, or propagated, and obtaining vulnerability assessment results and explanations of responsibility attribution.

Citation Information

Patent Citations

  • A method for validate and identifying unsafe sensitive input in Android system

    CN109299610A

  • Supply chain cross-packet vulnerability detection method and device, equipment and storage medium

    CN121051762A