A Fuzzy Testing Method and System Based on Symbolic Execution
By identifying mutual exclusion relationships among configuration items and establishing an association model, the combination of symbolic execution and fuzzing is optimized, thus solving the path coverage problem of configuration drivers and achieving efficient test coverage and resource utilization.
Patent Information
- Application Number
- CN202511204665.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-27
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2045-08-27
AI Technical Summary
Traditional fuzzing techniques struggle to overcome the path threshold effect when dealing with configuration-driven systems. The configuration item combination space explodes, and path-triggered results lack correlation modeling, leading to wasted test resources and coverage blind spots. Symbolic execution has high constraint solution complexity when configuration items participate in control flow, making it difficult to dynamically schedule test order.
By identifying the mutual exclusion relationships between configuration items, establishing the association between configuration combinations and paths, ignoring configuration constraints for symbolic execution, reconstructing path constraint expressions, dynamically adjusting the test order, introducing a coverage feedback mechanism to optimize the scheduling strategy, and combining symbolic execution with fuzz testing to optimize test resource allocation.
It improves the effectiveness of path exploration and overall testing efficiency, reduces redundant execution, enhances test coverage in complex configuration environments, dynamically schedules test resources, and improves the targeting and efficiency of path triggering.
Smart Images

Figure CN120780607B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of symbolic execution fuzzing technology, and specifically to a fuzzing method and system based on symbolic execution. Background Technology
[0002] As modern software system architectures become increasingly complex, configuration drivers, as a key technology for achieving system flexibility and scalability, are becoming increasingly important. These programs control their operational logic and the activation status of functional modules through numerous configuration parameters, causing the same program to exhibit significantly different execution behaviors under different configuration combinations. While this characteristic enhances software adaptability, it also presents significant challenges to testing: the triggering of program paths depends not only on input data but also on the complex relationships between configuration items.
[0003] Traditional fuzzing techniques exhibit significant limitations when dealing with configuration-driven systems. First, due to the path threshold effect created by configuration items, conventional input mutation strategies often fail to overcome configuration constraints, leading to premature pruning of many potential paths. Second, the mutual exclusion and dependency relationships between configuration items make it difficult for random or fixed configuration strategies to achieve balanced path coverage, resulting in wasted test resources and coverage blind spots. Furthermore, existing techniques lack effective modeling of the relationship between configuration combinations and path triggering, leading to significant redundant execution during testing.
[0004] Current improvement schemes, while employing multi-configuration fuzzing to enhance coverage through combinatorial testing strategies, still face the combinatorial explosion problem. The configuration combination space grows exponentially with the number of configuration items, making comprehensive exploration difficult with limited testing resources. Furthermore, existing methods fail to effectively establish a correlation model between path triggering results under different configuration combinations, leading to a lack of targeted test scheduling. Although symbolic execution techniques can assist in path analysis, their constraint-solving complexity increases dramatically in scenarios where configuration items participate in control flow, severely limiting their practical application effectiveness. Summary of the Invention
[0005] To address the aforementioned problems, the purpose of this invention is to provide a fuzz testing method and system based on symbolic execution, which can improve the effectiveness of path exploration and overall testing efficiency, and enhance test coverage in complex configuration environments.
[0006] This invention provides a fuzz testing method based on symbolic execution, comprising:
[0007] Retrieve multiple configuration items from the driver configuration;
[0008] Identify the mutual exclusion relationships between the configuration items, and combine the configuration items according to the mutual exclusion relationships to obtain multiple configuration combinations;
[0009] Perform fuzz testing on each of the configuration combinations and record the paths triggered by each configuration combination;
[0010] Establish an association relationship based on the configuration combination and the path;
[0011] Based on the association, identify mutually exclusive paths that appear only in a single configuration combination; the configuration combination in which the mutually exclusive path appears is the original configuration combination;
[0012] Obtain path constraints related to input features, ignore configuration constraints, perform symbolic execution on the mutually exclusive paths based on the path constraints, and reconstruct simplified path constraint expressions through the path constraints;
[0013] When the path constraint expression can be triggered in a non-original configuration combination, the expression input features of the path constraint expression are extracted, and a cross-configuration mapping relationship describing whether the input features can satisfy the path constraint expression under different configuration combinations is established.
[0014] When the path constraint expression cannot be triggered in a non-original configuration combination, the path is bound to the original configuration combination and a configuration constraint is generated;
[0015] The test order of configuration combinations is adjusted according to the cross-configuration mapping relationship or the configuration constraints in order to conduct the next round of fuzz testing.
[0016] In one possible implementation, obtaining the set of configuration items in the configuration driver includes:
[0017] Static analysis is performed on the source code of the configuration driver to identify the configuration parameters in the source code, the configuration parameters are transformed into the structure of configuration items, and the corresponding nodes of the configuration items are established in the control flow graph.
[0018] In one possible implementation, identifying the mutual exclusion relationship between the configuration items includes:
[0019] The logical expression is solved based on the Boolean variables on which the branch conditions in the control flow graph depend.
[0020] When multiple configuration items have conditions that cannot be satisfied simultaneously during program execution, these multiple configuration items are recorded as mutually exclusive to constrain the generation of configuration combinations.
[0021] In one possible implementation, performing fuzz testing on each of the configuration combinations includes:
[0022] Generate independent test contexts based on each configuration combination;
[0023] Fuzz testing is performed on each configuration combination based on the independent test context and input mutation strategy; the input mutation strategy is a method that modifies only the input data and not the configuration items of the configuration combination when generating fuzz test cases.
[0024] In one possible implementation, identifying mutually exclusive paths that appear only in a single configuration combination based on the association includes:
[0025] The number of times each path is triggered by different configuration combinations is calculated based on the aforementioned relationships;
[0026] When any path is triggered once by different configuration combinations, the path is determined to be a mutually exclusive path.
[0027] In one possible implementation, whether a path is configured to be triggered is determined based on the unique identifier of the path's entry basic block.
[0028] In one possible implementation, the unique identifier of the entry basic block of the path is the hash value or address code of the first basic block in the control flow graph.
[0029] In one possible implementation, adjusting the test order of configuration combinations based on the cross-configuration mapping relationship or the configuration constraints for the next round of fuzz testing includes:
[0030] After each round of fuzz testing is completed, the number of paths that are first triggered for each configuration combination in that round of fuzz testing is counted to obtain the number of new paths for each configuration combination.
[0031] The coverage feedback score of the configuration combination is obtained by the ratio of the number of newly added paths to the number of rounds in which the corresponding configuration combination is scheduled.
[0032] The coverage feedback scores of all configuration combinations are normalized to obtain the normalized scores of each configuration combination within the same scoring scale range.
[0033] Based on the normalized score, all configuration combinations to be tested are sorted in descending order to obtain the scheduling priority for the next round of fuzz testing.
[0034] The test order of each configuration combination is adjusted according to the next round of scheduling priority and the current scheduling priority of each configuration combination in order to conduct the next round of fuzz testing.
[0035] In one possible implementation, adjusting the test order of each configuration combination based on the next round's scheduling priority and the current scheduling priority of each configuration combination for the next round of fuzz testing includes:
[0036] When the ratio of the next round scheduling priority of a configuration combination to the current scheduling priority exceeds a preset ratio, the scheduling frequency of the next round of fuzz testing for the corresponding configuration combination is reduced.
[0037] When the decrease in the coverage feedback score of a configuration combination exceeds a preset decrease threshold, the interval period for the next round of fuzz testing for the corresponding configuration combination is extended.
[0038] When the coverage feedback score of a configuration combination exceeds a preset score threshold, the priority of the corresponding configuration combination in the next round of fuzz testing is increased.
[0039] This invention also provides a symbolic execution-based fuzzing system for performing any of the above-described fuzzing methods, comprising:
[0040] The first acquisition module is used to acquire multiple configuration items in the configuration driver;
[0041] The mutual exclusion relationship identification module is used to identify the mutual exclusion relationship between the configuration items, and combine the configuration items according to the mutual exclusion relationship to obtain multiple configuration combinations;
[0042] The testing module is used to perform fuzz tests on each of the configuration combinations and record the paths triggered by each configuration combination.
[0043] A module for establishing associations is used to establish associations based on the configuration combination and the path.
[0044] The mutual exclusion path identification module is used to identify mutual exclusion paths that appear only in a single configuration combination based on the association relationship; the configuration combination in which the mutual exclusion path appears is the original configuration combination.
[0045] The second acquisition module is used to acquire path constraints related to the input features, ignore configuration constraints, perform symbolic execution on the mutually exclusive paths according to the path constraints, and reconstruct simplified path constraint expressions through the path constraints.
[0046] When the path constraint expression can be triggered in a non-original configuration combination, the expression input features of the path constraint expression are extracted, and a cross-configuration mapping relationship describing whether the input features can satisfy the path constraint expression under different configuration combinations is established.
[0047] When the path constraint expression cannot be triggered in a non-original configuration combination, the path is bound to the original configuration combination and a configuration constraint is generated;
[0048] The adjustment module is used to adjust the test order of configuration combinations according to the cross-configuration mapping relationship or the configuration constraints, in order to conduct the next round of fuzz testing.
[0049] The symbolic execution-based fuzz testing method and system provided by this invention can improve the effectiveness of path exploration and overall testing efficiency, and enhance test coverage in complex configuration environments. Attached Figure Description
[0050] Figure 1 This is a flowchart illustrating the fuzz testing method provided in an embodiment of the present invention. Detailed Implementation
[0051] The embodiments of the present invention will be further described in detail below with reference to the accompanying drawings and examples. The following detailed description of the embodiments and the accompanying drawings are used to illustrate the principles of the present invention by way of example, but should not be used to limit the scope of the present invention. That is, the present invention is not limited to the described preferred embodiments, and the scope of the present invention is defined by the claims.
[0052] In the description of this invention, it should be noted that, unless otherwise stated, "a plurality of" means two or more; the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance; those skilled in the art can understand the specific meaning of the above terms in this invention as appropriate.
[0053] In existing technologies, the increasing complexity of software systems has led to the widespread application of configuration-driven programs. These programs control behavioral logic through configuration parameters, but traditional fuzzing methods face significant challenges in handling them. Since path triggering depends not only on input data but also on the relationships between configuration items, traditional fuzzing struggles to overcome configuration constraints, resulting in insufficient path coverage. While multi-configuration fuzzing attempts to test different configuration combinations, the configuration combination space grows exponentially, leading to significant waste of testing resources, and the path triggering results lack correlation modeling, failing to effectively guide subsequent testing. Symbolic execution, while aiding path analysis, suffers from inefficient constraint solving when configuration items participate in control flow, making dynamic scheduling of test sequences difficult.
[0054] To address the aforementioned issues, the inventors observed two core contradictions in traditional methods for handling path coverage in configuration-driven systems: first, the contradiction between configuration combination explosion and limited test resources; and second, insufficient decoupling between path triggering conditions and input mutation strategies. Further analysis revealed that different configuration combinations may activate the same path, while some paths depend only on specific configuration conditions. If the relationship between paths and configurations can be identified during testing and test priorities dynamically adjusted, redundant execution can be significantly reduced. Based on this, the inventors proposed combining symbolic execution with fuzzing testing. By reconstructing path constraints to skip configuration conditions, establishing cross-configuration mapping relationships, and introducing a coverage feedback mechanism to optimize the scheduling strategy, efficiency is improved while maintaining coverage depth.
[0055] Figure 1 A flowchart illustrating the fuzz testing method provided in an embodiment of the present invention is shown below. Figure 1 As shown, this invention provides a fuzz testing method based on symbolic execution, comprising:
[0056] Step S1: Obtain multiple configuration items from the configuration driver;
[0057] In one possible implementation, configuration items refer to parameters that affect the execution path during program runtime. Static analysis is performed on the source code of the configuration driver to identify configuration parameters, which are then transformed into configuration item structures, and corresponding nodes for these configuration items are established in the control flow graph. Configuration parameters include macro definitions, initialization parameters, and configuration file fields. The purpose of configuration items is to clarify the range of controllable variables that affect program behavior.
[0058] Static analysis refers to extracting configuration item information by parsing the syntax and semantic structure of the source code without executing the program. Specifically, it can be implemented using abstract syntax tree traversal or data flow analysis techniques to comprehensively identify the source of configuration parameters in the program.
[0059] Macro definitions are constants or conditional compilation blocks defined in the source code through preprocessor directives. Specifically, they can be implemented by parsing #define statements or #ifdef conditional branches in header files to extract compile-time variable configuration parameters.
[0060] Initialization parameters refer to configuration items assigned values through function parameters or global variables when the program starts. Specifically, they can be implemented by identifying the entry parameters of the main function or the global variable initialization code block, which is used to capture the initial configuration values at runtime.
[0061] Configuration file fields refer to key-value pair configuration data read by the program from external files. Specifically, this can be achieved by analyzing the structured parsing functions involved in file read and write operations, and is used to extract dynamically loaded configuration information.
[0062] Runtime loaded parameters refer to configuration items set through dynamic API calls or environment variables during program execution. Specifically, this can be achieved by tracing the call chain of the setenv function or specific configuration update interfaces, and is used to identify configuration parameters that can be adjusted at runtime.
[0063] The structure of configuration items refers to the unified encapsulation of configuration items from different sources into a standardized data structure that includes name, type, value range and dependency relationship. Specifically, it can be stored using a structure or hash table to eliminate the heterogeneity of configuration items and support subsequent combination and generation.
[0064] In a control flow graph, a node is a mapping relationship between each configuration item and its position of influence in the program control flow. This can be achieved by inserting configuration item markers into the basic blocks of the control flow graph, which are used to associate configuration items with constraints of the program execution path.
[0065] Specifically, during driver configuration testing, static analysis is first performed on the source code, traversing the abstract syntax tree to identify conditional compilation blocks in macro definitions, such as parsing the DEBUG flag in #ifdef DEBUG. Next, the global variable initialization code is examined to extract default parameters set at program startup, such as port numbers or log levels. Further analysis of configuration file parsing functions identifies field names and their corresponding key-value pairs read from JSON or INI files. Simultaneously, API calls that dynamically modify configurations at runtime are traced, for example, by capturing the parameter list passed to the set_config() function using Hook technology. These four types of configuration items are then uniformly transformed into a configuration item structure containing name, data type, and valid scope, eliminating differences in storage format between configuration items from different sources. Finally, the branch decision nodes affected by each configuration item are marked in the program control flow graph; for example, the log level configuration item is associated with its conditional statement in the log output function, thereby establishing an explicit mapping relationship between configuration items and execution paths.
[0066] Traditional methods typically handle configuration items from a single source, such as analyzing only configuration files while ignoring macro definitions or runtime parameters, leading to incomplete configuration combinations. Furthermore, existing technologies lack correlation analysis of the position of configuration items within the control flow, making it difficult to accurately identify mutual exclusion relationships between them. This invention covers all sources of configuration items through multi-dimensional static analysis and establishes a control flow graph node mapping, resulting in more comprehensive configuration item identification and clearer semantic relationships.
[0067] Through the above technical solution, this invention can accurately identify the sources of all potential configuration items in the configuration driver, including parameter settings at compile time, startup time, and runtime, avoiding incomplete test coverage due to missing configuration items. By transforming heterogeneous configuration items into a unified structure and associating them with control flow nodes, it provides an accurate dependency data foundation for subsequently generating valid configuration combinations, effectively preventing the generation of invalid configuration combinations that violate program semantics.
[0068] Step S2: Identify the mutual exclusion relationships between configuration items, and combine the configuration items according to the mutual exclusion relationships to obtain multiple configuration combinations;
[0069] In one possible implementation, logical expressions are solved based on Boolean variables on which the branch conditions in the control flow graph depend; when multiple configuration items have constraints that cannot be satisfied simultaneously during program execution, the multiple configuration items are recorded as mutually exclusive relationships to constrain the generation of configuration combinations.
[0070] Mutual exclusion refers to the constraint relationship between configuration items that logically cannot be effective simultaneously. This can be determined by solving the Boolean variables of the branch conditions in the control flow graph, and its purpose is to avoid generating invalid configuration combinations. Configuration combinations refer to legal parameter combinations that satisfy mutual exclusion constraints. These can be generated using constraint satisfaction algorithms, and their purpose is to reduce the test space and ensure the validity of the combinations.
[0071] The Boolean variables upon which branch conditions in a control flow graph depend refer to the variables involved in the conditional judgments affecting path selection during program execution. Specifically, this can be achieved by extracting the condition nodes in the program's control flow graph using static analysis tools and parsing the variable names and logical expressions they depend on. These variables characterize the control role of configuration items in the program logic.
[0072] Logical expression solving refers to the formal verification of possible logical conflicts between configuration items. Specifically, a satisfiability modular theory solver can be used to perform conflict detection on combinations of Boolean variables to determine whether different configuration items have conditions that cannot be satisfied simultaneously during program execution.
[0073] Semantic exclusion refers to the existence of constraints that prevent two or more configuration items from coexisting at the program logic level. Specifically, when these configuration items are activated simultaneously, there will inevitably be unsatisfactory path conditions in the program control flow graph, leading to the interruption of the execution flow.
[0074] Specifically, in the process of identifying mutual exclusion relationships among configuration items, firstly, all conditional branch nodes in the control flow graph are extracted through static analysis, and their dependent Boolean variable expressions are parsed. For each Boolean variable corresponding to a configuration item, its occurrence positions in the control flow graph are traversed, and all related logical constraints are collected. Subsequently, the logical constraints corresponding to different configuration items are input into the solver for conflict detection. If at least one input combination makes it impossible for the constraints of two configuration items to be satisfied simultaneously, then the two are determined to have a mutual exclusion relationship. Finally, the mutual exclusion relationship is stored in a mapping table in the form of key-value pairs, serving as a filtering rule in the configuration combination generation stage. For example, when the constraint condition corresponding to configuration item A is "x_0" and the constraint condition corresponding to configuration item B is "x≤0", the solver can verify that the two have a mutual exclusion relationship and record this relationship in the mapping table, ensuring that the subsequently generated configuration combinations will not contain both A and B simultaneously.
[0075] Traditional methods typically rely on static configuration file parsing or simple heuristic rules to infer configuration item relationships, failing to accurately capture logical conflicts arising during dynamic program execution. This invention, however, through program control flow analysis and formal verification, can precisely identify mutual exclusion relationships among configuration items determined by the program logic itself, avoiding redundancy or omissions in configuration combinations caused by incomplete manual rule definitions. For example, existing technologies may ignore implicit mutual exclusion relationships arising from nested conditional branches, while this invention, by traversing all conditional nodes in the control flow graph, can comprehensively cover potential conflicts in the program execution path.
[0076] Through the above technical solution, this invention can effectively identify the mutual exclusion relationships of configuration items determined by the program's internal logic, avoid generating invalid configuration combinations, and reduce path unreachability problems caused by configuration conflicts during fuzzing. Simultaneously, by establishing a mutual exclusion mapping table, it can provide clear combination generation rules for subsequent tests, reduce the consumption of test resources on invalid configurations, and improve the efficiency and effectiveness of test case generation.
[0077] Step S3: Perform fuzz testing on each configuration combination and record the paths triggered by each configuration combination;
[0078] In one possible implementation, an independent test context is generated based on each configuration combination; fuzz testing is performed on each configuration combination based on the independent test context and the input mutation strategy to ensure that the mutated input is not pruned before the path decision condition, thereby obtaining a set of valid paths.
[0079] The test context refers to the isolated runtime environment created for each configuration combination. This can be achieved through memory isolation or resource allocation mechanisms. Its purpose is to prevent the execution states of different configuration combinations from interfering with each other and to ensure the independence of fuzz test results.
[0080] The input mutation strategy is a method that modifies only the input data and not the configuration items of the configuration combination when generating fuzz test cases. Specifically, it can be implemented through random bit flipping or structure-aware mutation. Its purpose is to maintain the validity of the input data under the constraints of the configuration conditions and avoid the input being incorrectly pruned before the path determination due to the mutation of configuration items.
[0081] Specifically, during the fuzzing execution phase, the test context for each configuration combination is initialized independently, including memory stack, file handles, and thread resources, ensuring that the testing processes for different configuration combinations do not interfere with each other. The input mutation strategy modifies only the byte sequence or data structure of the input data when generating test cases, such as adjusting message length or replacing field values, while keeping the configuration parameters unchanged. Therefore, the mutated input data is not pre-filtered due to configuration parameter conflicts before entering the path condition judgment, thus triggering more effective path branches. The path set collected during the test will contain the complete execution trajectory under the combined effect of configuration combinations and input data.
[0082] Traditional fuzz testing often employs shared testing environments or hybrid mutation strategies in multi-configuration scenarios, leading to cross-contamination of test results from different configuration combinations. Furthermore, input mutation may unexpectedly alter configuration parameters, causing valid inputs to be incorrectly pruned. This invention decouples environment isolation from input mutation, ensuring the independence of the testing process while avoiding invalid path filtering, significantly improving the ability to discover valid paths.
[0083] Through the above technical solution, the present invention can accurately capture path branches triggered by configuration combinations and input data, avoiding the omission of effective paths due to interference from the test environment or misjudgment of configuration conditions, thereby improving the integrity of path coverage and test efficiency.
[0084] Step S4: Establish associations based on configuration combinations and paths;
[0085] In one possible implementation, the association refers to the mapping relationship between the configuration combination and the path it triggers. Specifically, it can be constructed by recording the execution results of fuzz tests, and its role is to reveal the dependency relationship between the path and the configuration.
[0086] Step S5: Identify mutually exclusive paths that appear only in a single configuration combination based on the association relationship;
[0087] In one possible implementation, the number of times each path is triggered by different configuration combinations is counted based on the association relationship; when the number of times any path is triggered by different configuration combinations is 1, the path is determined to be a mutually exclusive path.
[0088] In one possible implementation, whether a path is configured to be triggered is determined based on the unique identifier of the entry basic block of the path. The unique identifier is the hash value or address encoding of the first basic block in the control flow graph. Specifically, intermediate representation code generated by the compiler can be used to extract block features, which is used to accurately distinguish the starting execution point of different paths in complex control flows.
[0089] Among them, the configuration combination in which mutual exclusion paths appear is the original configuration combination; a mutual exclusion path is a path that can only be activated by a single configuration combination, which can be obtained by filtering by counting the number of path activations. Its purpose is to locate paths that are highly dependent on configuration conditions.
[0090] The number of activated configuration combinations refers to the statistical value of the number of times a path is triggered by different configuration combinations. Specifically, it can be achieved by traversing the length of the configuration combination list corresponding to each path in the configuration-path association relationship, which is used to quantify the degree of dependence of a path on a specific configuration combination.
[0091] Specifically, during driver configuration testing, the association between configuration combinations and paths is constructed as a hash table. The activation count of each path is counted by traversing the path set corresponding to each configuration combination in the hash table. When the activation count of a path is 1, it indicates that the path can only be triggered by a single configuration combination, and it is then added to the mutually exclusive path set. The unique identifier of the path entry block serves as the association key, ensuring that paths under different control flow branches can be accurately distinguished even if they have similar conditional branches. This mutually exclusive path filtering mechanism effectively identifies paths that highly depend on specific configurations, avoiding invalid path exploration caused by ignoring configuration constraints during subsequent testing.
[0092] Compared to existing technologies, traditional methods, when handling path analysis for configuration drivers, typically only focus on the correlation between input data and paths, ignoring the direct impact of configuration combinations on path activation. Existing path statistics often rely on hash value matching of complete execution trajectories, failing to distinguish path changes caused by configuration differences. This invention introduces a mechanism for statistically analyzing the number of activated configuration combinations, combined with precise identification of path entry blocks, enabling the establishment of a strong correlation model between paths and configurations at the configuration semantic level. This allows for the precise selection of critical paths that can only be triggered by a single configuration.
[0093] Through the above technical solution, this invention solves the problem that traditional fuzz testing cannot effectively identify configuration binding paths in configuration-driven scenarios, avoiding redundant operations of repeatedly attempting to execute mutually exclusive paths in invalid configuration combinations. Simultaneously, based on the identification mechanism of the path entry block, it can accurately distinguish path forks caused by configuration differences, improving the accuracy of path constraint analysis during subsequent symbolic execution.
[0094] Step S6: Obtain the path constraints related to the input features, ignore the configuration constraints, perform symbolic execution on the mutually exclusive paths according to the path constraints, and reconstruct the simplified path constraint expression through the path constraints;
[0095] In one possible implementation, during symbolic execution, configuration judgment conditions related to configuration items are skipped, path condition branches are ignored, and only symbolic conditions related to input data are retained. A simplified path constraint expression is constructed through path constraint reconstruction for subsequent path reachability judgment.
[0096] The configuration judgment condition refers to the branch logic related to the configuration item, which can be implemented through path constraint filtering. Its function is to extract path constraints that depend only on the input data.
[0097] Ignoring path condition branches related to configuration items means actively skipping condition judgment nodes triggered by configuration item values during symbolic execution. Specifically, this can be achieved by locating branch statements related to configuration items in the control flow graph through static analysis and marking them as met during symbolic execution, thereby avoiding interference from additional constraints introduced by configuration items that could hinder the path exploration of input data.
[0098] Path constraint reconstruction to build simplified path constraint expressions refers to filtering and reorganizing the path conditions collected during symbolic execution. Specifically, a constraint classifier can be used to divide the conditions into two categories: configuration-related and input-related. Only logical combinations of input-related constraints are retained, thereby generating constraint expressions that depend only on the input data, which are then used for subsequent path reachability verification across configuration combinations.
[0099] Specifically, during the symbolic execution phase, when analyzing the set of mutually exclusive paths, static analysis is first used to identify judgment statements related to configuration item values in the path conditions, such as macro definition switches or parameter initialization checks. Subsequently, during symbolic execution, these conditions are marked as satisfied and are no longer included in the path constraint collection scope. For the remaining condition branches related to input data, such as file parsing length verification or data format verification, their constraints are recorded according to the normal symbolic execution process. Further, a constraint expression optimization algorithm is used to logically simplify the collected input-related constraints, removing redundant conditions and merging equivalent expressions, ultimately generating a simplified path constraint expression. This expression only contains combinations of input data variables and operators and can be directly used for subsequent path triggering feasibility verification under different configuration combinations.
[0100] Traditional symbolic execution methods, when processing configuration-driven programs, must consider both configuration items and input data constraints simultaneously, leading to an exponential increase in the complexity of path constraint expressions and significantly increasing the computational burden of constraint solving. This invention, however, by removing configuration item-related conditions, focuses symbolic execution on input data-driven path constraints, effectively reducing the variable dimensionality and logical depth of constraint expressions, thus significantly improving the efficiency of subsequent path reachability determination.
[0101] Through the above technical solution, this invention can significantly reduce the number of constraints that need to be processed during symbolic execution without affecting the integrity of the path triggering logic, thereby shortening the path analysis time and improving the efficiency of path verification across configuration combinations. Furthermore, the simplified path constraint expressions are easier to decouple and match with different configuration combinations, providing accurate feasibility data for subsequent dynamic scheduling.
[0102] Step S7: When the path constraint expression can be triggered in a non-original configuration combination, extract the expression input features of the path constraint expression and establish a cross-configuration mapping relationship describing whether the input features can satisfy the path constraint expression under different configuration combinations.
[0103] In one possible implementation, cross-configuration mapping refers to the association between input features and configuration combinations that can trigger the path. Specifically, it can be constructed through constraint decoupling matching, and its role is to guide the collaborative optimization of input variation and configuration scheduling.
[0104] Among them, input feature terms refer to the variables and their constraints that are related to the input data and separated from the path constraint expression. Specifically, they can be realized by performing correlation analysis on the input field access operations and condition judgment nodes recorded during symbolic execution. Key influencing factors can be determined by symbolically tracing the propagation process of input data in the path.
[0105] Constraint decoupling matching refers to logically separating the input-related conditions in the path constraint from the mutually exclusive conditions in the configuration combination. Specifically, it can be achieved by using logical expression decomposition and variable substitution. By eliminating the interference of configuration item variables on the path constraint, an independent association between input features and configuration combinations is established.
[0106] Cross-configuration mapping refers to a two-dimensional relationship table describing whether input features can satisfy path constraints under different configuration combinations. Specifically, it can be implemented by filling matrix cells with Boolean values or probability weight values. For example, cross-validation can be performed between the value range of a specific field in the input feature item and the legal conditions of the configuration combination to generate a mapping relationship.
[0107] Specifically, the simplified path constraint expressions generated during symbolic execution are parsed into a combination of input data variables and logical operators. By traversing the variable references in the constraint expressions, variables related to the program input buffer, file reading interface, or network packet parsing are selected as input features. For each input feature, its corresponding constraint is matched against all possible configuration combinations. First, the mutual exclusion conditions of the configuration combinations are converted into logical expressions, and then the input constraints and configuration conditions are merged using a logical AND operation. The constraint solver is then used to verify whether the input feature under this combination satisfies path reachability. If a feasible solution exists, the correspondence between the input feature and the configuration combination is marked in the feasibility mapping relationship; if not, the possibility of associating the combination is excluded. The generated feasibility mapping relationship is used to guide the scheduling module to preferentially select configuration combinations compatible with the current input mutation direction. For example, when the input mutation engine generates data containing a specific header, the scheduler selects the configuration context that can activate the relevant path for loading based on the matrix matching result.
[0108] Traditional methods fail to decouple input constraints and configuration conditions during cross-configuration path mapping, making it impossible to distinguish the independent dependencies of path triggers on input features and configuration combinations. Existing technologies typically employ a combination testing strategy of fixed configurations and input variations, resulting in numerous invalid configuration switches and redundant executions. This invention, through constraint separation and feasibility mapping relationship modeling, clarifies the applicable boundaries of input features under different configurations, enabling the scheduler to dynamically select effective configuration combinations based on input variation features, avoiding the repeated execution of invalid tests under mutually exclusive configuration conditions.
[0109] Through the above technical solution, this invention solves the problem of low path coverage efficiency caused by the coupling of input and configuration in configuration driver testing. By establishing a feasibility mapping relationship between input features and configuration combinations, it can accurately match the direction of input variation and compatible configuration environments, reducing path exploration failures caused by configuration condition conflicts. Simultaneously, the scheduling mechanism driven by the feasibility mapping relationship avoids invalid attempts on irrelevant configuration combinations, thereby improving the path triggering success rate and execution efficiency of fuzz testing under limited resources.
[0110] Step S8: When the path constraint expression cannot be triggered in a non-original configuration combination, bind the path to the original configuration combination and generate configuration constraints.
[0111] Step S9: Adjust the test order of configuration combinations according to cross-configuration mapping relationships or configuration constraints in order to conduct the next round of fuzz testing.
[0112] In one possible implementation, after each round of fuzzing, the number of paths first triggered by each configuration combination in that round is counted to obtain the number of new paths for each configuration combination. The coverage feedback score of the configuration combination is obtained based on the ratio of the number of new paths to the number of rounds the corresponding configuration combination has been scheduled. The coverage feedback score is expressed as the ratio between the number of new paths and the total number of rounds the configuration combination has been scheduled for testing, reflecting the marginal coverage value of the combination at the current stage. The coverage feedback scores of all configuration combinations are normalized to obtain normalized scores for each configuration combination within the same scoring scale range; for example, the highest score is mapped to 1, the lowest score to 0, and the remaining scores are linearly mapped to this range. All configuration combinations to be tested are sorted in descending order according to the normalized scores to obtain the scheduling priority for the next round of fuzzing. The testing order of each configuration combination is adjusted according to the scheduling priority of the next round and the current scheduling priority of each configuration combination for the next round of fuzzing.
[0113] When the ratio of the next-round scheduling priority to the current scheduling priority of a configuration combination exceeds a preset ratio, it indicates a decrease in the marginal benefit of the configuration combination's coverage capability, and the scheduling frequency of the corresponding configuration combination in the next round of fuzzing is reduced. When the decrease in the coverage feedback score of a configuration combination exceeds a preset decrease threshold, the interval period for the next round of fuzzing for the corresponding configuration combination is extended. When the coverage feedback score of a configuration combination exceeds a preset score threshold, the priority of the corresponding configuration combination in the next round of fuzzing is increased. An updated configuration combination scheduling queue is generated and passed to the next-round fuzzing task scheduler as the basis for the order in which the fuzzing engine loads the configuration context.
[0114] The number of newly added paths refers to the number of paths that were not covered by historical tests and were triggered for the first time in the current test round for a specific configuration combination. This can be achieved by comparing path hash values or detecting sequence differences in basic block execution. This metric is used to quantify the immediate exploratory value of configuration combinations.
[0115] Coverage feedback score is a dynamic evaluation metric calculated by the ratio of the number of new paths to the number of scheduling rounds. Specifically, it can be implemented using sliding window statistics or exponential decay weighting methods. This metric is used to reflect the marginal coverage contribution rate of the configuration combination in the current testing phase.
[0116] Normalization refers to the process of converting coverage feedback scores of different dimensions into a unified scoring range. Specifically, it can be achieved by range standardization or quantile mapping methods. This process ensures that the priorities of different configuration combinations are comparable.
[0117] Scheduling frequency adjustment refers to dynamically controlling the test interval of the configuration combination based on changes in coverage feedback scores. Specifically, it can be implemented through priority queue reordering or time slice allocation algorithms. This mechanism is used to balance the allocation of test resources for exploration and utilization.
[0118] Specifically, after each round of fuzz testing, the newly added paths triggered by each configuration combination are first counted, and the first-time occurrence of a path is selected by comparing path identifiers. Then, a coverage feedback score is calculated based on the ratio of the number of newly added paths to the number of historical scheduling attempts. A higher ratio indicates that the configuration combination has a higher marginal coverage potential in the current stage. All scores are normalized to form a standardized priority score; for example, the highest-scoring configuration is mapped to priority 1, and the rest are distributed proportionally in the 0-1 range. A descending-order scheduling list is generated based on the normalized scores. Configuration combinations with significantly decreased scores have their subsequent test intervals extended, while high-scoring configurations are prioritized for testing. The final generated scheduling queue will guide the loading order of configuration contexts in the next round of testing, achieving dynamic optimization of test resource allocation.
[0119] Compared to existing technologies, traditional methods typically employ fixed-order or random scheduling strategies to perform multi-configuration fuzz testing, failing to adjust testing priorities based on real-time coverage performance, resulting in the failure to identify high-value configuration combinations in a timely manner. This invention, however, establishes a closed-loop mechanism of coverage feedback and dynamic scheduling, enabling rapid identification of configuration combinations with declining marginal returns and reducing their resource consumption. Simultaneously, it prioritizes the execution of configurations with high exploration potential, thereby improving the targeting of path triggering and testing efficiency.
[0120] Through the above technical solution, this invention can effectively reduce the number of redundant tests on inefficient configuration combinations and prioritize the scheduling of configuration combinations with high marginal coverage value, thereby significantly improving path exploration efficiency with limited testing resources. This mechanism adjusts the testing focus through real-time feedback, avoiding continuous resource investment in configurations with saturated coverage capabilities, while accelerating in-depth exploration of high-potential configurations, ultimately achieving efficient coverage of the configuration-driven path space.
[0121] This invention further proposes using the scheduling order update result to control the configuration context switching frequency of the input mutation engine in the next round of fuzzing, thereby improving the coupling efficiency of symbolic execution and fuzzing and the targeting of path triggering.
[0122] The scheduling order update result refers to the priority list of configuration combinations dynamically generated based on coverage feedback. Specifically, it can be implemented using a normalized score sorting algorithm. By calculating the ratio of the number of new paths added to the historical scheduling count for each configuration combination, it dynamically reflects the path exploration value at the current stage. This result reduces the repeated execution of invalid combinations by adjusting the loading order of the configuration context.
[0123] The configuration context switching frequency refers to the rate at which the input mutation engine switches between different configuration combinations. This can be controlled by adjusting the interval between configuration combinations in the scheduling queue. For example, the scheduling interval can be extended for configuration combinations with decreasing coverage feedback scores, while combinations with higher scores can be executed first. Optimizing this frequency avoids resource waste caused by frequent switching and improves the collaborative efficiency of symbolic execution and fuzzing.
[0124] Specifically, after each round of fuzzing, the system calculates a coverage feedback score based on the new path coverage of the configuration combinations, and generates a priority list after normalizing the score. The scheduling module dynamically adjusts the execution order of configuration combinations in the next round of testing based on this list. For configuration combinations with high scores, the system shortens their scheduling interval, allowing them to be loaded into the input mutation engine first; for combinations with continuously decreasing scores, the system extends their scheduling interval or reduces their execution frequency. By controlling the switching frequency of configuration contexts, the system can concentrate testing resources on configuration combinations that are more likely to trigger new paths, thereby improving the coupling efficiency between symbolic execution and fuzzing.
[0125] Existing fuzzing methods typically employ fixed or random scheduling strategies, resulting in high-value configuration combinations failing to execute promptly, while low-value combinations repeatedly consume resources. This invention, through a dynamic priority adjustment mechanism, optimizes the execution order of configuration combinations based on real-time coverage feedback, avoiding redundant testing. Simultaneously, by controlling the switching frequency, it reduces context switching overhead, enabling the path constraints generated by symbolic execution to more accurately guide the input mutation process.
[0126] Through the above technical solution, the present invention can effectively improve the collaborative efficiency of symbolic execution and input mutation in configuration driver fuzzing, reduce resource waste caused by invalid configuration switching, and enhance the targeting of path triggering to accelerate the discovery of new paths.
[0127] In summary, this invention first extracts configuration items and their mutual exclusion relationships through static analysis to generate legal configuration combinations, thereby reducing invalid tests. Then, fuzz testing is performed on each configuration combination, recording the triggered paths and establishing a configuration combination-path association. By analyzing these associations, mutually exclusive paths that can only be activated by a single configuration combination are selected, and symbolic execution is performed on these paths to remove configuration-related constraints, retaining only the input-related path conditions. If the simplified path constraint expression can be satisfied in other configuration combinations, a mapping relationship between input features and multiple configuration combinations is established; otherwise, the path is bound to the original configuration combination to avoid repeatedly trying invalid combinations in subsequent tests. Finally, the scheduling priority of configuration combinations is dynamically adjusted based on the newly added path coverage in each round of testing, prioritizing combinations with high marginal coverage value, thereby improving overall testing efficiency.
[0128] Compared to existing technologies, traditional multi-configuration fuzzing employs fixed or random scheduling strategies, failing to optimize test order based on path coverage feedback, leading to resource waste. Our method, however, dynamically identifies mutually exclusive paths and reconstructs constraints, deeply integrating symbolic execution with fuzzing to effectively distinguish between configuration-dependent and input-driven paths. Simultaneously, the coverage feedback mechanism updates scheduling priorities in real time, prioritizing the exploration of high-value configuration combinations and avoiding redundant execution. Compared to the inefficiency of existing symbolic execution methods in solving configuration-related path constraints, our method simplifies constraints by skipping configuration conditions, significantly reducing solution complexity.
[0129] Through the above technical solutions, this invention solves the problem of low path coverage efficiency in configuration drivers. It reduces the number of tests for invalid configuration combinations by establishing configuration-path associations and cross-configuration mappings; optimizes the test order through dynamic scheduling strategies, prioritizing the execution of configuration combinations with high coverage potential; and reduces symbolic execution complexity and improves path analysis efficiency through constraint refactoring. Ultimately, it significantly reduces test resource consumption while maintaining path coverage depth.
[0130] This invention also provides a symbolic execution-based fuzzing system for performing any of the above-described fuzzing methods, comprising:
[0131] The first acquisition module is used to acquire multiple configuration items in the configuration driver;
[0132] The mutual exclusion relationship identification module is used to identify the mutual exclusion relationships between configuration items and combine configuration items according to the mutual exclusion relationships to obtain multiple configuration combinations;
[0133] The testing module is used to perform fuzz tests on each configuration combination and record the paths triggered by each configuration combination.
[0134] The module for establishing associations is used to establish associations based on configuration combinations and paths;
[0135] The mutual exclusion path identification module is used to identify mutual exclusion paths that appear only in a single configuration combination based on the association relationship; the configuration combination in which the mutual exclusion path appears is the original configuration combination.
[0136] The second acquisition module is used to acquire path constraints related to the input features, ignore configuration constraints, perform symbolic execution on mutually exclusive paths according to the path constraints, and reconstruct simplified path constraint expressions through the path constraints.
[0137] When a path constraint expression can be triggered in a non-original configuration combination, extract the expression input features of the path constraint expression and establish a cross-configuration mapping relationship describing whether the input features can satisfy the path constraint expression under different configuration combinations.
[0138] When the path constraint expression cannot be triggered in a non-original configuration combination, the path is bound to the original configuration combination and a configuration constraint is generated.
[0139] The adjustment module is used to adjust the test order of configuration combinations based on cross-configuration mapping relationships or configuration constraints, in order to conduct the next round of fuzz testing.
[0140] Compared with existing technologies, the beneficial effects of the symbolic execution-based fuzzing method and system provided by this invention include:
[0141] 1) By introducing a mechanism for identifying mutually exclusive relationships of configuration items and determining the legality of configuration combinations, combined with path-configuration association modeling, redundant or invalid configuration combinations can be effectively screened out, avoiding the resource waste caused by the exponential growth of configuration combinations in traditional fuzzing, and improving the effectiveness of path exploration and overall testing efficiency.
[0142] 2) By statistically analyzing the activation status of paths in each configuration combination, mutually exclusive path sets are identified, and configuration conditions are skipped in symbolic execution, retaining only the path constraints related to the input. This establishes a cross-configuration mapping relationship between input features and path reachability, effectively solving the problem of high path constraint complexity in configuration-driven symbolic execution and improving path solving efficiency.
[0143] 3) Introduce a coverage feedback score mechanism based on the number of new paths, build a dynamic scheduling priority model, and perform real-time scheduling, sorting and frequency adjustment of configuration combinations to ensure that test resources are given priority to configuration combinations with high marginal coverage value, thereby improving the path discovery rate and reducing redundant test overhead.
[0144] 4) By establishing a cross-configuration mapping relationship between input features and configuration combinations, the input mutation strategy and configuration combination selection work together to avoid path pruning or trigger failure caused by configuration conflicts, effectively enhancing the targeting and efficiency of path triggering in fuzzing and improving test coverage in complex configuration environments.
[0145] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A fuzz testing method based on symbolic execution, characterized in that, include: Retrieve multiple configuration items from the driver configuration; Identify the mutual exclusion relationships between the configuration items, and combine the configuration items according to the mutual exclusion relationships to obtain multiple configuration combinations; Perform fuzz testing on each of the configuration combinations and record the paths triggered by each configuration combination; Establish an association relationship based on the configuration combination and the path; Based on the association, identify mutually exclusive paths that appear only in a single configuration combination; the configuration combination in which the mutually exclusive path appears is the original configuration combination; Obtain path constraints related to input features, ignore configuration constraints, perform symbolic execution on the mutually exclusive paths according to the path constraints, and reconstruct simplified path constraint expressions through the path constraints; When the path constraint expression can be triggered in a non-original configuration combination, the expression input features of the path constraint expression are extracted, and a cross-configuration mapping relationship describing whether the input features can satisfy the path constraint expression under different configuration combinations is established. When the path constraint expression cannot be triggered in a non-original configuration combination, the path is bound to the original configuration combination and a configuration constraint is generated; The test order of configuration combinations is adjusted according to the cross-configuration mapping relationship or the configuration constraints in order to conduct the next round of fuzz testing.
2. The fuzz testing method according to claim 1, characterized in that, The process of obtaining the set of configuration items in the configuration driver includes: Static analysis is performed on the source code of the configuration driver to identify the configuration parameters in the source code, the configuration parameters are transformed into the structure of configuration items, and the corresponding nodes of the configuration items are established in the control flow graph.
3. The fuzz testing method according to claim 2, characterized in that, The identification of mutual exclusion relationships between the configuration items includes: The logical expression is solved based on the Boolean variables on which the branch conditions in the control flow graph depend. When multiple configuration items have conditions that cannot be satisfied simultaneously during program execution, these multiple configuration items are recorded as mutually exclusive to constrain the generation of configuration combinations.
4. The fuzz testing method according to claim 1, characterized in that, Performing fuzz testing on each of the configuration combinations includes: Generate independent test contexts based on each configuration combination; Fuzz testing is performed on each of the configuration combinations based on the independent test context and input mutation strategy; the input mutation strategy is a method that modifies only the input data and not the configuration items of the configuration combination when generating fuzz test cases.
5. The fuzz testing method according to claim 1, characterized in that, The step of identifying mutually exclusive paths that appear only in a single configuration combination based on the association includes: The number of times each path is triggered by different configuration combinations is calculated based on the aforementioned relationships; When any path is triggered once by different configuration combinations, the path is determined to be a mutually exclusive path.
6. The fuzz testing method according to claim 5, characterized in that, Also includes: Whether a path is configured to be triggered is determined based on the unique identifier of the path's entry basic block.
7. The fuzz testing method according to claim 6, characterized in that, The unique identifier of the entry basic block of the path is the hash value or address code of the first basic block in the control flow graph.
8. The fuzz testing method according to claim 1, characterized in that, The step of adjusting the test order of configuration combinations according to the cross-configuration mapping relationship or the configuration constraints for the next round of fuzz testing includes: After each round of fuzz testing is completed, the number of paths that are first triggered for each configuration combination in that round of fuzz testing is counted to obtain the number of new paths for each configuration combination. The coverage feedback score of the configuration combination is obtained by the ratio of the number of newly added paths to the number of rounds in which the corresponding configuration combination is scheduled. The coverage feedback scores of all configuration combinations are normalized to obtain the normalized scores of each configuration combination within the same scoring scale range. Based on the normalized score, all configuration combinations to be tested are sorted in descending order to obtain the scheduling priority for the next round of fuzz testing; The test order of each configuration combination is adjusted according to the next round of scheduling priority and the current scheduling priority of each configuration combination in order to conduct the next round of fuzz testing.
9. The fuzz testing method according to claim 8, characterized in that, The step of adjusting the test order of each configuration combination according to the next round of scheduling priority and the current scheduling priority of each configuration combination for the next round of fuzz testing includes: When the ratio of the next round scheduling priority of a configuration combination to the current scheduling priority exceeds a preset ratio, the scheduling frequency of the next round of fuzz testing for the corresponding configuration combination is reduced. When the decrease in the coverage feedback score of a configuration combination exceeds a preset decrease threshold, the interval period for the next round of fuzz testing for the corresponding configuration combination is extended. When the coverage feedback score of a configuration combination exceeds a preset score threshold, the priority of the corresponding configuration combination in the next round of fuzz testing is increased.
10. A symbolic execution-based fuzzing system for performing the fuzzing method as described in any one of claims 1-9, characterized in that, include: The first acquisition module is used to acquire multiple configuration items in the configuration driver; The mutual exclusion relationship identification module is used to identify the mutual exclusion relationship between the configuration items, and combine the configuration items according to the mutual exclusion relationship to obtain multiple configuration combinations; The testing module is used to perform fuzz tests on each of the configuration combinations and record the paths triggered by each configuration combination. A module for establishing associations is used to establish associations based on the configuration combination and the path. The mutual exclusion path identification module is used to identify mutual exclusion paths that appear only in a single configuration combination based on the association relationship; the configuration combination in which the mutual exclusion path appears is the original configuration combination. The second acquisition module is used to acquire path constraints related to the input features, ignore configuration constraints, perform symbolic execution on the mutually exclusive paths according to the path constraints, and reconstruct simplified path constraint expressions through the path constraints. When the path constraint expression can be triggered in a non-original configuration combination, the expression input features of the path constraint expression are extracted, and a cross-configuration mapping relationship describing whether the input features can satisfy the path constraint expression under different configuration combinations is established. When the path constraint expression cannot be triggered in a non-original configuration combination, the path is bound to the original configuration combination and a configuration constraint is generated; The adjustment module is used to adjust the test order of configuration combinations according to the cross-configuration mapping relationship or the configuration constraints, in order to conduct the next round of fuzz testing.
Citation Information
Patent Citations
Intelligent fuzzy test method, device and system
CN112181833A
Fuzzy test system and method based on symbolic execution
CN116541294A