Vehicle-mounted computer test case automatic generation method and system based on path coverage
Through static analysis and instrumentation technology, the initial training data and instrumentation path are generated, and the training test cases automatically generates a model, solving the problem of manual writing and limited coverage in traditional test methods, and achieving efficient and automated test case generation.
Patent Information
- Application Number
- CN202510136434.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-06-03
AI Technical Summary
Traditional path coverage-based testing methods have problems such as time-consuming writing test cases manually, limited coverage, easy to ignore complex and boundary paths, and lack of effective optimization strategies.
The on-board computer control software is analyzed through static analysis tools, the data flow diagram is constructed, the initial training data is generated, and the instrumentation path is obtained through the instrumentation program. The model is automatically generated based on these path training test cases.
Improve the code coverage of test cases, find more potential defects, and the generated test cases are diverse and representative, reducing the time and cost of manually designed test cases, and achieving automated test case generation.
Smart Images

Figure CN120086135A_ABST
Abstract
Description
Technical Field
[0001] The present invention provides a method and system for automatically generating vehicle computer test cases based on path coverage, which relates to the technical field of software testing. Background Art
[0002] The test method based on path coverage is committed to ensuring the comprehensiveness of testing by generating test cases that cover all possible execution paths of the software. Path coverage enables testers to systematically analyze and generate test cases based on the control flow logic of the program, ensuring that all branch conditions and paths are fully tested. Compared with random testing or functional testing, path coverage testing can discover more hidden defects, especially complex conditions and edge cases.
[0003] However, traditional manual writing of test cases is time-consuming and has limited coverage; traditional test methods and random test case generation are prone to ignoring complex and boundary paths, resulting in incomplete coverage; traditional test methods only through static analysis fail to fully consider the dynamic characteristics of program execution, resulting in some paths being missed; traditional test methods lack effective optimization strategies and cannot fully explore large-scale complex path combinations. Summary of the Invention
[0004] The present invention provides a method and system for automatically generating vehicle computer test cases based on path coverage to solve the above-mentioned problems:
[0005] The method for automatically generating vehicle computer test cases based on path coverage proposed by the present invention includes:
[0006] Parsing the vehicle computer control software through a static analysis tool to construct a data flow graph, which takes the basic blocks in the control program as nodes and the changes in the control flow as edges, and generating initial training data based on the data flow graph;
[0007] Generating an instrumentation program, inputting the initial training data into the instrumentation program to obtain instrumentation paths, and training an automatic vehicle computer test case generation model based on the instrumentation paths.
[0008] Further, parsing the vehicle computer control software through a static analysis tool to construct a data flow graph, which takes the basic blocks in the control program as nodes and the changes in the control flow as edges, and generating initial training data based on the data flow graph includes:
[0009] Parsing the C code of the vehicle computer control software through Clang to generate an LLVM IR file;
[0010] Using the LLVM toolchain to parse the LLVM IR file into a data flow graph, which records variables and dependencies;
[0011] Extract the constraint relationships between variables according to the data flow diagram, where the constraint relationships between variables refer to the mutual dependencies and constraint conditions between different variables in the program;
[0012] Use a constraint solver and a random seed to generate different initial training data that satisfy the constraint relationships.
[0013] Further, generate an instrumentation program, input the initial training data into the instrumentation program, obtain instrumentation paths, and train an automatic generation model for in-vehicle computer test cases based on the instrumentation paths, including:
[0014] S31. Generate an instrumentation program for the in-vehicle computer control software program, and input the initial training data into the instrumentation program;
[0015] S32. Run and compile the instrumentation program, obtain instrumentation paths, screen the instrumentation paths based on path differences, and calculate the hierarchical depth of the screened instrumentation paths;
[0016] S33. Use the initial training data, instrumentation paths, and the hierarchical depth of the instrumentation paths as original samples, extract features from the original samples to generate test data, and select the pre-trained model Transformer;
[0017] S34. Input the test data into the pre-trained model, train the sub-model, and test whether the accuracy of the sub-model reaches the first preset threshold. If it reaches the first preset threshold, number the sub-model and save it to the trained sub-model library;
[0018] S35. If it does not reach the first preset threshold, sequentially add the node information in the instrumentation path, and then determine whether all sub-models have been trained. If not all sub-models have been trained, continue to train the sub-model;
[0019] S36. If all sub-models have been trained, check the number of sub-models, and determine whether the number of sub-models is equal to the number of nodes in the instrumentation path. If they are equal, link the sub-models according to the node order of the instrumentation path to obtain a completed automatic generation model for in-vehicle computer test cases;
[0020] S37. If the number of sub-models is not equal to the number of nodes in the instrumentation path, continue to select other pre-trained models, find the path nodes for which sub-models have not been constructed, and return to step S34 to continue execution for the path nodes for which sub-models have not been constructed.
[0021] Further, screening the instrumentation paths based on path differences includes:
[0022] List all the basic blocks that will appear in the program, and define the instrumentation path based on the basic blocks;
[0023] According to the order of the basic block set, record the number of times each basic block appears in the path based on the defined instrumentation path. The basic block serves as the dimension of the vector to obtain the vector representation of the instrumentation path;
[0024] Calculate the sparsity index of the path;
[0025] Calculate the weight of the path:
[0026] w = α·f + β·b
[0027] where w represents the weight, f represents the execution frequency of the path, b represents the number of basic blocks covered by the path, and α and β are the weights of the adjustment parameters;
[0028] Calculate the difference of the instrumentation path based on the path difference model. Specifically, the path difference model is:
[0029]
[0030] where D represents the difference of the path, n represents the number of instrumentation paths, w i and w j represent the weights of path p i and path p j respectively, d(p i , p j ) represents the distance metric function between path p i and path p j , and S represents the sparsity evaluation index;
[0031]
[0032] where d represents the number of data points in the instrumentation path, p i,k and p j,k represent the values of path p i and path p j at the k-th data point respectively, and δ 1 , δ 2 and δ 3 represent the weight coefficients;
[0033] If the calculated path difference is less than the second preset threshold, delete one of the paths whose distance calculated based on the distance metric function is less than the second preset threshold.
[0034] Furthermore, calculating the sparsity index of the path includes:
[0035] The situation of path - covered basic blocks is obtained based on the vector representation of the path, and the situation of path - covered basic blocks is converted into a path matrix;
[0036] The path matrix is summed by column to obtain the total number of times each basic block is covered by all paths, and a summation vector f is obtained;
[0037] The mean and variance of the summation vector are calculated, the standardized variance of the summation vector is calculated based on the mean and variance of the summation vector, and the sparsity index of the path is calculated based on the standardized variance:
[0038]
[0039] Among them, S represents the sparsity index, f is a vector composed of the coverage frequencies of each path for all basic blocks, represents the standardized variance, Var(f) is the variance of the summation vector, and μ(f) is the mean of the summation vector.
[0040] An automatic test case generation system for in - vehicle computers based on path coverage proposed by the present invention, the system includes:
[0041] A training data generation module, which is used to parse the in - vehicle computer control software through a static analysis tool to construct a data flow graph. The data flow graph takes the basic blocks in the control program as nodes and the changes in the control flow as edges, and generates initial training data based on the data flow graph;
[0042] A model training module, which is used to generate an instrumented program, input the initial training data into the instrumented program to obtain instrumented paths, and train an automatic test case generation model for in - vehicle computers based on the instrumented paths.
[0043] Furthermore, the training data generation module includes:
[0044] A code parsing module, which parses the C code of the in - vehicle computer control software through Clang to generate an LLVM IR file;
[0045] A data flow graph acquisition module, which is used to parse the LLVM IR file into a data flow graph using the LLVM tool chain. The data flow graph records variables and dependency relationships;
[0046] A constraint relationship extraction module, which is used to extract the constraint relationships between variables according to the data flow graph. The constraint relationships between variables refer to the mutual dependencies and constraint conditions between different variables in the program;
[0047] An initial training data generation module, which is used to generate different initial training data that satisfy the constraint relationships using a constraint solver and a random seed.
[0048] Further, the training model module includes:
[0049] A stub program generation module, which is used to generate a stub program for the in-vehicle computer control software program and input the initial training data into the stub program;
[0050] A screening module, which is used to run and compile the stub program, obtain the stub paths, screen the stub paths based on path differences, and calculate the hierarchical depth of the screened stub paths;
[0051] A test data generation module, which is used to use the initial training data, the stub paths, and the hierarchical depth of the stub paths as raw samples, extract the features in the raw samples to generate test data, and select the pre-trained model Transformer;
[0052] A sub-model number storage module, which is used to input the test data into the pre-trained model, train the sub-model, test whether the accuracy of the sub-model reaches the first preset threshold. If it reaches the first preset threshold, number the sub-model and save it to the trained sub-model library;
[0053] A node information addition module, which is used to, if the first preset threshold is not reached, sequentially add the node information in the stub paths, and then determine whether all sub-models have been trained. If not all sub-models have been trained, continue to train the sub-models;
[0054] A final model linking module, which is used to, if all sub-models have been trained, check the number of sub-models, determine whether the number of sub-models is equal to the number of nodes in the stub paths. If they are equal, link the sub-models according to the node order of the stub paths to obtain the built in-vehicle computer test case automatic generation model;
[0055] A continuous optimization module, which is used to, if the number of sub-models is not equal to the number of nodes in the stub paths, continue to select other pre-trained models, find the path nodes for which sub-models have not been built, and return to step S34 to continue execution for the path nodes for which sub-models have not been built.
[0056] Further, the screening module includes:
[0057] A path definition module, which is used to list all the basic blocks that will appear in the program and define the stub paths based on the basic blocks;
[0058] A vector representation path module, which is used to record the number of times each basic block appears in the path according to the order of the basic block set, use the basic blocks as the dimensions of the vector, and obtain the vector representation of the stub path;
[0059] A sparsity calculation module, which is used to calculate the sparsity index of the path;
[0060] A path weight calculation module for calculating the weight of a path:
[0061] w = α·f + β·b
[0062] Where w represents the weight, f represents the execution frequency of the path, b represents the number of basic blocks covered by the path, and α and β are the weights of the adjustment parameters;
[0063] A path difference calculation module for calculating the difference of the instrumented paths based on a path difference model. Specifically, the path difference model is:
[0064]
[0065] Where D represents the difference of the path, n represents the number of instrumented paths, w i and w j represent the weights of path p i and path p j respectively, d(p i , p j ) represents the distance metric function between path p i and path p j , and S represents the sparsity evaluation index;
[0066]
[0067] Where d represents the number of data points in the instrumented path, p i,k and p j,k represent the values of path p i and path p j at the k-th data point respectively, δ 1 , δ 2 and δ 3 represent the weight coefficients;
[0068] A path deletion module for deleting one of the paths with a distance less than a second preset threshold calculated based on the distance metric function if the calculated path difference is less than the second preset threshold.
[0069] Furthermore, the sparsity calculation module includes:
[0070] A path matrix conversion module for obtaining the situation of the basic blocks covered by the path based on the vector representation of the path and converting the situation of the basic blocks covered by the path into a path matrix;
[0071] A summation vector acquisition module for summing the path matrix by column to obtain the total number of times each basic block is covered by all paths and obtaining the summation vector f;
[0072] A sparsity index calculation module is used to calculate the mean and variance of the summation vector, calculate the standardized variance of the summation vector based on the mean and variance of the summation vector, and calculate the sparsity index of the path based on the standardized variance:
[0073]
[0074] where S represents the sparsity index, f is a vector composed of the coverage frequencies of each path for all basic blocks, represents the standardized variance, Var(f) is the variance of the summation vector, and μ(f) is the mean of the summation vector.
[0075] Advantages of the present invention: Improve the coverage rate. The test cases generated by the intelligent model can cover different paths, a large number of boundary cases and extreme cases of the program, thereby improving the code coverage rate of the test cases and discovering more potential defects; Instrumentation and screening make the generated test cases diverse and representative, preventing over-concentration on a few paths; Automated test case generation reduces the time and labor costs of manually designing test cases and speeds up the test process; The instrumentation program can efficiently capture the behavior data during program operation and provide sufficient information for subsequent automated analysis; The test case generation technology based on the deep learning model can dynamically generate effective test cases according to the program operation situation, realizing intelligence and adaptability; By using a high-precision training model, problems that may occur in the program can be predicted more accurately, improving the effectiveness and accuracy of the test; Wide and in-depth path coverage ensures comprehensive testing of the in-vehicle computer control software and increases the probability of discovering problems; The high reliability and stability of the in-vehicle computer control software are crucial. Comprehensive testing can effectively improve the system security and reduce potential failures. BRIEF DESCRIPTION OF THE DRAWINGS
[0076] Figure 1 It is a schematic diagram of the method for automatically generating test cases for in-vehicle computers based on path coverage according to the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0077] In order to more clearly understand the above objects, features and advantages of the present invention, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, without conflict, the embodiments of the present application and the features in the embodiments can be combined with each other.
[0078] Many specific details are set forth in the following description in order to fully understand the present invention. The described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0079] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which this invention belongs. The terms used in the specification of this invention herein are for the purpose of describing specific embodiments only and are not intended to limit the invention.
[0080] An embodiment of the present invention is an automatic generation method for in-vehicle computer test cases based on path coverage. The method includes:
[0081] Parsing the in-vehicle computer control software through a static analysis tool to construct a data flow graph, where the basic blocks in the control program are used as nodes and the changes in the control flow are used as edges, and generating initial training data based on the data flow graph;
[0082] Generating an instrumentation program, inputting the initial training data into the instrumentation program to obtain instrumentation paths, and training an automatic generation model for in-vehicle computer test cases based on the instrumentation paths.
[0083] The working principle and effects of the above technical solution are as follows: Use a static analysis tool to analyze the in-vehicle computer control software. The static analysis tool extracts the program's structure and control flow information by analyzing the source code or intermediate code (such as bytecode) without executing the code; through static analysis, a DataFlow Graph (DFG) is generated. In the DFG, basic blocks (one or more sequentially executed statements) serve as the nodes of the graph, and the control flow (the jump relationship between basic blocks) serves as the edges of the graph. The data flow graph intuitively represents the flow of the program between different basic blocks; rich initial training data is extracted from the data flow graph, and this data includes information about each basic block and the connection relationships between basic blocks. The initial training data includes variables on which the basic block depends or has an impact, the control structure of the basic block, etc.; the initial training data is input into the instrumentation program. Instrumentation is to add additional code (instrumentation points) at specified locations in the program to collect runtime information when the program runs; run the compiled instrumentation program to collect and record the instrumentation paths (i.e., the sequence of basic blocks passed through during the actual execution of the program); through instrumentation, the execution paths of the program and their related runtime data can be collected; screen the instrumentation paths, filter out redundant or invalid paths, and retain representative and diverse paths. This process is based on path diversity, that is, select instrumentation paths that cover different combinations of basic blocks and different control paths; use the screened instrumentation paths and related data to train an automatic generation model, using a deep learning model. The model can finally automatically generate test cases that cover different control paths by learning these path data. Improve the coverage rate. The test cases generated by the intelligent model can cover different paths, a large number of boundary conditions, and extreme conditions of the program, thereby increasing the code coverage rate of the test cases and discovering more potential defects; instrumentation and screening make the generated test cases diverse and representative, preventing over-concentration on a few paths; automated test case generation reduces the time and labor costs of manually designing test cases and speeds up the testing process; the instrumentation program can efficiently capture the behavior data during program operation and provide sufficient information for subsequent automated analysis; the test case generation technology based on the deep learning model can dynamically generate effective test cases according to the program's running conditions, realizing intelligence and adaptability; by using a high-precision training model, potential problems that may occur in the program can be predicted more accurately, improving the effectiveness and accuracy of testing; extensive and in-depth path coverage ensures comprehensive testing of the in-vehicle computer control software and increases the probability of discovering problems; the high reliability and stability of the in-vehicle computer control software are crucial, and comprehensive testing can effectively improve the system's security and reduce potential failures.
[0084] In one embodiment of the present invention, a static analysis tool is used to parse the in-vehicle computer control software to construct a data flow graph. The data flow graph takes the basic blocks in the control program as nodes and the changes in the control flow as edges. Based on the data flow graph, initial training data is generated, including:
[0085] Parse the C code of the in-vehicle computer control software through Clang to generate an LLVM IR file;
[0086] Use the LLVM toolchain to parse the LLVM IR file into a data flow graph, and the data flow graph records variables and dependencies;
[0087] Extract the constraint relationships between variables according to the data flow graph. The constraint relationships between variables refer to the mutual dependencies and constraint conditions between different variables in the program;
[0088] Use a constraint solver and a random seed to generate different initial training data that satisfy the constraint relationships.
[0089] The working principle and effects of the above technical solution are as follows: Use the Clang compiler to parse the C code of the in-vehicle computer control software. Clang is a highly extensible and modular compiler for the C language family (C / C++ / Objective-C / C++), and its design is based on the LLVM project; Clang compiles the C code into LLVM Intermediate Representation (LLVM IR), which is a low-level, language-agnostic intermediate representation. The LLVM IR file retains the control flow and data flow information of the program, facilitating further analysis; LLVM toolchain: Use the LLVM toolchain (such as the static analysis tool of LLVM) to parse the LLVM IR file and construct a Data Flow Graph (DFG); In the data flow graph, nodes represent variables, and edges represent the data dependency relationships and constraint conditions between variables. Recording variables and their corresponding dependency relationships helps to comprehensively understand the behavior and data flow of the program; Extract the constraint relationships between variables from the data flow graph, and the constraint relationships include dependencies, constraint conditions, assignment relationships, and comparison relationships between variables; Use a constraint solver tool (such as Z3, CVC4, etc.) to parse the extracted variable constraint relationships and generate specific values that satisfy these constraints. A constraint solver is a tool specifically used to solve logical constraint conditions. By given constraint conditions, it solves specific solutions that meet the conditions. Different initial training data are generated by random seeds to ensure the coverage and representativeness of the training data with diversity and richness. By different random seeds, different combinations of variable values can be generated, but all these combinations satisfy the extracted constraint relationships.Improve the diversity and coverage of test data. By using random seeds and constraint solvers to generate multiple different combinations of initial training data that satisfy the constraint relationships, the diversity of the training data is enhanced; the extensive variable combinations and their corresponding constraint relationships ensure the comprehensive coverage of the generated test data, thereby improving the test effect and defect discovery ability; the extracted variable constraint relationships accurately reflect the dependencies and restrictions between variables in the program, and the generated initial training data can accurately simulate the actual running situation of the program; the variable values generated through constraint solving conform to the logic and conditions during actual operation, ensuring that the test data has a high sense of reality; utilize toolchains such as Clang, LLVM, and constraint solvers to automate the generation of data flow graphs and data, reducing manual intervention and improving the efficiency of test data generation; the constraint solver can quickly solve complex logical constraint conditions and generate test data that meets the conditions, improving the efficiency of the entire process; improve the reliability and safety of in-vehicle computer control software. The diverse and comprehensive test data covers various possible running states and paths of the software, helping to discover potential defects and vulnerabilities. Precise constraint solving ensures the high quality of the test data, effectively enhancing the effectiveness and reliability of the test, and guaranteeing the stability and safety of the in-vehicle computer control software; the data flow graph and constraint relationships can handle complex variable dependencies and logical conditions in the program and are applicable to in-vehicle control software with high complexity; strong adaptability. The technical solution is based on abstract intermediate representations and toolchain analysis and can adapt to different types of control software and various complex application scenarios; through steps such as static analysis, LLVM IR parsing, data flow graph construction, constraint extraction, and constraint solving, high-quality initial training data is automatically generated, effectively improving the diversity, coverage, and accuracy of the test data. Its effects are significantly reflected in improving test efficiency, enhancing the reliability and safety of in-vehicle computer control software, and supporting the comprehensive analysis and testing of complex programs; by using toolchains and automation technologies, manual operations are significantly reduced, improving the overall test automation level and effect.
[0090] In one embodiment of the present invention, a stitched program is generated, and the initial training data is input into the stitched program to obtain stitched paths. Based on the stitched paths, a model for automatically generating in-vehicle computer test cases is trained, including:
[0091] S31. Generate a stitched program for the in-vehicle computer control software program, and input the initial training data into the stitched program;
[0092] S32. Run and compile the stitched program, obtain stitched paths, filter the stitched paths based on path differences, and calculate the hierarchical depth of the filtered stitched paths;
[0093] S33. Take the initial training data, the instrumentation path, and the hierarchical depth of the instrumentation path as the original samples, extract the features in the original samples to generate test data, and select the pre-trained model Transformer;
[0094] S34. Input the test data into the pre-trained model to train the sub-model, and test whether the accuracy of the sub-model reaches the first preset threshold. If it reaches the first preset threshold, number the sub-model and save it to the trained sub-model library;
[0095] S35. If it does not reach the first preset threshold, sequentially add the node information in the instrumentation path, and then determine whether all sub-models have been trained. If not all sub-models have been trained, continue to train the sub-model;
[0096] S36. If all sub-models have been trained, check the number of sub-models, and determine whether the number of sub-models is equal to the number of nodes in the instrumentation path. If they are equal, link the sub-models according to the node order of the instrumentation path to obtain the automatically generated model for in-vehicle computer test cases that has been constructed;
[0097] S37. If the number of sub-models is not equal to the number of nodes in the instrumentation path, continue to select other pre-trained models, find the path nodes for which sub-models have not been constructed, and return to step S34 to continue execution for the path nodes for which sub-models have not been constructed.
[0098] The working principle and effects of the above technical solution are as follows: Additional code (referred to as stubs) is inserted into the code of the in-vehicle computer control software. These stub codes are used to collect specific execution paths and other relevant information during program operation; the initial training data generated through the previous steps is input into the stub program so that these paths can be recorded during program execution, and the stubbed program is compiled to generate an executable file. The compiled stub program is run to obtain the stub paths, which record the specific path information during program execution; the obtained stub paths are screened based on path differences, and the representative and diverse paths are retained; the hierarchical depth of the screened stub paths is calculated. The hierarchical depth represents the nested level of nodes in the path; the initial training data, stub paths, and their hierarchical depths are used as original samples, and features are extracted from them to generate test data; a suitable pre-trained model (Transformer) is selected. This model is good at processing sequence data and can effectively capture the features of program paths; the generated test data is input into the selected pre-trained model to train the sub-model; check whether the accuracy of the sub-model reaches the first preset threshold. If it reaches this threshold, the sub-model is numbered and saved in the trained sub-model library; retrain by adding node information. If the accuracy of the sub-model does not reach the preset threshold, the node information in the stub path is gradually added to enhance the richness and diversity of the training data; check whether sub-models have been trained for all path nodes. If not, continue to train new sub-models. Check the number of sub-models and check whether all sub-models have been trained, that is, whether the number of sub-models is equal to the number of nodes in the stub path; if all sub-models have been trained and the numbers are equal, link the sub-models in the order of the nodes in the stub path to construct a completed in-vehicle computer test case automatic generation model; when sub-models cannot be generated for some path nodes, select other pre-trained models for new attempts; for the path nodes for which sub-models have not been constructed, return to step S34 to continue training until sub-models have been trained for all nodes and a final model is generated.Enhance the diversity and representativeness of test data. By screening the instrumented paths with different execution paths, various control flow possibilities of the program are covered, and the generated test data is diverse and has a wide coverage. When the sub-model training does not meet the standard, gradually increase the node information to ensure rich training data and improve the accuracy of the sub-model. Use the Transformer pre-trained model, which can effectively capture the sequence features of the program paths and train a sub-model with high accuracy. For sub-models with unqualified accuracy, repeatedly iterate the training by increasing the training data and node information to ensure the quality of the finally generated model. Automatically obtain the control flow information during program operation through instrumentation, reduce manual intervention, and achieve automated data collection. The entire training process is highly automated, including sub-model training, accuracy evaluation, model iterative improvement, etc., significantly improving the efficiency. Efficiently construct the final model, gradually link the sub-models, ensure that all sub-models are trained and the number of sub-models is consistent with the number of path nodes, and realize the construction of the final model through sequential linking to ensure the integrity and consistency of the model. For sub-models that cannot meet the standard, flexibly select other pre-trained models for trial to ensure that sub-models can be effectively generated under different scenarios and conditions and ensure the completion of the final model. Through automated and intelligent means, from static analysis, data flow graph construction, instrumentation path collection, pre-trained model selection and training, sub-model iterative improvement to the construction of the final model, the coverage rate, accuracy and efficiency of the test case generation for in-vehicle computer control software are significantly improved. Its working principle and implementation process ensure the diversity and representativeness of the test data, and finally construct a high-quality automatic test case generation model, greatly improving the security, reliability and performance of the software.
[0099] In one embodiment of the present invention, the instrumented paths are screened based on path differences, including:
[0100] List all basic blocks that will appear in the program, and define the instrumented paths based on the basic blocks;
[0101] According to the order of the basic block set, record the number of times each basic block appears in the path based on the defined instrumented paths. The basic blocks are used as the dimensions of the vector to obtain the vector representation of the instrumented paths;
[0102] Calculate the sparsity index of the path;
[0103] Calculate the weight of the path:
[0104] w = α·f + β·b
[0105] where w represents the weight, f represents the execution frequency of the path, b represents the number of basic blocks covered by the path, and α and β are the weights of the adjustment parameters;
[0106] Calculate the difference of the instrumentation path based on the path difference model. Specifically, the path difference model is as follows:
[0107]
[0108] where D represents the difference of the path, n represents the number of instrumentation paths, w i and w j represent the weights of path p i and path p j respectively, d(p i , p j ) represents the distance metric function between path p i and path p j , and S represents the sparsity evaluation index;
[0109]
[0110] where d represents the number of data points in the instrumentation path, p i,k and p j,k represent the values of path p i and path p j at the k-th data point respectively, and δ 1 , δ 2 and δ 3 represent the weight coefficients;
[0111] If the calculated path difference is less than the second preset threshold, delete one of the paths whose distance calculated based on the distance metric function is less than the second preset threshold.
[0112] The working principle and effect of the above technical solution are as follows: The path difference model measures the overall diversity of a set of paths. It evaluates the diversity of the entire path set, rather than simply comparing the differences between paths; the double summation ∑ in the model means traversing all combinations of paths. Since the diversity of paths reflects the group characteristics, all possible path pairs need to be considered. It represents the combination of traversing all paths; by combining weights and distances, the difference degree between different path pairs is evaluated. The weight reflects the importance of the path, and the distance metric function measures the difference between paths; the introduction of the sparsity S further comprehensively considers the uniformity of path coverage. The more uniform the coverage, the higher the sparsity, indicating higher diversity; 1 / [n(n - 1)] is a normalization factor used to average the results. Since the double summation traverses n(n - 1) / 2 path pairs, dividing by n(n - 1) normalizes the results and avoids the direct impact of the number of paths on the diversity calculation results; when performing static analysis on in-vehicle computer control software, all basic blocks in the program are identified. A basic block is a continuous instruction sequence without branch jumps in the program, usually bounded by the branches or convergences of the control flow; all the identified basic blocks are summarized to form a basic block set; the instrumented path is the sequence of basic blocks obtained through instrumentation during program execution, which reflects the execution path of the program; according to the order of the basic block set, the number of occurrences of each basic block in the path is recorded based on the defined instrumented path. These numbers serve as the dimensions of the vector, constituting the vector representation of the instrumented path; the sparsity represents the sparseness of the number of occurrences of basic blocks in the instrumented path, which helps to identify which paths have relatively sparse occurrences of basic blocks, thus reflecting the importance or particularity of the paths; if the calculated path difference is less than the second preset threshold, one of the paths with a distance less than the second preset threshold calculated based on the distance metric function is deleted. This is to maintain the diversity and representativeness of the paths and reduce redundant paths. By identifying and listing all basic blocks through static analysis, it is ensured that the instrumented paths cover all execution segments of the program; the number of occurrences of each basic block in the instrumented path is recorded in detail to accurately reflect the execution of the paths; by calculating the sparsity of the paths, paths with high representativeness and importance are identified, effectively avoiding data redundancy; based on the number of basic blocks and execution frequency, the weights of the paths are calculated to help screen out more important or frequent paths, which is helpful for optimizing the selection of test data; the difference of the instrumented paths is calculated, and by comparing the sparsity and distance metric of the paths, it is ensured that the selected paths have high diversity; redundant paths with low difference and close distance metric are deleted to further enhance the representativeness of the path set; improve the effectiveness and coverage of test data; use features such as basic blocks and the number of path occurrences to generate high-quality test data to ensure that the test covers all aspects of the program; input high-quality training data: the filtered path data provides high-quality initial training data, which helps to train a more accurate prediction model; reduce the interference of redundant information. By deleting low-difference paths, the negative impact of redundant information on model training is reduced, improving the accuracy and effectiveness of the model; through a series of methods such as static analysis, basic block identification, path recording, sparsity and weight calculation, and difference evaluation, the diversity, representativeness, and efficiency of the data are systematically guaranteed.The test data generated after path screening optimization has comprehensive coverage and high representativeness, further improving the effectiveness and accuracy of test case generation and ensuring the reliability and safety of in-vehicle computer control software; this process is not only fully automated but also well-considered, significantly enhancing the efficiency of the test process and the reliability of the results.
[0113] In one embodiment of the present invention, calculating the sparsity index of a path includes:
[0114] Obtaining the situation of path-covered basic blocks based on the vector representation of the path, and converting the situation of path-covered basic blocks into a path matrix;
[0115] Summing the path matrix by columns to obtain the total number of times each basic block is covered by all paths, and obtaining the summation vector f;
[0116] Calculating the mean and variance of the summation vector, calculating the standardized variance of the summation vector based on the mean and variance of the summation vector, and calculating the sparsity index of the path based on the standardized variance:
[0117]
[0118] Where S represents the sparsity index, f is a vector composed of the coverage frequencies of each path for all basic blocks, represents the standardized variance, Var(f) is the variance of the summation vector, and μ(f) is the mean of the summation vector.
[0119] The working principle and effect of the above technical solution are as follows: For each instrumented path, create a vector, each dimension of which corresponds to a basic block, and the value represents the number of times the basic block appears in the path; paths = {"data1": [1, 1, 1, 0, 0], "data2": [1, 1, 0, 1, 1], "data3": [1, 1, 0, 1, 0]} Convert the path coverage situation into a matrix: vectors = list(paths.values()) The path coverage matrix is as follows:
[0120]
[0121] Take the vector representation of each path as a row of a matrix to form a path matrix. The rows of the path matrix represent paths, the columns represent basic blocks, and the element values represent the number of occurrences of the corresponding basic block in the corresponding path; sum the path matrix by column to obtain a sum vector f; calculate the mean and variance of the sum vector, calculate the standardized variance, and evaluate the path sparsity. Through the path matrix, the coverage of all paths on basic blocks can be visually and quantitatively represented, helping to understand which basic blocks are covered by which paths; the sum vector f obtained by summing by column clearly represents the total number of times each basic block is covered in all paths, providing an overview of the comprehensive data coverage; calculating the mean and variance of the sum vector measures the balance of basic block coverage. The mean reflects the overall coverage level, and the variance reflects the degree of dispersion of the coverage; through the standardized variance, the influence of the absolute value of the number of coverages of different basic blocks is eliminated, and a more relative and fair measurement standard is obtained. Evaluate the path sparsity through the sparsity index. A higher sparsity index indicates a greater degree of imbalance in the path coverage of basic blocks, and a lower one indicates more balanced path coverage; based on the sparsity index, the test data generation process can be optimized, and paths with better coverage balance can be selected, thereby improving the test coverage rate and effectiveness; through a series of mathematical methods such as the path matrix, sum vector, mean, variance, and standardized variance, quantitatively evaluate and measure the basic block coverage and sparsity of the instrumented paths. The sparsity index, as an important measurement standard, helps to optimize path selection and test data generation, making the test process not only more comprehensive but also more balanced, improving the effectiveness and reliability of test case generation; the systematic and scientific path analysis method significantly improves the quality and efficiency of software testing.
[0122] An embodiment of the present invention, an automatic generation system for in-vehicle computer test cases based on path coverage, the system includes:
[0123] A training data generation module, configured to parse the in-vehicle computer control software through a static analysis tool, construct a data flow graph, where the data flow graph takes the basic blocks in the control program as nodes and the changes in the control flow as edges, and generate initial training data based on the data flow graph;
[0124] A training model module, configured to generate an instrumented program, input the initial training data into the instrumented program, obtain instrumented paths, and train an automatic generation model for in-vehicle computer test cases based on the instrumented paths.
[0125] The working principle and effects of the above technical solution are as follows: Use a static analysis tool to analyze the in-vehicle computer control software. The static analysis tool extracts the program's structure and control flow information by analyzing the source code or intermediate code (such as bytecode) without executing the code; through static analysis, a data flow graph (Data Flow Graph, DFG) is generated. In the DFG, basic blocks (one or more sequentially executed statements) are used as nodes of the graph, and the control flow (the jump relationship between basic blocks) is used as the edges of the graph. The data flow graph visually represents the flow of the program between different basic blocks; rich initial training data is extracted from the data flow graph, and this data includes information about each basic block and the connection relationships between basic blocks. The initial training data includes variables that a basic block depends on or affects, the control structure of the basic block, etc.; the initial training data is input into the instrumentation program. Instrumentation is adding additional code (instrumentation points) at specified positions in the program to collect runtime information when the program runs; run the compiled instrumentation program to collect and record the instrumentation paths (i.e., the sequence of basic blocks passed through during the actual execution of the program); through instrumentation, the execution paths of the program and their related runtime data can be collected; filter the instrumentation paths, filter out redundant or invalid paths, and retain representative and diverse paths. This process is based on path diversity, that is, select instrumentation paths that cover different combinations of basic blocks and different control paths; use the filtered instrumentation paths and related data to train an automatic generation model, using a deep learning model. The model can finally automatically generate test cases that cover different control paths by learning these path data. Improve the coverage rate. The test cases generated by the intelligent model can cover different paths, a large number of boundary conditions, and extreme conditions of the program, thereby improving the code coverage rate of the test cases and discovering more potential defects; instrumentation and filtering make the generated test cases diverse and representative, preventing over-concentration on a few paths; automated test case generation reduces the time and labor costs of manually designing test cases and speeds up the testing process; the instrumentation program can efficiently capture the behavior data during program operation and provide sufficient information for subsequent automated analysis; the test case generation technology based on the deep learning model can dynamically generate effective test cases according to the program's running situation, realizing intelligence and adaptability; by using a high-precision training model, problems that may occur in the program can be predicted more accurately, improving the effectiveness and accuracy of testing; extensive and in-depth path coverage ensures comprehensive testing of the in-vehicle computer control software and increases the probability of discovering problems; the high reliability and stability of the in-vehicle computer control software are crucial, and comprehensive testing can effectively improve the system's security and reduce potential failures.
[0126] In one embodiment of the present invention, the training data generation module includes:
[0127] Parsing code module, which parses the C code of in-vehicle computer control software through Clang to generate an LLVM IR file;
[0128] Data flow graph acquisition module, which is used to parse the LLVM IR file into a data flow graph by using the LLVM tool chain, and the data flow graph records variables and dependencies;
[0129] Constraint relation extraction module, which is used to extract the constraint relations between variables according to the data flow graph, and the constraint relations between variables refer to the mutual dependencies and constraint conditions between different variables in the program;
[0130] Initial training data generation module, which is used to generate different initial training data that meet the constraint relations by using a constraint solver and a random seed.
[0131] The working principle and effects of the above technical solution are as follows: Use the Clang compiler to parse the C code of the in-vehicle computer control software. Clang is a highly extensible and modular compiler for the C language family (C / C++ / Objective-C / C++), and its design is based on the LLVM project; Clang compiles the C code into LLVM intermediate representation (LLVM IR), which is a low-level, language-agnostic intermediate representation. The LLVM IR file retains the control flow and data flow information of the program, facilitating further analysis; LLVM toolchain: Use the LLVM toolchain (such as the static analysis tool of LLVM) to parse the LLVM IR file and construct a data flow graph (Data Flow Graph, DFG); In the data flow graph, nodes represent variables, and edges represent the data dependency relationships and constraint conditions between variables. Recording variables and their corresponding dependency relationships helps to comprehensively understand the behavior and data flow of the program; Extract the constraint relationships between variables from the data flow graph, and the constraint relationships include dependencies, constraint conditions, assignment relationships, and comparison relationships between variables; Use a constraint solver tool (such as Z3, CVC4, etc.) to parse the extracted variable constraint relationships and generate specific values that satisfy these constraints. A constraint solver is a tool specifically used to solve logical constraint conditions. By given constraint conditions, it solves specific solutions that meet the conditions. Different initial training data are generated with random seeds to ensure the coverage and representativeness of the training data with diversity and richness. Different variable value combinations can be generated through different random seeds, but all these combinations satisfy the extracted constraint relationships.Improve the diversity and coverage of test data. By using random seeds and constraint solvers, various different combinations of initial training data that satisfy the constraint relationships are generated, enhancing the diversity of the training data. The extensive variable combinations and their corresponding constraint relationships ensure the comprehensive coverage of the generated test data, thereby improving the test effect and defect discovery ability. The extracted variable constraint relationships accurately reflect the dependencies and restrictions between variables in the program, and the generated initial training data can accurately simulate the actual running situation of the program. The variable values generated through constraint solving conform to the logic and conditions during actual operation, ensuring that the test data has a high sense of reality. Utilize toolchains such as Clang, LLVM, and constraint solvers to automate the generation of data flow graphs and data, reducing manual intervention and improving the efficiency of test data generation. The constraint solver can quickly solve complex logical constraint conditions and generate test data that meets the conditions, improving the efficiency of the entire process. Improve the reliability and safety of in-vehicle computer control software. The diverse and comprehensive test data covers various possible running states and paths of the software, helping to discover potential defects and vulnerabilities. Precise constraint solving ensures the high quality of test data, effectively enhancing the effectiveness and reliability of testing, and guaranteeing the stability and safety of in-vehicle computer control software. The data flow graph and constraint relationships can handle complex variable dependencies and logical conditions in the program and are applicable to in-vehicle control software with high complexity. Strong adaptability. The technical solution is based on abstract intermediate representations and toolchain analysis and can adapt to different types of control software and various complex application scenarios. Through steps such as static analysis, LLVM IR parsing, data flow graph construction, constraint extraction, and constraint solving, high-quality initial training data is automatically generated, effectively improving the diversity, coverage, and accuracy of test data. Its effects are significantly reflected in improving test efficiency, enhancing the reliability and safety of in-vehicle computer control software, and supporting the comprehensive analysis and testing of complex programs. By using toolchains and automation technologies, manual operations are significantly reduced, improving the overall test automation level and effect.
[0132] In one embodiment of the present invention, the training model module includes:
[0133] A stub program generation module, which is used to generate a stub program for the in-vehicle computer control software program and input the initial training data into the stub program;
[0134] A screening module, which is used to run and compile the stub program, obtain stub paths, screen the stub paths based on path differences, and calculate the hierarchical depth of the screened stub paths;
[0135] A test data generation module, which is used to use the initial training data, stub paths, and the hierarchical depth of the stub paths as original samples, extract features from the original samples to generate test data, and select the pre-trained model Transformer;
[0136] The sub - model number storage module is used to input the test data into the pre - trained model, train the sub - model, and test whether the accuracy of the sub - model reaches the first preset threshold. If it reaches the first preset threshold, the sub - model is numbered and saved in the trained sub - model library;
[0137] The node information addition module is used to, if the first preset threshold is not reached, sequentially add the node information in the instrumentation path, and then determine whether all sub - models have been trained. If not all sub - models have been trained, continue to train the sub - models;
[0138] The final model linking module is used to, if all sub - models have been trained, check the number of sub - models, and determine whether the number of sub - models is equal to the number of nodes in the instrumentation path. If they are equal, link the sub - models according to the node order of the instrumentation path to obtain the automatically generated model for in - vehicle computer test cases that has been constructed;
[0139] The continuous optimization module is used to, if the number of sub - models is not equal to the number of nodes in the instrumentation path, continue to select other pre - trained models, find the path nodes for which sub - models have not been constructed, and return to step S34 to continue execution for the path nodes for which sub - models have not been constructed.
[0140] The working principle and effects of the above technical solution are as follows: Additional code (referred to as stubs) is inserted into the code of the in-vehicle computer control software, and these stub codes are used to collect specific execution paths and other relevant information during program operation; the initial training data generated through the previous steps is input into the stub program so that these paths can be recorded during program execution, and the stubbed program is compiled to generate an executable file. The compiled stub program is run to obtain the stub paths, which record the specific path information during program execution; the obtained stub paths are filtered based on path differences, and the representative and diverse paths are retained; the hierarchical depth of the filtered stub paths is calculated, and the hierarchical depth represents the nested level of the node in the path; the initial training data, stub paths, and their hierarchical depths are used as original samples, and features are extracted from them to generate test data; a suitable pre-trained model (Transformer) is selected, which is good at processing sequence data and can effectively capture the features of program paths; the generated test data is input into the selected pre-trained model to train the sub-model; the accuracy of the sub-model is tested to see if it reaches the first preset threshold. If it reaches this threshold, the sub-model is numbered and saved in the trained sub-model library; if the accuracy of the sub-model does not reach the preset threshold, the node information in the stub path is gradually added to enhance the richness and diversity of the training data; it is checked whether sub-models have been trained for all path nodes. If not, new sub-models continue to be trained. The number of sub-models is checked to see if all sub-models have been trained, that is, whether the number of sub-models is equal to the number of nodes in the stub path; if all sub-models have been trained and the numbers are equal, the sub-models are linked in the order of the nodes in the stub path to construct a completed in-vehicle computer test case automatic generation model; when sub-models cannot be generated for some path nodes, other pre-trained models are selected for new attempts; for the path nodes for which sub-models have not been constructed, return to step S34 to continue training until sub-models have been trained for all nodes and a final model is generated.Enhance the diversity and representativeness of test data. By screening the instrumentation paths with different execution paths, various control flow possibilities of the program are covered, and the generated test data is diverse and has a wide coverage. When the sub-model training is not up to the standard, gradually increase the node information to ensure rich training data and improve the accuracy of the sub-model. Use the Transformer pre-trained model, which can effectively capture the sequence features of the program paths and train a sub-model with high accuracy. For the sub-model with unqualified accuracy, repeatedly iterate the training by increasing the training data and node information to ensure the quality of the finally generated model. Automatically obtain the control flow information during program runtime through instrumentation, reduce manual intervention, and achieve automated data collection. The entire training process is highly automated, including sub-model training, accuracy evaluation, model iterative improvement, etc., significantly improving the efficiency. Efficiently construct the final model, gradually link the sub-models, ensure that all sub-models are trained and the number of sub-models is consistent with the number of path nodes, and achieve the construction of the final model through sequential linking to ensure the integrity and consistency of the model. For the sub-model that cannot meet the standard, flexibly select other pre-trained models for trial to ensure that sub-models can be effectively generated under different scenarios and conditions and ensure the completion of the final model. Through automated and intelligent means, from static analysis, data flow graph construction, instrumentation path collection, pre-trained model selection and training, sub-model iterative improvement to the construction of the final model, the coverage rate, accuracy and efficiency of the test case generation for in-vehicle computer control software are significantly improved. Its working principle and implementation process ensure the diversity and representativeness of the test data, and finally construct a high-quality automatic test case generation model, greatly improving the security, reliability and performance of the software.
[0141] In one embodiment of the present invention, the screening module includes:
[0142] A path definition module, configured to list all basic blocks that will appear in the program and define the instrumentation path based on the basic blocks;
[0143] A vector representation path module, configured to record the number of times each basic block appears in the path based on the defined instrumentation path in the order of the basic block set, use the basic block as the dimension of the vector, and obtain the vector representation of the instrumentation path;
[0144] A sparsity calculation module, configured to calculate the sparsity index of the path;
[0145] A path weight calculation module, configured to calculate the weight of the path:
[0146] w = α·f + β·b
[0147] where w represents the weight, f represents the execution frequency of the path, b represents the number of basic blocks covered by the path, and α and β are the weights of the adjustment parameters;
[0148] A calculation path difference module, which is used to calculate the difference of the instrumented paths based on a path difference model. Specifically, the path difference model is as follows:
[0149]
[0150] where D represents the difference of the paths, n represents the number of instrumented paths, w i and w j represent the weights of path p i and path p j respectively, d(p i , p j ) represents the distance metric function between path p i and path p j , and S represents the sparsity evaluation index;
[0151]
[0152] where d represents the number of data points in the instrumented paths, p i,k and p j,k represent the values of path p i and path p j at the k-th data point respectively, and δ 1 , δ 2 and δ 3 represent weight coefficients;
[0153] A path deletion module, which is used to delete one path from the paths whose distances calculated based on the distance metric function are less than a second preset threshold if the calculated path difference is less than the second preset threshold.
[0154] The working principle and effect of the above technical solution are as follows: The path difference model measures the overall diversity of a group of paths. It evaluates the diversity of the entire path set, rather than simply comparing the differences between paths; the double summation ∑ in the model means traversing all combinations of paths. Since the diversity of paths reflects the group characteristics, all possible path pairs need to be considered. It represents the combination of traversing all paths; by combining weights and distances, it evaluates the degree of difference between different path pairs. The weight reflects the importance of the path, and the distance metric function measures the difference between paths; the introduction of the sparsity S further comprehensively considers the uniformity of path coverage. The more uniform the coverage, the higher the sparsity, indicating higher diversity; 1 / [n(n - 1)] is a normalization factor used to average the results. Since the double summation traverses n(n - 1) / 2 path pairs, dividing by n(n - 1) normalizes the results and avoids the direct impact of the number of paths on the diversity calculation results; when performing static analysis on in-vehicle computer control software, all basic blocks in the program are identified. A basic block is a continuous instruction sequence in the program without branch jumps, usually bounded by the branches or convergences of the control flow; all the identified basic blocks are summarized to form a basic block set; the instrumented path is the sequence of basic blocks obtained through instrumentation during program execution, reflecting the execution path of the program; according to the order of the basic block set, the number of times each basic block appears in the path is recorded based on the defined instrumented path. These numbers serve as the dimensions of the vector, constituting the vector representation of the instrumented path; the sparsity represents the sparsity degree of the number of times basic blocks appear in the instrumented path, which helps to identify which paths have relatively sparse appearances of basic blocks, thus reflecting the importance or particularity of the paths; if the calculated path difference is less than the second preset threshold, one of the paths with a distance less than the second preset threshold calculated based on the distance metric function is deleted. This is to maintain the diversity and representativeness of the paths and reduce redundant paths. By identifying and listing all basic blocks through static analysis, it is ensured that the instrumented paths cover all execution segments of the program; the number of times each basic block appears in the instrumented path is recorded in detail to accurately reflect the execution situation of the path; by calculating the sparsity of the paths, paths with high representativeness and importance are identified, effectively avoiding data redundancy; based on the number of basic blocks and execution frequency, the weights of the paths are calculated to help screen out more important or frequent paths, which is helpful for optimizing the selection of test data; the difference of the instrumented paths is calculated, and by comparing the sparsity and distance metric of the paths, it is ensured that the selected paths have high diversity; redundant paths with low difference and close distance metric are deleted to further enhance the representativeness of the path set; improve the effectiveness and coverage of test data; use features such as basic blocks and the number of path appearances to generate high-quality test data to ensure that the test covers all aspects of the program; input high-quality training data: the filtered path data provides high-quality initial training data, which helps to train a more accurate prediction model; reduce the interference of redundant information. By deleting low-difference paths, the negative impact of redundant information on model training is reduced, improving the accuracy and effectiveness of the model; through a series of methods such as static analysis, basic block identification, path recording, sparsity and weight calculation, and difference evaluation, the diversity, representativeness, and efficiency of the data are systematically guaranteed.The test data generated after path screening optimization has comprehensive coverage and high representativeness, further improving the effectiveness and accuracy of test case generation and ensuring the reliability and safety of in-vehicle computer control software; this process is not only fully automated but also well-considered, significantly enhancing the efficiency of the test process and the reliability of the results.
[0155] In one embodiment of the present invention, the calculation sparsity module includes:
[0156] A conversion path matrix module, configured to obtain the situation of path-covered basic blocks based on the vector representation of the path, and convert the situation of path-covered basic blocks into a path matrix;
[0157] A summation vector acquisition module, configured to sum the path matrix by columns to obtain the total number of times each basic block is covered by all paths, and obtain a summation vector f;
[0158] A calculation sparsity index module, configured to calculate the mean and variance of the summation vector, calculate the standardized variance of the summation vector based on the mean and variance of the summation vector, and calculate the sparsity index of the path based on the standardized variance:
[0159]
[0160] Wherein, S represents the sparsity index, f is a vector composed of the coverage frequencies of each path for all basic blocks, represents the standardized variance, Var(f) is the variance of the summation vector, and μ(f) is the mean of the summation vector.
[0161] The working principle and effect of the above technical solution are as follows: For each instrumented path, create a vector, each dimension of which corresponds to a basic block, and the value represents the number of times the basic block appears in the path; paths = {"data1": [1, 1, 1, 0, 0], "data2": [1, 1, 0, 1, 1], "data3": [1, 1, 0, 1, 0]} Convert the path coverage situation into a matrix: vectors = list(paths.values()) The path coverage matrix is as follows:
[0162]
[0163] The vector representation of each path is used as a row of a matrix to form a path matrix. The rows of the path matrix represent paths, the columns represent basic blocks, and the element values represent the number of occurrences of the corresponding basic block in the corresponding path; sum the path matrix by column to obtain a summation vector f; calculate the mean and variance of the summation vector, calculate the standardized variance, and evaluate the path sparsity. Through the path matrix, the coverage of all paths to basic blocks can be visually and quantitatively represented, helping to understand which basic blocks are covered by which paths; the summation vector f obtained by summing by column clearly represents the total number of times each basic block is covered in all paths, providing an overview of the comprehensive data coverage; calculating the mean and variance of the summation vector measures the balance of basic block coverage. The mean reflects the overall coverage level, and the variance reflects the degree of dispersion of the coverage; through the standardized variance, the influence of the absolute value of the coverage quantity of different basic blocks is eliminated, and a more relative and fair measurement standard is obtained. The sparsity of the path is evaluated through the sparsity index. A higher sparsity index indicates a greater degree of imbalance in the path covering basic blocks, while a lower one indicates a more balanced path coverage; based on the sparsity index, the test data generation process can be optimized, and paths with better coverage balance can be selected, thereby improving the test coverage rate and effectiveness; through a series of mathematical methods such as the path matrix, summation vector, mean, variance, and standardized variance, the basic block coverage and sparsity of the instrumented paths are quantitatively evaluated and measured. The sparsity index, as an important measurement standard, helps to optimize path selection and test data generation, making the test process not only more comprehensive but also more balanced, and improving the effectiveness and reliability of test case generation; the systematic and scientific path analysis method significantly improves the quality and efficiency of software testing.
[0164] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these changes and modifications.
Claims
1. A method for automatically generating vehicle computer test cases based on path coverage, characterized in that: The method comprises: Analyze the vehicle computer control software by using a static analysis tool, build a data flow graph, the data flow graph uses basic blocks in the control program as nodes and changes in control flow as edges, and generate initial training data based on the data flow graph; Generate an instrumentation program, input the initial training data into the instrumentation program, obtain an instrumentation path, and train the vehicle computer test case based on the instrumentation path to automatically generate a model.
2. The method for automatically generating vehicle computer test cases based on path coverage according to claim 1 is characterized in that: The vehicle computer control software is parsed by a static analysis tool to construct a data flow graph, wherein the data flow graph uses the basic blocks in the control program as nodes and the changes in the control flow as edges, and generates initial training data based on the data flow graph, including: Use Clang to parse the C code of the onboard computer control software and generate LLVM IR files; Using the LLVM toolchain, the LLVM IR file is parsed into a data flow graph, wherein the data flow graph records variables and dependencies; Extracting constraint relationships between variables according to the data flow graph, wherein the constraint relationships between variables refer to mutual dependencies and constraint conditions between different variables in a program; A constraint solver and a random seed are used to generate different initial training data satisfying the constraint relationship.
3. The method for automatically generating vehicle computer test cases based on path coverage according to claim 1 is characterized in that: Generate an instrumentation program, input the initial training data into the instrumentation program, obtain an instrumentation path, and train the vehicle computer test case based on the instrumentation path to automatically generate a model, including: S31, generating a plug-in program for the vehicle-mounted computer control software program, and inputting the initial training data into the plug-in program; S32, running and compiling the instrumentation program, obtaining an instrumentation path, screening the instrumentation path based on path differences, and calculating the hierarchical depth of the screened instrumentation path; S33, taking the initial training data, the instrumentation path and the hierarchical depth of the instrumentation path as original samples, extracting features from the original samples to generate test data, and selecting a pre-trained model Transformer; S34, inputting the test data into the pre-trained model, training the sub-model, and testing whether the accuracy of the sub-model reaches a first preset threshold. If the first preset threshold is reached, numbering the sub-model and saving it to a trained sub-model library; S35, if the first preset threshold is not reached, then add the node information in the plugging path in sequence, and then determine whether all sub-models have been trained. If not, continue to train the sub-models; S36. If all sub-models have been trained, check the number of sub-models to determine whether the number of sub-models is equal to the number of nodes in the insertion path. If they are equal, link the sub-models according to the sequence of the insertion path nodes to obtain a constructed automatic generation model of the on-board computer test case; S37. If the number of sub-models is not equal to the number of nodes in the insertion path, continue to select other pre-trained models, and look for path nodes without constructed sub-models, and return to step S34 to continue execution for the path nodes without constructed sub-models.
4. The method for automatically generating vehicle computer test cases based on path coverage according to claim 3 is characterized in that: The instrumentation paths are screened based on path differences, including: List all basic blocks that will appear in the program, and define the stub path based on the basic blocks; According to the order of the basic block set, based on the defined instrumentation path, the number of times each basic block appears in the path is recorded, the basic block is used as the dimension of the vector, and the vector representation of the instrumentation path is obtained; Calculate the sparsity index of the path; Calculate the weight of the path: w=α·f+β·b Where w represents the weight, f represents the execution frequency of the path, b represents the number of basic blocks covered by the path, and α and β are the weights of the adjustment parameters; The difference of the instrumentation path is calculated based on a path difference model. Specifically, the path difference model is: Where D represents the difference of the path, n represents the number of instrumented paths, and w i and w j Respectively represent the path p i and path p j The weight of d(p i , p j ) represents the path p i and path p j The distance measurement function between them, S represents the sparsity evaluation index; Where d represents the number of data points in the instrumentation path, p i,k and p j,k Respectively represent the path p i and path p j The values at the kth data point, δ1, δ2, and δ3 represent weight coefficients; If the calculated path difference is less than the second preset threshold, one of the paths whose distance calculated based on the distance metric function is less than the second preset threshold is deleted.
5. The method for automatically generating vehicle computer test cases based on path coverage according to claim 4 is characterized in that: Calculate the sparsity metrics of the path, including: Obtain the coverage of basic blocks by the path based on the vector representation of the path, and convert the coverage of basic blocks by the path into a path matrix; Sum the path matrix by column to get the total number of times each basic block is covered by all paths, and obtain the sum vector f; The mean and variance of the summed vector are calculated, the standardized variance of the summed vector is calculated based on the mean and variance of the summed vector, and the sparsity index of the path is calculated based on the standardized variance: Among them, S represents the sparsity index, f is a vector composed of the coverage frequency of each path to all basic blocks, represents the standardized variance, Var(f) is the variance of the summed vector, and μ(f) is the mean of the summed vector.
6. The automatic generation system of vehicle computer test cases based on path coverage is characterized by: The system comprises: Generate training data module, used to parse the vehicle computer control software through static analysis tools, build a data flow graph, the data flow graph uses basic blocks in the control program as nodes and changes in control flow as edges, and generates initial training data based on the data flow graph; The training model module is used to generate an instrumentation program, input the initial training data into the instrumentation program, obtain an instrumentation path, and train the vehicle computer test case based on the instrumentation path to automatically generate a model.
7. The automatic generation system of vehicle computer test cases based on path coverage according to claim 6 is characterized in that: The generating training data module comprises: Parse the code module and use Clang to parse the C code of the onboard computer control software to generate LLVM IR files; Obtain a data flow graph module, for parsing the LLVM IR file into a data flow graph using an LLVM tool chain, wherein the data flow graph records variables and dependencies; A constraint relationship extraction module is used to extract the constraint relationship between variables according to the data flow graph, wherein the constraint relationship between variables refers to the mutual dependence and constraint conditions between different variables in the program; The initial training data generating module is used to generate different initial training data satisfying the constraint relationship by using a constraint solver and a random seed.
8. The automatic generation system of vehicle computer test cases based on path coverage according to claim 6 is characterized in that: The training model module includes: Generate a plug-in program module, which is used to generate a plug-in program for the vehicle-mounted computer control software program, and input the initial training data into the plug-in program; A screening module, used for running and compiling the instrumentation program, obtaining an instrumentation path, screening the instrumentation path based on path differences, and calculating the hierarchical depth of the screened instrumentation path; A test data generation module is used to use the initial training data, the instrumentation path and the hierarchical depth of the instrumentation path as original samples, extract features from the original samples to generate test data, and select a pre-trained model Transformer; A sub-model number storage module is used to input the test data into the pre-trained model, train the sub-model, test whether the accuracy of the sub-model reaches a first preset threshold, and if so, number the sub-model and save it in a trained sub-model library; A node information adding module is used to sequentially add the node information in the plugging path if the first preset threshold is not reached, and then determine whether all sub-models have been trained. If not, continue to train the sub-model; The final model linking module is used to check the number of sub-models if all sub-models have been trained, and determine whether the number of sub-models is equal to the number of nodes in the insertion path. If they are equal, the sub-models are linked according to the sequence of the insertion path nodes to obtain the constructed automatic generation model of the on-board computer test case; The continuous optimization module is used to continue to select other pre-trained models and find path nodes for which sub-models have not been constructed if the number of sub-models is not equal to the number of nodes in the insertion path, and return to step S34 for further execution of the path nodes for which sub-models have not been constructed.
9. The automatic generation system of vehicle computer test cases based on path coverage according to claim 8 is characterized in that: The screening module comprises: A path definition module is used to list all basic blocks that will appear in the program and define the stub path based on the basic blocks; A vector representation path module, for recording the number of occurrences of each basic block in the path based on the defined instrumentation path in the order of the basic block set, wherein the basic block is used as a dimension of the vector, and obtaining the vector representation of the instrumentation path; The sparsity calculation module is used to calculate the sparsity index of the path; The path weight calculation module is used to calculate the path weight: w=α·f+β·b Where w represents the weight, f represents the execution frequency of the path, b represents the number of basic blocks covered by the path, and α and β are the weights of the adjustment parameters; The path difference calculation module is used to calculate the difference of the plugging path based on the path difference model. Specifically, the path difference model is: Where D represents the difference of the path, n represents the number of instrumented paths, and w i and w j Respectively represent the path p i and path p j The weight of d(p i , p j ) represents the path p i and path p j The distance measurement function between them, S represents the sparsity evaluation index; Where d represents the number of data points in the instrumentation path, p i,k and p j,k Respectively represent the path p i and path p j The values at the kth data point, δ1, δ2, and δ3 represent weight coefficients; The path deletion module is used to delete a path among the paths whose distance calculated based on the distance metric function is less than the second preset threshold value if the calculated path difference is less than the second preset threshold value.
10. The automatic generation system of vehicle computer test cases based on path coverage according to claim 9 is characterized in that: The sparsity calculation module includes: A path matrix conversion module is used to obtain the situation of path coverage of basic blocks based on the vector representation of the path, and convert the situation of path coverage of basic blocks into a path matrix; A summation vector acquisition module is used to sum the path matrix by column to obtain the total number of times each basic block is covered by all paths, and obtain the summation vector f; A sparsity index calculation module is used to calculate the mean and variance of the summed vector, calculate the standardized variance of the summed vector based on the mean and variance of the summed vector, and calculate the sparsity index of the path based on the standardized variance: Among them, S represents the sparsity index, f is a vector composed of the coverage frequency of each path to all basic blocks, represents the standardized variance, Var(f) is the variance of the summed vector, and μ(f) is the mean of the summed vector.