White-box testing method and apparatus, program product and medium

By using code models to filter and evaluate risk functions in white box tests, the problem of low static analysis efficiency is solved, and a more efficient code risk assessment is achieved.

WO2025107535A1PCT designated stage expired Publication Date: 2025-05-30HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/092528
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-30
Filing Date
2024-05-11
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The static analysis scheme for existing white box tests requires traversing all code paths, resulting in low analysis efficiency.

Method used

By obtaining the code model of the tested code, filtering out the risk functions in the code path, using the evaluation model to analyze these risk functions, obtaining risk assessment results, thereby improving analysis efficiency.

Benefits of technology

The analysis of risk functions with unreachable paths is avoided, the analysis workload is reduced, and the code analysis efficiency is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024092528_30052025_PF_FP_ABST
    Figure CN2024092528_30052025_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a white-box testing method and apparatus, a program product and a medium in the field of computers, for use in carrying out risk assessment on a tested code, and obtaining a risk assessment result of the tested code by analyzing a path-reachable risk function, thereby improving the code analysis efficiency. The method comprises: identifying first risk functions comprised in a tested code; using a code model to screen for a first risk function in a code path from the first risk functions, wherein the code model comprises an upstream and downstream association relationship in the tested code, and the code path is an execution path of the tested code; and analyzing the first risk function in the code path by means of an assessment model to obtain a risk assessment result of the tested code.
Need to check novelty before this filing date? Find Prior Art

Description

A white box testing method, device, program product and medium

[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on November 21, 2023, with application number 202311558315.3, and with the invention name “A white box code security analysis method and device”, and claims priority to the Chinese patent application filed with the State Intellectual Property Office on November 30, 2023, with application number 202311628785.2, and with the invention name “A white box testing method, device, program product and medium”, all contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of computer technology, and in particular to a white box testing method, device, program product, and medium. Background Art

[0003] With the rapid development of information technology, software plays an increasingly important role in our daily lives and work. However, software security remains a serious concern. To ensure software security, white-box software security verification has become an important technical approach. White-box software security verification involves comprehensive analysis and testing of the software's internal structure and logic to identify potential security vulnerabilities and weaknesses. Compared to traditional black-box testing, white-box software security verification provides a deeper understanding of the software's internal operating mechanisms, enabling a more accurate assessment of its security.

[0004] Currently, various white-box software security verification methods and techniques have been proposed, including static analysis, dynamic analysis, symbolic execution, and fuzz testing. Common white-box testing methods include dynamic analysis and static analysis. Static analysis involves analyzing program source code without running the program to identify security vulnerabilities. Common methods use tools or manual methods to identify dangerous functions, obtain correlations between risky code, and analyze and flag risky code based on business practices and experience.

[0005] However, although the static analysis solution of white-box testing can automatically identify risky functions in the code, it needs to traverse all code paths, including risky paths that are reachable and unreachable, resulting in low code analysis efficiency.

[0006] Summary of the Invention

[0007] The present application provides a white box testing method, device, program product and medium for performing risk assessment on the code under test. By analyzing the path-reachable risk function, the risk assessment result of the code under test is obtained, thereby improving the efficiency of code analysis.

[0008] In view of this, on the first aspect, the present application provides a white box testing method, including: obtaining the code under test, and performing risk function identification on the obtained code under test to obtain a first risk function included in the code under test; in addition, a code model of the code under test should also be obtained, and the code model includes the association relationship between upstream and downstream in the code under test; then, the first risk function existing in the code path can be screened out through the obtained code model, and the code path is the execution path of the code under test; finally, the first risk function existing in the code path is analyzed through the evaluation model to obtain a risk assessment result of the code under test.

[0009] In the implementation method of the present application, the risk functions existing in the code path, i.e., the path-reachable risk functions, can be screened out based on the code model of the code under test. The risk functions existing in the code path are analyzed through the evaluation model to obtain the risk evaluation results, thereby avoiding the analysis of the risk functions of the unreachable path, thereby reducing the workload of risk function analysis and identification, and improving the analysis efficiency of the code under test.

[0010] In a possible implementation, the aforementioned analysis of the first risk function existing in the code path through the evaluation model to obtain the risk assessment result of the code under test may include: running the code under test to obtain one or more first associated item sets, the first associated item sets being included in the first risk function existing in the code path, the first associated item sets including multiple risk functions with associated relationships; performing association analysis on the one or more first associated item sets and the code path to obtain one or more second associated item sets, the second associated item sets including frequent item sets and upstream and downstream functions of the frequent item sets, the frequent item sets including first associated item sets with support higher than a preset value; analyzing the one or more second associated item sets through the evaluation model to obtain the risk assessment result of the code under test.

[0011] In an embodiment of the present application, further association analysis can be performed based on the obtained first associated item set and code path to obtain a frequent item set and upstream and downstream functions associated with the frequent item set, i.e., a second associated item set, and risk analysis can be performed on the second associated item set. By gradually searching for and analyzing more important risk functions and the code path of the risk function, the efficiency of code analysis can be improved.

[0012] In a possible implementation, the aforementioned code path may be obtained by analyzing the running code under test in a code instrumentation manner.

[0013] In the embodiment of the present application, the code path and the data flow path can be automatically obtained by code instrumentation, which can avoid missed detection when manually traversing the path.

[0014] In a possible implementation, the aforementioned running of the tested code to obtain one or more first associated item sets may include: obtaining one or more first associated item sets based on the first risk functions existing in the code path and the association relationship between the first risk functions existing in the code path.

[0015] In a possible implementation, the aforementioned performing association analysis on one or more first associated item sets and code paths to obtain one or more second associated item sets may include: performing association analysis on the one or more first associated item sets using an aggregation algorithm to obtain frequent item sets; and calling upstream and downstream functions associated with the frequent item sets according to the code paths to obtain one or more second associated item sets.

[0016] In an embodiment of the present application, an aggregation algorithm can be used to perform correlation analysis on the risk function to obtain frequent item sets of the risk function and functions associated with the frequent item sets. Key code paths can be analyzed based on the obtained frequent item sets, thereby improving the efficiency of code analysis and correcting false positives generated when the code recognition tool identifies the risk function.

[0017] In a possible implementation, the aforementioned use of an aggregation algorithm to perform association analysis on one or more first associated item sets to obtain frequent item sets may include: using an aggregation algorithm to calculate the support of the one or more first associated item sets; and taking the first associated item sets whose support is higher than a preset value as frequent item sets.

[0018] In one possible implementation, if the code under test includes uncovered code, and the uncovered code is a newly added code to be tested in the code under test, the method may further include: identifying a second risk function included in the baseline code; establishing a rule base for the use of risk functions based on the second risk function, and the rule base is used to indicate the constraints on the use of the second risk function and functions associated with the second risk function; risk grading the uncovered code based on the rule base; and updating the risk assessment result of the code under test based on the result of the risk grading to obtain an updated risk assessment result.

[0019] In the implementation mode of the present application, after performing risk analysis on the code under test, if new code to be tested is added to the code under test, only the newly added code to be tested can be automatically risk graded, and the risk assessment results of the code under test can be updated based on the risk grading results, thereby simplifying the risk analysis process and improving the efficiency of code risk analysis.

[0020] In one possible implementation, the aforementioned establishment of a rule base for risk function usage specifications based on the second risk function may include: using an aggregation algorithm to perform association analysis on the second risk function to obtain functions associated with the second risk function, calling paths of functions associated with the second risk function, and constraints; establishing a rule base based on the second risk function, functions associated with the second risk function, and calling paths of functions associated with the second risk function, as well as constraints.

[0021] In one possible implementation, the aforementioned risk grading of uncovered code based on the rule base may include: determining whether the constraints of the uncovered code meet the requirements of the rule base; if so, marking the uncovered code as low risk; if not, marking the uncovered code as high risk.

[0022] In the implementation manner of the present application, by automatically grading the risks of uncovered code, a key investigation direction is provided for subsequent investigation, which can assist testers in investigating high-risk code and improve the efficiency of code analysis.

[0023] In a second aspect, the present application provides a white box testing device, comprising:

[0024] an identification module, configured to identify a first risk function included in the code under test;

[0025] a screening module, configured to screen out the first risk function present in the code path from the first risk function using a code model, wherein the code model includes associations between upstream and downstream in the tested code, and the code path is an execution path of the tested code;

[0026] The analysis module is used to analyze the first risk function existing in the code path through the evaluation model to obtain a risk evaluation result of the tested code.

[0027] In one possible implementation, the analysis module is specifically configured to run the code under test to obtain one or more first associated item sets, where the first associated item set includes a first risk function present in the code path, and the first associated item set includes a plurality of risk functions having an associated relationship; perform association analysis on the one or more first associated item sets and the code path to obtain one or more second associated item sets, where the second associated item sets include frequent item sets and upstream and downstream functions of the frequent item sets, and the frequent item sets include first associated item sets with a support higher than a preset value; and analyze the one or more second associated item sets through an evaluation model to obtain a risk assessment result of the code under test.

[0028] In a possible implementation, the aforementioned code path may be obtained by analyzing the running code under test in a code instrumentation manner.

[0029] In a possible implementation, the analysis module is specifically configured to obtain one or more first associated item sets based on the first risk functions existing in the code paths and the association relationships between the first risk functions existing in the code paths.

[0030] In a possible implementation, the analysis module is specifically configured to perform association analysis on one or more first associated item sets using an aggregation algorithm to obtain frequent item sets; and to call upstream and downstream functions associated with the frequent item sets according to the code path to obtain one or more second associated item sets.

[0031] In a possible implementation, the analysis module is specifically configured to calculate the support of one or more first associated item sets using an aggregation algorithm; and to take first associated item sets with support higher than a preset value as frequent item sets.

[0032] In a possible implementation, if the code under test includes uncovered code, and the uncovered code is newly added code to be tested in the code under test, the apparatus further includes:

[0033] The identification module is further used to identify a second risk function included in the baseline code;

[0034] An establishment module is used to establish a rule base for risk function usage specifications based on the second risk function, where the rule base is used to indicate constraint conditions for the use of the second risk function and functions associated with the second risk function;

[0035] The classification module is used to classify the risks of uncovered codes according to the rule base;

[0036] The update module is used to update the risk assessment results of the tested code according to the risk classification results to obtain updated risk assessment results.

[0037] In one possible implementation, a module is established, specifically for using an aggregation algorithm to perform association analysis on the second risk function to obtain functions associated with the second risk function, calling paths of functions associated with the second risk function, and constraints; and a rule base is established based on the second risk function, functions associated with the second risk function, and calling paths of functions associated with the second risk function, as well as constraints.

[0038] In one possible implementation, the grading module is specifically used to determine whether the constraints of the uncovered code meet the requirements of the rule base; if so, the uncovered code is marked as low risk; if not, the uncovered code is marked as high risk.

[0039] In a third aspect, the present application provides a computing device cluster comprising at least one computing device, each computing device comprising a processor and a memory; the processor of at least one computing device is used to execute instructions stored in the memory of at least one computing device, so that the computing device cluster executes the method described in the first aspect above.

[0040] In a fourth aspect, the present application provides a computer program product comprising instructions, characterized in that when the instructions are executed by a computing device cluster, the computing device cluster executes the method described in the first aspect above.

[0041] In a fifth aspect, the present application provides a computer-readable storage medium, characterized in that it includes computer program instructions. When the computer program instructions are executed by a computing device cluster, the computing device cluster executes the method described in the first aspect above. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] FIG1 is a schematic diagram of a system framework provided by this application;

[0043] FIG2 is a flowchart of a code security analysis and verification provided by this application;

[0044] FIG3 is a schematic diagram of a white box testing method provided by the present application;

[0045] FIG4 is a flowchart of a code instrumentation analysis provided by this application;

[0046] FIG5 is a flow chart of another white box testing method provided by the present application;

[0047] FIG6 is a diagram showing the sequence of execution of a program under test provided by the present application;

[0048] FIG7 is a flow chart of another white box testing method provided by the present application;

[0049] FIG8 is a schematic structural diagram of a white box testing device provided by the present application;

[0050] FIG9 is a schematic diagram of the structure of a computing device according to an embodiment of the present application;

[0051] FIG10 is a schematic diagram of the structure of a computing device cluster according to an embodiment of the present application;

[0052] FIG11 is another structural diagram of a computing device cluster in an embodiment of the present application. DETAILED DESCRIPTION

[0053] The following will describe the technical solutions in the embodiments of this application in conjunction with the drawings in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making creative efforts shall fall within the scope of protection of this application.

[0054] First, let’s explain the terms involved in this application:

[0055] White box testing: White box testing, also known as structural testing, transparent box testing, logic-driven testing or code-based testing, is a test case design method. The box refers to the software being tested, and the white box means that the box is visible, that is, it is clear what is inside the box and how it works. The "white box" method can fully understand the internal logical structure of the program and test all logical paths.

[0056] Code stub: refers to inserting code into the program under test to track the execution process of the program under test, so as to obtain the execution of executable statements in the program and the execution path of the program.

[0057] Frequent itemsets: Frequent patterns refer to sets of items, sequences, or substructures that appear frequently in a data set. Frequent itemsets refer to sets whose support is greater than or equal to the minimum support.

[0058] Support: refers to the frequency of a set appearing in all transactions, that is, the probability that the item set {A, B} appears in the total item set.

[0059] Risky functions: These are functions that have potential risks or vulnerabilities in their code. When executed, these functions may introduce security risks, performance issues, memory leaks, or abnormal crashes.

[0060] The following is an introduction to the system framework on which the embodiments of the present application are based.

[0061] Referring to FIG. 1 , the present application provides a system framework 100 . As shown in system framework 100 , an automatic analysis module 110 can be directly used for risk assessment analysis of the code under test, and a risk assessment report for the code under test can be generated via a report generation module 120 . Automatic analysis module 110 can include a risk function identification module 111 , a code instrumentation analysis module 112 , a function sequence aggregation solution module 113 , a data storage module 114 , and a data comparison module 115 .

[0062] The risk function identification module 111 is primarily responsible for identifying risk functions or methods that may exist in the code under test. The code instrumentation analysis module 112 can analyze the execution of the business flow, which is the complete process of implementing a business or function. The function sequence aggregation solution module 113 can obtain the association relationship of functions and the frequent item sets of function sequences through aggregation algorithms. The data storage module 114 can be used to store the identified risk functions and their associated functions, code paths, etc. The data comparison module 115 is mainly used to perform comparative analysis of functions to determine the risk level of the functions.

[0063] The execution subject of the white-box testing method provided in this application can be a white-box testing tool, which can be program code software, or a medium storing relevant execution code, or the white-box testing tool can also be a physical device integrated or installed with relevant execution code, such as a chip, a microcontroller unit (MCU), a computer, or other electronic device. In addition, this application can be applied not only to the identification of risk functions of the tested code, but also to precision testing, scenario testing, and code performance testing, which are not specifically limited here.

[0064] The following introduces the code security analysis and verification process provided by this application in conjunction with the system framework described in Figure 1.

[0065] Referring to FIG2 , the code security analysis and verification flow chart provided in the present application, the baseline code can obtain the risk function of the baseline code, the function associated with the risk function, the call path of the risk function, the constraint conditions of the risk function, and the constraint conditions of the function associated with the risk function through the risk function identification module 111, the code instrumentation analysis module 112, and the function sequence aggregation solution module 113 in the automatic analysis module 110, and store them in the data storage module 114, thereby constructing a rule base for function usage specifications. The data comparison module 115 can be used to automatically perform risk classification on the newly added code, and judge the constraint conditions of the newly added code through the constructed rule base to obtain the risk level of the newly added code. The tested code can obtain a risk assessment report through the automatic analysis module 110 and the report generation module 120.

[0066] Referring to FIG3 , a flowchart of a white box testing method provided by the present application is shown as follows.

[0067] 301. Identify a first risk function included in the tested code;

[0068] Specifically, a code scanning tool can be used to identify the first risk function or method included in the tested code. The code scanning tool can select different tools according to the different code languages. For example, the bandit code scanning tool can be used for the Python language, and the spotbugs code scanning tool can be used for the Java language. The specific details are not limited here.

[0069] The code scanning tool may identify risk functions or methods by identifying keywords, or may identify risk functions or methods by one or more methods of syntax analysis or pattern matching, which are not specifically limited here.

[0070] 302. Filtering out first risk functions existing in the code path from the first risk functions using the code model;

[0071] Among them, the code model includes the association relationship between the upstream and downstream in the tested code, that is, it includes all the code paths of the tested code. Therefore, the code model can be used to screen out the first risk function (also called the path-reachable first risk function) existing in the code path. The code path is the execution path of the tested code, and the first risk function existing in the code path is a subset of the first risk function.

[0072] 303. Analyze the first risk function in the code path using the evaluation model to obtain a risk evaluation result of the tested code.

[0073] After obtaining the first risk function existing in the code path, it can be analyzed through the evaluation model to obtain the risk evaluation result of the tested code. In addition, the risk function of the uncovered code can also be evaluated.

[0074] Typically, code instrumentation is used to obtain information about a program during runtime and analyze its behavior during runtime, such as debugging, performance analysis, or security testing. In an embodiment of the present application, code instrumentation can be used to monitor the execution of the code under test, thereby obtaining the correlation between the risk functions of the code under test and the execution path of the code.

[0075] Optionally, the aforementioned code path can be obtained by analyzing the running code under test in a code instrumentation manner.

[0076] Optionally, after obtaining the path-reachable first risk function, one or more first associated item sets can be obtained based on the path-reachable risk function and the association relationship between the path-reachable risk functions, and the association relationship between the path-reachable risk functions can be obtained based on the code path. Each first associated item set is a subset of the path-reachable first risk function, and the first associated item set includes multiple path-reachable first risk functions with association relationships, wherein the flowchart using code instrumentation is shown in FIG4 ; subsequently, the obtained one or more first associated item sets and the code path can be subjected to association analysis to obtain one or more second associated item sets, wherein the second associated item set includes a frequent item set and upstream and downstream functions of the frequent item set, wherein the frequent item set is a first associated item set with a support higher than a preset value; after obtaining one or more second associated item sets, the one or more second associated item sets can be analyzed by the evaluation model to obtain the risk assessment result of the tested code.

[0077] Optionally, an aggregation algorithm may be used to perform association analysis on one or more first associated item sets to obtain frequent item sets; then, according to the code path, upstream and downstream functions first associated with the frequent item sets may be called to obtain one or more second associated item sets.

[0078] The aggregation algorithm may include any one of an association rule (Apriori) algorithm, an FP-growth (Frequent Pattern) algorithm, an Eclat algorithm, or a K-means clustering (K-means) algorithm, and the specifics are not limited here.

[0079] Optionally, the support of each first associated item set may be calculated by an aggregation algorithm, and first associated item sets with support higher than a preset value may be regarded as frequent item sets.

[0080] Furthermore, during the risk assessment of the code under test, a risk function can be identified for the baseline code (also referred to as historical code) to obtain a second risk function. The specific code scanning tools and methods are similar to those used to identify the risk function for the code under test, and are not further described here. Based on the obtained second risk function, a rule base for risk function usage specifications is established. The rule base includes multiple second risk functions, multiple functions associated with the second risk functions, call paths for code associated with the second risk functions, constraints on the use of the second risk function, and constraints on the use of functions associated with the second risk function.

[0081] Optionally, an aggregation algorithm can be used to perform association analysis on the second risk function to obtain a frequent item set of the second risk function. The frequent item set of the second risk function includes multiple second risk functions. By calling the upstream and downstream functions of the second risk function, the functions associated with the second risk function and the calling path of the second risk function and its related functions can be obtained, and the usage constraints of the second risk function and its related functions can be obtained to construct a rule base.

[0082] Optionally, the uncovered code can also be automatically risk graded. The uncovered code is the newly added tested code. It can be judged whether the constraints of the uncovered code meet the requirements of the rule base. If it meets the requirements, the uncovered code will be marked as low risk. Otherwise, the uncovered code will be marked as high risk. The code marked as high risk will be analyzed in detail to assist manual code risk analysis.

[0083] In the embodiment of the present application, the risk function of path reachability can be queried to perform risk analysis on the risk function of path reachability, thereby reducing the workload of code analysis and improving the efficiency of code analysis.

[0084] Referring to FIG5 , a flowchart of another white box testing method provided by the present application is described below.

[0085] 501. Identify a first risk function included in the tested code;

[0086] 502. Filtering out first risk functions existing in the code path from the first risk functions using the code model;

[0087] In the embodiment of the present application, step 501 is similar to step 301 described in FIG. 3 , and step 502 is similar to step 302 described in FIG. 3 , and the details are not repeated here.

[0088] 503. Run the code under test and obtain one or more first associated item sets according to the first risk function existing in the code path;

[0089] The code path can be obtained by analyzing the running code under test using code instrumentation. After obtaining the first risk function existing in the code path, one or more first associated item sets can be obtained based on the path-reachable risk functions and the association relationship between the path-reachable risk functions. The association relationship between the path-reachable risk functions can be obtained based on the code path.

[0090] It is also possible to obtain a business flow model of the code under test, which includes the association relationship between businesses. The code under test is the actual code used to implement the business. The business includes multiple risk functions of the code under test. The functions corresponding to the businesses in the business flow model can be regarded as the first set, and the first risk function existing in the code path can be regarded as the second set. The intersection of the first set and the second set is taken to obtain the third set. The first key union set includes one or more risk functions in the third set, thereby obtaining one or more first association item sets; then, the code under test is run to obtain the code path by using code instrumentation.

[0091] 504. Perform association analysis on the one or more first associated item sets and the code path using an aggregation algorithm to obtain one or more second associated item sets;

[0092] After obtaining the first associated item set and the code path, an aggregation algorithm may be used to calculate a frequent item set, and an association analysis may be performed based on the frequent item set to obtain one or more second associated item sets.

[0093] Specifically, the support of one or more first associated item sets can be calculated through an aggregation algorithm, and the first associated item sets with support greater than a preset value can be used as frequent item sets. Then, based on the frequent item sets and the code path, upstream and downstream functions can be called to obtain one or more second associated item sets. The second associated item sets can include one or more functions in the frequent item sets or upstream and downstream associated functions.

[0094] 505. Analyze one or more second associated item sets using the evaluation model to obtain a risk evaluation result of the tested code.

[0095] After obtaining one or more second associated item sets, the second associated item sets can be analyzed using the assessment model to obtain a risk assessment result for the code under test. This risk assessment result can be high risk, medium risk, or low risk, or can be a warning message directly issued in response to the risk function, the specific details of which are not limited here. The specific security analysis process of the code under test is shown in Figure 6.

[0096] In an embodiment of the present application, an aggregation algorithm can be used to identify functions associated with the risk function, and the input and end points of the associated code can be identified with the risk function as the center, thereby identifying the complete call path of the risk function. The risk function segment path can be aggregated into a logical long path, which is more in line with the real business scenario. By identifying the associated functions of the risk function, the code scanning tool's false positives for the risk function can be corrected to improve the efficiency of code analysis.

[0097] This application can not only perform risk assessment on the code under test through the evaluation model, but also automatically perform risk classification on the newly added code under test. The following is a detailed introduction to the method of risk classification of the newly added code under test.

[0098] Referring to FIG7 , a flowchart of another white box testing method provided by the present application is described below.

[0099] 701. Identify a second risk function for the baseline code;

[0100] The baseline code is the historical code for comparative analysis. Code scanning tools can be used to identify the baseline code to obtain the second risk function, laying the foundation for the subsequent construction of a standardized rule base for the use of risk functions.

[0101] 702. Establish a rule base for risk function usage specifications based on the second risk function;

[0102] After obtaining the second risk function, an aggregation algorithm may be used to perform correlation analysis on the second risk function, thereby constructing a rule base for the use of risk functions.

[0103] Specifically, the second risk function obtained can be used as the center to call upstream and downstream related functions to obtain the function and calling path that are first associated with the second risk function. The constraints of the second risk function and the function first associated with the second risk function can also be identified. The constraints are the limiting conditions when the risk function is used. The aforementioned second risk function, the function first associated with the second risk function, the calling path and the constraints are stored, and then a rule base for the risk function usage specification is constructed.

[0104] 703. Based on the rule base, risk classification is performed on uncovered codes.

[0105] For uncovered code, a comparative analysis can be performed using a standardized rule base based on the risk function to obtain the risk level of the uncovered code.

[0106] Specifically, the uncovered code can be risk-graded by determining whether its constraints meet the requirements of the rule base. If the constraints meet the requirements of the rule base, the uncovered code can be marked as low risk; otherwise, the uncovered code can be marked as high risk.

[0107] In addition, if there is no corresponding function in the rule base for the function in the uncovered code, the function is marked and analyzed, and after obtaining the risk grading result of the function, the risk grading and constraint conditions of the function are stored in the rule base, thereby completing the update of the rule base.

[0108] Therefore, in an embodiment of the present application, the risk of uncovered code can be automatically graded, so that the code path for key analysis can be identified according to the marked risk level, and the reachable code path and reachable data set can be automatically analyzed according to the risk level to assist testers in troubleshooting and improve code analysis efficiency.

[0109] The above describes the method flow provided by the present application. Based on the above method flow, the following describes the device provided by the present application.

[0110] Referring to FIG8 , a schematic diagram of the structure of a white box testing device provided by the present application includes:

[0111] Identification module 801, used to identify a first risk function included in the tested code;

[0112] A screening module 802 is configured to screen out first risk functions existing in a code path from the first risk functions using a code model, wherein the code model includes associations between upstream and downstream components in the tested code, and the code path is an execution path of the tested code;

[0113] The analysis module 803 is used to analyze the first risk function existing in the code path through the evaluation model to obtain a risk evaluation result of the tested code.

[0114] In one possible implementation, the analysis module 803 is specifically configured to run the code under test to obtain one or more first associated item sets, where the first associated item set includes a first risk function present in the code path, and the first associated item set includes a plurality of risk functions having an associated relationship; perform association analysis on the one or more first associated item sets and the code path to obtain one or more second associated item sets, where the second associated item sets include frequent item sets and upstream and downstream functions of the frequent item sets, and the frequent item sets include first associated item sets with a support higher than a preset value; and analyze the one or more second associated item sets through an evaluation model to obtain a risk assessment result of the code under test.

[0115] In a possible implementation, the aforementioned code path may be obtained by analyzing the running code under test in a code instrumentation manner.

[0116] In a possible implementation, the analysis module 803 is specifically configured to obtain one or more first associated item sets according to the first risk functions existing in the code paths and the association relationships between the first risk functions existing in the code paths.

[0117] In a possible implementation, the analysis module 803 is specifically configured to perform association analysis on one or more first associated item sets using an aggregation algorithm to obtain frequent item sets; and call upstream and downstream functions associated with the frequent item sets according to the code path to obtain one or more second associated item sets.

[0118] In a possible implementation, the analysis module 803 is specifically configured to calculate the support of one or more first associated item sets using an aggregation algorithm; and to take the first associated item sets with support higher than a preset value as frequent item sets.

[0119] In a possible implementation, if the code under test includes uncovered code, and the uncovered code is newly added code to be tested in the code under test, the apparatus further includes:

[0120] The identification module 801 is further configured to identify a second risk function included in the baseline code;

[0121] Establishing module 804, for establishing a rule base of risk function usage specifications based on the second risk function, where the rule base is used to indicate constraints on the use of the second risk function and functions associated with the second risk function;

[0122] A classification module 805 is used to classify the risk of uncovered code according to a rule base;

[0123] The updating module 806 is used to update the risk assessment result of the tested code according to the result of the risk classification to obtain an updated risk assessment result.

[0124] In one possible implementation, a module 804 is established, which is specifically used to use an aggregation algorithm to perform association analysis on the second risk function to obtain functions associated with the second risk function, calling paths of functions associated with the second risk function, and constraints; and a rule base is established based on the second risk function, functions associated with the second risk function, and calling paths of functions associated with the second risk function, as well as constraints.

[0125] In one possible implementation, the grading module 805 is specifically configured to determine whether the constraints of the uncovered code meet the requirements of the rule base; if so, the uncovered code is marked as low risk; if not, the uncovered code is marked as high risk.

[0126] The identification module, screening module, analysis module, establishment module, classification module, and update module can all be implemented via software or hardware. For example, the implementation of the identification module will be described below using the example of the identification module. Similarly, the implementation of the screening module, analysis module, establishment module, classification module, and update module can refer to the implementation of the identification module.

[0127] As an example of a software functional unit, the identification module may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Furthermore, the computing instance may be one or more. For example, the identification module may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region or in different regions. Furthermore, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including one data center or multiple geographically close data centers. Typically, a region may include multiple AZs.

[0128] Similarly, multiple hosts / virtual machines / containers running the code can be distributed within the same virtual private cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a region. Cross-region communication between two VPCs within the same region, or between VPCs in different regions, requires a communication gateway within each VPC to interconnect the VPCs.

[0129] As an example of a hardware functional unit, the identification module may include at least one computing device, such as a server. Alternatively, the identification module may be implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0130] The multiple computing devices included in the identification module can be distributed in the same region or in different regions. The multiple computing devices included in the identification module can be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in the identification module can be distributed in the same VPC or in multiple VPCs. The multiple computing devices can be any combination of servers, ASICs, PLDs, CPLDs, FPGAs, GALs, and other computing devices.

[0131] It should be noted that, in other embodiments, the identification module, screening module, analysis module, establishment module, grading module and update module can all be used to execute any step in the white box testing method. The steps that the screening module, analysis module, establishment module, grading module and update module are responsible for implementing can be specified as needed. The full functions of the white box testing device can be realized by respectively implementing different steps in the white box testing method through the screening module, analysis module, establishment module, grading module and update module.

[0132] This application also provides a computing device 900. As shown in Figure 9, computing device 900 includes a bus 902, a processor 904, a memory 906, and a communication interface 908. Processor 904, memory 906, and communication interface 908 communicate with each other via bus 902. Computing device 900 can be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in computing device 900.

[0133] Bus 902 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, among others. Buses may be classified as address buses, data buses, control buses, and the like. For ease of illustration, FIG9 illustrates a single bus line, but this does not imply a single bus or type of bus. Bus 904 may include a path for transmitting information between various components of computing device 900 (e.g., memory 906, processor 904, and communication interface 908).

[0134] The processor 904 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0135] The memory 906 may include volatile memory, such as random access memory (RAM). The processor 904 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).

[0136] Memory 906 stores executable program code, which processor 904 executes to implement the functions of the aforementioned identification module, screening module, analysis module, establishment module, classification module, and update module, thereby implementing the white-box testing method. In other words, memory 906 stores instructions for executing the white-box testing method.

[0137] Embodiments of the present application also provide a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.

[0138] As shown in Figure 10, the computing device cluster includes at least one computing device 900. The memory 906 in one or more computing devices 900 in the computing device cluster may store the same instructions for executing the white box testing method.

[0139] In some possible implementations, the memory 906 of one or more computing devices 900 in the computing device cluster may also store partial instructions for executing the white-box testing method. In other words, the combination of one or more computing devices 900 can jointly execute the instructions for executing the white-box testing method.

[0140] It should be noted that the memory 906 in different computing devices 900 in the computing device cluster can store different instructions, each for executing a portion of the functions of the white-box testing apparatus. In other words, the instructions stored in the memory 906 in different computing devices 900 can implement the functions of one or more of the identification module, screening module, analysis module, establishment module, grading module, and update module.

[0141] In some possible implementations, one or more computing devices in a computing device cluster may be connected via a network. The network may be a wide area network (WAN) or a local area network (LAN), etc. FIG. 11 illustrates a possible implementation. As shown in FIG. 11 , two computing devices 900A and 900B are connected via a network. Specifically, each computing device is connected to the network via a communication interface within the computing device. In this type of possible implementation, the memory 906 in the computing device 900A stores instructions for executing the functions of the identification module. Simultaneously, the memory 906 in the computing device 900B stores instructions for executing the functions of the screening module, the analysis module, the establishment module, the grading module, and the update module.

[0142] It should be understood that the functionality of the computing device 900A shown in FIG11 may also be implemented by multiple computing devices 900. Similarly, the functionality of the computing device 900B may also be implemented by multiple computing devices 900.

[0143] The present application also provides a computer program product comprising instructions. The computer program product may be software or a program product comprising instructions that can be run on a computing device or stored in any available medium. When the computer program product is run on at least one computing device, the computer program product causes the at least one computing device to perform a white-box testing method.

[0144] The present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device or a data storage device such as a data center that contains one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to perform the white box testing method.

[0145] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the protection scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A white box testing method, characterized in that: include: Identifying a first risk function included in the tested code; Filtering out a first risk function existing in a code path from the first risk function using a code model, wherein the code model includes an association relationship between upstream and downstream in the tested code, and the code path is an execution path of the tested code; The first risk function existing in the code path is analyzed by using an evaluation model to obtain a risk evaluation result of the tested code.

2. The method according to claim 1, characterized in that The step of analyzing the first risk function existing in the code path by using the evaluation model to obtain the risk evaluation result of the tested code includes: Running the code under test to obtain one or more first associated item sets, wherein the first associated item set includes a first risk function existing in the code path, and the first associated item set includes a plurality of risk functions having an associated relationship; Performing association analysis on the one or more first associated item sets and the code path to obtain one or more second associated item sets, wherein the second associated item sets include frequent item sets and upstream and downstream functions of the frequent item sets, and the frequent item sets include the first associated item sets with a support higher than a preset value; The one or more second associated item sets are analyzed by using the assessment model to obtain a risk assessment result of the tested code.

3. The method according to claim 2, characterized in that The step of running the code under test to obtain one or more first associated item sets includes: The one or more first associated item sets are obtained according to the first risk functions existing in the code paths and the association relationships between the first risk functions existing in the code paths.

4. The method according to any one of claims 2 or 3, characterized in that The performing association analysis on the one or more first associated item sets and the code path to obtain one or more second associated item sets includes: Performing association analysis on the one or more first associated item sets using an aggregation algorithm to obtain the frequent item sets; According to the code path, upstream and downstream functions associated with the frequent item sets are called to obtain the one or more second associated item sets.

5. The method according to claim 4, characterized in that The adopting an aggregation algorithm to perform association analysis on the one or more first associated item sets to obtain the frequent item sets includes: Using the aggregation algorithm to calculate the support of the one or more first associated item sets; The first associated item set whose support is higher than the preset value is used as the frequent item set.

6. The method according to any one of claims 1 to 5, characterized in that If the tested code includes uncovered code, and the uncovered code is a newly added code to be tested in the tested code, the method further includes: Identify secondary risk functions included in the baseline code; According to the second risk function, establishing a rule base of risk function usage specifications, wherein the rule base is used to indicate constraint conditions for use of the second risk function and functions associated with the second risk function; According to the rule base, risk classification is performed on the uncovered code; The risk assessment result of the tested code is updated according to the result of the risk grading to obtain an updated risk assessment result.

7. The method according to claim 6, characterized in that The step of establishing a rule base for using a risk function according to the second risk function includes: Performing association analysis on the second risk function using the aggregation algorithm to obtain a function associated with the second risk function, a calling path of the function associated with the second risk function, and the constraint condition; The rule base is established according to the second risk function, the function associated with the second risk function, the calling path of the function associated with the second risk function, and the constraint condition.

8. The method according to any one of claims 6 or 7, characterized in that The risk classification of uncovered codes according to the rule base includes: Determining whether the constraint condition of the uncovered code meets the requirement of the rule base; If the requirements are met, the uncovered code is marked as low risk; If the requirements are not met, the uncovered code is marked as high risk.

9. A white box testing device, characterized in that: include: An identification module, used for identifying a first risk function included in the tested code; A screening module, used to screen out the first risk function existing in the code path from the first risk function by using a code model, wherein the code model includes an association relationship between upstream and downstream in the tested code, and the code path is an execution path of the tested code; The analysis module is used to analyze the first risk function existing in the code path through an evaluation model to obtain a risk evaluation result of the tested code.

10. The device according to claim 9, characterized in that The analysis module is specifically used for: Running the code under test to obtain one or more first associated item sets, wherein the first associated item set includes a first risk function existing in the code path, and the first associated item set includes a plurality of risk functions having an associated relationship; Performing association analysis on the one or more first associated item sets and the code path to obtain one or more second associated item sets, wherein the second associated item sets include frequent item sets and upstream and downstream functions of the frequent item sets, and the frequent item sets include the first associated item sets with a support higher than a preset value; The one or more second associated item sets are analyzed by using the assessment model to obtain a risk assessment result of the tested code.

11. The device according to claim 10, characterized in that The analysis module is specifically used for: The one or more first associated item sets are obtained according to the first risk functions existing in the code paths and the association relationships between the first risk functions existing in the code paths.

12. The device according to any one of claims 10 or 11, characterized in that The analysis module is specifically used for: Performing association analysis on the one or more first associated item sets using an aggregation algorithm to obtain the frequent item sets; According to the code path, upstream and downstream functions associated with the frequent item sets are called to obtain the one or more second associated item sets.

13. The device according to claim 12, characterized in that The analysis module is specifically used for: Using the aggregation algorithm to calculate the support of the one or more first associated item sets; The first associated item set whose support is higher than the preset value is used as the frequent item set.

14. The device according to any one of claims 9 to 13, characterized in that If the code under test includes uncovered code, and the uncovered code is a newly added code to be tested in the code under test, the device further includes: The identification module is further used to identify a second risk function included in the baseline code; An establishing module, configured to establish a rule base for use specifications of a risk function according to the second risk function, wherein the rule base is used to indicate constraint conditions for use of the second risk function and functions associated with the second risk function; A classification module, used for classifying the risks of uncovered codes according to the rule base; An updating module is used to update the risk assessment result of the tested code according to the result of the risk classification to obtain an updated risk assessment result.

15. The device according to claim 14, characterized in that The establishment module is specifically used for: Performing association analysis on the second risk function using the aggregation algorithm to obtain a function associated with the second risk function, a calling path of the function associated with the second risk function, and the constraint condition; The rule base is established according to the second risk function, the function associated with the second risk function, the calling path of the function associated with the second risk function, and the constraint condition.

16. The device according to any one of claims 14 or 15, characterized in that The grading module is specifically used for: Determining whether the constraint condition of the uncovered code meets the requirement of the rule base; If the requirements are met, the uncovered code is marked as low risk; If the requirements are not met, the uncovered code is marked as high risk.

17. A computing device cluster, characterized in that: comprising at least one computing device, each computing device comprising a processor and a memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the method according to any one of claims 1 to 8.

18. A computer program product comprising instructions, characterized in that When the instructions are executed by a computing device cluster, the computing device cluster is caused to perform the method according to any one of claims 1 to 8.

19. A computer-readable storage medium, characterized in that: The method comprises computer program instructions, and when the computer program instructions are executed by a computing device cluster, the computing device cluster performs the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Program bug detection system and method

    CN102982282A

  • Vulnerability detection method and device based on assembly codes and electronic equipment

    CN112906004A

  • Software security code analyzer based on source code static analysis and detection method thereof

    CN116186705A

  • Code analysis method and device based on neural network, and electronic equipment

    CN116578980A

  • Generating containers for applications utilizing reduced sets of libraries based on risk analysis

    US20180025160A1