White box test method and device, program product and medium

By using code models to filter and evaluate the risk functions that can be accessed by paths in white box tests, the problem of low static analysis efficiency is solved, and a more efficient code risk assessment is achieved.

CN120029898APending Publication Date: 2025-05-23HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311628785.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-21
Filing Date
2023-11-30
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The static analysis scheme for existing white box tests requires traversing all code paths, resulting in low analysis efficiency.

Method used

By obtaining the code model of the tested code, filtering out the risk functions in the code path, using the evaluation model to analyze these risk functions, obtaining risk assessment results, thereby improving analysis efficiency.

Benefits of technology

The analysis of risk functions with unreachable paths is avoided, the analysis workload is reduced, and the code analysis efficiency is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029898A_ABST
    Figure CN120029898A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a white-box testing method and device, a program product and a medium in the field of computers, which are used for carrying out risk assessment on a tested code and analyzing a path reachable risk function to obtain a risk assessment result of the tested code, so that the code analysis efficiency is improved. The method comprises the following steps: identifying a first risk function included in a tested code; a first risk function existing in a code path is screened out from the first risk functions through a code model, the code model comprises the incidence relation between the upstream and the downstream in the tested code, and the code path is an execution path of the tested code; and analyzing a first risk function existing in the code path through an evaluation model to obtain a risk evaluation result of the tested code.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims the priority of the Chinese patent application filed with the State Intellectual Property Office on November 21, 2023, with application number 202311558315.3 and invention name “A white box code security analysis method and device”, the entire contents of which are incorporated by reference in this application. Technical Field

[0002] The present application relates to the field of computer technology, and in particular to a white box testing method, device, program product and medium. Background Art

[0003] With the rapid development of information technology, software plays an increasingly important role in our daily life and work. However, software security has always been a serious problem. In order to ensure the security of software, software white box security verification has become an important technical means. Software white box security verification refers to the comprehensive analysis and testing of the internal structure and logic of the software to discover potential security vulnerabilities and weaknesses. Compared with traditional black box testing, software white box security verification can provide a deeper understanding of the internal operating mechanism of the software, thereby more accurately evaluating its security.

[0004] At present, there are many methods and technologies for software white-box security verification, including static analysis, dynamic analysis, symbolic execution, fuzz testing, etc. Common white-box testing includes two solutions, dynamic analysis and static analysis. Among them, the static analysis solution refers to the analysis of program source code without running the program to find security vulnerabilities. The common method is to use tools or manual identification of dangerous functions, obtain the correlation relationship of risky codes, and analyze and mark risky codes based on business and experience.

[0005] However, although the static analysis solution of white-box testing can automatically identify risky functions in the code, it needs to traverse all code paths, including risky paths that are reachable and unreachable, resulting in low code analysis efficiency. Summary of the invention

[0006] The present application provides a white box testing method, device, program product and medium for performing risk assessment on the code under test, and obtaining the risk assessment result of the code under test by analyzing the path reachable risk function, thereby improving the efficiency of code analysis.

[0007] In view of this, on the first aspect, the present application provides a white box testing method, comprising: obtaining a code under test, and performing risk function identification on the obtained code under test to obtain a first risk function included in the code under test; in addition, a code model of the code under test should also be obtained, and the code model includes the association relationship between upstream and downstream in the code under test; then, the first risk function existing in the code path can be screened out through the obtained code model, and the code path is the execution path of the code under test; finally, the first risk function existing in the code path is analyzed through the evaluation model to obtain a risk assessment result of the code under test.

[0008] In the implementation method of the present application, the risk functions existing in the code path, i.e., the path-reachable risk functions, can be screened out according to the code model of the code under test. The risk functions existing in the code path are analyzed by the evaluation model to obtain the risk assessment results, thereby avoiding the analysis of the risk functions of the unreachable path, thereby reducing the workload of risk function analysis and identification and improving the analysis efficiency of the code under test.

[0009] In a possible implementation, the aforementioned analysis of the first risk function existing in the code path through the evaluation model to obtain the risk assessment result of the tested code may include: running the tested code to obtain one or more first associated item sets, the first associated item sets are included in the first risk function existing in the code path, and the first associated item sets include multiple risk functions with associated relationships; performing association analysis on the one or more first associated item sets and the code path to obtain one or more second associated item sets, the second associated item sets include frequent item sets and upstream and downstream functions of the frequent item sets, and the frequent item sets include first associated item sets with support higher than a preset value; analyzing the one or more second associated item sets through the evaluation model to obtain the risk assessment result of the tested code.

[0010] In an embodiment of the present application, association analysis can be further performed based on the obtained first associated item set and code path to obtain frequent item sets and upstream and downstream functions associated with the frequent item sets, i.e., second associated item sets, and risk analysis can be performed on the second associated item sets. By gradually searching and analyzing more important risk functions and the code paths of the risk functions, the efficiency of code analysis can be improved.

[0011] In a possible implementation, the aforementioned code path may be obtained by analyzing the running code under test in a code instrumentation manner.

[0012] In the embodiment of the present application, the code path and the data flow path can be automatically obtained by code insertion, which can avoid missed detection when manually traversing the path.

[0013] In a possible implementation, the aforementioned running of the tested code to obtain one or more first associated item sets may include: obtaining one or more first associated item sets based on first risk functions existing in the code path and the association relationship between the first risk functions existing in the code path.

[0014] In a possible implementation, the aforementioned performing association analysis on one or more first associated item sets and code paths to obtain one or more second associated item sets may include: performing association analysis on one or more first associated item sets using an aggregation algorithm to obtain frequent item sets; and calling upstream and downstream functions associated with the frequent item sets according to the code paths to obtain one or more second associated item sets.

[0015] In an embodiment of the present application, an aggregation algorithm can be used to perform association analysis on the risk function to obtain frequent item sets of the risk function and functions associated with the frequent item sets. Key code paths can be analyzed based on the obtained frequent item sets, thereby improving the efficiency of code analysis and correcting false positives generated when the code recognition tool identifies the risk function.

[0016] In a possible implementation, the aforementioned use of an aggregation algorithm to perform association analysis on one or more first associated item sets to obtain frequent item sets may include: using an aggregation algorithm to calculate the support of one or more first associated item sets; and taking first associated item sets whose support is higher than a preset value as frequent item sets.

[0017] In a possible implementation, if the code under test includes uncovered code, and the uncovered code is a newly added code to be tested in the code under test, the method may also include: identifying a second risk function included in the baseline code; based on the second risk function, establishing a rule base for the use of risk functions, wherein the rule base is used to indicate constraints on the use of the second risk function and functions associated with the second risk function; risk grading the uncovered code based on the rule base; and updating the risk assessment result of the code under test based on the result of the risk grading to obtain an updated risk assessment result.

[0018] In the implementation manner of the present application, after risk analysis is performed on the code under test, if new code to be tested is added to the code under test, only the newly added code to be tested can be automatically risk graded, and the risk assessment result of the code under test can be updated based on the risk grading result, thereby simplifying the risk analysis process and improving the efficiency of code risk analysis.

[0019] In one possible implementation, the aforementioned establishment of a rule base for risk function usage specifications based on the second risk function may include: using an aggregation algorithm to perform association analysis on the second risk function to obtain functions associated with the second risk function, calling paths of functions associated with the second risk function, and constraints; establishing a rule base based on the second risk function, functions associated with the second risk function, and calling paths of functions associated with the second risk function, and constraints.

[0020] In a possible implementation, the aforementioned risk grading of uncovered code based on the rule base may include: determining whether the constraints of the uncovered code meet the requirements of the rule base; if so, marking the uncovered code as low risk; if not, marking the uncovered code as high risk.

[0021] In the implementation manner of the present application, by automatically grading the risks of uncovered codes, a key investigation direction is provided for the subsequent investigation, which can assist testers in investigating high-risk codes and improve the efficiency of code analysis.

[0022] In a second aspect, the present application provides a white box testing device, comprising:

[0023] An identification module, used for identifying a first risk function included in the tested code;

[0024] A screening module, used to screen out the first risk function existing in the code path from the first risk function by using a code model, wherein the code model includes the association relationship between the upstream and downstream in the tested code, and the code path is the execution path of the tested code;

[0025] The analysis module is used to analyze the first risk function existing in the code path through the evaluation model to obtain the risk evaluation result of the tested code.

[0026] In a possible implementation, the analysis module is specifically used to run the code under test to obtain one or more first associated item sets, the first associated item set includes a first risk function existing in the code path, and the first associated item set includes multiple risk functions that have an associated relationship; perform association analysis on the one or more first associated item sets and the code path to obtain one or more second associated item sets, the second associated item sets include frequent item sets and upstream and downstream functions of the frequent item sets, and the frequent item sets include first associated item sets with support higher than a preset value; analyze the one or more second associated item sets through an evaluation model to obtain a risk assessment result of the code under test.

[0027] In a possible implementation, the aforementioned code path may be obtained by analyzing the running code under test in a code instrumentation manner.

[0028] In a possible implementation, the analysis module is specifically configured to obtain one or more first associated item sets according to the first risk functions existing in the code paths and the association relationships between the first risk functions existing in the code paths.

[0029] In a possible implementation, the analysis module is specifically configured to perform association analysis on one or more first associated item sets using an aggregation algorithm to obtain frequent item sets; and to call upstream and downstream functions associated with the frequent item sets according to the code path to obtain one or more second associated item sets.

[0030] In a possible implementation, the analysis module is specifically configured to calculate the support of one or more first associated item sets by using an aggregation algorithm; and to take the first associated item sets whose support is higher than a preset value as frequent item sets.

[0031] In a possible implementation manner, if the code under test includes uncovered code, the uncovered code is a newly added code to be tested in the code under test, the device further includes:

[0032] The identification module is further used to identify a second risk function included in the baseline code;

[0033] An establishment module is used to establish a rule base of risk function usage specifications according to the second risk function, and the rule base is used to indicate the constraint conditions for the use of the second risk function and functions associated with the second risk function;

[0034] The classification module is used to classify the risks of uncovered codes according to the rule base;

[0035] The update module is used to update the risk assessment results of the tested code according to the risk classification results to obtain updated risk assessment results.

[0036] In one possible implementation, a module is established, which is specifically used to use an aggregation algorithm to perform association analysis on the second risk function to obtain functions associated with the second risk function, calling paths of functions associated with the second risk function, and constraints; a rule base is established based on the second risk function, functions associated with the second risk function, and calling paths of functions associated with the second risk function, as well as constraints.

[0037] In a possible implementation, the grading module is specifically used to determine whether the constraints of the uncovered code meet the requirements of the rule base; if so, the uncovered code is marked as low risk; if not, the uncovered code is marked as high risk.

[0038] In a third aspect, the present application provides a computing device cluster, comprising at least one computing device, each computing device comprising a processor and a memory; the processor of at least one computing device is used to execute instructions stored in the memory of at least one computing device, so that the computing device cluster executes the method described in the first aspect above.

[0039] In a fourth aspect, the present application provides a computer program product comprising instructions, characterized in that when the instructions are executed by a computing device cluster, the computing device cluster executes the method described in the first aspect above.

[0040] In a fifth aspect, the present application provides a computer-readable storage medium, characterized in that it includes computer program instructions. When the computer program instructions are executed by a computing device cluster, the computing device cluster executes the method described in the first aspect above. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 A schematic diagram of a system framework provided for this application;

[0042] Figure 2 A code security analysis and verification flow chart provided for this application;

[0043] Figure 3 A flowchart of a white box testing method provided for this application;

[0044] Figure 4 A code instrumentation analysis flow chart provided for this application;

[0045] Figure 5 A flowchart of another white box testing method provided for this application;

[0046] Figure 6 A running sequence diagram of a tested program provided for this application;

[0047] Figure 7 A flowchart of another white box testing method provided for this application;

[0048] Figure 8 A schematic diagram of the structure of a white box testing device provided in this application;

[0049] Fig. 9 This is a schematic diagram of the structure of a computing device in an embodiment of the present application;

[0050] Fig.10 A schematic diagram of the structure of a computing device cluster in an embodiment of the present application;

[0051] Fig.11 Another structural diagram of a computing device cluster in an embodiment of the present application. DETAILED DESCRIPTION

[0052] The following will describe the technical solutions in the embodiments of the present application in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present application.

[0053] First, the terms involved in this application are explained:

[0054] White box testing: White box testing, also known as structural testing, transparent box testing, logic-driven testing or code-based testing, is a test case design method. The box refers to the software being tested, and the white box means that the box is visible, that is, it is clear what is inside the box and how it works. The "white box" method can fully understand the internal logical structure of the program and test all logical paths.

[0055] Code stub: refers to inserting code into the program under test to track the execution process of the program under test, so as to obtain the execution of executable statements in the program and the execution path of the program.

[0056] Frequent itemsets: Frequent patterns refer to sets of items, sequences, or substructures that appear frequently in a data set. Frequent itemsets refer to sets whose support is greater than or equal to the minimum support.

[0057] Support: refers to the frequency of a set appearing in all transactions, that is, the probability that the item set {A, B} appears in the total item set.

[0058] Risky function: refers to a function with potential risks or vulnerabilities in the code. When executing this function, it may introduce security risks, performance issues, memory leaks, or abnormal crashes.

[0059] The following is an introduction to the system framework on which the embodiments of the present application are based.

[0060] See also Figure 1 , the present application provides a system framework 100. As shown in the system framework 100, the automatic analysis module 110 can be directly used for risk assessment analysis of the tested code, and a risk assessment report of the tested code is generated through the report generation module 120. The automatic analysis module 110 may include a risk function identification module 111, a code instrumentation analysis module 112, a function sequence aggregation solution module 113, a data storage module 114, and a data comparison module 115.

[0061] Among them, the risk function identification module 111 is mainly responsible for identifying the risk functions or methods that may exist in the tested code, and the code stub analysis module 112 can analyze the situation when the business flow is executed, and the business flow is the complete process of realizing a business or function. The function sequence aggregation solution module 113 can obtain the association relationship of functions and the frequent item set of function sequences through the aggregation algorithm. The data storage module 114 can be used to store the identified risk functions and their associated functions, code paths, etc. The data comparison module 115 is mainly used to perform comparative analysis of functions to obtain the risk level of the function.

[0062] The execution subject of the white box testing method provided in this application can be a white box testing tool, which can be a program code software, or a medium storing relevant execution code, or the white box testing tool can also be a physical device integrated or installed with relevant execution code, such as a chip, a microcontroller unit (MCU), a computer, a computer and other electronic devices. In addition, this application can be applied not only to the identification of risk functions of the tested code, but also to precision testing, scenario testing and code performance testing, which are not limited here.

[0063] Combine the following Figure 1 The system framework described introduces the code security analysis and verification process provided by this application.

[0064] See also Figure 2 , the code security analysis and verification flow chart provided in this application, the baseline code can obtain the risk function of the baseline code, the function associated with the risk function, the call path of the risk function, the constraint conditions of the risk function and the constraint conditions of the function associated with the risk function through the risk function identification module 111, the code stub analysis module 112 and the function sequence aggregation solution module 113 in the automatic analysis module 110, and store them in the data storage module 114, so as to construct a rule base for function usage specifications. The data comparison module 115 can be used to automatically grade the risks of the newly added code, and judge the constraints of the newly added code through the constructed rule base to obtain the risk level of the newly added code. The tested code can obtain a risk assessment report through the automatic analysis module 110 and the report generation module 120.

[0065] See also Figure 3 , a flow chart of a white box testing method provided in this application is described as follows.

[0066] 301. Identify a first risk function included in the tested code;

[0067] Specifically, a code scanning tool can be used to identify the first risk function or method included in the tested code. The code scanning tool can select different tools according to different code languages. For example, the bandit code scanning tool can be used for the Python language, and the spotbugs code scanning tool can be used for the Java language. The specific details are not limited here.

[0068] The code scanning tool may identify the risk function or method by identifying keywords, or may identify the risk function or method by one or more methods of syntax analysis or pattern matching, which are not specifically limited here.

[0069] 302. Filter out the first risk function existing in the code path from the first risk function by using the code model;

[0070] Among them, the code model includes the association relationship between the upstream and downstream of the tested code, that is, it includes all the code paths of the tested code. Therefore, the code model can be used to screen out the first risk function (also called the path-reachable first risk function) existing in the code path. The code path is the execution path of the tested code, and the first risk function existing in the code path is a subset of the first risk function.

[0071] 303. Analyze the first risk function existing in the code path through the evaluation model to obtain a risk evaluation result of the tested code.

[0072] After obtaining the first risk function existing in the code path, it can be analyzed through the evaluation model to obtain the risk evaluation result of the tested code. In addition, the risk function of the uncovered code can also be evaluated.

[0073] Generally, by using code stubbing, it is possible to obtain information about the program when it is running and analyze the behavior of the program when it is running, such as debugging, performance analysis, or security testing. In the embodiment of the present application, by using code stubbing, it is possible to monitor the execution of the tested code to obtain the correlation relationship of the risk function of the tested code when it is running and the execution path of the code.

[0074] Optionally, the aforementioned code path may be obtained by analyzing the running code under test in a code instrumentation manner.

[0075] Optionally, after obtaining the path-reachable first risk function, one or more first associated item sets can be obtained based on the path-reachable risk function and the association relationship between the path-reachable risk functions, and the association relationship between the path-reachable risk functions can be obtained based on the code path. Each first associated item set is a subset of the path-reachable first risk function, and the first associated item set includes multiple path-reachable first risk functions with association relationships, wherein the flowchart using code stubs is as follows Figure 4 As shown; subsequently, an association analysis can be performed on the obtained one or more first associated item sets and the code path to obtain one or more second associated item sets, and the second associated item sets include frequent item sets and upstream and downstream functions of the frequent item sets, wherein the frequent item sets are first associated item sets whose support is higher than a preset value; after obtaining the one or more second associated item sets, the one or more second associated item sets can be analyzed through the evaluation model to obtain a risk assessment result of the tested code.

[0076] Optionally, an aggregation algorithm may be used to perform association analysis on one or more first associated item sets to obtain frequent item sets; then, according to the code path, upstream and downstream functions first associated with the frequent item sets may be called to obtain one or more second associated item sets.

[0077] The aggregation algorithm may include any one of an association rule (Apriori) algorithm, an FP-growth (Frequent Pattern) algorithm, an Eclat algorithm, or a K-means clustering (K-means) algorithm, and the specifics are not limited here.

[0078] Optionally, the support of each first associated item set may be calculated by an aggregation algorithm, and the first associated item sets with support higher than a preset value may be regarded as frequent item sets.

[0079] In addition, in the process of risk assessment of the tested code, the risk function of the baseline code (also called historical code) can also be identified to obtain a second risk function. The specific code scanning tools and methods are similar to the aforementioned risk function identification of the tested code, and will not be repeated here. According to the obtained second risk function, a rule base for the use of risk functions is established, and the rule base includes multiple second risk functions, multiple functions associated with the second risk function, the calling path of the code associated with the second risk function, the constraints on the use of the second risk function, and the constraints on the use of functions associated with the second risk function.

[0080] Optionally, an aggregation algorithm can be used to perform association analysis on the second risk function to obtain a frequent item set of the second risk function, which includes multiple second risk functions. By calling the upstream and downstream functions of the second risk function, the functions associated with the second risk function and the calling path of the second risk function and its related functions can be obtained, and the usage constraints of the second risk function and its associated functions can be obtained, thereby constructing a rule base.

[0081] Optionally, the risk of uncovered code can be automatically graded. The uncovered code is the newly added code under test. It can be determined whether the constraints of the uncovered code meet the requirements of the rule base. If it meets the requirements, the uncovered code will be marked as low risk; otherwise, the uncovered code will be marked as high risk. The code marked as high risk will be analyzed in detail to assist manual code risk analysis.

[0082] In the embodiment of the present application, the risk function of the path reachable can be queried to perform risk analysis on the path reachable risk function, thereby reducing the workload of code analysis and improving the efficiency of code analysis.

[0083] See also Figure 5 , a flow chart of another white box testing method provided by this application is described as follows.

[0084] 501. Identify a first risk function included in the tested code;

[0085] 502. Filter out the first risk function existing in the code path from the first risk function by using the code model;

[0086] In the embodiment of the present application, step 501 is the same as the above Figure 3 The step 301 is similar to the above step 502. Figure 3 The step 302 is similar and will not be described in detail here.

[0087] 503. Run the code under test, and obtain one or more first associated item sets according to the first risk function existing in the code path;

[0088] The code path can be obtained by analyzing the running code under test by code instrumentation. After obtaining the first risk function existing in the code path, one or more first associated item sets can be obtained according to the path-reachable risk functions and the association relationship between the path-reachable risk functions, and the association relationship between the path-reachable risk functions can be obtained according to the code path.

[0089] A business flow model of the code under test can also be obtained, and the business flow model includes the association relationship between businesses. The code under test is the actual code used to implement the business. The business includes multiple risk functions of the code under test. The functions corresponding to the businesses in the business flow model can be regarded as the first set, and the first risk function existing in the code path can be regarded as the second set. The intersection of the first set and the second set is taken to obtain the third set. The first key union set includes one or more risk functions in the third set, thereby obtaining one or more first association item sets; then, the code under test is run by code instrumentation to obtain the code path.

[0090] 504. Perform association analysis on one or more first associated item sets and code paths using an aggregation algorithm to obtain one or more second associated item sets;

[0091] After obtaining the first associated item set and the code path, an aggregation algorithm may be used to calculate a frequent item set, and an association analysis may be performed based on the frequent item set to obtain one or more second associated item sets.

[0092] Specifically, the support of one or more first associated item sets can be calculated through an aggregation algorithm, and the first associated item sets with support greater than a preset value can be used as frequent item sets. Then, according to the frequent item sets and the code path, the upstream and downstream functions can be called to obtain one or more second associated item sets. The second associated item sets can include one or more functions in the frequent item sets or functions associated with the upstream and downstream.

[0093] 505. Analyze one or more second associated item sets through the evaluation model to obtain a risk evaluation result of the tested code.

[0094] After obtaining one or more second associated item sets, the second associated item sets can be analyzed by the evaluation model to obtain a risk assessment result of the tested code, which can be a high risk, a medium risk, or a low risk, or a warning message directly issued to the risk function, which is not limited here. Figure 6 shown.

[0095] In an embodiment of the present application, an aggregation algorithm can be used to identify functions associated with a risk function, and the input and end points of the associated code can be identified with the risk function as the center, thereby identifying the complete calling path of the risk function. The risk function segment paths can be aggregated into a logical long path, which is more in line with the actual business scenario. By identifying the associated functions of the risk function, the code scanning tool's false positives for the risk function can be corrected to improve the efficiency of code analysis.

[0096] This application can not only perform risk assessment on the code under test through an evaluation model, but also automatically classify the risk levels of the newly added code to be tested. The following details the method for classifying the risk levels of the newly added code to be tested.

[0097] Refer to Figure 7 , the process schematic diagram of another white box testing method provided by this application is described as follows.

[0098] 701. Identify the second risk function of the baseline code;

[0099] The baseline code is the historical code for comparative analysis. A code scanning tool can be used to identify the baseline code to obtain the second risk function, laying a foundation for subsequent construction of a rule library for the usage specification of the risk function.

[0100] 702. Establish a rule library for the usage specification of the risk function according to the second risk function;

[0101] After obtaining the second risk function, an aggregation algorithm can be used to perform correlation analysis on the second risk function, thereby constructing a rule library for the usage specification of the risk function.

[0102] Specifically, taking the obtained second risk function as the center, upstream and downstream related functions can be called to obtain the functions associated with the second risk function and the call paths. Additionally, the constraint conditions of the second risk function and the functions associated with the second risk function can be identified. These constraint conditions are the limiting conditions when using the risk function. Store the aforementioned second risk function, the functions associated with the second risk function, the call paths, and the constraint conditions, and then construct a rule library for the usage specification of the risk function.

[0103] 703. Classify the risk levels of the uncovered code according to the rule library.

[0104] For the uncovered code, comparative analysis can be performed according to the rule library for the usage specification of the risk function to obtain the risk level of the uncovered code.

[0105] Specifically, the risk level of the uncovered code can be classified by determining whether the constraint conditions of the uncovered code meet the requirements of the rule library. If the constraint conditions of the uncovered code meet the requirements of the rule library, the uncovered code can be marked as low risk; otherwise, the uncovered code can be marked as high risk.

[0106] In addition, if there is no corresponding function in the rule library for the function in the uncovered code, the function is marked and analyzed. After obtaining the risk classification result of the function, the risk classification and constraint conditions of the function are stored in the rule library to complete the update of the rule library.

[0107] Therefore, in an embodiment of the present application, the risk of uncovered code can be automatically graded, so that the code path for key analysis can be identified according to the marked risk level, and the reachable code path and reachable data set can be automatically analyzed according to the risk level to assist testers in troubleshooting and improve code analysis efficiency.

[0108] The above is an introduction to the method flow provided by the present application. Based on the above method flow, the following is an introduction to the device provided by the present application.

[0109] See also Figure 8 , a schematic diagram of the structure of a white box testing device provided by the present application, comprising:

[0110] An identification module 801 is used to identify a first risk function included in the tested code;

[0111] A screening module 802 is used to screen out the first risk function existing in the code path from the first risk function by using a code model, wherein the code model includes the association relationship between the upstream and downstream in the tested code, and the code path is the execution path of the tested code;

[0112] The analysis module 803 is used to analyze the first risk function existing in the code path through the evaluation model to obtain the risk evaluation result of the tested code.

[0113] In a possible implementation, the analysis module 803 is specifically used to run the code under test to obtain one or more first associated item sets, the first associated item set includes a first risk function existing in the code path, and the first associated item set includes multiple risk functions that have an associated relationship; perform association analysis on the one or more first associated item sets and the code path to obtain one or more second associated item sets, the second associated item sets include frequent item sets and upstream and downstream functions of the frequent item sets, and the frequent item sets include first associated item sets with support higher than a preset value; analyze the one or more second associated item sets through an evaluation model to obtain a risk assessment result of the code under test.

[0114] In a possible implementation, the aforementioned code path may be obtained by analyzing the running code under test in a code instrumentation manner.

[0115] In a possible implementation, the analysis module 803 is specifically configured to obtain one or more first associated item sets according to the first risk functions existing in the code paths and the association relationships between the first risk functions existing in the code paths.

[0116] In a possible implementation, the analysis module 803 is specifically configured to perform association analysis on one or more first associated item sets using an aggregation algorithm to obtain frequent item sets; and to call upstream and downstream functions associated with the frequent item sets according to the code path to obtain one or more second associated item sets.

[0117] In a possible implementation, the analysis module 803 is specifically configured to calculate the support of one or more first associated item sets by using an aggregation algorithm; and take the first associated item sets whose support is higher than a preset value as frequent item sets.

[0118] In a possible implementation manner, if the code under test includes uncovered code, the uncovered code is a newly added code to be tested in the code under test, the device further includes:

[0119] The identification module 801 is further used to identify a second risk function included in the baseline code;

[0120] Establishing module 804, used to establish a rule base of risk function usage specifications according to the second risk function, the rule base is used to indicate the constraint conditions for the use of the second risk function and functions associated with the second risk function;

[0121] A classification module 805 is used to classify the risks of uncovered codes according to a rule base;

[0122] The updating module 806 is used to update the risk assessment result of the tested code according to the result of the risk classification to obtain an updated risk assessment result.

[0123] In one possible implementation, a module 804 is established, which is specifically used to use an aggregation algorithm to perform association analysis on the second risk function to obtain functions associated with the second risk function, calling paths of functions associated with the second risk function, and constraints; a rule base is established based on the second risk function, functions associated with the second risk function, and calling paths of functions associated with the second risk function, as well as constraints.

[0124] In a possible implementation, the grading module 805 is specifically used to determine whether the constraints of the uncovered code meet the requirements of the rule base; if so, the uncovered code is marked as low risk; if not, the uncovered code is marked as high risk.

[0125] Among them, the identification module, screening module, analysis module, establishment module, classification module and update module can be implemented by software, or can be implemented by hardware. Exemplarily, the implementation of the identification module is introduced below by taking the identification module as an example. Similarly, the implementation of the screening module, analysis module, establishment module, classification module and update module can refer to the implementation of the identification module.

[0126] As an example of a software functional unit, the identification module may include code running on a computing instance. Among them, the computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Furthermore, the above-mentioned computing instance may be one or more. For example, the identification module may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region (region) or in different regions. Furthermore, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including one data center or multiple data centers with close geographical locations. Among them, usually a region may include multiple AZs.

[0127] Similarly, multiple hosts / virtual machines / containers used to run the code can be distributed in the same virtual private cloud (VPC) or in multiple VPCs. Usually, a VPC is set up in a region. For cross-region communication between two VPCs in the same region and between VPCs in different regions, a communication gateway needs to be set up in each VPC to achieve interconnection between VPCs through the communication gateway.

[0128] As an example of a hardware functional unit, the identification module may include at least one computing device, such as a server, etc. Alternatively, the identification module may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL) or any combination thereof.

[0129] The multiple computing devices included in the identification module can be distributed in the same region or in different regions. The multiple computing devices included in the identification module can be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in the identification module can be distributed in the same VPC or in multiple VPCs. The multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.

[0130] It should be noted that, in other embodiments, the identification module, the screening module, the analysis module, the establishment module, the grading module and the update module can all be used to execute any step in the white-box testing method, and the steps that the screening module, the analysis module, the establishment module, the grading module and the update module are responsible for implementing can be specified as needed. The full functions of the white-box testing device can be realized by respectively implementing different steps in the white-box testing method through the screening module, the analysis module, the establishment module, the grading module and the update module.

[0131] The present application also provides a computing device 900. Fig. 9 As shown, the computing device 900 includes: a bus 902, a processor 904, a memory 906, and a communication interface 908. The processor 904, the memory 906, and the communication interface 908 communicate through the bus 902. The computing device 900 can be a server or a terminal device. It should be understood that the present application does not limit the number of processors and memories in the computing device 900.

[0132] The bus 902 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Fig. 9 The bus 904 is represented by only one line, but does not mean that there is only one bus or one type of bus. The bus 904 may include a path for transmitting information between various components of the computing device 900 (eg, the memory 906, the processor 904, and the communication interface 908).

[0133] The processor 904 may include any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0134] The memory 906 may include a volatile memory, such as a random access memory (RAM). The processor 904 may also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid state drive (SSD).

[0135] The memory 906 stores executable program codes, and the processor 904 executes the executable program codes to respectively implement the functions of the aforementioned identification module, screening module, analysis module, establishment module, classification module, and update module, thereby implementing the white box testing method. That is, the memory 906 stores instructions for executing the white box testing method.

[0136] The embodiment of the present application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smart phone.

[0137] like Fig.10 As shown, the computing device cluster includes at least one computing device 900. The memory 906 in one or more computing devices 900 in the computing device cluster may store the same instructions for executing the white box testing method.

[0138] In some possible implementations, the memory 906 of one or more computing devices 900 in the computing device cluster may also store partial instructions for executing the white-box testing method. In other words, the combination of one or more computing devices 900 may jointly execute instructions for executing the white-box testing method.

[0139] It should be noted that the memory 906 in different computing devices 900 in the computing device cluster can store different instructions, which are respectively used to execute part of the functions of the white box testing apparatus. That is, the instructions stored in the memory 906 in different computing devices 900 can implement the functions of one or more modules among the identification module, the screening module, the analysis module, the establishment module, the classification module and the update module.

[0140] In some possible implementations, one or more computing devices in the computing device cluster may be connected via a network, which may be a wide area network or a local area network. Fig.11 A possible implementation is shown. Fig.11 As shown, two computing devices 900A and 900B are connected via a network. Specifically, the network is connected via a communication interface in each computing device. In this type of possible implementation, the memory 906 in the computing device 900A stores instructions for executing the functions of the identification module. At the same time, the memory 906 in the computing device 900B stores instructions for executing the functions of the screening module, the analysis module, the establishment module, the classification module, and the update module.

[0141] It should be understood that Fig.11 The functions of the computing device 900A shown in FIG. 9A may also be completed by multiple computing devices 900. Similarly, the functions of the computing device 900B may also be completed by multiple computing devices 900.

[0142] The embodiment of the present application also provides a computer program product including instructions. The computer program product may be software or a program product including instructions that can be run on a computing device or stored in any available medium. When the computer program product is run on at least one computing device, the at least one computing device performs a white box testing method.

[0143] The embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state hard disk). The computer-readable storage medium includes instructions that instruct the computing device to perform the white box testing method.

[0144] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present invention.

Claims

1. A white box testing method, It is characterized in that include: Identifying a first risk function included in the tested code; Filtering out a first risk function existing in a code path from the first risk function using a code model, wherein the code model includes an association relationship between upstream and downstream in the tested code, and the code path is an execution path of the tested code; The first risk function existing in the code path is analyzed by using an evaluation model to obtain a risk evaluation result of the tested code.

2. The method according to claim 1, It is characterized in that The step of analyzing the first risk function existing in the code path by using the evaluation model to obtain the risk evaluation result of the tested code includes: Running the code under test to obtain one or more first associated item sets, wherein the first associated item set includes a first risk function existing in the code path, and the first associated item set includes a plurality of risk functions having an associated relationship; Performing association analysis on the one or more first associated item sets and the code path to obtain one or more second associated item sets, wherein the second associated item sets include frequent item sets and upstream and downstream functions of the frequent item sets, and the frequent item sets include the first associated item sets with a support higher than a preset value; The one or more second associated item sets are analyzed by using the assessment model to obtain a risk assessment result of the tested code.

3. The method according to claim 2, It is characterized in that The step of running the code under test to obtain one or more first associated item sets includes: The one or more first associated item sets are obtained according to the first risk functions existing in the code paths and the association relationships between the first risk functions existing in the code paths.

4. The method according to any one of claims 2 or 3, It is characterized in that The performing association analysis on the one or more first associated item sets and the code path to obtain one or more second associated item sets includes: Performing association analysis on the one or more first associated item sets using an aggregation algorithm to obtain the frequent item sets; According to the code path, upstream and downstream functions associated with the frequent item sets are called to obtain the one or more second associated item sets.

5. The method according to claim 4, It is characterized in that The adopting an aggregation algorithm to perform association analysis on the one or more first associated item sets to obtain the frequent item sets includes: Using the aggregation algorithm to calculate the support of the one or more first associated item sets; The first associated item set whose support is higher than the preset value is used as the frequent item set.

6. The method according to any one of claims 1 to 5, It is characterized in that If the tested code includes uncovered code, and the uncovered code is a newly added code to be tested in the tested code, the method further includes: Identify secondary risk functions included in the baseline code; According to the second risk function, establishing a rule base of risk function usage specifications, wherein the rule base is used to indicate constraint conditions for use of the second risk function and functions associated with the second risk function; According to the rule base, risk classification is performed on the uncovered code; The risk assessment result of the tested code is updated according to the result of the risk grading to obtain an updated risk assessment result.

7. The method according to claim 6, It is characterized in that The step of establishing a rule base for using a risk function according to the second risk function includes: Performing association analysis on the second risk function using the aggregation algorithm to obtain a function associated with the second risk function, a calling path of the function associated with the second risk function, and the constraint condition; The rule base is established according to the second risk function, the function associated with the second risk function, the calling path of the function associated with the second risk function, and the constraint condition.

8. The method according to any one of claims 6 or 7, It is characterized in that The risk classification of uncovered codes according to the rule base includes: Determining whether the constraint condition of the uncovered code meets the requirement of the rule base; If the requirements are met, the uncovered code is marked as low risk; If the requirements are not met, the uncovered code is marked as high risk.

9. A white box testing device, It is characterized in that include: An identification module, used for identifying a first risk function included in the tested code; A screening module, used to screen out the first risk function existing in the code path from the first risk function by using a code model, wherein the code model includes an association relationship between upstream and downstream in the tested code, and the code path is an execution path of the tested code; The analysis module is used to analyze the first risk function existing in the code path through an evaluation model to obtain a risk evaluation result of the tested code.

10. The device according to claim 9, wherein the analysis module is specifically used for: Running the code under test to obtain one or more first associated item sets, wherein the first associated item set includes a first risk function existing in the code path, and the first associated item set includes a plurality of risk functions having an associated relationship; Performing association analysis on the one or more first associated item sets and the code path to obtain one or more second associated item sets, wherein the second associated item sets include frequent item sets and upstream and downstream functions of the frequent item sets, and the frequent item sets include the first associated item sets with a support higher than a preset value; The one or more second associated item sets are analyzed by using the assessment model to obtain a risk assessment result of the tested code.

11. The device according to claim 10, It is characterized in that The analysis module is specifically used for: The one or more first associated item sets are obtained according to the first risk functions existing in the code paths and the association relationships between the first risk functions existing in the code paths.

12. The device according to any one of claims 10 or 11, It is characterized in that The analysis module is specifically used for: Performing association analysis on the one or more first associated item sets using an aggregation algorithm to obtain the frequent item sets; According to the code path, upstream and downstream functions associated with the frequent item sets are called to obtain the one or more second associated item sets.

13. The device according to claim 12, It is characterized in that The analysis module is specifically used for: Using the aggregation algorithm to calculate the support of the one or more first associated item sets; The first associated item set whose support is higher than the preset value is used as the frequent item set.

14. The device according to any one of claims 9 to 13, It is characterized in that If the code under test includes uncovered code, and the uncovered code is a newly added code to be tested in the code under test, the device further includes: The identification module is further used to identify a second risk function included in the baseline code; An establishing module, configured to establish a rule base for use specifications of a risk function according to the second risk function, wherein the rule base is used to indicate constraint conditions for use of the second risk function and functions associated with the second risk function; A classification module, used for classifying the risks of uncovered codes according to the rule base; An updating module is used to update the risk assessment result of the tested code according to the result of the risk classification to obtain an updated risk assessment result.

15. The device according to claim 14, It is characterized in that The establishment module is specifically used for: Performing association analysis on the second risk function using the aggregation algorithm to obtain a function associated with the second risk function, a calling path of the function associated with the second risk function, and the constraint condition; The rule base is established according to the second risk function, the function associated with the second risk function, the calling path of the function associated with the second risk function, and the constraint condition.

16. The device according to any one of claims 14 or 15, It is characterized in that The grading module is specifically used for: Determining whether the constraint condition of the uncovered code meets the requirement of the rule base; If the requirements are met, the uncovered code is marked as low risk; If the requirements are not met, the uncovered code is marked as high risk.

17. A computing device cluster, It is characterized in that comprising at least one computing device, each computing device comprising a processor and a memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the method according to any one of claims 1 to 8.

18. A computer program product comprising instructions, It is characterized in that When the instructions are executed by a computing device cluster, the computing device cluster is caused to perform the method according to any one of claims 1 to 8.

19. A computer-readable storage medium, It is characterized in that The method comprises computer program instructions, and when the computer program instructions are executed by a computing device cluster, the computing device cluster performs the method according to any one of claims 1 to 8.