A method and system for evaluating the completion of crowdsourcing test tasks
By using a segmented reliability model framework in crowdsourcing testing, combining testing labor costs and test report costs, estimating the total number of potential defects of the software and calculating the detection defect coverage, the problem of difficulty in evaluating the completion degree of crowdsourcing testing tasks in the existing technology is solved, and effective assessment of task completion degree and reasonable management of resources are achieved.
Patent Information
- Application Number
- CN202011619793.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-30
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2040-12-30
AI Technical Summary
It is difficult for prior art to effectively evaluate the completion of crowdsourcing testing tasks, especially when considering the impact of test reports and tester costs on defect discovery efficiency.
A piecewise reliability model framework based on testing labor costs and test report costs is proposed. Through correlation analysis and least squares maximum likelihood estimation methods, the total number of potential defects of software is estimated, and the detection defect coverage is calculated to evaluate the task completion degree.
Real-time evaluation of the completion degree of crowdsourcing test tasks is realized, helping the platform reasonably manage the task process, avoid wasting test resources, and improve software quality and platform performance.
Smart Images

Figure CN114691479B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a task completion evaluation method, in particular to a crowdsourcing test task completion evaluation method. Background Art
[0002] Crowdsourcing testing is an emerging testing method. Its working mode is usually that the crowdsourcing testing platform releases tasks to different test workers within a specified time, and integrates and evaluates the workers' test reports after the test is completed. The unique advantages of this method, such as personnel diversity and complementarity, can greatly improve the efficiency of defect discovery. However, compared with traditional testing methods, this distributed testing method lacks scientific and reasonable task division and task allocation, and its testing process is also invisible. Therefore, the crowdsourcing testing platform cannot effectively monitor and manage the testing process, and project managers can only plan the end of crowdsourcing testing tasks based on personal experience. However, these experience-based task completion evaluation strategies may lead to unsatisfactory software quality and may also cause waste of testing resources. In response to this problem, there is an existing technology that evaluates the completion of test tasks based on software reliability. First, the capture-recapture model is used to predict the defect scale in the software based on the overlap of defects in the test report, and then the reliability and completion of the software are measured by analyzing the percentage of the number of detected defects in the defect scale.
[0003] However, the above studies ignore the impact of test reports and tester costs on defect discovery efficiency. Generally speaking, the more active the personnel involved in the test task are, the more test reports the platform receives, and the greater the probability of defects being discovered. To address this problem, the Software Reliability Growth Model (SRGM) based on testing-effort (TE) was proposed and widely used. SRGM models reliability from the perspective of software failure, and uses mathematical methods based on differential equations (groups) to establish a quantitative function model between several random parameters in the software testing process, such as test time, cumulative number of defects detected, test workload, etc., to solve the cumulative number of defects detected function expression. Then the progress of the test task is evaluated by comparing the number of defects detected and the estimated total number of defects.
[0004] Although the existing TE-based SRGM can estimate the number of potential defects, its research is mainly focused on the traditional software testing model. Crowdsourcing testing is different from traditional testing methods. The random diversity of its participants and the huge number of test reports make the software products more reasonably and fully tested. This large amount of manpower and report cost investment may lead to low efficiency of testing. Therefore, this task completion evaluation method needs to be not only applicable to the traditional software testing model, but also needs to be improved in combination with the test cost data characteristics under the crowdsourcing testing model. Summary of the invention
[0005] The technical problem to be solved by the present invention is to overcome the defects of the prior art and provide a crowdsourcing test report automatic evaluation method and system.
[0006] In order to solve the above technical problems, the present invention provides a crowdsourcing test task completion evaluation method, which is characterized by:
[0007] Obtaining a test cost element data set according to a preset test process condition, wherein the test cost element data set includes a test labor cost data set and a test report cost data set, and the preset test process condition enables the task completion evaluation method to meet the crowdsourcing test behavior;
[0008] Correlation analysis between test labor cost dataset and test report cost dataset;
[0009] Obtain a defect data set of the test report results, select a corresponding modeling equation from a pre-built piecewise reliability model framework based on the degree of correlation between the test labor cost data set and the test report cost data set according to the result of the correlation analysis, select a workload function with the smallest fitting error with the actual workload and substitute it into the modeling equation, and estimate the parameters of the modeling equation based on the defect data set using the least squares and maximum likelihood estimation methods to obtain the total number of potential defects in the software;
[0010] The detection defect coverage is calculated based on the actual number of defects detected in the defect data set and the total number of potential defects in the software to evaluate the task completion.
[0011] Furthermore, the test process conditions include:
[0012] (1) During the test, the defect detection process follows a non-homogeneous Poisson process over time;
[0013] (2) All defects are independent and can be detected;
[0014] (3) All test times are refined, calendar time is converted into timestamp format, and the test start time is ensured to be zero;
[0015] (4) After each test report is submitted, the defects found will be immediately and perfectly fixed without introducing new defects;
[0016] (5) Repeated defects in the test report are accumulated only for the first time;
[0017] (6) The average number of defects detected within the time (t, t+Δt) is proportional to the average number of defects remaining in the software under the combined effect of the current defect detection rate b(t), the tester occupancy rate w1(t) at time t, and the test report consumption rate w2(t) at time t.
[0018] Furthermore, the correlation analysis is performed using a correlation coefficient formula and a heat map.
[0019] Furthermore, the segmented reliability model framework is:
[0020]
[0021] Where w1(t) is the tester occupancy rate at time t, w2(t) is the test report consumption rate at time t, a represents the total number of defects, λ and μ are the correlation harmonic coefficients, b(t) represents the defect detection rate, a represents the total number of defects in the software, and m(t) represents the expected mean function of the number of detected defects in the time interval [0, t];
[0022] The three correlation cases of the piecewise reliability model framework are expressed as:
[0023] When ρ(P, R) ≥ 0.7, P and R are highly correlated, indicating that the average number of defects detected in the time (t, t + Δt) is proportional to the average number of defects remaining in the software under a certain element workload consumption rate. P is the test labor cost data set, and R is the test report cost data set.
[0024] When 0.4≤ρ(P,R)<0.7, the correlation between P and R is weak, which means that the average number of defects detected in the time (t, t+Δt) is proportional to the average number of defects remaining in the software under the weighted sum of the two element consumption rates;
[0025] When ρ(P,R)≤0.4, P and R are basically independent, which means that the average number of defects detected within the time (t, t+Δt) is proportional to the average number of defects remaining in the software under the common consumption rate of the two elements.
[0026] Furthermore, the workload function includes:
[0027] Weibull-Type TEF, Logistic TEF and Log-Logistic TEF;
[0028] By using the above three workload functions to fit the personnel and report cost data and comparing the fitting results, the workload function with the smallest fitting error with the actual workload is determined.
[0029] A crowdsourcing test task completion evaluation system, comprising:
[0030] A first determination module is used to determine a test cost element data set according to a preset test process condition, wherein the test cost element data set includes a test labor cost data set and a test report cost data set, and the preset test process condition enables the task completion evaluation method to meet the crowdsourcing test behavior;
[0031] An analysis module is used to analyze the correlation between the test labor cost data set and the test report cost data set;
[0032] The second determination module is used to obtain the defect data set of the test report results, select the corresponding modeling equation from the pre-built piecewise reliability model framework based on the correlation degree of the test labor cost data set and the test report cost data set according to the result of the correlation analysis, select the workload function with the smallest fitting error with the actual workload and substitute it into the modeling equation, and estimate the parameters of the modeling equation based on the defect data set using the least squares and maximum likelihood estimation methods to determine the total number of potential defects in the software;
[0033] The calculation module is used to calculate the detection defect coverage rate according to the actual number of defects detected in the defect data set and the total number of potential defects in the software, and evaluate the task completion.
[0034] Furthermore, the first determination module includes a condition determination module, which is used to determine the following test process conditions:
[0035] (1) During the test, the defect detection process follows a non-homogeneous Poisson process over time;
[0036] (2) All defects are independent and can be detected;
[0037] (3) All test times are refined, calendar time is converted into timestamp format, and the test start time is ensured to be zero;
[0038] (4) After each test report is submitted, the defects found will be immediately and perfectly fixed without introducing new defects;
[0039] (5) Repeated defects in the test report are accumulated only for the first time;
[0040] (6) The average number of defects detected within the time (t, t+Δt) is proportional to the average number of defects remaining in the software under the combined effect of the current defect detection rate b(t), the tester occupancy rate w1(t) at time t, and the test report consumption rate w2(t) at time t.
[0041] Furthermore, the analysis module uses a correlation coefficient formula and a heat map to analyze the correlation between the test labor cost data set and the test report cost data set.
[0042] Furthermore, the second determination module includes a model determination module, which is used to determine the following segmented reliability model framework:
[0043]
[0044] Where w1(t) is the tester occupancy rate at time t, w2(t) is the test report consumption rate at time t, a represents the total number of defects, λ and μ are the correlation harmonic coefficients, b(t) represents the defect detection rate, a represents the total number of defects in the software, and m(t) represents the expected mean function of the number of detected defects in the time interval [0, t];
[0045] The three correlation cases of the piecewise reliability model framework are expressed as:
[0046] When ρ(P, R) ≥ 0.7, P and R are highly correlated, indicating that the average number of defects detected in the time (t, t + Δt) is proportional to the average number of defects remaining in the software under a certain element workload consumption rate. P is the test labor cost data set, and R is the test report cost data set.
[0047] When 0.4≤ρ(P,R)<0.7, the correlation between P and R is weak, which means that the average number of defects detected in the time (t, t+Δt) is proportional to the average number of defects remaining in the software under the weighted sum of the two element consumption rates;
[0048] When ρ(P,R)≤0.4, P and R are basically independent, which means that the average number of defects detected within the time (t, t+Δt) is proportional to the average number of defects remaining in the software under the common consumption rate of the two elements.
[0049] Furthermore, the second determination module includes a workload function determination module, which is used to fit the personnel and report cost data by using Weibull-Type TEF, Logistic TEF and Log-Logistic TEF workload functions, and after comparing the fitting results, determine the workload function with the smallest fitting error with the actual workload.
[0050] The beneficial effects achieved by the present invention are:
[0051] Based on SRGM, the present invention proposes a crowdsourcing test task completion evaluation method, which can estimate the total number of potential software defects according to the impact of test reports and tester costs on defect discovery efficiency, and evaluate the test task completion by calculating the detection defect coverage. It can help the platform to monitor the task process in real time and make reasonable task closing decisions, thereby ensuring the quality of the test software while avoiding a large amount of test time and test resources. Waste, and improve the overall performance of the crowdsourcing test platform. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 It is the overall framework diagram of the implementation method of the present invention;
[0053] Figure 2 It is the correlation analysis result between tester cost and test report cost in task sample. DETAILED DESCRIPTION
[0054] The present invention will be further described below in conjunction with the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and cannot be used to limit the protection scope of the present invention.
[0055] like Figure 1 As shown, the input is the tester cost dataset P, the test report cost dataset R, and the defect dataset D of the test report results. The steps are as follows:
[0056] Step 1: Model the crowdsourcing testing process and pre-set the prerequisites so that the task completion evaluation method meets the characteristics of crowdsourcing testing behavior.
[0057] Since the crowdsourcing testing process only involves defect detection and does not include repair work, this method needs to set up a virtual repair process, which is a perfect repair. Secondly, the scale of crowdsourcing testing software is usually small and not suitable for system-level analysis, so the time needs to be refined. Third, considering that the test report has the characteristics of high repetition and high complementarity, the platform needs to merge and remove duplicate reports. Therefore, the task completion evaluation method based on SRGM needs to meet the following conditions:
[0058] (1) During the test, the defect detection process follows the NHPP process (Non-homogeneous Poisson Process) over time;
[0059] (2) All defects are independent and equally detectable;
[0060] (3) All test times are refined, calendar time is converted into timestamp format, and the test start time is ensured to be zero;
[0061] (4) After each test report is submitted, a virtual repair will be performed, that is, the defects found will be immediately and perfectly repaired without introducing new defects, so the total number of defects in the software is a constant a;
[0062] (5) Repeated defects in the test report are accumulated only for the first time. After the defect is virtually repaired, if it is still reported in the test report, this defect will be ignored;
[0063] (6) The average number of defects detected within the time (t, t+Δt) is proportional to the average number of defects remaining in the software under the combined consumption rate of the current defect detection rate b(t) and the two test cost workloads w1(t) and w2(t).
[0064] Step 2: Model the crowdsourcing testing process based on the assumptions in step (1) and propose a piecewise reliability modeling framework that can include two testing cost elements.
[0065] TE-based SRGM usually only considers the case of a single test cost element, and its general software reliability modeling framework is as follows:
[0066]
[0067] Where m(t) represents the expected mean function of the number of detected defects in the time interval [0, t]; b(t) represents the defect detection rate at that time, and a represents the total number of defects in the software. w(t) represents the test workload consumption rate function, which is the derivative of the test workload function (denoted as W(t)) with respect to the test time t.
[0068] However, considering the huge number of crowdsourcing testing labor costs and test report costs, which have a great impact on defect discovery efficiency, this paper takes both the test manpower and the number of submitted test reports into account in the test workload. Formula (1) cannot directly substitute two test cost elements. Therefore, based on formula (1), the model introduces two test cost elements to calculate the average number of defects finally detected. Taking into account the different correlations between the test cost elements, the model establishes the following segmented reliability model framework:
[0069]
[0070] Where w1(t) is the tester occupancy rate at time t, w2(t) is the test report consumption rate at time t, a represents the total number of defects, λ and μ are the correlation harmonic coefficients, and b(t) represents the defect detection rate.
[0071] The three correlation situations of the two test cost elements are described as follows:
[0072] (1) When ρ(P, R) ≥ 0.7, P and R are highly correlated. The influence of the two elements can be replaced by any one of them, that is, the average number of defects detected in the time (t, t + Δt) is proportional to the average number of defects remaining in the software under the workload consumption rate of a certain element;
[0073] (2) When 0.4≤ρ(P,R)<0.7, the correlation between P and R is weak. At this time, it is determined that P and R affect the number of defect detections in a certain proportion, and the average number of defects detected in the time (t, t+Δt) is proportional to the average number of defects remaining in the software under the weighted sum of the two element consumption rates;
[0074] (3) When ρ(P, R)≤0.4, P and R are basically unrelated. At this time, it is assumed that P and R are completely independent of each other and jointly affect the number of defect detections. The average number of defects detected in the time (t, t+Δt) is proportional to the average number of defects remaining in the software under the common consumption rate of the two elements.
[0075] During the entire testing process, the overall defect detection rate will eventually increase and tend to be constant as the number of defects is detected and eliminated. Therefore, this paper uses the existing b(t) formula:
[0076]
[0077] Among them, r is called the bending factor, which represents the proportion of irrelevant defects in the software. When r = 1, the obtained m(t) is a concave model, and the other cases are S-shaped models.
[0078] Step 3: Input the tester cost dataset P and the test report cost dataset R and perform correlation analysis on them.
[0079] In the actual test environment, the test cost elements are not completely independent of each other, nor are they completely consistent. They are more or less related. Therefore, the model uses the correlation coefficient to represent the correlation between two elements. The correlation coefficient is a statistical indicator correlation first designed by statistician Karl Pearson. Its calculation formula is as follows:
[0080]
[0081] Among them, P is the participating test cost dataset P, R is the submitted test report cost dataset R, Var represents variance, and E represents expected value.
[0082] This paper uses the above correlation coefficient calculation formula to randomly sample the cost data of the existing data set for correlation analysis and draws the analysis results into a heat map as shown in the figure. Figure 2As shown, the correlation between personnel cost and reporting cost for this data set is 0.65, which indicates that there is a certain positive correlation between the two test cost elements, but the correlation is weak.
[0083] Step 4: Input the defect data set of the test report results. According to the correlation results in step (3), select the corresponding modeling equation from the piecewise reliability model framework in step (2), and substitute the appropriate workload function to solve the equation. Then, use the least squares and maximum likelihood estimation methods to estimate the parameters of the equation based on the defect data set to obtain the total number of potential defects in the software.
[0084] According to the results of the correlation analysis of cost elements from multiple samplings, the correlation coefficients are generally greater than 0.4, so the relationship between two cost elements is mostly extremely strong or weakly correlated. Therefore, we usually choose the first or second differential equation in formula (2) to describe the test process.
[0085] Secondly, for the workload function fitting the number of tester costs and the number of test report costs, we selected the following three commonly used TEFs:
[0086] (1) Weibull-Type TEF: Throughout the testing phase, the TE per unit time is not constant. In fact, the instantaneous TE will eventually decrease during the testing life cycle, so the cumulative TE approaches a finite limit. This analysis is reasonable because no software company will spend unlimited resources on software testing. Therefore, Yamada et al. proposed the Weibull distribution, whose cumulative TE consumed in (0, t] is as follows:
[0087]
[0088] The instantaneous TE consumed at the test time t is:
[0089]
[0090] Where W is the total amount of TE expenditure, β is the scale parameter, and δ and θ are shape parameters.
[0091] There are several special cases of Weibull-Type curves: when θ=1, δ=1, the cumulative TE is an exponential curve; when θ=1, δ=2, the cumulative TE is a Rayleigh curve; when θ=1, the cumulative TE is a Weibull curve.
[0092] (2) Logistic TEF: When δ>3, the Weibull-Type curve has an obvious peak phenomenon. This phenomenon does not conform to the actual software development / testing process. Therefore, Huang et al. proposed to use the logistic TE function to describe the test working mode. The cumulative TE consumed in (0, t] is as follows:
[0093]
[0094] The instantaneous TE consumed at the test time t is:
[0095]
[0096] (3) Log-Logistic TEF: In some defect data sets, the rate of unit TE loss may also increase or decrease as the test progresses, and the existing model cannot capture this trend. Therefore, Gokhale and Trivedi proposed the Log-Logistic TE function to describe this trend. The cumulative TE consumed in (0, t] is:
[0097]
[0098] The instantaneous TE consumed at the test time t is:
[0099]
[0100] Step 5: Calculate the detection defect coverage rate based on the number of detected defects and the estimated total number of potential software defects obtained from the test report to evaluate the task completion. The detection defect coverage rate formula is as follows:
[0101] Detection defect coverage = number of defects actually detected / total estimated potential software defects.
[0102] The implementation effect of the present invention is verified as follows. This embodiment selects four sets of real defect data sets DS1 to DS4 from the MoocTest crowdsourcing test platform. These four sets of crowdsourcing test data sets cover different numbers of defects and testers, and the number of participating testers is greater than 200, and the number of reports is greater than 1000. Its cost data far exceeds that of traditional testing. After these projects have been crowdsourced tested, the statistics and integration of test reports, test manpower and defect information are carried out according to their complementary characteristics to form a defect data set. The specific content is shown in Table 1.
[0103] Table 1 Statistics of crowdsourcing test defect dataset
[0104]
[0105] This embodiment selects the baseline method - TE-based SRGM as the control group for comparison, where the TEF used in the test cost workload is the optimal cost function curve selected based on the comparison of fitting effects after implementation. The parameter estimation of TEF only uses LSE, while the parameter estimation of the cumulative detection defect number function expression m(t) uses MLE and LSE.
[0106] In the performance evaluation phase, Accuracy of Estimation (AE) is used to evaluate the accuracy and effectiveness of the method. It is calculated by the initial estimated number of defects in the software and the final cumulative number of defects actually detected. The error between the estimated number of potential software defects and the detection result reflects the accuracy of the estimated value of the number of software defects.
[0107]
[0108] Among them, M a is the cumulative number of defects actually detected after testing, and a is the estimated total number of potential software defects. The estimated value a can eventually be used to compare with the number of defects detected to measure the completion of the test task.
[0109] In step (4), when three TEFs are used to fit the two cost quantities of test personnel and test reports respectively, the optimal TEF for each cost quantity is obtained by estimating the function parameters using LSE. The specific parameters of the optimal TEF for the cost quantity of each data set are as follows.
[0110] Table 2 Optimal TEF parameters for cost data of each dataset
[0111]
[0112] It can be seen from the above table that in this embodiment, on the four crowdsourcing test data sets, all test manpower quantity fitting results are close to the Logistic TEF curve, and most of the test report quantity fitting results are close to the Weibull-Type TEF curve.
[0113] Then, according to the correlation degree of the cost elements of each data set, the corresponding differential equation is selected in the piecewise reliability model framework, and then the optimal TEF of the above experimental data sets is substituted into the differential equation for solution. For unknown parameters, the MLE and LSE methods are used to estimate the unknown parameters of m(t) and obtain the total number of potential defects in the software parameter a. The comparative experimental results of each data set are listed in Table 3 below.
[0114] Table 3 Parameter estimation results and performance comparison of each data set
[0115]
[0116] Among them, the equation selected for the DS3 data set is the same as that of the control group. We only need to focus on its performance results. Its AE values are 8.31% and 12.24% respectively, which can estimate the potential number of defects more accurately. In DS1 and DS4, the AE values of the estimated results in the MLE and LSE methods are smaller than those of the control group, indicating that the prediction accuracy of this embodiment is higher than that of the baseline method before improvement. In DS2, its AE value is similar to that of the control group, and the advantages of this embodiment are not obvious.
[0117] In general, the accuracy of the estimation results of the present invention is higher than that of the baseline method, and the AE values using the MLE method in all data sets are lower than those using the LSE method, indicating that the performance of the present invention using the MLE method is better.
[0118] From the experimental data, the prediction result of the total number of potential defects in the software in this embodiment is very accurate. Therefore, the accuracy of calculating the detection defect coverage based on this data will be higher than that of the baseline method, so the ability to evaluate the completion of the test task will be stronger. When monitoring the crowdsourcing test process, the method of the present invention can accurately predict the total number of potential defects in the software, effectively evaluate the task completion, help the platform to monitor the task process in real time, and make reasonable task closing decisions, thereby ensuring the quality of the test software while avoiding a large amount of test time and test resources. Waste, and improve the overall performance of the crowdsourcing test platform.
[0119] Correspondingly, the present invention also provides a crowdsourcing test task completion evaluation system, comprising:
[0120] A first determination module is used to determine a test cost element data set according to a preset test process condition, wherein the test cost element data set includes a test labor cost data set and a test report cost data set, and the preset test process condition enables the task completion evaluation method to meet the crowdsourcing test behavior;
[0121] An analysis module is used to analyze the correlation between the test labor cost data set and the test report cost data set;
[0122] The second determination module is used to obtain the defect data set of the test report results, select the corresponding modeling equation from the pre-built piecewise reliability model framework based on the correlation degree of the test labor cost data set and the test report cost data set according to the result of the correlation analysis, select the workload function with the smallest fitting error with the actual workload and substitute it into the modeling equation, and estimate the parameters of the modeling equation based on the defect data set using the least squares and maximum likelihood estimation methods to determine the total number of potential defects in the software;
[0123] The calculation module is used to calculate the detection defect coverage rate according to the actual number of defects detected in the defect data set and the total number of potential defects in the software, and evaluate the task completion.
[0124] The first determination module includes a condition determination module, which is used to determine the following test process conditions:
[0125] (1) During the test, the defect detection process follows a non-homogeneous Poisson process over time;
[0126] (2) All defects are independent and can be detected;
[0127] (3) All test times are refined, calendar time is converted into timestamp format, and the test start time is ensured to be zero;
[0128] (4) After each test report is submitted, the defects found will be immediately and perfectly fixed without introducing new defects;
[0129] (5) Repeated defects in the test report are accumulated only for the first time;
[0130] (6) The average number of defects detected within the time (t, t+Δt) is proportional to the average number of defects remaining in the software under the combined effect of the current defect detection rate b(t), the tester occupancy rate w1(t) at time t, and the test report consumption rate w2(t) at time t.
[0131] The analysis module uses a correlation coefficient formula and a heat map to analyze the correlation between the test labor cost data set and the test report cost data set.
[0132] The second determination module includes a model determination module, which is used to determine the following segmented reliability model framework:
[0133]
[0134] Among them, w1(t) is the test personnel occupancy rate at time t, w2(t) is the test report consumption rate at time t, a represents the total number of defects, λ and μ are the correlation harmonic coefficients, b(t) represents the defect detection rate, a represents the total number of defects in the software, and m(t) represents the expected mean function of the number of detected defects in the time interval [0, t];
[0135] The three correlation cases of the piecewise reliability model framework are expressed as:
[0136] When ρ(P, R) ≥ 0.7, P and R are highly correlated, indicating that the average number of defects detected in the time (t, t + Δt) is proportional to the average number of defects remaining in the software under a certain element workload consumption rate. P is the test labor cost data set, and R is the test report cost data set.
[0137] When 0.4≤ρ(P,R)<0.7, the correlation between P and R is weak, which means that the average number of defects detected in the time (t, t+Δt) is proportional to the average number of defects remaining in the software under the weighted sum of the two element consumption rates;
[0138] When ρ(P,R)≤0.4, P and R are basically independent, which means that the average number of defects detected within the time (t, t+Δt) is proportional to the average number of defects remaining in the software under the common consumption rate of the two elements.
[0139] The second determination module includes a workload function determination module, which is used to fit the personnel and report cost data by using Weibull-Type TEF, Logistic TEF and Log-Logistic TEF workload functions, and after comparing the fitting effects, determine the workload function with the smallest fitting error with the actual workload.
[0140] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.
[0141] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0142] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0143] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0144] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A crowdsourcing test task completion evaluation method, characterized in that: Obtaining a test cost element data set according to a preset test process condition, wherein the test cost element data set includes a test labor cost data set and a test report cost data set, and the preset test process condition enables the task completion evaluation method to meet the crowdsourcing test behavior; Correlation analysis between test labor cost dataset and test report cost dataset; Obtain a defect data set of the test report results, select a corresponding modeling equation from a pre-built piecewise reliability model framework based on the degree of correlation between the test labor cost data set and the test report cost data set according to the result of the correlation analysis, select a workload function with the smallest fitting error with the actual workload and substitute it into the modeling equation, and estimate the parameters of the modeling equation based on the defect data set using the least squares and maximum likelihood estimation methods to obtain the total number of potential defects in the software; The detection defect coverage is calculated based on the number of defects actually detected in the defect data set and the total number of potential defects in the software to evaluate the task completion; The test process conditions include: (1) During the test, the defect detection process follows a non-homogeneous Poisson process over time; (2) All defects are independent and can be detected; (3) All test times are refined, calendar time is converted into timestamp format, and the test start time is ensured to be zero; (4) After each test report is submitted, the defects found will be immediately and perfectly fixed, and no new defects will be introduced; (5) Repeated defects in the test report only accumulate the first discovery; (6) In ( t , t +Δ t ) The average number of defects detected within a certain period of time and the current defect detection rate b ( t )and t Test personnel occupancy rate at all times w 1( t )and t Time test report consumption rate w 2( t ) is proportional to the average number of defects remaining in the software at the combined consumption rate of The piecewise reliability model framework is: (2); in, w 1( t )for t Test personnel occupancy rate at all times, w 2( t )for t Test report consumption rate at all times, a represents the total number of defects, λ and μ is the correlation harmonic coefficient, b ( t ) represents the defect detection rate, m ( t ) represents the time interval [0, t ] is the expected mean function of the number of detected defects within ; The three correlation cases of the piecewise reliability model framework are expressed as: when ρ When (P, R) ≥ 0.7, P and R are highly correlated, indicating that ( t , t +Δ t ) The average number of defects detected in the time is proportional to the average number of defects remaining in the software under a certain element workload consumption rate, P is the test labor cost data set, and R is the test report cost data set; When 0.4≤ ρ When (P, R) < 0.7, the correlation between P and R is weak, indicating that ( t , t +Δ t ) The average number of defects detected in the time is proportional to the average number of defects remaining in the software under the weighted sum of the two element consumption rates; when ρ When (P,R) ≤0.4, P and R are basically irrelevant, indicating that ( t , t +Δ t ) time is proportional to the average number of defects remaining in the software at the combined consumption rate of the two elements.
2. The crowdsourcing test task completion evaluation method according to claim 1, characterized in that: The correlation analysis is performed using a correlation coefficient formula and a heat map.
3. The crowdsourcing test task completion evaluation method according to claim 1, characterized in that: The workload function includes: Weibull-Type TEF, Logistic TEF and Log-Logistic TEF; By using the above three workload functions to fit the personnel and report cost data and comparing the fitting results, the workload function with the smallest fitting error with the actual workload is determined.
4. A crowdsourcing test task completion evaluation system, characterized in that: include: A first determination module is used to determine a test cost element data set according to a preset test process condition, wherein the test cost element data set includes a test labor cost data set and a test report cost data set, and the preset test process condition enables the task completion evaluation method to meet the crowdsourcing test behavior; An analysis module is used to analyze the correlation between the test labor cost data set and the test report cost data set; The second determination module is used to obtain the defect data set of the test report results, select the corresponding modeling equation from the pre-built piecewise reliability model framework based on the correlation degree of the test labor cost data set and the test report cost data set according to the result of the correlation analysis, select the workload function with the smallest fitting error with the actual workload and substitute it into the modeling equation, and estimate the parameters of the modeling equation based on the defect data set using the least squares and maximum likelihood estimation methods to determine the total number of potential defects in the software; A calculation module, used to calculate the detection defect coverage rate according to the number of defects actually detected in the defect data set and the total number of potential defects in the software, and evaluate the task completion; The first determination module includes a condition determination module, which is used to determine the following test process conditions: (1) During the test, the defect detection process follows a non-homogeneous Poisson process over time; (2) All defects are independent and can be detected; (3) All test times are refined, calendar time is converted into timestamp format, and the test start time is ensured to be zero; (4) After each test report is submitted, the defects found will be immediately and perfectly fixed, and no new defects will be introduced; (5) Repeated defects in the test report only accumulate the first discovery; (6) In ( t , t +Δ t ) The average number of defects detected within a certain period of time and the current defect detection rate b ( t )and t Test personnel occupancy rate at all times w 1( t )and t Time test report consumption rate w 2( t ) is proportional to the average number of defects remaining in the software at the combined consumption rate of The piecewise reliability model framework is: (2); in, w 1( t )for t Test personnel occupancy rate at all times, w 2( t )for t Test report consumption rate at all times, a represents the total number of defects, λ and μ is the correlation harmonic coefficient, b ( t ) represents the defect detection rate, m ( t ) represents the time interval [0, t ] is the expected mean function of the number of detected defects within ; The three correlation cases of the piecewise reliability model framework are expressed as: when ρ When (P, R) ≥ 0.7, P and R are highly correlated, indicating that ( t , t +Δ t ) The average number of defects detected in the time is proportional to the average number of defects remaining in the software under a certain element workload consumption rate, P is the test labor cost data set, and R is the test report cost data set; When 0.4≤ ρ When (P, R) < 0.7, the correlation between P and R is weak, indicating that ( t , t +Δ t ) The average number of defects detected in the time is proportional to the average number of defects remaining in the software under the weighted sum of the two element consumption rates; when ρ When (P,R) ≤0.4, P and R are basically irrelevant, indicating that ( t , t +Δ t ) time is proportional to the average number of defects remaining in the software at the combined consumption rate of the two elements.
5. The crowdsourcing test task completion evaluation system according to claim 4, characterized in that: The analysis module uses a correlation coefficient formula and a heat map to analyze the correlation between the test labor cost data set and the test report cost data set.
6. The crowdsourcing test task completion evaluation system according to claim 4, characterized in that: The second determination module includes a workload function determination module, which is used to fit the personnel and report cost data by using Weibull-Type TEF, Logistic TEF and Log-Logistic TEF workload functions, and after comparing the fitting effects, determine the workload function with the smallest fitting error with the actual workload.
Citation Information
Patent Citations
Software reliability assessment method and device based on hybrid testing
CN102063375A
Mobile application crowdsourcing test report automatic evaluation method and computer storage medium
CN110928764A