A software maintenance scale prediction method based on multi-layer iteration

By adopting multi-layer iterative prediction model and normalized data preprocessing technology in software development, the problem of inaccurate prediction of software maintenance scale in the existing technology is solved, and higher prediction accuracy and better data processing capabilities are achieved.

CN114154730BActive Publication Date: 2025-05-06HANGZHOU DIANZI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111491975.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-08
Publication Date
2025-05-06
Estimated Expiration
2041-12-08

AI Technical Summary

Technical Problem

The prior art is difficult to accurately predict the scale of software maintenance in the early stages of software development, and lacks effective processing methods for unbalanced data, resulting in poor prediction results.

Method used

A multi-layer iteration-based prediction model is adopted, combined with normalized data preprocessing technology, a multi-layer iteration prediction model is constructed through the GMDH method, and 10-fold cross-validation and normalization processing are used to improve prediction accuracy.

Benefits of technology

It significantly improves the prediction accuracy of the software maintenance scale, reduces the MMRE value, increases the Pred(25) value, and can more accurately predict the software maintenance scale.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114154730B_ABST
    Figure CN114154730B_ABST
Patent Text Reader

Abstract

The present invention discloses a software maintenance scale prediction method based on multi-layer iteration, which firstly performs data acquisition: obtain all available classes in an existing public software system and record them as training data sets, then select measurement factors: obtain the original measurement values ​​of each class in the training data set; sample and measure the normalized training data set, and then perform training data sampling: divide the normalized training data set into a validation set and a training set; then construct and train a multi-layer iterative prediction model: finally, the prediction of the software maintenance scale is completed through the trained multi-layer iterative prediction model. The present invention continuously eliminates the intermediate layer models with poor prediction effects by designing a multi-layer iterative model, and finally obtains a relatively accurate prediction model of the software maintenance scale. The method proposed by the present invention is used to predict the software maintenance scale, and finally obtains a better result, which can accurately predict the software maintenance scale.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of software engineering, and in particular relates to a method for predicting the maintenance scale of a software system based on multi-layer iteration. Background Art

[0002] Software maintenance is an integral phase in any software development life cycle, which begins once the software is delivered to the customer. Software maintenance is the most expensive phase of software development, sometimes consuming nearly 60-70% of the total development cost, depending on the complexity of the system under consideration. Software maintenance scale refers to the scale of code modifications during the software maintenance process, such as correcting errors, removing obsolete code, adding new code, etc., and is one of the most important attributes for evaluating software quality. Software quality can be improved by predicting the maintenance scale in the early stages of development. However, predicting software maintenance scale in the early stages is very difficult, and research in this area still has not yielded good results due to the complexity of software system maintenance behavior.

[0003] In recent years, researchers in this field have studied and developed many software maintenance scale prediction models, including the use of traditional machine learning methods such as decision trees, Bayesian networks, SVM, etc., as well as various deep neural network algorithms such as recurrent neural networks, fuzzy neural networks, etc. Generally, 10-fold cross validation is used to improve the accuracy of the model, and MRE, MMRE, Pred (25) (the probability that the MRE value is less than 25%), etc. are used to evaluate the accuracy of the model. These methods do not produce good results in predicting the maintenance scale, including the acquisition and preprocessing of the data set are not ideal, and there is no good way to deal with unbalanced data. Summary of the invention

[0004] In order to overcome the above-mentioned deficiencies of the prior art, the present invention proposes a software maintenance scale prediction method based on multi-layer iteration, which, combined with normalized data preprocessing technology, can greatly improve the model's software maintenance scale prediction ability.

[0005] The technical solutions specifically adopted in the present invention are as follows:

[0006] A software maintenance scale prediction method based on multi-layer iteration includes the following steps:

[0007] S1. Data acquisition: obtain all available classes in an existing public software system, denoted as training data set TR;

[0008] S2. Metric factor selection: Get the original metric values ​​of each category in the training data set TR,

[0009] S3. Metric normalization processing:

[0010] S4. Training data sampling: The normalized training data set LM′ train The data is randomly divided into 10 parts, and one of them is taken as the validation set The remaining 9 sets are used as training sets

[0011] S5. Build and train a multi-layer iterative prediction model:

[0012] Construct a multi-layer iterative prediction model based on the GMDH method. In the process of training the prediction model, the output of each layer is selected as the input of the next layer, and the training is stopped according to the termination criterion to obtain the final trained prediction model.

[0013] S6. Complete the prediction of software maintenance scale through the trained multi-layer iterative prediction model.

[0014] Furthermore, the metrics include: WMC, NOC, LCOM, DIT, RFC, MPC, NOM, DAC, SIZE1, SIZE2, and the number of code modification lines CHANGE (including the number of added, deleted and changed lines) between different versions is selected as the indicator of software maintenance scale, where the meaning of each metric is as follows:

[0015] WMC: the sum of the cyclomatic complexity of all methods in the class;

[0016] NOC: the number of direct subclasses of the class;

[0017] LCOM: the number of disjoint method sets in a class;

[0018] DIT: The position of the class in the inheritance hierarchy;

[0019] RFC: The sum of the number of class methods and the number of class call methods;

[0020] MPC: the number of times methods defined in a class are called by other classes;

[0021] NOM: the number of methods in the class;

[0022] DAC: the number of abstract data types defined in the class;

[0023] SIZE1: The number of lines of pure code in the class;

[0024] SIZE2: The sum of the number of attributes and methods in the class;

[0025] CHANGE: The number of lines of code changed between two versions of the class, including the number of new, deleted, and modified lines of code;

[0026] Let LM be the class metric set of the training data set TR. train={M1, M2, ..., M l , .., M n}, n represents the number of classes, M l =<WMC, NOC, LCOM, DIT, RFC, MPC, NOM, DAC, SIZE1, SIZE2, CHANGE>;

[0027] Furthermore, the S3 specific method is as follows:

[0028] Class metric set LM for training dataset TR train Perform normalization to obtain the normalized training data set LM′ train =[M1′, M2′, ..., M l ′, ..., M n '],in Without loss of generality, remember is the measure x j The normalized value, the normalization formula is as follows:

[0029]

[0030] where x j represents the original measurement value, x max 、x min They are x j Corresponding LM train The maximum and minimum values ​​of this type of measurement;

[0031] Furthermore, the specific steps of S5 are as follows:

[0032] S5-1. First-layer prediction model training:

[0033] The normalized training set Every M l Elements in ' As an independent variable, As the actual value of the dependent variable software maintenance scale (i.e., CHANGE); the training method is: Combine the two elements of each combination and Substitute the following prediction formula and Substitute the actual value into the prediction formula y and calculate it using the widrow-hoff learning rule Group weights:

[0034]

[0035] Where y is the output value, indicating the software maintenance scale (i.e., CHANGE), and a0, a1, a2, a3, a4, and a5 are weights;

[0036] Then Substituting the group weights into the above prediction formula, we get candidate intermediate models;

[0037] S5-2. Tier 1 prediction model validation:

[0038] The validation set Input the candidate intermediate model obtained by S5-1, output the predicted value of the software maintenance scale (i.e., CHANGE), and calculate the average relative error value MAE of each candidate intermediate model by comparing it with the corresponding true value of the output software maintenance scale (i.e., CHANGE), obtain the MAE set E, and use the minimum MAE of the current layer as the criterion value W1 of the current layer. The calculation formula of the average relative error value MAE of each candidate intermediate model is as follows:

[0039]

[0040] Among them, m is the validation set The number of classes, y i′ The true value of the validation set software maintenance scale (i.e., CHANGE), is the corresponding model prediction value;

[0041] S5-3. Optimal selection of the first-layer prediction model: Select the candidate intermediate models corresponding to the 10 smallest MAE values ​​from the MAE set E as the optimal models and connect them to the next layer of the prediction model;

[0042] S5-4. Multi-layer iterative training prediction model: The training is continuously repeated according to the method of steps S5-1 to S5-3, that is, in the training of the kth (k=2, 3..., indicating the number of layers currently trained) layer prediction model, the optimal model output by the k-1th layer is taken as input, and the inputs are combined in pairs according to S5-1, and the weight of each combination is calculated, and the weight is substituted into the prediction formula to obtain the candidate intermediate model; according to step S5-2, the verification set is input into the candidate intermediate model to obtain the corresponding prediction value, and the MAE of each candidate intermediate model is calculated according to the formula of S5-2, and the minimum MAE value of the current layer is used as the criterion value W of the current layer. k ; Perform model optimization on the candidate intermediate models. If the MAE value of the candidate intermediate model in the k-th layer is less than W k-1 If the number of candidates is more than 10, the 10 candidate intermediate models with the smallest MAE values ​​are selected as the optimal models and connected to the next layer of the prediction model. Otherwise, the 10 candidate intermediate models with MAE values ​​less than W are selected as the optimal models. k-1 All candidate intermediate models are connected to the next layer of the prediction model as optimal models;

[0043] S5-5. Training termination criteria for prediction models:

[0044] When the minimum MAE value of the candidate intermediate model in the current layer is greater than the criterion value of the previous layer, the training process is stopped, and the model corresponding to the criterion value in the previous layer of optimal models is selected as the final layer of the prediction model;

[0045] When the current layer has only one candidate intermediate model and the MAE value of the candidate intermediate model is less than or equal to the criterion value of the previous layer, the training process is stopped and the current candidate intermediate model is selected as the final layer of the prediction model;

[0046] S5-6. Train to obtain the final multi-layer iterative prediction model;

[0047] Furthermore, the specific method of S6 is as follows:

[0048] For a new software system, obtain the class metric set LM according to step S2 test , perform normalization processing according to step S3 to obtain the normalized training data set LM′ test , and then the normalized training data set LM′ test The prediction model trained by S5 is input to obtain the predicted value of the software maintenance scale, and the original value of the predicted value is reversed according to the normalization formula of S3. The obtained original value is the predicted value of the number of code modification lines of the next version of each class included in the software system.

[0049] Beneficial effects of the present invention:

[0050] The present invention designs a multi-layer iterative model to continuously eliminate the intermediate layer models with poor prediction effects, and finally obtains a relatively accurate prediction model for the software maintenance scale. The method proposed by the present invention is used to predict the software maintenance scale, and finally obtains better results. Among them, for the imbalance problem of data between classes of the same attribute and between various attributes of a class, the present invention uses normalization to reduce the gap, and uses 10-fold cross validation for the training set. Finally, this method predicts the maintenance scale of the software more accurately, that is, the MMRE value of the prediction result is smaller and the Pred(25) value is larger. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 This is a flow chart of an embodiment of the present invention;

[0052] Figure 2 A specific example diagram of a network based on a multi-layer iterative software maintenance scale prediction model;

[0053] Figure 3 This is a comparison chart of the prediction and actual results of the present invention;

[0054] Figure 4 A comparison diagram of the present invention and other methods; DETAILED DESCRIPTION

[0055] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings.

[0056] like Figure 1 As shown, a software maintenance scale prediction method based on multi-layer iteration includes the following steps:

[0057] S1. Data acquisition: Use the CKJM-extended tool to obtain all available classes in an existing public JAVA software system, denoted as the training dataset TR;

[0058] S2. Metric factor selection: Use CKJM-extended tool to obtain the original metric values ​​in each category, including: WMC, NOC, LCOM, DIT, RFC, MPC, NOM, DAC, SIZE1,

[0059] SIZE2, and use the Beyond Compare tool to obtain the number of code modification lines CHANGE (including the number of added, deleted, and changed lines) between different versions as an indicator of the software maintenance scale, where the meanings of each metric are as follows:

[0060] WMC: The sum of the cyclomatic complexity of all methods in the class. The methods are plotted as control flow graphs. The cyclomatic complexity is the number of linearly independent paths in the graph. The calculation formula is:

[0061] V(G)=e-n+2p

[0062] Where e represents the number of edges in the control flow graph, n represents the number of nodes in the control flow graph, and p represents the number of connected graphs in the control flow graph. Since all control flow graphs are connected, p is 1.

[0063] NOC: the number of direct subclasses of the class;

[0064] LCOM: The number of disjoint method sets in a class. Any two methods in the same method set share at least one variable.

[0065] DIT: The position of a class in the inheritance hierarchy. Map all classes in the software into a tree structure of classes. The root node class has no parent class, and its DIT value is 1. The DIT value of the class at the nth level is n.

[0066] RFC: The sum of the number of class methods and the number of class call methods;

[0067] MPC: the number of times methods defined in a class are called by other classes;

[0068] NOM: the number of methods in the class;

[0069] DAC: the number of abstract data types defined in the class;

[0070] SIZE1: The number of pure code lines in the class. Since a line of code ends with a semicolon, the number of semicolons in non-comments is used to calculate the number of code lines.

[0071] SIZE2: The sum of the number of attributes and methods in the class;

[0072] CHANGE: The number of lines of code that have changed between two versions of a class, including the number of lines of code that have been added, deleted, and modified. A deletion or addition operation is counted as 1, and a modification operation is counted as 2 because it includes both deletion and addition operations.

[0073] Let LM be the class metric set of the training data set TR. train ={M1, M2, ..., M l , ..., M n}, n represents the number of classes, M l =<WMC, NOC, LCOM, DIT, RFC, MPC, NOM, DAC, SIZE1, SIZE2, CHANGE>;

[0074] S3. Metric normalization processing:

[0075] Class metric set LM for training dataset TR train Perform normalization to obtain the normalized training data set LM′ train =[M1′, M2′, ..., M l ′, ..., M n '],in Without loss of generality, remember is the measure x j The normalized value, the normalization formula is as follows:

[0076]

[0077] where x j represents the original measurement value, x max 、x min They are x j Corresponding LM train The maximum and minimum values ​​of this type of measurement;

[0078] S4. Training data sampling: Using the 10-fold cross-validation method, the training data set LM′ train The data is randomly divided into 10 parts, and one of them is taken as the validation set each time , and the remaining 9 are used as training sets

[0079] S5. Training prediction model: Construct a multi-layer iterative prediction model. In the process of training the prediction model, the output of each layer is selected as the input of the next layer, and the training is stopped according to the termination criteria to obtain the final trained prediction model. The detailed process includes the following sub-steps:

[0080] S5-1. First-layer prediction model training:

[0081] The normalized training set Every M l Elements in ' As an independent variable, As the actual value of the dependent variable software maintenance scale (i.e., CHANGE); the training method is: Combine the two elements of each combination and Substitute the following prediction formula and Substitute the actual value into the prediction formula y and calculate it using the widrow-hoff learning rule Group weights:

[0082]

[0083] Where y is the output value, indicating the software maintenance scale (i.e., CHANGE), and a0, a1, a2, a3, a4, and a5 are weights;

[0084] Then Substituting the group weights into the above prediction formula, we get candidate intermediate models;

[0085] S5-2. Tier 1 prediction model validation:

[0086] The validation set Input the candidate intermediate model obtained by S5-1, output the predicted value of the software maintenance scale (i.e., CHANGE), and calculate the average relative error value MAE of each candidate intermediate model by comparing it with the corresponding true value of the output software maintenance scale (i.e., CHANGE), obtain the MAE set E, and use the minimum MAE of the current layer as the criterion value W1 of the current layer. The calculation formula of the average relative error value MAE of each candidate intermediate model is as follows:

[0087]

[0088] Among them, m is the validation set The number of classes, y i′The true value of the validation set software maintenance scale (i.e., CHANGE), is the corresponding model prediction value;

[0089] S5-3. Optimal selection of the first-layer prediction model: Select the candidate intermediate models corresponding to the 10 smallest MAE values ​​from the MAE set E as the optimal models and connect them to the next layer of the prediction model;

[0090] S5-4. Multi-layer iterative training prediction model: The training is continuously repeated according to the method of steps S5-1 to S5-3, that is, in the training of the kth (k=2, 3..., indicating the number of layers currently trained) layer prediction model, the optimal model output by the k-1th layer is taken as input, and the inputs are combined in pairs according to S5-1, and the weight of each combination is calculated, and the weight is substituted into the prediction formula to obtain the candidate intermediate model; according to step S5-2, the verification set is input into the candidate intermediate model to obtain the corresponding prediction value, and the MAE of each candidate intermediate model is calculated according to the formula of S5-2, and the minimum MAE value of the current layer is used as the criterion value W of the current layer. k ; Perform model optimization on the candidate intermediate models. If the MAE value of the candidate intermediate model in the k-th layer is less than W k-1 If the number of candidates is more than 10, the 10 candidate intermediate models with the smallest MAE values ​​are selected as the optimal models and connected to the next layer of the prediction model. Otherwise, the 10 candidate intermediate models with MAE values ​​less than W are selected as the optimal models. k-1 All candidate intermediate models are connected to the next layer of the prediction model as optimal models;

[0091] S5-5. Training termination criteria for prediction models:

[0092] When the minimum MAE value of the candidate intermediate model in the current layer is greater than the criterion value of the previous layer, the training process is stopped, and the model corresponding to the criterion value in the previous layer of optimal models is selected as the final layer of the prediction model;

[0093] When the current layer has only one candidate intermediate model and the MAE value of the candidate intermediate model is less than or equal to the criterion value of the previous layer, the training process is stopped and the current candidate intermediate model is selected as the final layer of the prediction model;

[0094] S5-6. Train to obtain the final multi-layer iterative prediction model;

[0095] S6. For a new software system, obtain the class metric set LM according to step S2 test , perform normalization processing according to step S3 to obtain the normalized training data set LM′ test , and then the normalized training data set LM′ testThe prediction model trained by S5 is input to obtain the predicted value of the software maintenance scale, and the original value of the predicted value is reversed according to the normalization formula of S3. The obtained original value is the predicted value of the number of code modification lines of the next version of each class included in the software system.

[0096] Based on the above method flow, the technical effect is further demonstrated through examples.

[0097] Example

[0098] The steps of this embodiment are the same as those of the specific implementation method, and will not be repeated here. The following is a demonstration of some implementation processes and implementation results:

[0099] Data source acquisition: A total of 313 classes were obtained from a software system. For each class, the CKJM-extended tool was used to obtain the ten metrics of WMC, NOC, LCOM, DIT, RFC, MPC, NOM, DAC, SIZE1, and SIZE2. The Beyond Compare tool was used to collect the number of lines of code added, deleted, and modified between two versions of each class to obtain the training data set. These data were preprocessed according to steps S2-S4.

[0100] First-level prediction model training: The training data set is divided into a training set and a validation set in a ratio of 9:1 according to the 10-fold cross-validation method. Figure 2 The figure shows the complete process of training a multi-layer iterative software maintenance scale prediction model, that is, first use the training set data to train the first layer of the model, and then use the independent variable As input, As the actual software maintenance scale (i.e., CHANGE) value, substitute it into the formula of S51, and calculate the weight of each candidate intermediate model according to the weight calculation method of S5-1, and obtain 45 first-level intermediate models with different weights.

[0101] First-level prediction model verification and optimization: Input the verification set into the intermediate model obtained above The predicted value of software maintenance scale (i.e., CHANGE) is obtained, and the MAE value of each candidate intermediate model is calculated using the mean relative error formula according to its corresponding actual value. The ten candidate intermediate models with the smallest MAE are left as the optimal models and named w1~w 10 As the input of the second layer, the minimum MAE value is assigned to W1, which is the first layer criterion value.

[0102] Iterative training process: the second layer input w1~w 10 According to S5-4, the second layer candidate intermediate model is obtained by training , select the ten candidate intermediate models with the smallest MAE as the optimal models, named z1~z 10 As the input of the third layer, the model training and optimization process is repeated continuously. According to the S5-5 model termination criterion, since there is only one candidate intermediate model f in the end, and its MAE value is less than the criterion value of the previous layer, f is used as the last layer of the prediction model, and the training is terminated to obtain the final multi-layer iterative software maintenance scale prediction model.

[0103] Model usage evaluation: A new software system is predicted using a prediction model to obtain a predicted value of the software maintenance scale (i.e., CHANGE). Compared with the actual value y and the accuracy is as follows Figure 3 As shown. The indicators used are MRE, MMRE and Pred(25) values ​​to evaluate the prediction results. The formula is as follows:

[0104]

[0105]

[0106]

[0107] Where y is the actual value, is the predicted value; m is the number of classes in the test dataset; K is the number of observations with MRE less than or equal to 25%.

[0108] Where MRE represents relative error, MMRE represents the mean value of relative error, and Pred(25) represents the ratio of relative error ≤ 25%. Figure 3 It can be seen that the model has a good fit, with the result that MMRE reaches 15.23%, which is 13.2% lower than other methods, and Pred(25) reaches 95.13%, that is, most of the relative error MRE values ​​are below 25%. Figure 4 The comparison chart shows the prediction accuracy of the proposed model with the other four models: GRNN, SVM, Decision-Tree, and Bayesian. The proposed model shows higher performance when evaluated by the above three indicators: MRE, MMRE, and PRED (25).

Claims

1. A software maintenance scale prediction method based on multi-layer iteration, characterized in that: The steps include: S1. Data acquisition: obtain all available classes in an existing public software system, denoted as training data set TR; S2. Metric factor selection: Get the original metric values ​​of each category in the training data set TR, S3. Metric normalization processing: S4. Training data sampling: The normalized training data set LM' train The data is randomly divided into 10 parts, and one of them is taken as the validation set The remaining 9 sets are used as training sets S5. Build and train a multi-layer iterative prediction model: Construct a multi-layer iterative prediction model based on the GMDH method. In the process of training the prediction model, the output of each layer is selected as the input of the next layer, and the training is stopped according to the termination criterion to obtain the final trained prediction model. S6. Complete the prediction of software maintenance scale through the trained multi-layer iterative prediction model; The specific steps of S5 are as follows: S5-1. First-layer prediction model training: The normalized training set Every M l 'Medium element As an independent variable, Actual value of software maintenance scale as the dependent variable; The training method is: Combine the two elements of each combination and Substitute the following prediction formula and Substitute the actual value into the prediction formula y and calculate it using the widrow-hoff learning rule Group weights: Where y is the output value, indicating the scale of software maintenance, and a0, a1, a2, a3, a4, a5 are weights; Then Substituting the group weights into the above prediction formula, we get candidate intermediate models; S5-2. Tier 1 prediction model validation: The validation set Input the candidate intermediate model obtained by S5-1, output the predicted value of the software maintenance scale, and calculate the average relative error value MAE of each candidate intermediate model by comparing it with the corresponding true value of the output software maintenance scale, obtain the MAE set E, and use the minimum MAE of the current layer as the criterion value W1 of the current layer. The calculation formula of the average relative error value MAE of each candidate intermediate model is as follows: Among them, m is the validation set The number of classes, y i' Maintaining scale truth for validation set software, is the corresponding model prediction value; S5-3. Optimal selection of the first-layer prediction model: Select the candidate intermediate models corresponding to the 10 smallest MAE values ​​from the MAE set E as the optimal models and connect them to the next layer of the prediction model; S5-4. Multi-layer iterative training prediction model: Repeat the training according to the method of steps S5-1 to S5-3, that is, in the training of the k-th layer prediction model, k = 2, 3..., indicating the number of layers currently trained; take the optimal model output by the k-1th layer as input, combine the inputs in pairs according to S5-1, calculate the weight of each combination, substitute the weight into the prediction formula to obtain the candidate intermediate model; input the verification set into the candidate intermediate model according to step S5-2 to obtain the corresponding prediction value, calculate the MAE of each candidate intermediate model according to the formula of S5-2, and use the minimum MAE value of the current layer as the criterion value W of the current layer k ; Perform model optimization on the candidate intermediate models. If the MAE value of the candidate intermediate model in the k-th layer is less than W k-1 If the number of candidates is more than 10, the 10 candidate intermediate models with the smallest MAE values ​​are selected as the optimal models and connected to the next layer of the prediction model. Otherwise, the 10 candidate intermediate models with MAE values ​​less than W are selected as the optimal models. k-1 All candidate intermediate models are connected to the next layer of the prediction model as optimal models; S5-5. Training termination criteria for prediction models: When the minimum MAE value of the candidate intermediate model in the current layer is greater than the criterion value of the previous layer, the training process is stopped, and the model corresponding to the criterion value in the previous layer of optimal models is selected as the final layer of the prediction model; When the current layer has only one candidate intermediate model and the MAE value of the candidate intermediate model is less than or equal to the criterion value of the previous layer, the training process is stopped and the current candidate intermediate model is selected as the final layer of the prediction model; S5-6. Train to obtain the final multi-layer iterative prediction model.

2. The software maintenance scale prediction method based on multi-layer iteration according to claim 1 is characterized in that: Metrics include: WMC, NOC, LCOM, DIT, RFC, MPC, NOM, DAC, SIZE1, SIZE2, and select the number of code modification lines CHANGE between different versions as the indicator of software maintenance scale, where the meaning of each metric is as follows: WMC: the sum of the cyclomatic complexity of all methods in the class; NOC: the number of direct subclasses of the class; LCOM: the number of disjoint method sets in a class; DIT: The position of the class in the inheritance hierarchy; RFC: The sum of the number of class methods and the number of class call methods; MPC: the number of times the methods defined in a class are called by other classes; NOM: the number of methods in the class; DAC: the number of abstract data types defined in the class; SIZE1: The number of lines of pure code in the class; SIZE2: The sum of the number of attributes and methods in the class; CHANGE: The number of lines of code changed between two versions of the class, including the number of new, deleted, and modified lines of code; Let LM be the class metric set of the training data set TR. train ={M1,M2,…,M l ,…,M n }, n represents the number of classes, M l =<WMC,NOC,LCOM,DIT,RFC,MPC,NOM,DAC,SIZE1,SIZE2,CHANGE> .

3. The software maintenance scale prediction method based on multi-layer iteration according to claim 2 is characterized in that: The specific method of S3 is as follows: Class metric set LM for training dataset TR train Perform normalization to obtain the normalized training data set LM' train =[M1',M2',…,M l ',…,M n '],in Without loss of generality, remember is the measure x j The normalized value, the normalization formula is as follows: where x j represents the original measurement value, x max 、x min They are x j Corresponding LM train The maximum and minimum values ​​of this type of metric.

4. The software maintenance scale prediction method based on multi-layer iteration according to claim 3 is characterized in that: S6 specific methods are as follows: For a new software system, obtain the class metric set LM according to step S2 test , perform normalization processing according to step S3 to obtain the normalized training data set LM' test , and then the normalized training data set LM' test The prediction model trained by S5 is input to obtain the predicted value of the software maintenance scale, and the original value of the predicted value is reversed according to the normalization formula of S3. The obtained original value is the predicted value of the number of code modification lines of the next version of each class included in the software system.

Citation Information

Patent Citations

  • Method for evaluating and predicting maintenance work load of open source software (OSS) based on code quality

    CN104809066A

  • A method and a system for predicting defects of object-oriented software

    CN106991047A