Software Refactoring Prediction Method Based on Code Smells
By applying machine learning integration technology in software development, combining code structure and odor information, predicting the necessity of code reconstruction, the problem of difficult to predict the timing of software code reconstruction in the existing technology is solved, and the maintainability and scalability of the software are improved.
Patent Information
- Application Number
- CN202111468006.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-03
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2041-12-03
AI Technical Summary
The existing technology is difficult to effectively predict the timing of software code reconstruction, resulting in a long-term code odor, increasing the difficulty and cost of software maintenance and expansion.
Using a machine learning integration technology method, the structural information of the code file and the intensity and historical information of the odor are used to predict the necessity of code reconstruction through feature selection and model integration.
It significantly improves the accuracy and effectiveness of code reconstruction prediction, helps developers refactor them at the right time, and improves the maintainability and scalability of the software.
Smart Images

Figure CN114138328B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of software maintenance, and particularly to a software refactoring prediction method based on code smells. Background Art
[0002] In the process of software maintenance and evolution, software systems need to be continuously changed by developers in order to implement new requirements, enhance existing features, or fix important bugs. Due to time pressure or other relevant information, developers do not always have the time or willingness to control the complexity of the system and find good design solutions before applying their modifications. They may simply write code to implement functions, ignoring the structure and readability of the program. The code writing becomes increasingly chaotic, making the entire code structure bloated. At this time, code smells are often introduced.
[0003] Code smells indicate suboptimal design or implementation choices in the source code that are prone to causing changes and errors, which will result in a decline in code quality. At the same time, it will cause trouble for software developers in understanding and maintaining project code, and is an important factor leading to technical debt, thus generating unnecessary maintenance costs.
[0004] In the development and life cycle of software, maintenance is considered one of the most difficult and expensive tasks. Especially if there are too many smells in the software system, it often leads to extremely costly refactoring work.
[0005] Recent research shows that the phenomenon of code smells existing in existing software systems is very common. A large number of smell instances are often introduced when the files where they are located are created, and they tend to exist in the software system for a long time and survive, which often brings potential problems for subsequent maintenance and expansion of the software system.
[0006] Therefore, it is very important to choose a suitable time to perform refactoring to eliminate smells, which will improve the maintainability and extensibility of the software.
[0007] Previous research often focused on the detection of code smells or the priority of software refactoring, and there were few studies on predicting the timing of code refactoring. Summary of the Invention
[0008] From the perspective of code refactoring prediction, the present invention uses the structural information of code files (such as size, complexity, coupling degree, cohesion, etc.), the intensity of smells, and historical information (such as the duration of smells in the file, the number of changes, etc.) to provide a code refactoring prediction method based on machine learning integration technology, which can effectively solve the above problems. The specific technical solutions adopted by the present invention are as follows:
[0009] A software refactoring prediction method based on code smells, specifically including the following steps:
[0010] Step (1) Given a set F of m*n source code file versions in the system to be analyzed, F = (F 1,1 , F 1,2 , …, F i,j , …, F m,n ), where F i,j represents the j-th version of the source code file F i . Use a code parsing tool to parse each source code file, and represent the structure and smell information metric of each source code file version F i,j in the form of S i,j = <className, classVersion, structure, hasSmell>. Here, i = 1, 2, …, m, j = 1, 2, …, n. Among them, className represents the class name of the source code file version F i,j (assuming a source code file contains one class), classVersion represents the version number of the source code file version F i,j in the project history, structure represents the set of structural feature W of the source code file version F i,j = <w LOC , w NOA , w CBO , w MPC , w TCC , w McCabe , w WMC >, where w LOC represents the number of lines of code in this file, w NOA represents the number of attributes in this file, w CBO represents the number of target classes coupled with this file (coupling means that the methods in this file call the methods or variables of the target class), w MPC represents the number of methods in this file that call other methods, w TCC represents the number of methods directly related by accessing the same attribute, w McCabe represents the complexity calculated by the McCabe metric for this file, w WMC represents the sum of the cyclomatic complexities of the methods in this file, and hasSmell represents whether there is a certain code smell in the source code file version F i,j (1 means there is a smell, 0 means there is no smell);
[0011] Step (2) If the source code file version F i,j is identified as having a certain code smell, and the set of feature thresholds for determining whether it has this code smell is T = <w 1 |b 1,...,w g |b g , ..., w t |b t >, where w g is a feature in W, b g To identify the feature w of this code smell g The corresponding threshold, g = 1, 2, ..., t, is calculated by the following formula to obtain the odor intensity:
[0012]
[0013] Among them, m(w g ) represents the existence of the system to be analyzed due to the feature w g The maximum or minimum value that causes a certain code smell (when w g More than b g When some code smell occurs, the maximum value is selected. g Less than b g When it causes some code smell, choose the minimum value);
[0014] Source code file version F after adding strength information i,j The structure and smell information measurement of is expressed as:
[0015] S′ i,j =<className,classVersion,structure,hasSmell,intensity>
[0016] Step (3) Obtain source code file history information metrics:
[0017] Suppose source code file F i A code smell was introduced in a historical version p, the source code file F i The current version is j, then the source code file version is F i,j The historical information measurement can be expressed as:
[0018] H i,j =<className,classVersion,diffDays,diffVersions,action>
[0019] Where diffDays represents the number of natural days between versions p and j, diffVersions represents the number of versions between versions p and j, and action represents the number of file F between versions p and j. i The number of times the modification occurred;
[0020] In step (4), we find all the source code file versions with code smells eliminated in the source code file set F and classify them into the following two categories:
[0021] 1) For F i,j where hasSmell = 1, while for F i,j+1 hasSmell = 0;
[0022] 2) For F i,j where hasSmell = 1, and j is the last version of the source code file F i Considering that the file deletion may also be the result of excessive code smells, this situation will also be considered;
[0023] According to the strategies in 1) and 2), all the last source code file versions F i,j with a certain code smell in the source code file set F are obtained, forming the source code file set with code smells eliminated, and it is expanded through the oversampling technique SMOTE to form ζ P , which is the positive sample in the dataset. The remaining source code file versions in the source code file set F are used as the negative sample ζ N Finally, the complete sample dataset is obtained: ζ = ζ P ∪ζ N , where the reconstruction label value y corresponding to each source code file version in ζ P is 1, and the reconstruction label value y corresponding to each source code file version in ζ N is 0;
[0024] In step (5), the feature recursive elimination technique (RFE), Random Forest Classifier, and LGBMClassifier are used to retain the most important z features, denoted as W * =(w 1 , w 2 , …, w z ); The source code file version F i,j after feature selection can be represented by (S * i,j , H * i,j ), where S * i,j represents the retained structural and smell information metric, and H * i,j represents the retained historical information metric;
[0025] In step (6), a part of the data in the dataset is used as the training set ζ train ;
[0026] Step (7) For the information representation (S * i,j , H * i,j ) of each source code file version, input S * i,j into LGBM to obtain the output h 1 , and input H * i,j into Logistic Regression to obtain the output h 2 . The output of the final model represents the predicted probability of reconstruction where a 1 and a 2 represent weights, and a 1 + a 2 = 1;
[0027] Step (8) Use the Cross-entropy Loss Function to calculate the loss between the reconstructed label value y and the output . The definition of the loss function is as follows:
[0028]
[0029] where d represents the number of training samples;
[0030] Step (9) Use the training set ζ train to train the model parameters of LGBM and Logistic Regression until the maximum number of iterations MaximumIter is reached, and obtain the LGBM and Logistic Regression models with the best parameters after training;
[0031] Step (10) For a source code file, first obtain the file's code smell information metric and historical information metric according to steps (1), (2), and (3), then extract the best features according to step (5), input these features into the models obtained in step (9), and finally obtain the predicted probability of whether this file needs to be reconstructed. If the probability is greater than or equal to 0.5, it means reconstruction is required; if it is less than 0.5, it means reconstruction is not required.
[0032] The present invention has the following benefits:
[0033] 1. This method demonstrates the feasibility and effectiveness of machine learning techniques in the field of code refactoring prediction.
[0034] 2. According to the definition of code smells, it reasonably utilizes the multi-dimensional information contained in source code files. Relying on the strong complementarity between different information metrics, it significantly improves the effectiveness of the code refactoring prediction method.
[0035] 3. Compared with GBDT, LGBM has better accuracy and can process large-scale data in parallel, and can better utilize the structural features in source files. The good interpretability of the Logistic Regression model can help better observe the impact of historical metric information on the final result. With the help of ensemble techniques, we improved the prediction accuracy by using the LGBM + Logistic Regression approach. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 It is a flowchart of the software refactoring prediction method based on code smells of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0037] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be described in detail below with reference to the accompanying drawings.
[0038] For the convenience of description, the following relevant symbols are defined:
[0039] F: The set of source code file versions.
[0040] F i,j : The j-th version of the source code file F i
[0041] W: The set of structural features of the source code file version F i,j
[0042] T: The set of feature thresholds for detecting code smells.
[0043] S i,j : The structural and smell information metric representation of the source code file version F i,j
[0044] intensity: The intensity of smells in the source code file version F i,j
[0045] H i,j : The historical information metric representation of the source code file version F i,j
[0046] S * i,j : The structural and smell information metric retained after feature selection for the source code file version F i,j
[0047] H * i,j : Source code file version F i,j Historical information metrics retained after feature selection.
[0048] MaximumIter: Size of the number of iterations.
[0049] a 1 and a 2 : Learning weight.
[0050] The following combines the attached Figure 1 , and details the software refactoring prediction method based on code smells provided by this invention patent, including the following steps:[[]]
[0051] Step (1) Given a set F = (F 1,1 , F 1,2 , …, F i,j , …, F m,n ) of m*n source code file versions in the system to be analyzed, where F i,j represents the j-th version of the source code file F i . Use a code parsing tool to parse each source code file, and represent the structure and smell information metrics of each source code file version F i,j in the form of S i,j = <className, classVersion, structure, hasSmell>, i = 1, 2, …, m, j = 1, 2, …, n, where className represents the class name of the source code file version F i,j (assuming a source code file contains one class), classVersion represents the version number of the source code file version F i,j in the project history, structure represents the set of structural features W = <w i,j , w LOC , w NOA , w CBO , w MPC , w TCC , w McCabe , w WMC > of the source code file version F LOC where w NOA represents the number of lines of code in the file, w CBO represents the number of attributes in the file, w MPC represents the number of target classes coupled to the file (coupling means that the methods in the file call the methods or variables of the target class), w TCC represents the number of methods that call other methods in the file, w McCabeIndicates the complexity of the file calculated by the McCabe metric, w WMC Indicates the sum of the cyclomatic complexities of the methods in the file, and hasSmell indicates whether there is a certain code smell in the source code file version F i,j (1 indicates the presence of a smell, 0 indicates the absence);
[0052] Step (2) If the source code file version F i,j is identified as having a certain code smell, the set of characteristic thresholds for determining whether it has this code smell is T = <w 1 |b 1 ,..., w g |b g ,..., w t |b t >, where w g is a feature in W, and b g is the threshold corresponding to the feature w g recognized as this code smell, g = 1, 2,..., t, and the smell intensity is calculated by the following formula:
[0053]
[0054] where m(w g ) represents the maximum or minimum value of a certain code smell caused by the feature w g existing in the system to be analyzed (when w g exceeds b g and causes a certain code smell, the maximum value is selected; when w g is less than b g and causes a certain code smell, the minimum value is selected);
[0055] After adding the intensity information, the structure and smell information metric of the source code file version F i,j are expressed as:
[0056] S′ i,j = <className, classVersion, structure, hasSmell, intensity>
[0057] Step (3) Obtain the historical information metric of the source code file:
[0058] Considering that the time when the smell exists in the file and the change frequency of the source file are also important factors to consider when implementing refactoring, we need to obtain the characteristic metric of the smell historical information through the software version history. Assume that the source code file F i introduced a code smell in a certain historical version p, and the source code file F iThe current version is j, and the source code file version is F i,j The measurement of historical information in
[0059] H i,j =<className, classVersion, diffDays, diffVersions, action>
[0060] where diffDays represents the number of natural days between versions p and j, diffVersions represents the number of versions between versions p and j, and action represents the number of times the file F i has been modified between versions p and j;
[0061] In step (4), we find all the source code file versions with code smells eliminated in the source code file set F and divide them into the following two categories:
[0062] 1) For F i,j hasSmell = 1, while for F i,j+1 hasSmell = 0;
[0063] 2) For F i,j hasSmell = 1, and j is the last version of the source code file F i Considering that the file being deleted may also be the result of excessive code smells, this situation will also be considered;
[0064] According to the strategies in 1) and 2), all the last source code file versions F i,j with a certain code smell are obtained in the source code file set F, forming a source code file set with code smells eliminated. Given that this application scenario is essentially a classification task with an imbalanced number of positive and negative samples, we use the oversampling technique SMOTE to expand it. Specifically, SMOTE randomly selects an instance from the minority class and identifies its nearest neighbors, creates new instances between them, and finally obtains ζ P , which are the positive samples in the dataset. The remaining source code file versions in the source code file set F are used as the negative samples ζ N of the dataset. Finally, a complete sample dataset is obtained: ζ = ζ P ∪ζ N , where the reconstruction label value y corresponding to each source code file version in ζ P is 1, and the reconstruction label value y corresponding to each source code file version in ζ N is 0;
[0065] Step (5) To further extract important features from the dataset, the Recursive Feature Elimination (RFE) technique, Random Forest Classifier, and LGBM Classifier are used to retain the most important z features, denoted as W * =(w 1 ,w 2 ,…,w z ); The source code file version F i,j after feature selection can be represented by (S * i,j ,H * i,j ), where S * i,j represents the retained structure and smell information metric, and H * i,j represents the retained historical information metric;
[0066] Step (6) 90% of the data in the dataset is used as the training set ζ train , and the remaining 10% is used as the test set ζ test .
[0067] Step (7) For the information representation (S * i,j ,H * i,j ) of each source code file version, input S * i,j into LGBM to obtain the output h 1 , input H * i,j into Logistic Regression to obtain the output h 2 , and the output of the final model represents the reconstructed prediction probability where a 1 and a 2 represent weights, and a 1 +a 2 =1;
[0068] Step (8) Use the Cross-entropy Loss Function to calculate the loss between the reconstructed label value y and the output . The definition of the loss function is as follows (where d represents the number of training samples):
[0069]
[0070] Step (9) Utilize the training set ζ trainTrain the LGBM and Logistic Regression model parameters until the maximum number of iterations MaximumIter is reached. We set it to 500, and finally obtain the LGBM and Logistic Regression models with the best parameters after training;
[0071] In step (10), for a source code file, first obtain the file's smell information metrics and historical information metrics according to steps (1), (2), and (3). Then, extract the best features according to step (5) and input these features into the model obtained in step (9). Finally, obtain the predicted probability of whether this file needs to be refactored. If the probability is greater than or equal to 0.5, it means refactoring is required; if it is less than 0.5, it means refactoring is not required.
[0072] The present invention can be used for the refactoring prediction of industrial software to find a suitable time to refactor the software to eliminate smells.
Claims
1. Software refactoring prediction method based on code smells, characterized in that it includes the following steps: Step 1: Given a set F = (F 1,1 , F 1,2 , …, F i,j , …, F m,n ) of m * n source code file versions in the system to be analyzed, where F i,j represents the j-th version of the source code file F i . Use a code parsing tool to parse each source code file, and represent the structure and smell information metric of each source code file version F i,j in the form of S i,j = <className, classVersion, structure, hasSmell>, where i = 1, 2, …, m and j = 1, 2, …, n. Here, className represents the class name of the source code file version F i,j . Assume that a source code file contains one class. classVersion represents the version number of the source code file version F i,j in the project history. structure represents the set of structural features W of the source code file version F i,j . hasSmell indicates whether there is a certain code smell in the source code file version F i,j . 1 means there is a smell, and 0 means there is no smell; Feature set W = <w LOC , w NOA , w CBO , w MPC , w TCC , w McCabe , w WMC >, where w LOC represents the number of lines of code of the file, w NOA represents the number of attributes in the file, w CBO represents the number of target classes coupled to the file, where coupling means that the methods in the file call the methods or variables of the target class; w MPC represents the number of methods in the file that call other methods, W TCC represents the number of methods that are directly related by accessing the same attribute, w McCabe represents the complexity of the file calculated by the McVabe metric, w WMC represents the sum of the cyclomatic complexities of the methods in the file; Step 2: If the source code file version F i,j is identified as having a certain code smell, and the set of feature thresholds for determining whether it has this code smell is T = <w 1 |b 1 ,..., w g |b g ,..., w t |b t >, where w g is a feature in W, and b g is the threshold corresponding to the feature w g identified as this code smell, g = 1, 2,..., t, and the smell intensity is calculated through the following formula: Among them, m(w g ) represents the maximum or minimum value of a certain code smell existing in the system to be analyzed due to the feature w g : when w g exceeds b g and causes a certain code smell, the maximum value is selected; when w g is less than b g and causes a certain code smell, the minimum value is selected; Source code file version F after adding intensity information i,j The structure and smell information metric are represented as: S′ i,j =<className, classVersion, structure, hasSmell, intensity>; Step 3: Obtain the measurement of the historical information of the source code file: Let the source code file be F i In a certain historical version p, code smells were introduced in the source code file F i If the current version is j, then the source code file version F i,j The measurement of historical information in it is expressed as: H i,j =<className, classVersion, diffDays, diffVersions, action>, where diffDays represents the number of natural days between versions p and j, diffVersions represents the number of versions between versions p and j, and action represents the number of times the file F i has been modified between versions p and j; Step 4: We find all the source code file versions with code smells eliminated in the source code file set F and divide them into the following two categories: 1)F i,j hasSmell of which is 1, while for F i,j+1 hasSmell of which is 0; 2)F i,j whose hasSmell = 1, and j is the last version of the source code file F i ; According to the strategies in 1) and 2), all the last source code file versions F with a certain code smell are obtained from the set of source code files F i,j , forming a set of source code files for code smell elimination, and expanding it through the oversampling technique SMOTE to form ζ P , which are the positive samples in the dataset. The remaining source code file versions in the set of source code files F are used as the negative samples ζ of the dataset N , and finally a complete sample dataset is obtained: ζ = ζ P ∪ζ N , where the refactoring label value y corresponding to each source code file version in ζ P is 1, and the refactoring label value y corresponding to each source code file version in ζ N is 0; Step 5: Use feature recursive elimination technique, Random Forest Classifier, and LGBM Classifier to retain the most important z features, denoted as W * =(w 1 , w 2 , …, w z ); The source code file version F after feature selection i,j is represented by (S * i,j , H * i,j ), where S * i,j represents the retained structure and smell information metric, and H * i,j represents the retained historical information metric; Step 6: Use a part of the data in the dataset as the training set ζ train ; Step 7: For the information representation (S * i,j , H * i,j ) of each source code file version, input S * i,j into LGBM to obtain the output h 1 , and input H * i,j into Logistic Regression to obtain the output h 2 . The output of the final model represents the predicted probability of reconstruction where a 1 and a 2 represent weights, and a 1 + a 2 = 1; Step 8: Use the cross-entropy loss function to calculate the loss between the reconstructed label value y and the output The loss function is defined as follows: where d represents the number of training samples; Step 9: Use the training set ζ train to train the model parameters of LGBM and Logistic Regression until the maximum number of iterations MaximumIter is reached, and obtain the LGBM and Logistic Regression models with the best parameters after training; Step 10: For a source code file, first obtain the file's smell information measurement and historical information measurement according to Step 1, Step 2, and Step 3. Then extract the best features according to Step 5 and input these features into the model obtained in Step 9. Finally, obtain the prediction probability of whether this file needs to be refactored. If the probability is greater than or equal to 0.5, it means refactoring is required; if it is less than 0.5, it means refactoring is not required.
Citation Information
Patent Citations
Code bad smell detection method based on BP neural network
CN110502277A
Software taste detection method based on machine learning
CN111813442A