Method and device for improving performance of software defect prediction model

The harmony search-based cost-sensitive decision tree optimizes parameters across preprocessing and classification stages to enhance software defect prediction models, improving accuracy and resource allocation.

WO2025150605A1PCT designated stage expired Publication Date: 2025-07-17IND COOP FOUND CHONBUK NAT UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/001382
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-08
Filing Date
2024-01-30
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

Existing software defect prediction models suffer from significant performance deterioration when trained with suboptimal parameters, particularly in preprocessing and classification stages, leading to inefficient resource allocation and prediction accuracy issues.

Method used

A method and device that utilize a harmony search-based cost-sensitive decision tree to simultaneously optimize parameters across preprocessing and classification stages, including normalization, feature selection, class imbalance learning, and decision tree hyperparameters, to enhance model performance.

Benefits of technology

The approach achieves substantial performance improvement by automatically allocating optimal parameter sets, enhancing defect prediction accuracy and resource efficiency through improved G-measure, defect detection rate, and reduced false alarm rates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024001382_17072025_PF_FP_ABST
    Figure KR2024001382_17072025_PF_FP_ABST
Patent Text Reader

Abstract

A method for improving the performance of a software defect prediction model according to an embodiment of the present invention comprises: a software defect prediction model provision step of providing a software defect prediction model that identifies a module in which the occurrence of a software defect is possible; a parameter optimization step of simultaneously optimizing at least one parameter in respective steps of a software defect prediction process by using an optimization algorithm, in order to improve the performance of the software defect prediction model, wherein a search space of the optimization algorithm simultaneously considers a preprocessing step and a classification model construction step; and a performance evaluation step of evaluating the performance of the software defect prediction model.
Need to check novelty before this filing date? Find Prior Art

Description

Performance improvement method and device for software defect prediction model

[0001] The present invention relates to the field of software defect prediction (SDP), and more particularly, to a performance improvement method and device for a software defect prediction model.

[0002] Software quality assurance (QA) is a key topic in software engineering. Software defect prediction (SDP) is a technique used to ensure software quality.

[0003] Software defect prediction models identify modules prone to defects. Setting appropriate parameters in software defect prediction models is crucial because it impacts model performance. Software defect prediction aims to identify as many defective modules as possible in a software system before testing. Furthermore, software defect prediction helps effectively allocate limited human and material resources in software projects.

[0004] For software defect prediction, previous studies have applied various machine learning models to identify defect-prone software modules.

[0005] Machine learning models have various configurable parameters that must be set based on the developer's experience. For example, in a support vector machine (SVM) model, the kernel type is specified before training.

[0006] Similarly, the K value of the K-nearest neighbor model is specified before training. However, training with suboptimal parameters significantly degrades the predictive model's performance. Therefore, optimizing model parameters by tuning them is crucial to achieve significant performance improvements.

[0007] The present invention is intended to solve the problems of the past, and provides a performance improvement method and device for a software defect prediction model that seeks to substantially improve performance in software defect prediction by simultaneously optimizing parameters not only in the preprocessing stage but also in the classification model construction stage in software defect prediction.

[0008] The tasks of the present invention are not limited to the tasks mentioned above, and other tasks not mentioned can be clearly understood from the description below.

[0009] In a performance improvement method for a software defect prediction model according to one embodiment of the present invention,

[0010] A software defect prediction model providing step for providing a software defect prediction model that identifies a module in which a software defect may occur; and

[0011] A parameter optimization step for simultaneously optimizing at least one parameter in each step of a software defect prediction process using an optimization algorithm to improve the performance of the software defect prediction model, wherein the search space of the optimization algorithm simultaneously considers the preprocessing step and the classification model construction step.

[0012]

[0013] Preferably,

[0014] The above optimization algorithm uses a cost-sensitive decision tree based on harmony search (HS-CSDT).

[0015] The above harmony search-based cost-sensitive decision tree is characterized by using the harmony search algorithm (HS), which is a meta-heuristic algorithm.

[0016]

[0017] Preferably,

[0018] The above preprocessing step includes normalization, feature selection, and class imbalance learning, and the classification model construction step is characterized by including a decision tree (DT) model.

[0019]

[0020] Preferably,

[0021] The above parameter optimization step is

[0022] A parameter extraction step for extracting parameters in the normalization, feature selection, and class imbalance learning and hyperparameters of the decision tree model by executing the cost-sensitive decision tree based on the harmony search with training data; and

[0023] A performance evaluation step for evaluating the performance of the software defect prediction model using the extracted parameters in the normalization, feature selection, and class imbalance learning and the hyper parameters of the decision tree model;

[0024]

[0025] Preferably,

[0026] The performance evaluation of the above software defect prediction model is characterized in that it is performed by calculating the probability of detection, probability of false alarm, G-measure, and file inspection reduction (FIR) using validation data, and calculating the average value by averaging the calculated results.

[0027]

[0028] Preferably,

[0029] The above parameter optimization step is,

[0030] It includes adjusting the parameters in the normalization, the feature selection and the class imbalance learning and the hyperparameters of the decision tree model to increase the G-measure.

[0031]

[0032] In a performance improvement device for a software defect prediction model according to one embodiment of the present invention,

[0033] A software defect prediction model providing unit that provides a software defect prediction model that identifies modules in which software defects may occur; and

[0034] A parameter optimization unit that simultaneously optimizes at least one parameter in each stage of a software defect prediction process using an optimization algorithm to improve the performance of the software defect prediction model, wherein the search space of the optimization algorithm simultaneously considers the preprocessing stage and the classification model construction stage.

[0035]

[0036] Preferably,

[0037] The above optimization algorithm uses a cost-sensitive decision tree based on harmony search (HS-CSDT).

[0038] The above harmony search-based cost-sensitive decision tree is characterized by using the harmony search algorithm (HS), which is a meta-heuristic algorithm.

[0039]

[0040] Preferably,

[0041] The above preprocessing step includes normalization, feature selection, and class imbalance learning, and the classification model construction step is characterized by including a decision tree (DT) model.

[0042]

[0043] Preferably,

[0044] The above parameter optimization section,

[0045] A parameter extraction unit that extracts parameters in the normalization, feature selection, and class imbalance learning and hyperparameters of the decision tree model by executing the cost-sensitive decision tree based on the harmony search with training data; and

[0046] It includes a performance evaluation unit that evaluates the performance of the software defect prediction model using the extracted parameters in the normalization, feature selection, and class imbalance learning and the hyper parameters of the decision tree model.

[0047]

[0048] Preferably,

[0049] The performance evaluation of the software defect prediction model in the above performance evaluation section is characterized in that it is performed by calculating the probability of detection, the probability of false alarm, the G-measure, and the code inspection effort (FIR) using validation data, and calculating the average value by averaging the calculated results.

[0050]

[0051] Preferably,

[0052] The above parameter optimization section,

[0053] It includes adjusting the parameters in the normalization, the feature selection and the class imbalance learning and the hyperparameters of the decision tree model to increase the G-measure.

[0054]

[0055] Specific details of other embodiments are included in the detailed description and drawings.

[0056] The performance improvement method and device for a software defect prediction model according to the present invention has the advantage of being able to achieve substantial performance improvement by simultaneously optimizing parameters in all process stages of software defect prediction.

[0057] According to the performance improvement method and device for a software defect prediction model according to the present invention, an optimal parameter set is automatically allocated at all stages of a software project, thereby demonstrating excellent defect prediction performance.

[0058] However, the effects of the present invention are not limited to the effects mentioned above, and other effects not mentioned can be clearly understood from the description below.

[0059] FIG. 1 is a flowchart illustrating a performance improvement method for a software defect prediction model according to one embodiment of the present invention.

[0060] FIG. 2 is a diagram showing the overall process of software defect prediction according to one embodiment of the present invention.

[0061] FIG. 3 is a diagram schematically illustrating the execution process of a cost-sensitive decision tree based on harmony search according to one embodiment of the present invention.

[0062] FIG. 4 is a diagram summarizing an algorithm of a harmony search-based cost-sensitive decision tree (HS-CSDT) according to one embodiment of the present invention.

[0063] FIG. 5 is a diagram schematically illustrating the configuration of a performance improvement device for a software defect prediction model according to one embodiment of the present invention.

[0064] FIG. 6 is a diagram schematically illustrating the configuration of a parameter optimization unit according to one embodiment of the present invention.

[0065] FIG. 7 is a diagram summarizing the results of statistical analysis between the algorithm of the Harmony Search-based Cost-Sensitive Decision Tree (HS-CSDT) and other comparison methods in software defect prediction according to one embodiment of the present invention.

[0066] FIG. 8 is a diagram illustrating an exemplary computing device that may implement devices and / or systems according to various embodiments of the present invention.

[0067] A performance improvement method for software defect prediction according to an embodiment of the present invention is characterized by including: a software defect prediction model providing step of providing a software defect prediction model that identifies a module in which a software defect may occur; and a parameter optimization step of simultaneously optimizing at least one parameter in each step of a software defect prediction process using an optimization algorithm to improve the performance of the software defect prediction model, wherein the search space of the optimization algorithm simultaneously considers a preprocessing step and a classification model construction step.

[0068] A performance improvement device for a software defect prediction model according to an embodiment of the present invention comprises: a software defect prediction model providing unit that provides a software defect prediction model that identifies a module in which a software defect may occur; and a parameter optimization unit that simultaneously optimizes at least one parameter in each step of a software defect prediction process using an optimization algorithm to improve the performance of the software defect prediction model, wherein the search space of the optimization algorithm simultaneously considers a preprocessing step and a classification model construction step.

[0069] The advantages and features of the present invention, and the methods for achieving them, will become clearer with reference to the embodiments described in detail below together with the accompanying drawings. However, the present invention is not limited to the embodiments disclosed below and may be implemented in various different forms. These embodiments are provided only to ensure that the disclosure of the present invention is complete and to fully inform those skilled in the art of the scope of the invention, and the present invention is defined only by the scope of the claims. Like reference numerals designate like elements throughout the specification.

[0070] Embodiments described herein will be described with reference to cross-sectional and / or plan views, which are ideal illustrations of the present invention. In the drawings, the thicknesses of components are exaggerated for the purpose of effectively explaining the technical contents. Accordingly, the components illustrated in the drawings have a schematic nature, and the shapes of the components illustrated in the drawings are intended to illustrate specific forms of the components and are not intended to limit the scope of the invention. Although terms such as first, second, and third are used to describe various components in various embodiments of the present specification, these components should not be limited by such terms. These terms are used only to distinguish one component from another. The embodiments described and illustrated herein also include complementary embodiments thereof.

[0071] The terminology used herein is for the purpose of describing embodiments only and is not intended to limit the present invention. In this specification, the singular also includes the plural unless specifically stated otherwise. As used herein, the terms "comprises" and / or "comprising" do not exclude the presence or addition of one or more other components, steps, operations, and / or elements to the mentioned components, steps, operations, and / or elements.

[0072] Unless otherwise defined, all terms (including technical and scientific terms) used herein may be used in their common sense to those of ordinary skill in the art to which the present invention pertains. Furthermore, terms defined in commonly used dictionaries are not to be interpreted ideally or excessively unless explicitly and specifically defined otherwise.

[0073] Hereinafter, with reference to the drawings, the concept of the present invention and embodiments thereof will be described in detail.

[0074]

[0075] FIG. 1 is a flowchart illustrating a performance improvement method for a software defect prediction model according to one embodiment of the present invention.

[0076] A performance improvement method for a software defect prediction model according to one embodiment of the present invention includes a software defect prediction model providing step (S110) and a parameter optimization step (S120).

[0077] The software defect prediction model provision step (S110) provides a software defect prediction model that identifies a module in which a software defect may occur.

[0078] The parameter optimization step (S120) simultaneously optimizes at least one parameter in each step of the software defect prediction process using an optimization algorithm to improve the performance of the software defect prediction model.

[0079] In the parameter optimization step (S120), the search space of the optimization algorithm simultaneously considers the preprocessing step and the classification model construction step.

[0080] The software defect prediction model of the present invention is characterized by improving the performance of the model by simultaneously considering parameters in both the preprocessing stage and the classification model construction stage.

[0081] The optimization algorithm of the present invention utilizes a cost-sensitive decision tree based on harmony search (HS-CSDT). HS-CSDT utilizes the Harmony Search algorithm (HS), a metaheuristic algorithm. HS-CSDT uses the Harmony Search algorithm to simultaneously identify optimal feature selection, regularization techniques, class weights, and decision tree hyperparameters.

[0082] In one embodiment, the preprocessing step includes normalization, feature selection, and class imbalance learning, and the classification model building step includes a decision tree (DT) model.

[0083] Feature selection selects a subset of features that effectively characterize the given data (input data). In the present invention, the input data is generated using software metrics. These metrics include information such as lines of code (LOC) and response for a class (RFC).

[0084] Normalization is a technique for adjusting the values ​​of each feature within a specified range so that multiple features have equal weighting. Defect prediction performance can vary depending on the normalization technique used, such as z-score or min-max.

[0085] Class imbalance learning is a factor that affects defect prediction performance, implying that addressing this issue can improve performance. Therefore, in class imbalance learning, the weights (or ratios) of defective and non-defective classes are important parameters in the preprocessing stage. Software defect data suffers from class imbalance, where the number of non-defective instances exceeds the number of defective instances. In most machine learning applications, a skewed proportion of instances in a particular class negatively impacts defect prediction performance.

[0086] The present invention utilizes cost-sensitive learning. Cost-sensitive learning addresses class imbalance problems at the algorithmic level. In software QA activities, misclassifying a defective instance is more costly than misclassifying a non-defective instance. Therefore, cost-sensitive learning methods aim to build a predictive model with the lowest misclassification cost.

[0087] In one embodiment, the parameter optimization step includes a parameter extraction step of running a harmony search-based cost-sensitive decision tree with training data to extract parameters in regularization, feature selection, and class imbalance learning and hyperparameters of the decision tree model, and a performance evaluation step of evaluating the performance of the software defect prediction model using the extracted parameters in regularization, feature selection, and class imbalance learning and hyperparameters of the decision tree model.

[0088] In machine learning, hyperparameters are variables set in a model to implement an optimal training model. They can determine factors such as the learning rate, number of epochs (the number of training iterations), and weight initialization. Furthermore, hyperparameter tuning techniques can be applied to find optimal values ​​for the trained model.

[0089] In one embodiment, the performance evaluation of the software defect prediction model is characterized in that it is performed by calculating the probability of detection, the probability of false alarm, the G-measure, and the code inspection effort (FIR) using validation data, and calculating the average value by averaging the calculated results.

[0090] Predictive performance for binary classification is typically evaluated using the confusion matrix in Table 1. Specifically, in software defect prediction, the goal is to increase the defect detection rate (PD) and reduce the false alarm rate (PF). In this invention, performance is assessed using the following four evaluation metrics.

[0091] The confusion matrix for the classification results of the software defect prediction model is shown in Table 1.

[0092] [Table 1l Confusion matrix

[0093] _DEFECTIVE: TP(TRUE POSITIVE), FP(FALSE POSITIVE)

[0094] -CLEAN: FN(FALSE NEGATIVE), TN(TRUE NEGATIVE)

[0095] _ Probability of detection (PD)

[0096] The defect detection rate (PD) is an indicator of the number of actual defects predicted by the model to be defects, and is calculated as in Equation (1).

[0097] _ Formula (1)

[0098] PD = TP / (TP + FN)

[0099]

[0100] _Probability of False Alarm (PF)

[0101] The defect false alarm rate represents the ratio of the number of non-defects incorrectly classified as defects to the total number of non-defects, and is calculated as in Equation (2).

[0102] _ Formula (2)

[0103] PF = FP / (FP + TN)

[0104] _G-measure

[0105] The G-measure is calculated as the harmonic mean of the fault detection rate (PD) and the fault false alarm rate (PF), and is an indicator suitable for evaluating class imbalance data sets, as shown in Equation (3).

[0106] _ Formula (3)

[0107] G-measure = 2 × [ {PD ×(1- pf)} / {PD + (1-pf)}]

[0108] _ Code Inspection Effort (FIR)

[0109] The File Inspection Effort (FIR) is a metric that indicates the extent to which a software defect prediction model reduces code inspection effort. This evaluation metric represents the rate at which the number of files to be inspected decreases to achieve the same defect detection rate (PD). A higher FIR performance indicates that smaller files are easier to detect defects. In the FIR equation, FI (File Inspection) is the ratio of the number of files to be inspected to the total number of files. FIR is calculated as shown in Equation (4).

[0110]

[0111] _ Formula (4)

[0112] FIR = (PD - FI) / PD

[0113] In one embodiment, the parameter optimization step includes adjusting parameters in regularization, feature selection, and class imbalance learning, as well as hyperparameters of the decision tree model, to increase the G-measure. The parameter optimization step aims to improve the performance of the software defect prediction model by adaptively adjusting parameters based on the dataset input to the software defect prediction model.

[0114]

[0115] FIG. 2 is a diagram showing the overall process of software defect prediction according to one embodiment of the present invention.

[0116] Figure 2 illustrates a typical file-level defect prediction process.

[0117] In the first step ((1) labeling / counting), files (instances) are collected from a software archive. The software archive (210) includes a version control system, a bug tracking system, and an email archive. Then, if an instance contains one or more defects, it is labeled "buggy," otherwise it is labeled "clean."

[0118] In the second step ((2) feature extraction), software metrics, including object-oriented metrics (230), are extracted from instances (220) and turned into features. These metrics can be extracted using the CKJM tool.

[0119] The third step ((3) Training Corpus Creation) is to create training instances (240) used to train a machine learning-based model. The labels and features generated in the previous two steps are then used as training data.

[0120] The fourth step ((4) preprocessing) builds a defect prediction model. Before using training data as input data for the prediction model, preprocessing methods (250) are considered. Research on software defect prediction models (SDPs) uses preprocessing methods such as normalization, feature selection, and class imbalance learning.

[0121] Finally ((6) Prediction & Evaluation) the trained model is used to predict and evaluate whether a new instance (270) has a bug or is clean.

[0122]

[0123] FIG. 3 is a diagram schematically illustrating the execution process of a cost-sensitive decision tree based on harmony search according to one embodiment of the present invention.

[0124] Figure 3 schematically illustrates the execution process of a cost-sensitive decision tree based on harmony search, which includes a preprocessing step (310), a learning step (320), and an evaluation step (330).

[0125] The preprocessing step (310) includes normalizing (310) the data input to the software defect prediction model and then feature selection (320).

[0126] In the present invention, class imbalance learning is additionally included in the preprocessing step (310).

[0127] The learning step (320) loads parameters (321) to build a classifier (322) and trains the built classifier (323).

[0128] The classifier or classification model building step includes a decision tree (DT) model.

[0129] Finally, the evaluation step (330) inputs a new instance into the classifier (331), determines whether the input new instance is defective (332) (defective / clean), and evaluates the performance.

[0130] The performance of the software defect prediction model is evaluated by calculating the probability of detection, probability of false alarm, G-measure, and FIR, and then calculating the average value by averaging the calculated results.

[0131] In addition, parameters in the preprocessing stage and hyperparameters in the classification model construction stage are optimized to increase the G-measure.

[0132]

[0133] FIG. 4 is a diagram summarizing an algorithm of a harmony search-based cost-sensitive decision tree (HS-CSDT) according to one embodiment of the present invention.

[0134] The algorithm illustrated in Fig. 4 represents the pseudocode of a Harmony Search-based Cost-Sensitive Decision Tree (HS-CSDT). The defect prediction data (DATA) is divided into training data (Xtrain) and test data (Xtest) through a hierarchical K-fold validator (lines 1-4). Based on the training data, the optimal parameter set (hi) with the optimal fitness value is obtained through Harmony Search (HS).

[0135] Hyperparameter optimization is performed using Harmony Search (HS) (line 5).

[0136] Assign regularization parameters (vn), feature selection parameters (vf), class weight parameters (vw), and model hyperparameters (vh) (lines 6-9).

[0137] Next, the model is trained by providing model parameters through a preprocessing step that converts the training data into a feature subset based on information about each parameter (lines 10-12).

[0138] Next, the trained model is used to predict defects in the test data (line 13). Finally, lines 2-15 are repeated to average the performance of the Harmony Search-based Cost-Sensitive Decision Tree (HS-CSDT).

[0139]

[0140] FIG. 5 is a diagram schematically illustrating the configuration of a performance improvement device for a software defect prediction model according to one embodiment of the present invention.

[0141] A performance improvement device (500) for a software defect prediction model according to one embodiment of the present invention includes a software defect prediction model providing unit (510) and a parameter optimization unit (520).

[0142] The software defect prediction model providing unit (510) provides a software defect prediction model that identifies modules in which software defects may occur.

[0143] The parameter optimization unit (520) simultaneously optimizes at least one parameter at each stage of the software defect prediction process using an optimization algorithm to improve the performance of the software defect prediction model. In the optimization process, the search space of the optimization algorithm simultaneously considers the preprocessing stage and the classification model construction stage.

[0144] In one embodiment, the optimization algorithm utilizes a cost-sensitive decision tree based on harmony search (HS-CSDT). An optimization algorithm is a computer algorithm that finds solutions to various engineering problems that minimize or maximize a given cost function by adjusting the values ​​of optimization variables within each search range.

[0145] In one embodiment, the harmony search based cost sensitive decision tree is characterized by using a metaheuristic algorithm, the harmony search algorithm (HS).

[0146] In one embodiment, the preprocessing step includes normalization, feature selection, and class imbalance learning, and the classification model building step includes a decision tree (DT) model.

[0147] The parameter optimization unit (520) includes a parameter extraction unit (610) that extracts parameters in normalization, feature selection, and class imbalance learning and hyperparameters of a decision tree model by executing a cost-sensitive decision tree based on harmony search with training data, and a performance evaluation unit (620) that evaluates the performance of a software defect prediction model using the extracted parameters in normalization, feature selection, and class imbalance learning and hyperparameters of a decision tree model.

[0148] The performance evaluation of the software defect prediction model in the performance evaluation unit (620) is characterized by being performed by calculating the probability of detection, probability of false alarm, G-measure, and FIR (Finally Inspected Intensity Ratio) using validation data, and calculating the average value by averaging the calculated results.

[0149]

[0150] FIG. 6 is a diagram schematically illustrating the configuration of a parameter optimization unit according to one embodiment of the present invention.

[0151] The parameter optimization unit (600) of the present invention includes a parameter extraction unit (610) and a performance evaluation unit (620).

[0152] The parameter extraction unit (610) executes a cost-sensitive decision tree based on harmony search with training data to extract parameters and hyperparameters of the decision tree model in normalization, feature selection, and class imbalance learning.

[0153] The performance evaluation unit (620) evaluates the performance of the software defect prediction model using the extracted parameters in normalization, feature selection, and class imbalance learning and the hyperparameters of the decision tree model.

[0154] The parameter optimization unit (600) includes adjusting parameters in regularization, feature selection, and class imbalance learning and hyperparameters of the decision tree model to increase the G-measure.

[0155]

[0156] FIG. 7 is a diagram summarizing the results of statistical analysis between the algorithm of the Harmony Search-based Cost-Sensitive Decision Tree (HS-CSDT) and other comparison methods in software defect prediction according to one embodiment of the present invention.

[0157] Comparisons of HSOCS-US-SVM (710) and SOGA-LR (720) demonstrate above-average performance. The proposed method differs from the proposed method in all evaluation metrics: PD, PF, G-measure, and FIR. This demonstrates that the Harmony Search-based Cost-Sensitive Decision Tree (HS-CSDT) algorithm of the present invention offers statistically significant performance improvements.

[0158] Compared to SOGA-DT (730) and COSTE-MLP (740), some metrics are below average, but the G-measure significantly surpasses them. Since SOGA-DT (730) and the HS-CSDT of the present invention differ only in their search spaces, the search space proposed in the present invention is more effective in terms of performance enhancement.

[0159] The performance difference between DT(750) and the HS-CSDT method of the present invention shows a large performance difference in the remaining evaluation indices except PF.

[0160] Additionally, the HS-CSDT of the present invention shows a large level of performance improvement compared to the performance of the LR (760) and SVM models (770).

[0161] The comparative results in Figure 7 confirm the effectiveness of metaheuristic algorithms in terms of performance improvement. In particular, the HS-CSDT of the present invention demonstrates a significant performance advantage over related methods in terms of G-measure. Considering that data used in software defect prediction (SDP) suffers from class imbalance, G-measure, which comprehensively considers defect detection rate (PD) and false alarm rate (PF), demonstrates significant performance improvements.

[0162]

[0163] FIG. 8 is a diagram illustrating an exemplary computing device that may implement devices and / or systems according to various embodiments of the present invention.

[0164] An exemplary computing device (800) capable of implementing devices according to some embodiments of the present disclosure will now be described in more detail with reference to FIG. 8.

[0165] A computing device (800) may include one or more processors (810), a bus (850), a communication interface (870), a memory (830) for loading a computer program (891) to be executed by the processor (810), and a storage (890) for storing the computer program (891). However, only components related to the embodiment of the present disclosure are illustrated in FIG. 8.

[0166] Accordingly, a person skilled in the art will appreciate that the present disclosure may further include general components other than those illustrated in FIG. 8.

[0167] The processor (810) controls the overall operation of each component of the computing device (800). The processor (810) may include a central processing unit (CPU), a microprocessor unit (MPU), a microcontroller unit (MCU), a graphic processing unit (GPU), or any other form of processor (810) well known in the art of the present disclosure. In addition, the processor (810) may perform operations for at least one application or program for executing a method according to embodiments of the present disclosure. The computing device (800) may include one or more processors (810). The computing device (800) may refer to artificial intelligence (AI).

[0168] The memory (830) stores various data, commands, and / or information. The memory (830) can load one or more programs (891) from the storage (890) to execute methods according to embodiments of the present disclosure. The memory (830) may be implemented as a volatile memory such as RAM, but the technical scope of the present disclosure is not limited thereto.

[0169] The bus (850) provides communication between components of the computing device (800). The bus (850) may be implemented as various types of buses, such as an address bus, a data bus, and a control bus.

[0170] The communication interface (870) supports wired and wireless Internet communication of the computing device (800). Furthermore, the communication interface (870) may support various communication methods other than Internet communication. To this end, the communication interface (870) may be configured to include a communication module well known in the technical field of the present disclosure.

[0171] According to some embodiments, the communication interface (870) may be omitted.

[0172] Storage (890) can non-temporarily store one or more programs (891) and various data.

[0173] Storage (890) may be configured to include non-volatile memory such as Read Only Memory (ROM), Erasable Programmable ROM (EPROM), Electrically Erasable Programmable ROM (EEPROM), flash memory, a hard disk, a removable disk, or any form of computer-readable recording medium well known in the art to which the present disclosure pertains.

[0174] The computer program (891) may include one or more instructions that, when loaded into the memory (830), cause the processor (810) to perform methods / operations according to various embodiments of the present disclosure. That is, the processor (810) may perform the methods / operations according to various embodiments of the present disclosure by executing the one or more instructions.

[0175]

[0176] Although the preferred embodiments of the present invention have been illustrated and described above, the present invention is not limited to the specific embodiments described above, and various modifications may be made by a person skilled in the art without departing from the gist of the present invention as claimed in the claims. Furthermore, such modifications should not be understood individually from the technical idea or prospect of the present invention.

Claims

1. A software defect prediction model providing step for providing a software defect prediction model that identifies modules where software defects may occur; and A method for improving the performance of a software defect prediction model, comprising: a parameter optimization step of simultaneously optimizing at least one parameter at each step of a software defect prediction process using an optimization algorithm to improve the performance of the software defect prediction model, wherein the search space of the optimization algorithm simultaneously considers the preprocessing step and the classification model construction step.

2. In claim 1, The above optimization algorithm uses a cost-sensitive decision tree based on harmony search (HS-CSDT). A method for improving performance for a software defect prediction model, characterized in that the above harmony search-based cost-sensitive decision tree uses a harmony search algorithm (HS), which is a meta-heuristic algorithm.

3. In claim 2, A method for improving performance for a software defect prediction model, wherein the preprocessing step includes normalization, feature selection, and class imbalance learning, and the classification model construction step includes a decision tree (DT) model.

4. In claim 3, The above parameter optimization step is A parameter extraction step for extracting parameters in the normalization, feature selection and class imbalance learning and hyper parameters of the decision tree model by executing the cost-sensitive decision tree based on the harmony search with training data; and A performance improvement method for a software defect prediction model, comprising: a performance evaluation step of evaluating the performance of the software defect prediction model using the extracted parameters in the normalization, feature selection, and class imbalance learning and the hyper parameters of the decision tree model.

5. In claim 4, A performance improvement method for a software defect prediction model, characterized in that the evaluation of the performance of the software defect prediction model is performed by calculating the probability of detection, the probability of false alarm, the G-measure, and the file inspection reduction (FIR) using validation data, and calculating the average value by averaging the calculated results.

6. In claim 5, The above parameter optimization step is, A method for improving performance for a software defect prediction model, comprising adjusting parameters in the regularization, the feature selection and the class imbalance learning and the hyperparameters of the decision tree model to increase the G-measure.

7. A software defect prediction model providing unit that provides a software defect prediction model that identifies modules in which software defects may occur; and A performance improvement device for a software defect prediction model, comprising: a parameter optimization unit that simultaneously optimizes at least one parameter at each stage of a software defect prediction process using an optimization algorithm to improve the performance of the software defect prediction model, wherein the search space of the optimization algorithm simultaneously considers the preprocessing stage and the classification model construction stage.

8. In claim 7, The above optimization algorithm uses a cost-sensitive decision tree based on harmony search (HS-CSDT). A performance improvement device for a software defect prediction model, characterized in that the above harmony search-based cost-sensitive decision tree uses a harmony search algorithm (HS), which is a meta-heuristic algorithm.

9. In claim 8, A performance improvement device for a software defect prediction model, characterized in that the preprocessing step includes normalization, feature selection, and class imbalance learning, and the classification model construction step includes a decision tree (DT) model.

10. In claim 9, The above parameter optimization section, A parameter extraction unit for executing the cost-sensitive decision tree based on the harmony search with training data to extract parameters in the normalization, feature selection and class imbalance learning and hyper parameters of the decision tree model; and A performance improvement device for a software defect prediction model, comprising: a performance evaluation unit for evaluating the performance of the software defect prediction model using the extracted parameters in the normalization, feature selection, and class imbalance learning, and the hyper parameters of the decision tree model.

11. In claim 10, A performance improvement device for a software defect prediction model, characterized in that the evaluation of the performance of the software defect prediction model in the above performance evaluation section is performed by calculating the probability of detection, the probability of false alarm, the G-measure, and the file inspection reduction (FIR) using validation data, and calculating the average value by averaging the calculated results.

12. In claim 11, The above parameter optimization section, A performance improvement device for a software defect prediction model, comprising adjusting parameters in the regularization, the feature selection and the class imbalance learning and the hyper parameters of the decision tree model to increase the G-measure.

Citation Information

Patent Citations

  • Wafer processing method

    KR1020200038416A

  • prefab sofa and combined IoT equipment

    KR1020220071157A

  • Camera modure for vehicle

    KR1020240073692A

  • User-selectable meta verse space combination design system incorporating the concept of unit space

    KR102523515B1