A software defect prediction method integrating genetic algorithm and deep neural network

By integrating genetic algorithms and deep neural networks, extracting and mapping subsets of features of software projects, the problem that existing software defect prediction methods rely on manual design features is solved, and more efficient and accurate software defect prediction is achieved.

CN115185732BActive Publication Date: 2025-05-02NANTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210849578.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-19
Publication Date
2025-05-02
Estimated Expiration
2042-07-19

AI Technical Summary

Technical Problem

Existing software defect prediction methods rely on manual design characteristics, with uncertainty and high costs, making it difficult to effectively predict software defects.

Method used

Using a method of fusion genetic algorithm and deep neural network, a subset of features of source and target projects are extracted through selection, crossover and Bayesian optimization variations, and mapped them to the same implicit space through a variational autoencoder for unsupervised training to extract high-level feature representations.

Benefits of technology

Improves the accuracy and efficiency of software defect prediction, reduces manual intervention and costs, and achieves better performance on the same dataset than other methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115185732B_ABST
    Figure CN115185732B_ABST
Patent Text Reader

Abstract

The present invention provides a software defect prediction method integrating genetic algorithm and deep neural network, which belongs to the field of computer technology and solves the technical problem that new features in automatic defect prediction are uncertain and may differ from the prediction results; its technical solution is: using a result-optimized genetic algorithm to select the features of a data set, combining a variational autoencoder and a maximum mean difference distance, learning the common features of the source project and the target project, to train a reliable defect prediction model. The beneficial effects of the present invention are: the genetic algorithm of the present invention combines the Bayesian algorithm to replace the random mutation process of the traditional genetic algorithm, designs a new fitness function, reduces unnecessary features, and compares with the traditional cross-project defect prediction method on multiple data sets, showing that the present invention can improve the effectiveness of software prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a software defect prediction method integrating a genetic algorithm and a deep neural network. Background Art

[0002] Software is a collection of computer data and instructions organized in a specific order. It has great advantages for daily life and social development. In the past period of time, with the rapid development of the economy, the market for various software needs has also expanded. With the rapid increase in the number of software, its quality has also become the focus of people's increasing concern. Most people believe that code defects in software implementation are the main cause of crashes. In order to solve this problem and improve the reliability of software, software testing is introduced to discover and solve the code that causes crashes. However, with the rapid development of various aspects at home and abroad, more and more software is being developed, and code defects are still a point that is not given much attention.

[0003] With the increasing number of accidents caused by defects every year, it will cause huge unexpected losses in business, making the relationship between companies and customers tense. Therefore, software testing is gradually gaining attention. Many companies use manual inspection of code segments and units. Although this can find bugs and defects, manual inspection is definitely unrealistic for future software. The labor cost is getting higher and higher, and the cost is limited. Therefore, the introduction of automatic defect prediction is necessary and beneficial to the limited cost.

[0004] Recent studies divide automatic defect prediction into two parts: first, extract features from source files, and then use various machine learning algorithms to develop classifiers. In the past, in order to improve the accuracy of prediction, new identification features or feature combinations were manually designed in the first stage. However, the new features are uncertain and may differ from the predicted results. Therefore, it is necessary to explore a convenient and effective method to solve the defect prediction problem in software.

[0005] How to solve the above technical problems becomes the subject faced by the present invention. Summary of the invention

[0006] The purpose of the present invention is to provide a software defect prediction method integrating genetic algorithm and deep neural network, which can perform defect prediction based on feature learning of source project and target project.

[0007] The idea of ​​the present invention is: the present invention proposes a software defect prediction method that integrates genetic algorithm and deep neural network, that is, through selection, crossover and variation after Bayesian optimization in the genetic algorithm, feature subsets of the source project and the target project are extracted, and then, through the variational autoencoder, the two projects are mapped to the same implicit space for unsupervised training to extract high-level feature representation. The method of the present invention achieves better performance than other methods in the same data set.

[0008] The present invention is implemented by the following technical solution: a software defect prediction method integrating genetic algorithm and deep neural network, which includes the following steps:

[0009] (1) Encode the public PROMIS, NASA, ReLink, and AEEEM datasets, initialize the population, and design a suitable fitness function to evaluate the quality of individuals. The specific processing operations include the following steps:

[0010] (1-1) First, the relatively simple and commonly used binary encoding method is used to encode the features;

[0011] (1-2) Based on Huffman coding, the fitness function is designed in combination with the branch distance method. The K value is 0.01 to ensure that the branch function returns a non-negative value.

[0012] (1-3) When the branch predicate is false, the value of the branch function is positive, and when the branch predicate is true, the branch function is 0.

[0013] (2) The processed population is subjected to selection, crossover and Bayesian optimization-improved mutation operations through a genetic algorithm improved by Bayesian optimization to screen out a suitable feature subset, which specifically includes the following steps:

[0014] (2-1) The classic roulette wheel selection method is used to select excellent individuals. The higher the fitness of an individual, the greater the probability of its being selected.

[0015] (2-2) Using the single-point crossover method, a point is randomly selected and the two chromosomes are crossovered at this point;

[0016] (2-3) When the genotypes of the individuals selected for crossover are the same, the genotype of the first half of the crossover point of the first individual remains unchanged, and the genotype of the last half remains consistent with the flashback coding after the crossover point of the second individual;

[0017] (2-4) The individuals remaining after crossover are used as known points and brought into the Bayesian optimization model to obtain two parameter values. These two parameter values ​​are added to the population as new individuals to improve the diversity of the population and increase the convergence speed.

[0018] (3) The feature subset after feature selection is used as the features of the source project and the target project, and the parts that have an adverse effect on the results are screened out;

[0019] (4) Using a variational autoencoder, the feature subsets selected from the source and target items are mapped into a robust implicit feature space for unsupervised pre-training, which specifically includes the following steps:

[0020] (4-1) First, the source items and target items after feature selection are filled with zeros so that they have the same feature dimension;

[0021] (4-2) Put the data of the two projects into a variational autoencoder so that they can learn implicit features;

[0022] (4-3) Sampling from the posterior distribution, the new encoding input is generated into the network to generate a new x.

[0023] (5) Obtaining the common features of the source project and the target project, by introducing the maximum mean difference constraint to constrain the distance between the mean parameters of the implicit feature distribution of the source project and the target project, which is corresponded by the mean vector output by the hidden layer, and their difference features are corresponded by the variance vector, which specifically includes the following steps:

[0024] (5-1) Half of the samples are taken from the source project and the other half are sampled from the target project. The variational autoencoder can learn the distribution information of both at the same time;

[0025] (5-2) Obtain the maximum mean difference between the mean parameters of the source and target items according to their implicit feature distributions;

[0026] (5-3) The implicit features are resampled from a Gaussian distribution with a similar mean to obtain the true difference between the source and target items.

[0027] (6) In order to ensure that the implicit space has good reliability, we added a discriminant network to extract high-level feature representations and judge the performance.

[0028] Furthermore, in step (2), the Bayesian optimization algorithm is used to guide the individual mutation step in the genetic algorithm, the selected probability proxy model is the Gaussian process of the non-parametric model, the confidence upper bound strategy UCB function is selected as the acquisition function, and the fitness is judged after the crossover. Only when the highest fitness is lower than that of the previous generation, the mutation operation is performed, otherwise, the next iteration is directly performed. Compared with the original genetic algorithm, the convergence speed is accelerated, new individuals are introduced, and the diversity of the population is improved.

[0029] Furthermore, in step (4), the variational autoencoder in the deep neural network is used to further learn the features of the source project and the target project, that is, unsupervised training is performed through an implicit feature space, the maximum mean difference is introduced to obtain the common features, and finally a layer of discriminant network is added to ensure the reliability of the implicit space, thereby improving the accuracy of the prediction of the target project.

[0030] The parameter settings of the software defect prediction method integrating genetic algorithm and deep neural network are as follows:

[0031] K in the branching function: 0.01.

[0032] Compared with the prior art, the beneficial effects of the present invention are as follows: a software defect prediction method integrating genetic algorithm and deep neural network proposed in the present invention takes into account four aspects: characteristic factors that have an inverse effect on the results, conditions for variation, differences in characteristic dimensions between projects, and reliability of the model. Good features are screened out by using the genetic algorithm with the optimal solution. Compared with the traditional genetic algorithm, the Gaussian process of the non-parametric model in Bayesian optimization is adopted in terms of variation. This method will increase the diversity of the population and effectively improve its convergence speed, and can obtain higher quality features. When the dimensions between projects are different, the feature dimensions with small dimensions are padded with zeros, and a layer of nonlinear perceptron is added at the end, which helps to improve its reliability. The software defect prediction method integrating genetic algorithm and deep neural network proposed in the present invention is simpler, more reliable, and more effective than the cross-project defect prediction method. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] The accompanying drawings are used to provide further understanding of the present invention and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention and do not constitute a limitation of the present invention.

[0034] Figure 1 A system framework diagram of a hybrid software defect prediction method based on Bayesian optimization genetic algorithm and deep neural network provided by the present invention. DETAILED DESCRIPTION

[0035] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. Of course, the specific embodiments described here are only used to explain the present invention and are not used to limit the present invention.

[0036] Example 1

[0037] See also Figure 1 As shown, this embodiment provides a software defect prediction method integrating genetic algorithm and deep neural network, which specifically includes the following contents:

[0038] (1) Collect public PROMIS, NASA, ReLink, and AEEEM datasets, which contain 18 projects including CM1, MW1, and PC1, with 37, 29, 26, and 61 metrics, respectively.

[0039] (2) Perform preprocessing operations on the data set, including deleting the previous introduction, deleting the annotation information, and changing whether there is a defect to 0 or 1; Table 1 shows the detailed statistical information of the sample number, measurement number and defect rate of the experimental data set.

[0040] Table 1 Dataset information

[0041]

[0042] (3) To ensure the accuracy of the method, we selected one project from the 18 projects as the target and then used each project in the other groups as the source project in turn.

[0043] (4) Encode the selected data set, initialize the population, and design a suitable fitness function to evaluate the quality of individuals;

[0044] (4-1) First, a relatively simple and commonly used binary encoding method is used to encode the features;

[0045] (4-2) Based on Huffman coding, the fitness function is designed in combination with the branch distance method. The K value is 0.01 to ensure that the branch function returns a non-negative value; when the branch predicate is false, the value of the branch function is positive, and when the branch predicate is true, the branch function is 0.

[0046] (5) The processed population is subjected to selection, crossover and Bayesian optimization-improved genetic algorithm to screen out a suitable feature subset.

[0047] (5-1) The classic roulette wheel selection method is used to select excellent individuals. The higher the fitness of an individual, the greater the probability of its selection.

[0048] (5-2) Using the single-point crossover method, a point is randomly selected and the two chromosomes are crossovered at this point;

[0049] (5-3) When the genotypes of the individuals selected for crossover are the same, the genotype of the first half of the crossover point of the first individual remains unchanged, and the genotype of the last half remains consistent with the flashback coding after the crossover point of the second individual;

[0050] (5-4) The individuals remaining after crossover are used as known points and brought into the Bayesian optimization model to obtain two parameter values. These two parameter values ​​are added to the population as new individuals to improve the diversity of the population and increase the convergence speed.

[0051] (6) The feature subsets selected by the source and target items are mapped into a robust implicit feature space using a variational autoencoder for unsupervised pre-training.

[0052] (6-1) First, the source items and target items after feature selection are filled with zeros so that they have the same feature dimension;

[0053] (6-2) Put the data of the two projects into a variational autoencoder so that they can learn implicit features;

[0054] (6-3) Sampling from the posterior distribution, the new encoding input is generated into the network to generate a new x.

[0055] (7) The common features of the source and target items are obtained by introducing the maximum mean difference constraint to constrain the distance between the mean parameters of the implicit feature distribution between the source and target items, which is corresponded by the mean vector output by the hidden layer, and their difference features are corresponded by the variance vector.

[0056] (8) In order to ensure that the implicit space has good reliability, we added a discriminant network to extract high-level feature representations and judge the performance.

[0057] (9) The optimal parameter settings of this method are as follows:

[0058] K in the branching function: 0.01.

[0059] (10) The method of this embodiment and the existing cross-project defect prediction method are evaluated on the same data set, and the quality of the method is evaluated based on the accuracy.

[0060] Table 2 Comparison of results between the method of this embodiment and other methods

[0061]

[0062] Experiments have shown that the software defect prediction method that integrates genetic algorithm and deep neural network proposed in this embodiment can perform more reliable cross-project defect prediction compared with the baseline method; specifically, the method of this embodiment first performs feature selection by an optimized genetic algorithm and then performs defect prediction, which can surpass these baseline methods; for Accuracy, the method of this embodiment can at least improve the performance by 5.1% respectively; these results show that the method proposed in this embodiment is highly competitive.

[0063] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. A software defect prediction method integrating genetic algorithm and deep neural network, characterized in that: The following steps are involved: S1: Encode the public PROMIS, NASA, ReLink, and AEEEM datasets, initialize the population, and design a suitable fitness function to evaluate the quality of individuals; S2: The population is selected through the genetic algorithm improved by Bayesian optimization, and the operations of crossover and mutation improved by Bayesian optimization are performed to screen out the appropriate feature subset; S3: Map the selected feature subsets of the source and target items into an implicit feature space and perform unsupervised pre-training on them using a variational autoencoder; In step S3, the source items and the target items are mapped into the implicit feature space for unsupervised pre-training, which includes the following steps: S31: Fill the source items and target items after feature selection with zeros so that they have the same feature dimension; S32: Put the data of the two projects into a variational autoencoder so that they learn the implicit feature space; S33: Sampling from the posterior distribution, the new encoding input is used to generate a new feature representation x in the generative network; S4: The maximum mean difference is introduced to constrain the distance between the mean parameters of the implicit feature distribution of the source item and the target item, capturing the common features of the source item and the target item, which are represented by the mean vector output by the hidden layer, and the feature difference between the two is represented by the variance vector; The step S4 of learning the common features of the source project and the target project includes the following steps: S41: Using half of the samples from the source project and the other half from the target project, the variational autoencoder can learn the distribution information of both at the same time; S42: obtaining the maximum mean difference between the mean parameters of the source item and the target item according to their implicit feature distributions; S43: The implicit features are resampled from Gaussian distributions with similar means to obtain the true differences between the source and target items; S5: By adding a discriminant network to extract high-level feature representation, the implicit feature space is guaranteed to have good reliability; The step S5 extracts high-level feature representations using an additional layer of nonlinear perceptron as a discriminant network: The hybrid software defect prediction method based on the genetic algorithm improved by Bayesian optimization and deep neural network adopts cross entropy calculation, and takes the probability of defects included and whether the defect marking is true as loss terms.

2. The software defect prediction method integrating genetic algorithm and deep neural network according to claim 1, characterized in that: The step S2 uses the Bayesian improved genetic algorithm to select the feature subset, including the following steps: S21: Use the classic selection method, roulette wheel to select excellent individuals. The higher the fitness of an individual, the greater the probability of its being selected. S22: Using the single-point crossover method, a point is randomly selected and the two chromosomes are crossovered at this point; S23: When the genotypes of the individuals selected for crossover are consistent, the genotype of the first half of the crossover point of the first individual is kept unchanged, and the genotype of the last half is kept consistent with the flashback coding after the crossover point of the second individual; S24: Take the individuals remaining after crossover as known points and bring them into the Bayesian optimization model to obtain two parameter values, which are then added to the population as new individuals.

Citation Information

Patent Citations

  • Software defect prediction method based on genetic algorithm and random forest

    CN109977028A

  • Deep neural network structure optimization method based on fusion of prediction mechanism and genetic algorithm

    CN110490320A