A software defect category prediction method based on big data
By adopting big data and deep learning technologies in software defect prediction, multi-source domain data features are extracted and attention mechanism weighted, combined with Flort template fusion features, and finally multi-source domain fusion classifier model prediction is finally solved, which solves the problem that traditional methods are difficult to accurately predict software defect categories, and improves the accuracy and stability of prediction.
Patent Information
- Application Number
- CN202411001090.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-25
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-07-25
AI Technical Summary
Traditional software defect prediction methods are difficult to meet the requirements of accuracy and generalization capabilities in actual production, and cannot effectively predict software defect categories.
A software defect category prediction method based on big data is proposed. By collecting software running data from different source domains, calculating the weight of each source domain data, extracting data features using feature complete extraction technology, weighting features using attention mechanism, fusion features through Flort template, and finally inputting the multi-source domain fusion classifier model for prediction.
It improves the accuracy and stability of software defect category prediction, can better capture the key features of software defects, and enhances the prediction performance of the model.
Smart Images

Figure CN118965105B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of deep learning, and in particular relates to a software defect category prediction method based on big data. Background Art
[0002] In the context of the development of industrial automation and intelligence, software plays an important role in industrial production. With the widespread application of software in modern life and work, the timely identification and prediction of software defects has become crucial. The stable operation of software directly affects the efficiency and safety of the production line. Therefore, accurate prediction of software defect categories is crucial to improve production efficiency and reduce production interruptions.
[0003] Traditional software defect prediction methods often rely on empirical rules or simple statistical methods, which are difficult to meet the requirements of accuracy and generalization ability in actual production. Summary of the invention
[0004] In order to improve software quality and stability, and to effectively predict software defect categories, and to help developers discover and solve problems in a timely manner, the present invention proposes a software defect category prediction method based on big data, which specifically includes the following steps:
[0005] S1. Collect software operation data from different source domains and perform preprocessing operations on the data;
[0006] S2, calculate the weight of each source domain data, and use the feature complete extraction technology to extract the data features of each source domain from the preprocessed data;
[0007] S3, using the attention mechanism to weight the features of different source domains to obtain the first fused feature, and inputting the different domain data and their corresponding attention weights into the Flort template to obtain the second fused feature;
[0008] S4, fusing the first fusion feature and the second fusion feature to obtain a final fusion feature, and using the final fusion feature as the input of the multi-source domain fusion classifier model;
[0009] S5. The multi-source domain fusion classifier model considers the relationship between the source domain and the target source, and between the source domains to predict whether the input features contain the defect type in the target domain.
[0010] Furthermore, the weight calculation of a source domain includes:
[0011]
[0012] Among them, w i is the weight parameter learned from source domain i, n is the number of sample source domains, m is the number of samples in a source domain, and x is,j is the jth sample of source domain i, y i s,j is the label of the jth sample in source domain i; f w (·) is a neural network model with a network parameter W, which is used to input samples into the neural network to predict the defect label of the input sample; L s is the source domain loss function, L d is the domain adversarial loss function, X t represents the sample of the target domain t, and μ is the weight parameter of the loss function.
[0013] Furthermore, the process of extracting data features of each source domain using feature complete extraction technology includes:
[0014]
[0015] in, represents the data features of source domain i extracted using feature complete extraction technology, σ is the activation function, and m is the number of samples in a source domain; is the attention weight of the jth data feature in a source domain, represents the jth data feature in a source domain, for The vector representation of ; represents the attention score of the jth data feature in a source domain; w represents the learnable weight vector, w T represents the transpose of w; b1 is the bias term.
[0016] Furthermore, step S3 specifically includes the following steps:
[0017] S31, constructing similarity vectors between features of different source domains and calculating the correlation between features;
[0018] S32. Weight the features of different source domains through the attention mechanism to obtain weighted feature representation;
[0019] S33, sending the weighted feature representation to the Flort module for feature fusion to obtain a comprehensive feature representation.
[0020] Furthermore, the process of weighting the features of different source domains through the attention mechanism includes:
[0021] S i =(f i s ) T Y
[0022] S i ′=w i ×S i
[0023]
[0024] Among them, S i Represents the data feature f of source domain i i s The correlation score between the source domain and the label Y in the target domain. In the present invention, the label of the defect type in each divided source domain is obtained as the label Y; S i ′ is the weight parameter w learned from the source domain i i For S i Weighted relevance score; α i represents the attention weight of the data features in source domain i; N represents the number of source domains; Z represents the data features after attention weighting.
[0025] Furthermore, the process of fusing the different domain data and their corresponding attention weights through the Flort template to obtain the second fusion feature includes:
[0026]
[0027] Among them, F fuesd is the comprehensive feature representation after fusion; W i and b i is the weight and bias term of the linear transformation corresponding to the source domain i, σ is the activation function, and tanh represents the hyperbolic tangent function.
[0028] Furthermore, the process of finally fusing the first fusion feature and the second fusion feature to obtain the final fusion feature includes:
[0029] F combined =tanh(W Z ×Z+W f ×F fused +b2)
[0030] Among them, F combined To obtain the final fusion feature; W Z , W f is a learnable weight matrix, which is used to integrate the first fusion feature Z and the second fusion feature F fused Mapped into the same feature space; b2 is the bias term, which is used to adjust the baseline of the final fusion feature.
[0031] Furthermore, the classification process of the multi-source domain fusion classifier model includes:
[0032]
[0033] Among them, Predicition (X t) represents the input feature X on the predicted target domain t t defect category; σ is the activation function, which is used to convert the prediction result into a probability distribution function; Weighted Attention (f i s ,f j s ,Y) represents the data feature similarity between source domain i and source domain j weighted by the attention weight between source domain i and target domain data; F combined It represents the final fusion feature obtained by fusion; b3 is the bias term in the multi-source domain fusion classifier model.
[0034] Furthermore, the data feature similarity between source domain i and target domain j is weighted by the attention weight between source domain i and target domain data. Attention (f i s ,f j s ,Y) is expressed as:
[0035] Weighted Attention (f i s ,f j s ,Y)=Attention(f i s ,Y)×Similarity(f i s ,f j s )
[0036] Among them, Similarity(f i s ,f j s ) represents the data feature f in the source domain i i s and the data features f in the source domain j j s Similarity vector between i s ,Y) represents the data feature f in i i s The attention weight between the feature Y in the target domain and the feature Y in the target domain is weighted.
[0037] The present invention proposes an innovative method based on big data and deep learning for the problem of predicting software defect categories. The core of the method is to accurately predict software defect categories through different data collected from multiple source domains and model training processes. The present invention first uses CDANN and PTR modules, which helps to extract data features from different source domains, thereby better capturing the key features of software defects. This feature extraction is crucial for accurately predicting software defect categories. Then, the attention mechanism further weights the features of different source domains, which helps to better capture key features, thereby improving the prediction performance of the model. Subsequently, the training of the multi-source domain fusion classifier model further enhances the prediction of defect categories on the target domain. Finally, Flort monitors and dynamically adjusts the threshold of the classifier in real time, which helps to maintain the prediction performance of the model and enables it to adapt to the changing environment. In summary, the present invention provides a beneficial effect for identifying software category defects, thereby improving the accuracy and stability of prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 A method flow chart of a method for predicting software defect categories based on big data according to an embodiment of the present invention;
[0039] Figure 2 This is a model training flowchart of a software defect category prediction method based on big data in an embodiment of the present invention. DETAILED DESCRIPTION
[0040] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0041] The present invention proposes a software defect category prediction method based on big data, such as Figure 1 , specifically including the following steps:
[0042] S1. Collect software operation data from different source domains and perform preprocessing operations on the data;
[0043] S2. Calculate the weight of each source domain data, and use the Property Thorough Refine (PTR) technology to extract the data features of each source domain from the preprocessed data;
[0044] S3, using the attention mechanism to weight the features of different source domains to obtain the first fused feature, and inputting the different domain data and their corresponding attention weights into the Flort template to obtain the second fused feature;
[0045] S4, fusing the first fusion feature and the second fusion feature to obtain a final fusion feature, and using the final fusion feature as the input of the multi-source domain fusion classifier model;
[0046] S5. The multi-source domain fusion classifier model considers the relationship between the source domain and the target source, and between the source domains to predict whether the input features contain the defect type in the target domain.
[0047] In the present invention, different source domains refer to different defect types of software. For example, in this embodiment, different source domains can be performance defects, security defects, compatibility defects and configuration defects of the software, among which: performance defects refer to poor performance of the software during operation, such as long response time, excessive resource consumption and other problems; security defects refer to security vulnerabilities in the software, which may be maliciously exploited to cause losses; compatibility defects refer to inconsistent performance or inability to run software in different environments or platforms; configuration defects refer to the inability of the software to work properly under a specific configuration. Therefore, the number of source domains in this embodiment is 4. Those skilled in the art can divide the defect types according to actual needs, and use the data obtained by the division to perform defect detection through the method of the present invention. This embodiment only takes the above four types as examples.
[0048] In this embodiment, a method for predicting software defect categories based on big data includes the following steps:
[0049] (I) First, it is necessary to collect the software's operating data and calculate the CDANN weight (i.e., the weight of each source domain data) through the software's operating data. The weight acquisition process includes:
[0050]
[0051] Among them, w i is the weight parameter learned from source domain i, n is the number of sample source domains, m is the number of samples in a source domain, and x i s,j is the jth sample of source domain i, y i s,j is the label of the jth sample in source domain i; f w (·) is a neural network model with a network parameter W, which is used to input samples into the neural network to predict the defect label of the input sample; L s is the source domain loss function, L d is the domain adversarial loss function, X t represents the sample of the target domain t, and μ is the weight parameter of the loss function.
[0052] (ii) While implementing step (i), data features corresponding to each source domain are obtained from the historical data of each source domain based on the PTR technology. The process of obtaining the features includes the following steps:
[0053]
[0054] Among them, f i s represents the data features of source domain i extracted using feature complete extraction technology, σ is the activation function, and m is the number of samples in a source domain; is the attention weight of the jth data feature in a source domain, represents the jth data feature in a source domain, for The vector representation of ; represents the attention score of the jth data feature in a source domain; w represents the learnable weight vector, w T represents the transpose of w; b1 is the bias term.
[0055] (III) The features of different source domains are weighted through the attention mechanism to obtain weighted feature representation; the weighted feature representation is sent to the Flort module for feature fusion to obtain a comprehensive feature representation, where:
[0056] The process of weighting the features of different source domains through the attention mechanism includes:
[0057] S i =(f i s ) T Y
[0058] S i ′=w i ×S i
[0059]
[0060] Among them, S i Represents the data feature f of source domain i i s The correlation score between the label Y in the target domain, S i ′ is the weight parameter w learned from the source domain i i For S i Weighted relevance score; α i represents the attention weight of the data features in source domain i; N represents the number of source domains; Z represents the data features after attention weighting.
[0061] The process of fusing different domain data and their corresponding attention weights through the Flort template to obtain the second fusion feature includes:
[0062]
[0063] Among them, F fuesd is the comprehensive feature representation after fusion; f i s is the data feature of source domain i, α i represents the data feature attention weight in source domain i, W i and b i are the weights and bias terms of the linear transformation corresponding to the source domain i, σ is the activation function, and tanh represents the hyperbolic tangent function.
[0064] (IV) The process of finally fusing the first fusion feature and the second fusion feature to obtain the final fusion feature includes:
[0065] F combined =tanh(W Z ×Z+W f ×F fused +b2)
[0066] Among them, F combined is the final fusion feature obtained by fusion; tanh represents the hyperbolic tangent function; W Z , W f is a learnable weight matrix, which is used to integrate the first fusion feature Z and the second fusion feature F fused Mapped into the same feature space; b2 is the bias term, which is used to adjust the baseline of the final fusion feature.
[0067] (V) The final fusion features are input into the multi-source domain fusion classifier model for classification. The classification process includes:
[0068]
[0069] Among them, Predicition (X t ) represents the input feature X on the predicted target domain t t defect category; σ is the activation function, which is used to convert the prediction result into a probability distribution function; Weighted Attention (f i s ,f j s ,Y) represents the data feature similarity between source domain i and source domain j weighted by the attention weight between source domain i and target domain data; F combined It represents the final fusion feature obtained by fusion; b3 is the bias term in the multi-source domain fusion classifier model.
[0070] The similarity of data features between source domain i and source domain j weighted by the attention weight between source domain i and target domain data Attention (f i s ,f j s ,Y) is expressed as:
[0071] Weighted Attention (f i s ,f j s ,Y)=Attention(f i s ,Y)×Similarity(f i s ,f i s )
[0072] Among them, Similarity(f i s ,f j s ) represents the data feature f in the source domain i i s and the data features f in the source domain j j s Similarity vector between i s ,Y) represents the data feature f in the source domain i i s The attention weight between the feature Y in the target domain and the feature Y in the target domain is weighted.
[0073] In this embodiment, the similarity vector between features of different source domains can be calculated by calculating the cosine similarity, Euclidean distance, etc. between two feature vectors. This embodiment also provides a specific process for calculating the similarity vector between features of different source domains, including:
[0074]
[0075] Among them, Similarity(X i ,X j ) represents the data feature f in the source domain i i s and the data features f in the source domain j j s The similarity vector between ij represents the angle between the feature vectors of source domain i and source domain j; r represents the imaginary unit.
[0076] In this embodiment, the data feature f in the source domain i i s The attention weight between the data feature Y in the target domain (f i s ,Y) is expressed as:
[0077]
[0078] Where Y is the feature vector of the target domain label, which is obtained by one-hot encoding the target domain label; |·| represents the modulus of the vector.
[0079] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A software defect category prediction method based on big data, characterized in that: The specific steps include: S1. Collect software operation data from different source domains and perform preprocessing operations on the data; S2. Calculate the weight of each source domain data, and use the feature complete extraction technology to extract the data features of each source domain from the preprocessed data. The process includes: Among them, f i s represents the data features of source domain i extracted using feature complete extraction technology, σ is the activation function, and m is the number of samples in a source domain; is the attention weight of the jth data feature in a source domain, represents the jth data feature in a source domain, for The vector representation of ; represents the attention score of the jth data feature in a source domain; w represents the learnable weight vector, w T represents the transpose of w, w i is the weight parameter learned from the source domain i; b1 is the bias term; S3. Use the attention mechanism to weight the features of different source domains to obtain the first fused feature, input the different domain data and their corresponding attention weights into the Flort template for fusion to obtain the second fused feature. The process of obtaining the second fused feature includes: Among them, F fuesd is the comprehensive feature representation after fusion, N represents the number of source domains; f i s is the data feature of source domain i, α i represents the data feature attention weight in source domain i, W i and b i is the weight and bias term of the linear transformation corresponding to the source domain i, σ is the activation function, and tanh represents the hyperbolic tangent function; S4, fusing the first fusion feature and the second fusion feature to obtain a final fusion feature, and using the final fusion feature as the input of the multi-source domain fusion classifier model; S5. The multi-source domain fusion classifier model considers the relationship between the source domain and the target source, and between the source domains to predict whether the input features contain the defect type in the target domain.
2. The method for predicting software defect categories based on big data according to claim 1, characterized in that: The weight calculation of a source domain includes: Among them, w i is the weight parameter learned from source domain i, m is the number of samples in a source domain, x i s,j is the jth sample of source domain i, y i s,j is the label of the jth sample in source domain i; f w (·) is a neural network model with a network parameter W, which is used to input samples into the neural network to predict the defect label of the input sample; L s is the source domain loss function, L d is the domain adversarial loss function, X t represents the sample of the target domain t, and μ is the weight parameter of the loss function.
3. The method for predicting software defect categories based on big data according to claim 1, characterized in that: The process of weighting the features of different source domains through the attention mechanism include: S i =(f i s ) T Y S i ′=w i ×S i Among them, S i Represents the data feature f of source domain i i s The correlation score between the label Y in the target domain, S i ′ is the weight parameter w learned from the source domain i i For S i The weighted relevance score; Z represents the data feature after attention weighting.
4. The method for predicting software defect categories based on big data according to claim 1, characterized in that: The process of finally fusing the first fusion feature and the second fusion feature to obtain the final fusion feature includes: F combined =tanh(W Z ×Z+W f ×F fused +b2) Among them, F combined To obtain the final fusion feature; W Z , W f is a learnable weight matrix, which is used to integrate the first fusion feature Z and the second fusion feature F fused Mapped into the same feature space; b2 is the bias term used to adjust the baseline of the final fusion feature.
5. The method for predicting software defect categories based on big data according to claim 1, characterized in that: The classification process of the multi-source domain fusion classifier model includes: Among them, Predicition (X t ) represents the input feature X on the predicted target domain t t defect category; σ is the activation function, which is used to convert the prediction results into a probability distribution function; represents the data feature similarity between source domain i and source domain j weighted by the attention weight between source domain i and target domain data; F combined It represents the final fusion feature obtained by fusion; b3 is the bias term in the multi-source domain fusion classifier model.
6. The method for predicting software defect categories based on big data according to claim 1, characterized in that: The similarity of data features between source domain i and source domain j weighted by the attention weight between source domain i and target domain data It is expressed as: in, Represents the data feature f in the source domain i i s and the data features f in the source domain j j s Similarity vector between i s ,Y) represents the data feature f in i i s The attention weight between the feature Y in the target domain and the feature Y in the target domain is weighted.
Citation Information
Patent Citations
Cross-domain small sample defect target detection method
CN115587969A
Bearing cross-domain fault diagnosis method based on multi-modal attention adaptive network
CN117972307A