Multi-feature fusion software defect positioning method and device, equipment and medium
By adopting a multi-feature fusion method in software defect positioning and using multi-gated network to control feature weights, the problem of inefficient defect positioning in the prior art is solved, and more efficient and accurate defect positioning is achieved.
Patent Information
- Application Number
- CN202510169059.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art is difficult to efficiently handle large-scale code bases in software defect location, traditional methods are inefficient and error-prone, and deep learning models are insufficient in complex defect location scenarios, especially in processing multiple features and multimodal data.
A multi-feature fusion software defect positioning method based on defect categories is adopted to construct a multi-feature defect positioning model through data acquisition, unsupervised clustering and multi-task learning, and a multi-gated network is added to control the weight size of features under different defect categories.
It significantly improves the accuracy and efficiency of defect positioning, can locate defects more precisely, reduce the dependence of manual analysis, and improve debugging and maintenance efficiency during software development.
Smart Images

Figure CN120216336A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of software defect localization problems, and particularly relates to a multi-feature fusion software defect localization method, device, equipment and medium. Background Art
[0002] In recent years, the application of machine learning and deep learning technologies in software defect localization has attracted wide attention. Software defect localization, as a key link in software quality assurance, involves complex tasks such as code analysis, defect identification, and intelligent recommendation. With the expansion of software system scale and the increase of code complexity, traditional defect localization methods are difficult to efficiently meet the defect detection requirements in large-scale code libraries. They often rely on manual analysis, which is inefficient and error-prone. The progress of deep learning, especially natural language processing and graph neural network technologies, provides new intelligent solutions for software defect localization, which can enhance the accuracy of localization while improving efficiency.
[0003] Although the application of deep learning in software defect localization has broad prospects, it also faces some challenges, such as the complexity of defect report and source code matching, the generalization ability of the model, and the need to efficiently process large-scale code libraries. In practical applications, how to improve the performance of the model in complex defect localization scenarios, especially the ability to process multiple features and multi-modal data, remains an important research direction. Through continuous technological innovation and practical exploration, deep learning and natural language processing are expected to further improve the intelligent level of software defect localization, provide more efficient solutions for debugging and maintenance in the software development process, and promote the quality improvement and R & D efficiency optimization of the software industry. Summary of the Invention
[0004] To solve at least one of the technical problems existing in the prior art to a certain extent, an object of the present invention is to provide a multi-feature fusion software defect localization method, device, equipment and medium based on defect categories.
[0005] The first technical solution adopted by the present invention is:
[0006] A multi-feature fusion software defect localization method includes the following steps:
[0007] Perform data collection, and construct a defect localization data set based on all source code files in defect reports, corresponding defect files, and defect versions;
[0008] Perform unsupervised clustering on defect reports to divide defects into different categories;
[0009] Construct a multi-feature defect localization model based on multi-task learning, and add a multi-gating network to the defect localization model to control the weight sizes of different features under different defect categories;
[0010] Based on the clustering division results, the defect localization model is supervised trained using the defect localization data set;
[0011] Obtain the defect reports and project source code files with labeled categories, and input them into the trained defect localization model to obtain the defect probability ranking results of all files.
[0012] Furthermore, the unsupervised clustering of the defect reports to divide the defects into different categories includes:
[0013] Based on the text information of the defect reports, use the text clustering algorithm to perform unsupervised clustering on the defect reports to divide the defect categories.
[0014] Furthermore, the use of the text clustering algorithm to perform unsupervised clustering on the defect reports to divide the defect categories includes:
[0015] Use the Latent Dirichlet Allocation for text topic modeling and perform clustering through K-Means; the core idea of the Latent Dirichlet Allocation is to represent each document as a probability distribution of several topics, and each topic is again a probability distribution of words:
[0016] Construct a joint distribution by sampling the document topic distribution and sampling the topic word distribution:
[0017]
[0018] In the formula, P(·) represents probability; α and β respectively represent sampling the document topic distribution and sampling the topic word distribution; W represents the set of all words in all documents, Z represents the topic distribution of each word, θ represents the topic distribution of the document, represents the word distribution of the topic; z d,n represents the topic of the nth word in the dth document, w d,n represents the nth word in the dth document; θ d represents the topic distribution of document d, represents the word distribution of topic k; M represents the total number of documents, N d represents the number of words in document d, and K represents the total number of topics;
[0019] Estimate θ through variational inference or Gibbs sampling d and Obtain the topic distribution θ d as the feature vector of the document, representing the probability distribution of the document on K topics:
[0020] θ d =[θ d,1 ,θ d,2 ,···,θ d,k (2)
[0021] where θ d,k represents the probability distribution of the d-th document on the k-th topic;
[0022] Based on the topic distribution θ of each document d , use the K-Means clustering algorithm to divide it into C clusters, initialize C centroids, and iteratively assign the remaining documents to the nearest centroid. The goal is to minimize the sum of squared errors within the clusters:
[0023]
[0024] where is the indicator function, y i is the cluster label of document i, μ j is the centroid of the j-th cluster, and θ i is the topic distribution of the i-th document;
[0025] Finally, label the defect reports as C different categories.
[0026] Furthermore, for the defect localization model constructed based on multi-task learning, a multi-gating network is added to control the weight sizes of different features under different defect categories, including:
[0027] Build a multi-feature extraction network, and use corresponding graph neural networks to extract features for multiple graph representations; the graph representations include the Abstract Syntax Tree (AST), Control Flow Graph (CFG), and Program Dependency Graph (PDG) of the code. Through the multi-feature extraction network, features of the code in terms of syntax structure, execution order (functional semantics), and inter-statement dependency relationships can be obtained respectively. These features can not only help the model identify syntax errors in the code and defects in functional semantics such as loop assignments, but also provide semantic information at the code abstraction level, enabling the model to learn high-level semantic features of the code and further alleviating the semantic gap between defect reports and code files;
[0028] For the extraction of code text semantics and defect report text semantics, perform semantic embedding and combine with CNN to extract local semantic information of the text;
[0029] Design a multi-gating network corresponding to each defect category and train the weight parameters of multiple features under different categories.
[0030] Furthermore, the design of the multi-gating network corresponding to each defect category and the training of the weight parameters of multiple features under different categories include:
[0031] For each category t, the gating network generates a weight vector g based on the input features (t) to adjust the feature importance of this category;
[0032] Based on the weight vector g obtained from the gating network (t) calculate feature weighting to generate the feature vector h for category t (t) ;
[0033] For each category t, input the weighted feature vector h (t) into the corresponding task network to calculate the output y of the task (t) .
[0034] Furthermore, the calculation formula for the weight vector g (t) is:
[0035] g (t) = σ(W (t) br + b (t) ) (4)
[0036] In the formula, is the weight matrix of the gating network for category t, is the bias vector, br represents the feature vector of the defect report text, and σ(·) represents the activation function, which can be the Sigmoid or Softmax function;
[0037] The calculation formula for the feature vector h (t) is:
[0038] h (t) = g (t) ⊙ F (5)
[0039] In the formula, represents an n - row and d - column multi - feature matrix composed of graph features and semantic features, and ⊙ represents the element - by - element multiplication operation;
[0040] The calculation formula for the output y (t) is:
[0041] y (t) = f(h (t) ) (6)
[0042] where f(·) represents the final defect probability prediction function.
[0043] Furthermore, the input of the defect localization model is the text of the defect report and the source code file, and the output is the defect probability of the source code file;
[0044] During the model training process, set the loss function to the FocalLoss function:
[0045] L loc =-α t (1 - p t ) γ log(p t ) (7)
[0046] Where p t is the predicted probability, α t represents the balance parameter, which is used to control the weights of positive and negative samples in the loss, and γ is the focusing parameter, which controls the weights of easy and hard samples.
[0047] The second technical solution adopted by the present invention is:
[0048] A multi - feature fusion software defect localization device, comprising:
[0049] A data collection and pre - processing module, which is used for data collection and constructs a defect localization data set based on all source code files in the defect report, the corresponding defect files and the defect version;
[0050] A defect clustering module, which is used for unsupervised clustering of defect reports to divide the defects into different categories;
[0051] A model construction module, which is used for constructing a multi - feature defect localization model based on multi - task learning, and adding a multi - gate network to the defect localization model to control the weight sizes of different features under different defect categories;
[0052] A model training module, which is used for supervised training of the defect localization model using the defect localization data set based on the clustering division result;
[0053] A defect file sorting module, which is used for obtaining defect reports and project source code files with labeled categories, and inputting them into the trained defect localization model to obtain the defect probability sorting results of all files.
[0054] The third technical solution adopted by the present invention is:
[0055] An electronic device, the electronic device includes a processor and a memory, and at least one instruction, at least one program, a code set or an instruction set is stored in the memory, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement a multi - feature fusion software defect localization method as described above.
[0056] The fourth technical solution adopted by the present invention is:
[0057] A computer-readable storage medium stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement the multi-feature fusion software defect localization method as described above.
[0058] The fifth technical solution adopted by the present invention is:
[0059] A computer program product or a computer program includes computer instructions stored in a computer-readable storage medium. A processor of a computer device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions to enable the computer device to execute the above method.
[0060] The beneficial effects of the present invention are as follows: By subdividing defect categories to control feature weights, the present invention can effectively improve the accuracy of the defect localization method; and by fusing various code structure representations, more sufficient feature information is provided for the software defect localization model. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following introduces the drawings of the related technical solutions in the embodiments of the present invention or the prior art. It should be understood that the drawings in the following introduction are only for conveniently and clearly presenting some embodiments of the technical solutions in the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative efforts.
[0062] Figure 1 It is a flowchart of the steps of a multi-feature fusion software defect localization method in an embodiment of the present invention;
[0063] Figure 2 It is a schematic flowchart of a multi-feature fusion software defect localization method based on defect categories in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0064] The embodiments of the present invention are described in detail below. Examples of the embodiments are shown in the drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the drawings are exemplary only for explaining the present invention and should not be construed as a limitation of the present invention. For the step numbers in the following embodiments, they are only set for convenience of description and illustration, and no limitation is imposed on the order between the steps. The execution order of each step in the embodiments can be adjusted adaptively according to the understanding of those skilled in the art.
[0065] In the description of the present invention, it should be understood that for the orientation description, such as the orientation or positional relationship indicated by up, down, front, back, left, right, etc., is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to the present invention.
[0066] In the description of the present invention, the meaning of "several" is one or more, the meaning of "multiple" is more than two, and understandings such as "greater than", "less than", "exceeding", etc. do not include the present number, and understandings such as "above", "below", "within", etc. include the present number. If there is a description of "first" and "second", it is only for the purpose of distinguishing technical features, and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or implicitly indicating the sequence relationship of the indicated technical features.
[0067] In the description of the present invention, unless otherwise clearly defined, words such as "set", "install", "connect", etc. should be understood in a broad sense, and those skilled in the art can reasonably determine the specific meanings of the above words in the present invention in combination with the specific content of the technical solution.
[0068] Aiming at the existing technical problems, the present invention provides a multi-feature fusion software defect localization solution based on defect categories. First, data collection is carried out to construct a defect localization data set based on defect reports, defect files, and source code. Then, a defect clustering model is constructed to perform unsupervised clustering on defect reports and divide the defects into different categories. Next, a multi-feature defect localization model is constructed based on multi-task learning, and a multi-gate network is added to this model to control the weights of features under different categories. In the defect localization model, supervised learning is performed on the classified defect reports and source code in combination with clustering information. Finally, the clustering and localization models are integrated to construct an end-to-end defect localization model, providing users with accurate defect localization and defect probability ranking results. This method significantly improves the accuracy and user-friendliness of defect localization by subdividing defect categories and adopting an end-to-end design.
[0069] Embodiment 1
[0070] As Figure 1 and Figure 2 shown, this embodiment provides a multi-feature fusion software defect localization method based on defect categories, including the following steps:
[0071] S1. Perform data collection to construct a defect localization data set based on defect reports, corresponding defect files, and all source code files in the defect version.
[0072] Specifically, collect the required dataset from the publicly available defect tracking system, such as defect reports, defect files, and all source code files in the corresponding version. Filter, clean, and preprocess the collected dataset, including text tokenization, stop word removal, stemming, etc.
[0073] Exemplarily, the defect report is the defect data provided by any publicly available defect tracking system; the defect file is the file modified in the corresponding commit of the software project described in the defect report; all source code files of the defect corresponding version are all source code files in the corresponding commit of the software project described in the defect report.
[0074] S2. Perform unsupervised clustering on the defect reports to divide the defects into different categories.
[0075] Based on the text information of the defect reports, use a text clustering algorithm to perform unsupervised clustering on the defect reports to divide the defect categories, which can be clustered based on traditional information retrieval methods or deep neural networks. Specifically, Latent Dirichlet Allocation (LDA) can be used for text topic modeling and K-Means for clustering. The core idea of LDA is to represent each document as a probability distribution of several topics, and each topic is again a probability distribution of words:
[0076] Construct a joint distribution by sampling the document topic distribution and sampling the topic word distribution:
[0077]
[0078] where P(·) represents probability; α and β represent the sampling document topic distribution and sampling topic word distribution respectively; W represents the set of all words in all documents, Z represents the topic distribution of each word, θ represents the topic distribution of the document, represents the word distribution of the topic; z d,n represents the topic of the nth word in the dth document, w d,n represents the nth word in the dth document; θ d represents the topic distribution of document d, represents the word distribution of topic k; M represents the total number of documents, N d represents the number of words in document d, and K represents the total number of topics.
[0079] Estimate θ d and obtain the topic distribution θ d as the feature vector of the document, representing the probability distribution of the document on K topics:
[0080] θ d =[θd,1 , θ d,2 , ···, θ d,k (2)
[0081] Based on the topic distribution θ of each document d , the K-Means clustering algorithm is used to divide them into C clusters, initialize C centroids, and iteratively assign the remaining documents to the nearest centroid. The goal is to minimize the sum of squared errors within the clusters:
[0082]
[0083] where is the indicator function, y i is the cluster label of document i, μ j is the centroid of the j-th cluster, and θ i is the topic distribution of the i-th document.
[0084] Through the above text clustering method, the defect reports are finally labeled as C different categories.
[0085] S3. Build a multi-feature defect localization model based on multi-task learning, and add a multi-gating network to the defect localization model to control the weight sizes of different features under different defect categories.
[0086] Build a multi-feature extraction network, and use corresponding graph neural networks for feature extraction for various graph representations;
[0087] For the abstract syntax tree (AST), use the graph attention network (GAT) for feature extraction to obtain the feature vector of the code at the AST level;
[0088] For heterogeneous graphs such as the control flow graph (CFG), use the relational graph convolutional network (RGCN) for feature extraction to obtain the feature vector of the code at the CFG level;
[0089] For the extraction of the semantic information of code text and defect report text, use methods such as Word2Vec for semantic embedding, and combine CNN to extract the local semantic information of the text;
[0090] Design a multi-gating network corresponding to each defect category, and train the weight parameters of multi-features under different categories. Specifically:
[0091] For each category t, the gating network generates a weight vector g (t) based on the input features to adjust the feature importance of this category. The gating weight vector can be calculated by the following formula:
[0092] g (t) = σ(W (t) br + b (t)) (4)
[0093] Among them is the gating network weight matrix for class t, is the bias vector, br represents the feature vector of the defect report text, and σ(·) represents the activation function, which can adopt the Sigmoid or Softmax function.
[0094] Calculate the feature weighting based on the weight parameters obtained from the gating network to generate the feature vector h for class t (t) :
[0095] h (t) = g (t) ⊙F (5)
[0096] Among them, g (t) represents the weight parameter for class t, represents the multi-feature matrix of n rows and d columns composed of graph features and semantic features, and ⊙ represents the element-wise multiplication operation.
[0097] For each class t, input the weighted feature h (t) into the corresponding task network to calculate the output y of the task (t) :
[0098] y (t) = f(h (t) ) (6)
[0099] Among them, f(·) represents the final defect probability prediction function.
[0100] S4. Based on the clustering division results, use the defect localization dataset to perform supervised training on the defect localization model.
[0101] On the defect localization model, perform supervised learning on the defect reports and project source code datasets with labeled categories. The input of the defect localization model is the text of the defect report and the source code file, and the output is the defect probability of the source code file. To overcome the problem of class imbalance, set the loss function to the FocalLoss function:
[0102] L loc = -α t (1 - p t ) γ log(p t ) (7)
[0103] Among them, p t is the predicted probability, α t represents the balance parameter, which is used to control the weights of positive and negative samples in the loss, and γ is the focusing parameter, which controls the weights of easy and difficult samples.
[0104] S5. Obtain defect reports with labeled categories and project source code files, and input them into the trained defect localization model to obtain the defect probability ranking results of all files.
[0105] Provide defect reports with labeled categories and project source code files for the software defect localization model described above to obtain the defect probability ranking results of all files.
[0106] To verify the effectiveness of the method of the present invention, the following experiments were designed: comparative experiments on defect localization performance were carried out on the classic defect localization datasets Tomcat and AspectJ, and the performance was compared with traditional defect localization method tools DNNLOC and BugLocator to verify the effectiveness and advancement of the method of this article. This experiment uses multi-fold cross-validation, aiming to obtain the average localization performance of the method under different training data distributions, making the experimental results more reliable and persuasive.
[0107] Table 1 Comparison of localization performance of each method on Tomcat and AspectJ datasets
[0108]
[0109] Table 1 shows the performance comparison results with traditional defect localization methods on Tomcat and AspectJ datasets. From the data in Table 1, it can be seen that the method proposed in the present invention is superior to widely recognized software defect localization methods such as DNNLOC and BugLocator in all indicators. In terms of the Top1 indicator of the Tomcat dataset, the method of the present invention is 5.4% and 16.4% higher than DNNLOC and BugLocator respectively, which means that this method has a higher probability of directly ranking the defective file as the first in the list of defect suspects. At the same time, the higher MRR indicator also shows that this method has a higher average ranking for defective files, which can greatly reduce the time and energy consumption for developers to further investigate; from the MAP indicator, this method can find more defective files among the top k sorted files, with a higher recall rate; and on the AspectJ dataset, the Top1 and MRR and other indicators of the method of the present invention are relatively better, with strong competitiveness.
[0110] Generally speaking, in view of the problems of diverse defect types, large code scale, and insufficient positioning accuracy in traditional defect localization methods, the present invention proposes a multi-feature fusion software defect localization method based on defect categories. First, by subdividing the defect categories of defect reports, the supplementary category information can effectively guide model training, and the most suitable feature combination is selected from multiple features for defect localization. This method reduces the interference of the remaining samples on the model parameters, enabling the model to achieve more refined and specialized defect localization. Secondly, this method fuses multiple graph representations to supplement the features of the code from multiple perspectives, helping the defect localization model learn more comprehensive features of the code, improving the accuracy of defect localization, and verifying the effectiveness and advancement of this method through experiments. Finally, the present invention explicitly fuses two models of defect clustering and defect localization to construct an end-to-end software defect localization model, which can better meet the actual application requirements.
[0111] Embodiment 2
[0112] This embodiment provides a multi-feature fusion software defect localization device, including:
[0113] A data collection and preprocessing module, configured to perform data collection and construct a defect localization data set based on all source code files in defect reports, corresponding defect files, and defect versions;
[0114] A defect clustering module, configured to perform unsupervised clustering on defect reports to divide defects into different categories;
[0115] A model construction module, configured to construct a multi-feature defect localization model based on multi-task learning, and add a multi-gating network to the defect localization model to control the weight sizes of different features under different defect categories;
[0116] A model training module, configured to perform supervised training on the defect localization model using the defect localization data set based on the clustering division result;
[0117] A defect file sorting module, configured to obtain defect reports and project source code files with labeled categories, and input them into the trained defect localization model to obtain the defect probability sorting results of all files.
[0118] Since this device is a multi-feature fusion software defect localization device of an embodiment of the present invention, and the principle of this device to solve problems is similar to that of this method, the implementation of this device can refer to the implementation process of the above method embodiment, and the repeated parts will not be described again.
[0119] Embodiment 3
[0120] An embodiment of the present invention further provides an electronic device, which includes a processor and a memory. At least one instruction, at least one program, a code set, or an instruction set is stored in the memory and is loaded and executed by the processor to implement Figure 1 a multi-feature fusion software defect localization method as shown.
[0121] It can be understood that the memory may include a random access memory (RAM) and may also include a read-only memory. Optionally, the memory includes a non-transitory computer-readable storage medium. The memory is used to store instructions, programs, codes, code sets, or instruction sets. The memory may include a program storage area and a data storage area. Among them, the program storage area may store instructions for implementing an operating system, instructions for at least one function, instructions for implementing the above method embodiments, etc.; the data storage area may store data created according to the use of the server, etc.
[0122] The processor may include one or more processing cores. The processor uses various interfaces and lines to connect various parts within the entire server. By running or executing instructions, programs, code sets, or instruction sets stored in the memory, and by calling data stored in the memory, the processor executes various functions of the server and processes data. Optionally, the processor may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor may integrate a combination of one or more of a central processing unit (CPU) and a modem. Among them, the CPU mainly processes the operating system and application programs, etc.; the modem is used to process wireless communication. It can be understood that the above modem may not be integrated into the processor and may be implemented separately by a chip.
[0123] Since this electronic device is the electronic device corresponding to the multi-feature fusion software defect localization method of the embodiment of the present invention, and the principle of the electronic device for solving problems is similar to that of the method, the implementation of this electronic device can refer to the implementation process of the above method embodiment, and the repeated parts will not be described again.
[0124] Embodiment 4
[0125] An embodiment of the present invention further provides a computer-readable storage medium, in which at least one instruction, at least one program, a code set or an instruction set is stored, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement Figure 1 a multi-feature fusion software defect localization method as shown
[0126] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program, and this program can be stored in a computer-readable storage medium. The storage medium includes a read-only memory (ROM), a random access memory (RAM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electrically-erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc memories, magnetic disc memories, tape memories, or any other computer-readable medium capable of carrying or storing data.
[0127] Since this storage medium is the storage medium corresponding to the multi-feature fusion software defect localization method of the embodiment of the present invention, and the principle of solving problems by this storage medium is similar to that of this method, the implementation of this storage medium can refer to the implementation process of the above method embodiment, and the repeated parts will not be described again.
[0128] Embodiment 5
[0129] In some possible embodiments, various aspects of the method of the embodiments of the present invention can also be implemented in the form of a program product, which includes program code. When the program product runs on a computer device, the program code is used to cause the computer device to execute the steps of a multi-feature fusion software defect localization method according to various exemplary embodiments described above in this specification. Among them, the executable computer program code or "code" for executing each embodiment can be written in a high-level programming language such as C, C++, C#, Smalltalk, Java, JavaScript, Visual Basic, structured query language (e.g., Transact-SQL), Perl, or in various other programming languages.
[0130] It should be understood that each part of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits with logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits with appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0131] In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0132] The above embodiments are only for illustrating the technical concept and characteristics of the present invention, and their purpose is to enable those of ordinary skill in the art to understand the content of the present invention and implement it accordingly. It should not be used to limit the protection scope of the present invention. Any equivalent changes or modifications made according to the essence of the content of the present invention should be covered within the protection scope of the present invention.
Claims
1. A multi-feature fusion software defect location method, characterized in that: The following steps are involved: Conduct data collection and build a defect location dataset based on defect reports, corresponding defect files, and all source code files in the defective version; Perform unsupervised clustering of defect reports to classify defects into different categories; A multi-feature defect localization model is built based on multi-task learning. A multi-gating network is added to the defect localization model to control the weights of different features under different defect categories. Based on the clustering results, the defect localization data set is used to conduct supervised training on the defect localization model; Obtain the defect reports and project source code files of the labeled categories, input them into the trained defect localization model, and obtain the defect probability ranking results of all files.
2. A multi-feature fusion software defect location method according to claim 1, characterized in that: The unsupervised clustering of defect reports is performed to classify defects into different categories, including: Based on the text information of defect reports, a text clustering algorithm is used to perform unsupervised clustering of defect reports to divide defect categories.
3. A multi-feature fusion software defect location method according to claim 2, characterized in that: The text clustering algorithm is used to perform unsupervised clustering on the defect report to classify the defect categories, including: The latent Dirichlet distribution is used for text topic modeling and clustering through K-Means. The core idea of the latent Dirichlet distribution is to represent each document as a probability distribution of several topics, and each topic is a probability distribution of words: a joint distribution is constructed by sampling document topic distribution and sampling topic word distribution: Where P(·) represents probability; α and β represent the sampled document topic distribution and sampled topic word distribution respectively; W represents the word set of all documents, Z represents the topic distribution of each word, and θ represents the topic distribution of the document. The word distribution representing the topic; z d,n represents the topic of the nth word in the dth document, w d,n represents the nth word in the dth document; θ d represents the topic distribution of document d, represents the word distribution of topic k; M represents the total number of documents, N d represents the number of words in document d, and K represents the total number of topics; Estimate θ via variational inference or Gibbs sampling d and Get the topic distribution θ d As the feature vector of the document, it represents the probability distribution of the document on K topics: i d =[θ d,1 ,i d,2 ,···,θ d,k ] (2) In the formula, θ d,k represents the probability distribution of the kth topic in the dth document; Based on the topic distribution θ of each document d , use the K-Means clustering algorithm to divide into C clusters, initialize C centroids, and iteratively assign the remaining documents to the nearest centroids, with the goal of minimizing the sum of squared errors within the cluster: In the formula, is the indicator function, y i is the cluster label of document i, μ j is the centroid of the jth cluster, θ i is the topic distribution of the i-th document; Finally, the defect reports are labeled into C different categories.
4. The multi-feature fusion software defect location method according to claim 1, characterized in that: The multi-feature defect localization model is constructed based on multi-task learning, and a multi-gating network is added to the defect localization model to control the weights of different features under different defect categories, including: Build a multi-feature extraction network and use the corresponding graph neural network to extract features for various graph representations; Semantic embedding is performed to extract the semantics of code text and defect report text, and local semantic information of the text is extracted by combining CNN. Based on the above semantic information, the model can predict the defect probability of the code file by measuring the semantic similarity between the defect report and the code file. A multi-gated network corresponding to each defect category is designed, and the weight parameters of multiple features under different categories are trained to control the importance of different features.
5. A multi-feature fusion software defect location method according to claim 4, characterized in that: The multi-gated network designed to correspond one-to-one to the defect categories is used to train weight parameters of multiple features under different categories, including: For each category t, the gating network generates a weight vector g based on the input features. (t) , to adjust the feature importance of this category; The weight vector g obtained based on the gating network (t) Calculate feature weights to generate feature vector h for category t (t) ; For each category t, the weighted feature vector h (t) Input into the corresponding task network and calculate the output y of the task (t) .
6. A multi-feature fusion software defect location method according to claim 5, characterized in that: The weight vector g (t) The calculation formula is: g (t) =σ(W (t) br+b (t) ) (4) Where W (t) is the gating network weight matrix of category t, b (t) is the bias vector, br represents the feature vector of the defect report text, and σ(·) represents the activation function; The feature vector h (t) The calculation formula is: h (t) =g (t) ⊙F (5) Where F represents a multi-feature matrix of n rows and d columns consisting of graph features and semantic features, and ⊙ represents the bit-by-bit multiplication operation of the elements; The output y (t) The calculation formula is: y (t) =f(h (t) ) (6) Where f(·) represents the final defect probability prediction function.
7. The multi-feature fusion software defect location method according to claim 1, characterized in that: The input of the defect localization model is the text of the defect report and the source code file, and the output is the defect probability of the source code file; During model training, the loss function is set to the FocalLoss function: L loc =-a t (1-p t ) γ log(p t ) (7) In the formula, p t is the predicted probability, α t represents the balance parameter, which is used to control the weight of positive and negative samples in the loss, and γ is the focusing parameter.
8. A multi-feature fusion software defect location device, characterized in that: include: The data collection and preprocessing module is used to collect data and build a defect localization dataset based on defect reports, corresponding defect files, and all source code files in the defective version; Defect clustering module, which is used to perform unsupervised clustering of defect reports to classify defects into different categories; The model building module is used to build a multi-feature defect localization model based on multi-task learning. A multi-gated network is added to the defect localization model to control the weights of different features under different defect categories. A model training module is used to perform supervised training on the defect location model using the defect location dataset based on the clustering results; The defect file sorting module is used to obtain defect reports and project source code files of labeled categories, and input the trained defect localization model to obtain the defect probability sorting results of all files.
9. An electronic device, characterized in that: The electronic device includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the method described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The storage medium stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the method according to any one of claims 1 to 7.