An Automatic Recovery Method for Traceability Links Based on the Combination of Multi-Source Information
Through the method of combining multi-source information, the cross-attention mechanism and pre-trained model are used to integrate the metadata, text and code information of defect reports, and the problem of inaccurate link recovery in the existing technology is solved, achieving more efficient link recovery and software development traceability.
Patent Information
- Application Number
- CN202310254576.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-16
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2043-03-16
AI Technical Summary
When building links between defect reports and code submissions, the existing automated link recovery methods have shortcomings such as insufficient data mining, insufficient semantic understanding ability, insufficient code understanding ability, and large data training requirements, resulting in inaccurate link recovery and a lot of manpower.
Using a multi-source information combination method, the metadata, text information and code information of defect reports are fused through feature representation and cross-attention mechanism, and vectorized representation is used for pre-trained models such as RoBERTa and Unixcoder, and classified with the help of a full connection layer to achieve traceability link recovery between defect reports and code submission.
It realizes more accurate semantic understanding and code understanding, improves the accuracy of link recovery, reduces the demand for data volume, and improves traceability and observability in the software development process.
Smart Images

Figure CN116401289B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent software engineering, and particularly to an automatic recovery method for traceability links based on the combination of multi-source information. Background Art
[0002] The software repository contains a large amount of data on software system development, which contains rich information on software updates and evolution. Developers use the software repository data to carry out various software maintenance tasks to promote software iteration. At the same time, researchers make full use of the software repository data to carry out empirical research to provide valuable guiding experience for software engineering development practice and help software development teams improve software development quality.
[0003] The traceability links between various artifacts in the software repository play an important role in the software maintenance process. Among them, the traceability link between defect reports (Issues) and code commits (Commits) (i.e., Issue-Commit link) closely associates code with defects, functions, etc., improving developers' observability of software performance and software status and strengthening the traceability of software functions and modules. Therefore, the traceability link between defect reports and code commits is widely analyzed in various software maintenance tasks to prepare data for various software maintenance tasks and build software maintenance models (such as defect localization, defect assignment, defect prediction, function localization, etc.). However, since defect reports and code commits are stored separately in defect tracking systems (such as JIRA, Bugzilla, etc.) and version control systems (such as SVN, Git, etc.), the link between them is often manually added with the specified defect report ID in the code commit log by developers, which results in a large number of traceability links between defect reports and code commits being lost during the software maintenance process due to reasons such as developers' negligence. Since the relationship between defect reports and code commits is many-to-many, that is, a certain defect report may be fixed by multiple code commits, and a certain code commit may fix multiple defect reports at the same time, this also makes it not only require a lot of effort from domain experts but also be error-prone to manually complete the link recovery between defect reports and code commits, which poses a huge obstacle to the data preparation and model construction of various software maintenance tasks.
[0004] In order to improve the traceability of software system development and thus improve software development efficiency, many automatic link recovery methods have been proposed to generate traceability links between defect reports and code commits. According to the technologies involved, these methods can be divided into link recovery technologies based on heuristic methods, link recovery technologies based on machine learning, and link recovery technologies based on deep learning.
[0005] Wu et al. proposed a heuristic rule-based method ReLink, which is a link recovery technique based on heuristic methods. It determines whether there is a link between a given defect report and a code commit by formulating some heuristic rules. Specifically, they believe that the defect repair time should be close to the commit time of the code modification, the change log of the code commit and the description information of the error report have similar text similarity, and the defect report creator is usually the submitter of the defect repair, etc. The specific threshold criteria are determined by learning in the correct link set. If a given link pair meets three rules simultaneously during the link recovery process, it is recognized as a correct link. However, the heuristic rule-based methods represented by ReLink have serious deficiencies such as insufficient accuracy and inadequate utilization of text data.
[0006] Mazrae et al. proposed a method hybrid-linker that combines text data and non-text data, which is a link recovery technique based on machine learning. It constructs a hybrid model composed of a linear combination of a text component, a non-text component, and a classifier. Specifically, this method uses TF-IDF to convert text into feature vectors and inputs them into a GB classifier for classification. The non-text components are classified by means of an XGBoost classifier and a GB classifier, and finally the classification results are integrated to determine whether the link pair is a correct link. However, the machine learning-based methods represented by hybrid-linker only judge similarity by text similarity for text data and cannot deeply mine semantic associations. At the same time, the feature information of the text component and the non-text component cannot be fully interactively combined, and the classification processes of each component are relatively independent.
[0007] Ruan et al. proposed a semantic-enhanced link recovery method DeepLink, which is a link recovery technique based on deep learning. It uses word embedding technology (Word2vec) and a recurrent neural network (RNN) model to obtain the vector representation of the link pair and uses cosine similarity to determine whether there is a link between the defect report and the code commit. However, the existing deep learning-based methods represented by DeepLink still have deficiencies such as insufficient semantic understanding ability and difficulty in ensuring the model effect with the amount of training data.
[0008] Generally speaking, the existing automated link recovery methods provide a good research foundation for building links between defect reports and code commits, but there are still deficiencies in the actual development process, such as insufficient data mining, insufficient semantic understanding ability, insufficient code understanding ability, and high data training requirements, which cannot meet the actual application needs. Summary of the Invention
[0009] Objective of the Invention: The objective of the present invention is to provide a traceability link automatic recovery method based on the combination of multi-source information, which can fully exploit and utilize the multi-source information in defect reports and code submissions, automatically construct the missing links between defect reports and code submissions in the software development process, thereby improving the traceability and observability of open-source projects and providing support for various software maintenance tasks.
[0010] Technical Solution: The traceability link automatic recovery method of the present invention includes the following steps:
[0011] S1, Select a target project and obtain defect report data and code submission data, and construct a target data set according to the links manually completed by developers;
[0012] S2, For the metadata information in the defect report, use the feature representation method to complete the feature representation of the metadata information through feature engineering; for the text information and code information in the defect report, obtain the vectorized embedding representation through the existing pre-trained model;
[0013] S3, For the text data feature representation and the code data feature representation, respectively use the cross-attention mechanism for fusion and pooling to obtain the text feature vector y and the code feature vector z; for the metadata feature representation, first use the principal component analysis method for dimensionality reduction to obtain the metadata feature vector x, and then fuse the metadata feature vector x, the text feature vector y and the code feature vector z to obtain the multi-source fusion feature vector for link recovery;
[0014] S4, Utilize the multi-source fusion feature vector, use the fully connected layer as the classification layer of the network model, learn the classification feature relationship through convolution, and then input the classification feature relationship into the Softmax function for normalization as the final binary classification result, thereby realizing the determination of whether there is a traceability link between the defect report and the code submission.
[0015] Further, in step S1, the detailed implementation steps of constructing the target data set are as follows:
[0016] S11, Screen out software systems with different data scales, different development languages and different popularity levels from the collected target projects;
[0017] S12, Use the JQL language provided by the JIRA platform to extract the multi-source information in the required defect reports from the defect tracking system;
[0018] S13, Use the software repository mining tool to mine the multi-source information of the required code submissions in the version control system;
[0019] S14. Construct the correct link set and the incorrect link set according to the heuristic link construction method, and add classification labels to obtain the required link pair data set.
[0020] Furthermore, in step S14, the construction of the correct link set is as follows: Analyze whether there is a defect report ID manually added by the developer in the code submission log. If there is a manually added defect report ID, then use the code submission associated with the defect report as a correct link and add it to the correct link set.
[0021] The construction of the incorrect link set is as follows: For each code submission explicitly linked to a defect report by the developer, select all defect reports whose closing time differs from the submission time of the code submission by no more than 7 days, and form potential links with the code submission. If the potential link does not exist in the correct link set, then add it to the incorrect link set.
[0022] Furthermore, in step S2, the implementation steps of the feature representation of the metadata information are as follows:
[0023] S211. Represent the defect report priority and defect report type according to one-hot encoding.
[0024] S212. Represent the defect report creation time, update time, closing time, and code submission time in the form of the RBF (Radial Basis Function) for time information.
[0025] S213. Represent the defect report reporter, code submission author, and submitter according to the feature hashing method.
[0026] Furthermore, in step S2, for the text (it1, it2,... it n ) in the defect report and the text (ct1, ct2,... ct m ) in the code submission, perform vectorized embedding representation on the preprocessed text data in the form of a text sequence. The implementation process is as follows:
[0027] Concatenate the text in the defect report and the text in the code submission in the input format of RoBERTa, in the following form:
[0028] ([CLS], it1, it2,... it n , [SEP], ct1, ct2,... ct m , [SEP]), where m and n satisfy m + n + 3 = 512;
[0029] Among them, [CLS] is a special token in the pre-trained model, used to complete subsequent classification; [SEP] is used to separate different sentence sequences;
[0030] Take the vector output of the last layer of the RoBERTa model as the feature representation of the text information.
[0031] Furthermore, in step S2, splice the code in the defect report and the code in the code submission according to the input method of the UnixCoder model, in the following form: ([CLS], it1, it2,... itn, [SEP], ct1, ct2,... ctm, [SEP]). Set the maximum sequence length to 512, and obtain the vector output of the last layer of the UnixCoder model. The dimension of the vector is 512×768, and obtain the feature representation of the code data.
[0032] Furthermore, in step S2, the metadata information in the defect report mainly includes: defect report type, priority, reporter, defect report creation time, update time, and closing time;
[0033] The text information in the defect report includes the defect report title and the defect report description;
[0034] The code information in the defect report is extracted from the defect report description by regular expressions according to the conventional code identifier extraction pattern.
[0035] Furthermore, in step S3, for the text data feature representation, the implementation process of fusing and pooling using the cross-attention mechanism to obtain their respective feature vectors is as follows:
[0036] Based on the original feature representation, obtain the query Q matrix. After performing matrix multiplication with the key K matrix in the code feature representation, divide each element in the Q matrix and the K matrix by d k which is the dimension size of matrix K; then input the Softmax function to normalize the result. At this time, the dimension of the matrix is 512×512. Multiply the matrix with dimension 512×512 by the value V matrix in the code feature representation to obtain the feature vector Yz’ which is the text feature representation that integrates the code feature information;
[0037] Concatenate the feature vector Yz’ with the original text feature representation Y and perform max pooling to obtain the text feature vector y;
[0038] For the code data feature representation, combine the code feature representation and the text feature representation through the cross-attention mechanism to obtain Zy’, concatenate it with the original code feature representation Z and perform max pooling to obtain the code feature vector z.
[0039] Further, in step S4, the binary classification result indicates whether a given defect report-code commit pair is a correct link; if it is determined that there is a correct link between the defect report and the code commit, a traceability link is added between the given defect report and the code commit.
[0040] Compared with the prior art, the present invention has the following remarkable effects:
[0041] 1. The present invention makes full use of multi-source information from defect reports and code commits, namely metadata information, text information, and code information, and details the acquisition and utilization methods of each type of information, enabling more accurate mining and utilization of the information in the link pair; at the same time, based on the understanding of sequence data such as text and code by the pre-trained model, only a small amount of data fine-tuning is required to achieve more accurate semantic understanding and code understanding.
[0042] 2. In the field of traceability link recovery, the present invention uses advanced pre-trained models such as RoBERTa and Unixcoder to complete the vectorization representation of sequence data, and proposes a multi-source information feature fusion method, that is, based on the cross-attention mechanism to complete the fusion of text feature representation and code feature representation, enhancing the correlation between text data and code data, and further fusing with the metadata feature vector to achieve more accurate link recovery. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 is the model structure diagram of the present invention;
[0044] Figure 2 is the schematic diagram of the code submission code processing method in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0045] The following further describes the present invention in detail with reference to the accompanying drawings of the specification and the specific embodiments.
[0046] The traceability link automatic recovery method of the present invention extracts the information in the Issue-Commit link pair from three perspectives of metadata, text, and code, and uses a pre-trained model to characterize serialized information such as code and text. For the obtained multi-source feature information, first, the cross-attention mechanism is used to complete the fusion of text features and code features, enhancing the correlation between code and text, and then the obtained features are fused with the metadata feature vector after dimensionality reduction by the principal component analysis method to complete classification. With the help of text features, code features, and metadata features, this method can more accurately mine and utilize the information in the link pair. At the same time, based on the understanding of sequence data such as text and code by the pre-trained model, only a small amount of data fine-tuning is required to achieve more accurate semantic understanding and code understanding. The model structure of the traceability link automatic recovery method based on the combination of multi-source information is asFigure 1 As shown below. The specific implementation steps are as follows:
[0047] Step 1, Multi-source data preparation
[0048] Select target projects from a high-quality project platform and obtain the software artifacts required for the experiment, namely defect reports and code commits, and then build the target dataset according to the link construction method manually completed by developers.
[0049] Step 11, Open-source software project selection
[0050] Collect high-quality target projects from Github and screen software systems with different data scales, different development languages, and different popularity levels. Specifically, in this embodiment, software projects using the JIRA platform as the defect tracking system and Git as the version control system are selected. These projects usually have relatively high quality and provide convenient interfaces to facilitate data acquisition. In addition, comprehensively considering information such as the number of favorites in the software repository, software development cycle, software size, and software development language, the final selection of projects and the construction of the dataset are completed.
[0051] Step 12, Defect report data acquisition
[0052] Use the JQL (JIRA Query Language) provided by the JIRA platform to extract multi-source information from the defect reports required for the experiment from the defect tracking system.
[0053] The metadata information in the required defect reports mainly includes: defect report type, priority, reporter, defect report creation time, update time, and closure time. The defect report type is used to describe the main objective of the defect report, mainly including faults, tasks, new features, and improvements, etc. The defect report priority indicates to a certain extent the importance of the task described in the defect report, mainly including critical, important, minor, and trivial, etc.
[0054] The text data in the required defect reports mainly includes the defect report title and the defect report description. The defect report title concisely expresses the main purpose and content of the defect report, while the defect report description elaborately describes the main content of the defect report, which not only contains text information but also includes code blocks, code lines, and identifiers, etc.
[0055] The code data in the required defect reports is mainly extracted from the defect report description according to common code identifier extraction patterns (such as camel case naming, underscore naming, etc.) using regular expressions. The identifiers extracted are concatenated to form the code information in the defect report.
[0056] Step 13, Code commit data acquisition
[0057] Use the software repository mining tool Pydriller to mine multi-source information of code commits required for experiments in the software repository from the version control system.
[0058] The metadata information in the required code commits mainly includes the commit time, the author of the code commit, and the committer. Among them, the author of the code commit is the developer responsible for completing the code modification, and the committer is the developer who applies the code commit to the software project.
[0059] The text information in the required code commits is the commit log message written by the committer when the code is committed. According to the developer specification of the Apache community, this log information usually clearly expresses the modified content of the code commit. If the code commit solves the problem in the defect report, the ID of the defect report (e.g., ACCUMULO-1234) should be explicitly included.
[0060] In order to reasonably utilize the code information in the code commits and improve the understanding of the code change process and the purpose of code modification, the present invention proposes a more accurate code utilization method compared to code pattern matching to construct the code modification information in the code commits required for the method. The code processing method is as Figure 2 shown. Specifically, for each modified file in a given code commit, the code data is constructed according to the following method: 1) Use Pydriller to determine the modification type of the file in this code commit (such as MODIFY, DELETE, and ADD) and splice it with the file name to obtain the file name modification string s; 2) Determine whether the file is a code file. If it is a code file, further mine which functions are modified in the file, and splice all the modified function names with the file name modification string s obtained in 1). Otherwise (that is, if it is non-code file), directly use the file name modification string as the code modification information; Finally, splice the code modification information of each extracted file to obtain the code modification information of this code commit.
[0061] Step 14, Construction of Issue-Commit link pairs
[0062] Construct the correct link set and the wrong link set according to the traditional heuristic link construction method.
[0063] Construction of the correct link set: Analyze whether the defect report ID is manually added by the developer in the code commit log. If there are such explicit links where the defect report ID is manually added, then take this code commit and its associated defect report as the correct link and add it to the correct link set.
[0064] Construction of the incorrect link set: For each code commit explicitly linked by the developer to a defect report, select all defect reports whose closing time and the commit time of the code commit differ by no more than 7 days, and form potential links with the code commit. If the potential link does not exist in the correct link set, add it to the incorrect link set.
[0065] After constructing the correct link set and the incorrect link set and adding classification labels, the Issue-Commit link pair dataset required by the present invention can be obtained.
[0066] Step 2, multi-source data feature representation
[0067] The present invention mainly utilizes the metadata information, text data, and code data in the Issue-Commit link pair, where the text data and code data are both processed in sequence form. For the metadata information, feature engineering is completed through a feature representation method to represent the metadata information. For the text data and code data, the method uses existing pre-trained models to obtain vectorized embedding representations to construct feature representations.
[0068] Step 21, metadata feature representation
[0069] As described in Steps 12 and 13, for the metadata in the defect report, information such as the defect report type, priority, reporter, defect report creation time, update time, and closing time are mainly extracted. For the metadata in the code commit, information such as the commit time, code commit author, and committer are mainly extracted.
[0070] According to the existing feature representation technology, the defect report priority and defect report type are represented by one-hot encoding; the defect report creation time, update time, closing time, and code commit time are represented by the RBF radial basis function for time information; finally, information such as the defect report reporter, code commit author, and committer are represented by the feature hashing method. Thus, a complete metadata feature representation of an Issue-Commit link pair sample, that is, a feature vector with a dimension of 1×m, can be obtained.
[0071] Step 22, text data feature representation
[0072] For the text (it1, it2,... it n ) in the defect report and the text (ct1, ct2,... ct m ) in the code commit, common preprocessing strategies are adopted to complete text cleaning, that is, removing URLs, removing stop words, word segmentation, stemming, etc.
[0073] Subsequently, the preprocessed text data is vectorized and embedded in the form of a text sequence, mainly by means of the currently most commonly used RoBERTa model. As a pre-trained model more robust than the BERT model, the RoBERTa model only needs to use a small amount of data on the basis of pre-training to fine-tune the model, and has stronger text semantic representation ability.
[0074] Next, the text in the defect report and the text in the code submission are concatenated according to the input method of the RoBERTa model, in the following form: ([CLS], it1, it2,... it n , [SEP], ct1, ct2,... ct m , [SEP]), and the maximum sequence length is 512, that is, m and n satisfy m + n + 3 = 512; where [CLS] is a special token in the RoBERTa model, used to complete subsequent classification, and [SEP] is used to separate different sentence sequences. Take the vector output of the last layer network in the RoBERTa model output (with a dimension of 512×768, that is, the vector representation of each token in the input is obtained) as the feature representation of the text data.
[0075] Step 23, Feature representation of code data
[0076] Similar to the feature representation of text data, the extracted code information is processed in a sequence manner. For the code (ic1, ic2,... ic n ) in the defect report extracted in step 12 and the code (cc1, cc2,... cc m ) in the code submission extracted in step 13, the Unixcoder model, which currently has the best effect in the field of code understanding, is used for vectorized representation of the code. As a unified cross-modal pre-trained model, the Unixcoder model can be compatible with three modes: Encoder-Only, Decoder-Only, and Encoder-Decoder, and it has achieved the most advanced results in tasks such as code clone detection, code translation, and code search.
[0077] Similar to the processing of text data, the code in the defect report and the code in the code submission are concatenated according to the input method of the Unixcoder model, in the form of ([CLS], it1, it2,... itn, [SEP], ct1, ct2,... ctm, [SEP]), and the maximum sequence length is also set to 512, and the vector output of the last layer network of the Unixcoder model (with a dimension of 512×768, that is, the vector representation of each token in the input is obtained) is obtained to get the feature representation of the code data.
[0078] Step 3, Multi-source data feature fusion
[0079] Perform feature fusion on the metadata feature representation, text data feature representation, and code data feature representation, where both text data and code data are processed as sequence data. For the text data feature representation and code data feature representation, this method uses the cross-attention mechanism for fusion and pooling to obtain their respective feature vectors; for the metadata feature representation, the principal component analysis method is first used for dimensionality reduction and then concatenated with the text feature vector and code feature vector.
[0080] Step 31, Text and code feature fusion based on the cross-attention mechanism
[0081] The cross-attention mechanism is an attention mechanism that mixes two different embedded sequences. In the case where the two sequences have the same dimension, one sequence is used to obtain the input query Q, and the other sequence provides the key K and value V. The cross-attention mechanism has good effects in the field of feature fusion. Using the cross-attention mechanism to achieve the fusion of text and code features can further enhance the feature correlation between code and text.
[0082] As Figure 1 shown, taking the text feature representation as an example, on the basis of the original feature representation, a query Q matrix is obtained. After performing matrix multiplication with the key K matrix in the code feature representation, each element in the matrix (i.e., the Q matrix with dimensions 512×768 and the K matrix with dimensions 512×768) is divided by d k which is the dimension size of the K matrix. Subsequently, the result is input into the Softmax function for normalization. At this time, the dimension of the matrix is 512×512. The matrix with dimensions 512×512 is multiplied by the value V matrix in the code feature representation to obtain the feature vector Yz’, which is the text feature representation that has fused the code feature information; Yz’ is concatenated with the original text feature representation Y and then max-pooled to obtain the text feature vector y. The attention mechanism formula is as follows:
[0083]
[0084] Similarly, the code feature representation also uses the cross-attention mechanism to fuse with the text feature representation and pool to obtain the code feature vector z. The code feature representation and the text feature representation are combined through the cross-attention mechanism to obtain Zy’, which is concatenated with the original code feature representation Z and then max-pooled to obtain the code feature vector z, as Figure 1 shown.
[0085] Step 32, Combination of multi-source feature vectors
[0086] For the metadata feature representation obtained in step 21, the principal component analysis method is used to reduce the dimension of the original feature representation, convert the original feature indicators into a few comprehensive indicators, so that the low-dimensional components retain the most important aspects of the data, and finally obtain the metadata feature vector x with a dimension of 1×m'.
[0087] The metadata feature vector x is concatenated with the text feature vector y and the code feature vector z obtained in step 31 to obtain the final multi-source fusion feature vector for link recovery.
[0088] Step 4, automated link recovery
[0089] Using the multi-source fusion feature vector obtained in step 3, a binary classification is completed with the help of a fully connected layer, thereby determining whether there is a traceability link between the defect report and the code submission.
[0090] Specifically, the fully connected layer is used as the classification layer of the network model, and the convolutional layer is used to fully learn the classification feature relationship. Subsequently, the classification feature relationship is input into the Softmax function for normalization as the final binary classification result, which indicates whether the given defect report-code submission pair is a correct link. That is, if it is determined that there is a correct link between the defect report and the code submission, a traceability link is added between the given defect report and the code submission.
Claims
1. A method for automatically restoring traceability links based on the combination of multi-source information, characterized in that, The steps are as follows: S1. Select a target project, obtain defect report data and code submission data, and construct a target dataset according to the links manually completed by developers; S2. For the metadata information in the defect report, use a feature representation method to perform feature engineering to complete the feature representation of the metadata information; For the text information and code information in the defect report, obtain vectorized embedding representations through existing pre-trained models; S3. For the text data feature representation and code data feature representation, use the cross-attention mechanism for fusion and pooling respectively to obtain the text feature vector y and the code feature vector z; for the metadata feature representation, first use the principal component analysis method to reduce the dimension to obtain the metadata feature vector x, and then fuse the metadata feature vector x, the text feature vector y, and the code feature vector z to obtain a multi-source fusion feature vector for link recovery; S4. Use the multi-source fusion feature vector, use the fully connected layer as the classification layer of the network model, learn the classification feature relationship through convolution, and then input the classification feature relationship into the Softmax function for normalization as the final binary classification result, so as to determine whether there is a traceable link between the defect report and the code submission.
2. The automatic recovery method for traceability links based on the combination of multi-source information according to claim 1, characterized in that In step S1, the detailed implementation steps for constructing the target dataset are as follows: S11. Screen software systems with different data scales, different development languages, and different popularity levels from the collected target projects; S12. Use the JQL language provided by the JIRA platform to extract multi-source information in the required defect reports from the defect tracking system; S13. Use a software repository mining tool to mine multi-source information of required code submissions in the version control system; S14. Construct a correct link set and an incorrect link set according to the heuristic link construction method, and add classification labels to obtain the required link pair dataset.
3. The automatic recovery method for traceability links based on multi-source information combination according to claim 2, characterized in that, In step S14, the construction of the correct link set is as follows: Analyze whether the defect report ID is manually added in the code submission log. If there is a manually added defect report ID, add the code submission and its associated defect report as a correct link to the correct link set; The construction of the incorrect link set is as follows: For each code submission explicitly linked to a defect report by a developer, select all defect reports whose closing time differs from the submission time of the code submission by no more than 7 days, and form potential links with the code submission; If the potential link does not exist in the correct link set, add it to the incorrect link set.
4. The automatic recovery method of traceability link based on multi-source information combination according to claim 1, characterized in that In step S2, the implementation steps for the feature representation of the metadata information are as follows: S211. Represent the defect report priority and defect report type according to one-hot encoding; S212. Represent the defect report creation time, update time, closing time, and code submission time according to the RBF radial basis function for time information; S213. Represent the defect report reporter, code submission author, and submitter according to the feature hashing method.
5. The automatic recovery method for traceability links based on the combination of multi-source information according to claim 1, wherein In step S2, for the text in the defect report and the text in the code submission, the preprocessed text data is vectorized and embedded in the form of a text sequence; The implementation process is as follows: The text in the defect report and the text in the code submission are concatenated according to the input method of RoBERTa, in the following form: ([CLS], it1, it2,... it n , [SEP], ct1, ct2,... ct m , [SEP]), m and n satisfy m + n + 3 = 512; the text in the defect report is: it1, it2,... it n ; the text in the code submission is ct1, ct2,... ct m ; Among them, [CLS] is a special token in the pre-trained model, used to complete subsequent classification; [SEP] is used to separate different sentence sequences; Take the vector output of the last layer of the RoBERTa model as the feature representation of the text information.
6. The automatic recovery method of traceability link based on multi-source information combination according to claim 1, characterized in that In step S2, the code in the defect report and the code in the code submission are concatenated according to the input method of the UnixCoder model, in the following form: ([CLS], it1, it2,... itn, [SEP], ct1, ct2,... ctm, [SEP]). Set the maximum sequence length to 512, and obtain the vector output of the last layer of the UnixCoder model. The dimension of the vector is 512×768, and the feature representation of the code data is obtained.
7. The automatic recovery method for traceability links based on multi-source information combination according to claim 1, characterized in that In step S2, the metadata information in the defect report mainly includes: defect report type, priority, reporter, defect report creation time, update time, and closing time; The text information in the defect report includes the defect report title and the defect report description; The code information in the defect report is extracted from the defect report description according to the regular code identifier extraction pattern using regular expressions.
8. The automatic recovery method of traceability link based on multi-source information combination according to claim 1, characterized in that, In step S3, for the feature representation of the text data, the implementation process of using the cross-attention mechanism for fusion and pooling to obtain their respective feature vectors is as follows: Based on the original feature representation, the query Q matrix is obtained. After performing matrix multiplication with the key K matrix in the code feature representation, each element in the Q matrix and the K matrix is divided by , where d k is the dimension size of matrix K. Subsequently, the result is normalized by inputting it into the Softmax function. At this time, the matrix dimension is 512×512. The matrix with dimension 512×512 is multiplied with the value V matrix in the code feature representation to obtain the feature vector which is the text feature representation that integrates the code feature information; Concatenate the feature vector with the original text feature representation Y and perform max pooling to obtain the text feature vector y; For the code data feature representation, the code feature representation and the text feature representation are combined through a cross-attention mechanism to obtain , which is concatenated with the original code feature representation Z and undergoes max pooling to obtain the code feature vector z.
9. The automatic recovery method for traceability links based on multi-source information combination according to claim 1, wherein In step S4, the binary classification result indicates whether a given defect report-code submission pair is a correct link; if it is determined that there is a correct link between the defect report and the code submission, a traceability link is added between the given defect report and the code submission.
Citation Information
Patent Citations
Link generation method based on BERT pre-training model
CN114816497A
Deep Learning Source Code Analyzer and Repairer
US20170212829A1