Project defect code determination method and device, equipment, medium and product

By performing word embedding and language encoding on defect reports and code change information, and using a defect localization model to determine the association probability of defect codes, the problems of high cost of manual localization and high resource requirements for automatic localization in existing technologies are solved, achieving efficient and accurate defect code localization.

CN120950370APending Publication Date: 2025-11-14AGRICULTURAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511050736.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-29
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

In existing technologies, manual location of software defect codes is costly and inefficient, while automatic location methods have limited efficiency and high resource requirements, resulting in high defect code location costs.

Method used

By acquiring current defect reports and code change information, and using a pre-trained defect localization model for word embedding and language encoding processing, the association probability of defective code is determined, and the degree of matching between defect reports and code blocks is automatically analyzed.

Benefits of technology

It achieves low-cost, high-efficiency, and high-accuracy defect code location, saving manpower and resource costs and improving location efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120950370A_ABST
    Figure CN120950370A_ABST
Patent Text Reader

Abstract

The invention discloses a project defect code determination method and device, equipment, a medium and a product, and the method comprises the steps: obtaining a current defect report and code change information of a current project, carrying out the word embedding coding processing of the current defect report and the code change information, and obtaining a defect report input vector and at least one code block input vector; performing language coding processing on the defect report input vector and each code block input vector by utilizing a pre-trained defect positioning model to obtain a defect report feature vector and at least one code block feature vector; and performing defect positioning classification processing on the defect report feature vector and each code block feature vector by using a pre-trained defect positioning model to obtain at least one association probability, and determining a defect code of the current defect report according to each association probability. According to the method, the current defect report can be automatically analyzed, and the defect code of the current defect report is accurately and efficiently determined according to the association probability of the current defect report and each code block.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a method, apparatus, equipment, medium, and product for determining project defect codes. Background Technology

[0002] During software development, various defects will arise. Compiling these defects yields a defect report. Defect localization refers to identifying the code segment corresponding to the problem recorded in the defect report—that is, the code segment that caused the problem described in the defect report.

[0003] Currently, defect code localization methods include manual and automated methods. Manual localization requires developers to read defect reports, understand the defects, and then filter out relevant defective code from the codebase. Automated localization involves repeatedly running test cases, collecting information such as execution paths and results, and then using this information to identify the defective code in the software project. However, manual localization requires developers to have a thorough understanding of the software project's architecture and business logic, and involves repeated debugging, resulting in high labor and time costs and low efficiency and speed. Automated localization, requiring repeated test case runs, has limited efficiency and high computational resource requirements, making it less economical and costly.

[0004] Therefore, proposing a low-cost, high-efficiency, and high-accuracy defect code location method that saves on defect code location costs while ensuring the effectiveness of defect code location is an urgent problem to be solved. Summary of the Invention

[0005] This invention provides a method, apparatus, device, medium, and product for determining project defect codes, which can be applied to the financial technology field. It aims to automatically analyze the current defect report and accurately and efficiently determine the defect code of the current defect report based on the correlation probability between the current defect report and each code block, ensuring the effectiveness of defect code location while saving the cost of defect code location.

[0006] According to one aspect of the present invention, a method for determining project defect codes is provided, the method comprising:

[0007] Obtain the current defect report and code change information of the current project, and perform word embedding encoding on the current defect report and code change information to obtain the defect report input vector and at least one code block input vector; wherein, the code change information includes at least one code block, and the code block and the code block input vector correspond one-to-one;

[0008] Using a pre-trained defect localization model, language encoding is performed on the defect report input vector and at least one code block input vector to obtain the defect report feature vector and at least one code block feature vector; wherein, the code block feature vector and the code block input vector correspond one-to-one.

[0009] Using a pre-trained defect localization model, defect localization and classification are performed on the defect report feature vector and at least one code block feature vector to obtain at least one association probability, and the defect code of the current defect report is determined based on the at least one association probability; wherein, the association probability corresponds one-to-one with the code block feature vector.

[0010] According to another aspect of the present invention, a project defect code determination apparatus is provided. The project defect code determination apparatus is used to implement the project defect code determination method in any embodiment of the present invention. The apparatus includes:

[0011] The information acquisition module is used to acquire the current defect report and code change information of the current project, and to perform word embedding encoding on the current defect report and code change information to obtain the defect report input vector and at least one code block input vector; wherein, the code change information includes at least one code block, and the code block and the code block input vector correspond one-to-one.

[0012] The feature vector determination module is used to perform language encoding processing on the defect report input vector and at least one code block input vector using a pre-trained defect localization model to obtain the defect report feature vector and at least one code block feature vector; wherein, the code block feature vector and the code block input vector correspond one-to-one.

[0013] The defect code localization module is used to perform defect localization and classification processing on the defect report feature vector and at least one code block feature vector using a pre-trained defect localization model, obtain at least one association probability, and determine the defect code of the current defect report based on the at least one association probability; wherein, the association probability corresponds one-to-one with the code block feature vector.

[0014] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:

[0015] At least one processor; and a memory communicatively connected to the at least one processor;

[0016] The memory stores a computer program that can be executed by at least one processor, which enables the at least one processor to execute the method for determining project defect codes in any embodiment of the present invention.

[0017] According to another aspect of the present invention, a computer-readable storage medium is provided that stores computer instructions for causing a processor to execute a method for determining project defect codes in any embodiment of the present invention.

[0018] According to another aspect of the present invention, a computer program product is provided, the computer program product including a computer program, which, when executed by a processor, implements a method for determining project defect codes according to any embodiment of the present invention.

[0019] The method for determining the defect code of the project according to the present invention includes: obtaining the current defect report and code change information of the current project, and performing word embedding encoding processing on the current defect report and code change information to obtain a defect report input vector and at least one code block input vector; the code change information includes at least one code block, and the code block and the code block input vector correspond one-to-one; using a pre-trained defect localization model, performing language encoding processing on the defect report input vector and at least one code block input vector to obtain a defect report feature vector and at least one code block feature vector; the code block feature vector and the code block input vector correspond one-to-one; using a pre-trained defect localization model, performing defect localization and classification processing on the defect report feature vector and at least one code block feature vector to obtain at least one association probability, and determining the defect code of the current defect report based on the at least one association probability; the association probability corresponds one-to-one with the code block feature vector. The technical solution of this invention automatically analyzes each code block in the current defect report and code change information to obtain the defect report input vector and the code block input vector of each code block. A pre-trained defect localization model is then used to perform language encoding processing on the defect report input vector and the code block input vector of each code block to obtain the defect report feature vector and the code block feature vector of each code block. The feature vector is an optimized input vector that reduces the amount of data while retaining vector features. Defect code localization is performed based on the correlation probability between the defect report feature vector and the code block feature vector of each code block. This allows for a quantitative measurement of the matching degree between each code block and the defect report, thereby improving the accuracy and efficiency of defect code localization. It is a low-cost, high-efficiency, and high-accuracy defect code localization method that can save on defect code localization costs while ensuring the effectiveness of defect code localization. It solves the problems of manual defect location methods, which require developers to have a thorough understanding of the software project's architecture and business logic and repeatedly debug the code, resulting in high labor and time costs and low efficiency and speed in locating defective code. It also solves the problems of automatic defect location methods, which require repeated running of test cases, resulting in limited efficiency in locating defective code, high demand for computing resources, low economic efficiency, and high cost in locating defective code.

[0020] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in this invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 This is a flowchart illustrating a method for determining project defect codes provided by the present invention;

[0023] Figure 2 This is a schematic diagram of the architecture of a defect localization model provided by the present invention;

[0024] Figure 3 This is a schematic diagram of the information processing process of a defect location model provided by the present invention;

[0025] Figure 4 This is a flowchart illustrating another method for determining project defect codes provided by the present invention;

[0026] Figure 5 This is a schematic diagram of the structure of a device for determining project defect codes provided by the present invention;

[0027] Figure 6 This is a schematic diagram of the structure of an electronic device provided by the present invention. Detailed Implementation

[0028] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are merely some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort should fall within the scope of protection of the present invention.

[0029] It should be noted that the terms "first," "second," "initial," "intermediate," "candidate," "alternate," "target," etc., used in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0030] The acquisition, storage, use, and processing of data in the technical solution of this invention all comply with relevant laws and regulations. Specifically, the user information collected in this invention is information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of related data all comply with relevant laws, regulations, and standards of relevant countries and regions, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entry points for users to choose to authorize or reject automated decision results; if the user chooses to reject, the process proceeds to the expert decision-making process.

[0031] Figure 1 This is a flowchart illustrating a method for determining project defect codes provided by the present invention. This embodiment is applicable to low-cost, high-efficiency, and high-accuracy defect code location. The method can be executed by the project defect code determination device provided by the present invention. This device can be implemented in hardware and / or software. In a specific embodiment, the device can be integrated into an electronic device. The following embodiments will illustrate this using the integration of the device into an electronic device as an example. (Refer to...) Figure 1 The method specifically includes the following steps:

[0032] S101. Obtain the current defect report and code change information of the current project, and perform word embedding encoding on the current defect report and code change information to obtain the defect report input vector and at least one code block input vector.

[0033] The current project can be understood as a project under development and testing, such as a software project undergoing requirements clarification, coding, testing, integration, deployment preparation, and project maintenance. During development and maintenance, software projects require continuous testing, and the code is fine-tuned based on test reports to improve performance. The current defect report is a structured document that records anomalies or errors encountered during testing, serving as both a problem log and a basis for repair. Code change information is a record of code changes within the software project, indicating changes to code blocks, including but not limited to the time and content of those changes. It's a collection of all code block changes during development; generally, any project defect will be accompanied by a corresponding code block change. Word embedding encoding can be understood as a data processing method that converts linguistic or code information into vectors. It transforms discrete symbols (words, phrases, etc.) in natural language and instructions in code into continuous, low-dimensional, dense real-valued vectors. It is a distributed representation that captures the latent relationships between words in a computable form, ensuring that semantically or syntactically similar words are close to each other in the vector space, thereby determining the degree of matching between various types of information (defect reports and individual code blocks). The defect report input vector can be understood as the vector obtained after word embedding encoding of the current defect report. The code block input vector is the vector obtained after word embedding encoding of the code block. Code change information includes at least one code block, and there is a one-to-one correspondence between the code block and its input vector.

[0034] Optionally, word embedding encoding is performed on the current defect report and code change information to obtain a defect report input vector and at least one code block input vector. This includes: extracting the title, report content, and reporter of the current defect report to obtain the defect report feature text, and performing word embedding encoding on the defect report feature text using a byte pair encoding algorithm to obtain the defect report input vector; and splitting the code change information according to the code change record to obtain at least one code block, and performing word embedding encoding on at least one code block using a byte pair encoding algorithm to obtain at least one code block input vector.

[0035] The title of the current defect report can be understood as the name of the current defect report or a summary of its key features, distinguishing it from other defect reports. The report content can be understood as the project anomaly recorded in the report; anomaly information is essential for locating defective code. The reporter can be understood as the person who filled out the report or the tester of the software project; it can be used for data backtracking and to analyze the reporter's reporting habits, aiding in defect code location. The defect report feature text can be understood as extracting key information from the report content and then combining it with the current defect report title and reporter's name to obtain the final defect report text. This setup ensures that project anomalies are fully recorded while reducing the content of the current defect report, improving defect code location efficiency and saving resource requirements for defect code location work.

[0036] Byte-pair encoding is a word segmentation method that identifies frequently occurring characters or character combinations (called "sub-words") in a text. Following linguistic logic, these sub-words are progressively merged to generate a text that conforms to the sub-word concatenation rules and contains the report's content. Vector embedding encoding is then applied to this text to obtain the defect report input vector. The defect report input vector is of type rotation position encoding + word embedding encoding. This setup ensures that the input vector fully expresses the meaning of the defect report. It's worth noting that the defect report input vector can also be obtained by directly performing vector embedding encoding on the segmented sub-words; similarly, this defect report input vector is also of type rotation position encoding + word embedding encoding.

[0037] Code change logs can be understood as a record of code block changes, including the time and content of the changes. Based on these logs, code block changes occurring in different testing cycles can be distinguished. Specifically, the same code block may have undergone the same changes in different testing cycles. This invention deduplicates code blocks when splitting code change information, preventing the processor from processing duplicate code blocks and saving processor resources. The code block input vector can be understood as the vector obtained after word embedding encoding of the code block. The word embedding encoding process for code blocks is the same as that for defect reports, and this invention does not limit it.

[0038] S102. Using a pre-trained defect localization model, language encoding is performed on the defect report input vector and at least one code block input vector to obtain the defect report feature vector and at least one code block feature vector.

[0039] The pre-trained defect localization model can be understood as training an initial defect localization model using change information (code change blocks and code change comments) from similar projects (projects with the same programming language and technology stack) of the current project, and then fine-tuning the trained defect localization model using historical code block change information and historical defect reports of the current project. The result is a model capable of extracting high-dimensional feature vectors from the defect report input vector and the code block input vector. The initial defect localization model includes, but is not limited to, general defect localization models and defect localization models from related projects. It consists of a pre-trained language model with semantic understanding capabilities and a defect localization module that can learn a random distribution of parameters. The pre-trained language model is used to parse the feature vectors of code blocks, code change comments, and defect reports. The defect localization module is used to determine the similarity probability between the code block feature vector and the code change comment (or defect report) feature vector. Language encoding processing is used to parse the high-dimensional feature vectors of the defect report input vector and the input vectors of each code block. The defect report feature vector can be understood as a high-dimensional feature vector of the defect report input vector, and the code block feature vector can be understood as a high-dimensional feature vector of the code block input vector. There is a one-to-one correspondence between the code block feature vector and the code block input vector.

[0040] The change information is obtained from similar projects of the current project using code management tools. This information includes code change blocks and code change comments, which share the same programming language and technology stack as the current project. Following the code change information processing method recorded in S101, the above code change information is processed to obtain vector information, including but not limited to code change block vectors and code change comment vectors. This invention utilizes test information (code change blocks and code change comments) from similar projects of the current project, as well as historical code block change information and historical defect reports of the current project, to train a defect localization model. The purpose of this training is to train a model with high matching degree and strong feature condensation effect with a smaller sample size, reducing model training cost and time, improving the model's adaptability, and thus improving the defect code localization effect. Secondly, the defect localization model can learn the commonalities of similar tasks (task pairing of code-natural language description defects) for subsequent defect code localization. It can also provide knowledge reserves for the model, thereby improving the parsing accuracy of defect reports and code blocks, and enabling vector analysis and defect matching from the perspective of code changes.

[0041] The defect localization model consists of two parts: a feature vector parsing module and an association probability determination module. Figure 2 This is a schematic diagram of the architecture of a defect localization model provided by the present invention, showing the architecture of the feature vector parsing module of the defect localization model. Figure 2The input text is a defect report (code change comment) and / or code block. After processing such as word segmentation and embedding encoding, the defect report (code change comment) and / or code block are input into the language model for processing, resulting in a vector of the defect report (code change comment) and / or code block. The feature vector parsing module of the defect localization model of this invention will also perform supplementary processing on the vector, that is, use a regularization (Dropout) layer and a fully connected layer to convert the vector output by the language model into a vector of a specific dimension, and standardize the vector for subsequent data processing. The dimension can be set and adjusted according to the data processing requirements, and this invention does not limit it.

[0042] Optionally, determining the defect localization model includes: obtaining at least one code change block and at least one code change comment from similar projects of the current project, with a one-to-one correspondence between the code change block and the code change comment; performing word embedding encoding on each code change block and each code change comment to obtain at least one code change block input vector and at least one code change comment input vector; determining positive training samples and negative training samples of the defect localization model based on the correspondence between the at least one code change block input vector and at least one code change comment input vector, wherein the number of positive training samples and negative training samples of the defect localization model is the same, the code change block input vector and the code change comment input vector are corresponding in the positive training samples of the defect localization model, and the code change block input vector and the code change comment input vector are not corresponding in the negative training samples of the defect localization model; training the initial defect localization model using the positive training samples and negative training samples of the defect localization model with the cross-entropy function as the loss function to obtain a backup defect localization model; and fine-tuning the backup defect localization model using at least one code block and at least one historical defect report of the current project to obtain the defect localization model.

[0043] In this context, "similar projects" can be understood as software projects similar to or with the same project logic as the current project. "Code change blocks" can be understood as code changes in similar projects of the current project. "Code change comments" can be understood as defect descriptions corresponding to code change blocks in similar projects. Generally, the code change blocks and code change comments of the current project belong to the same project. The processing procedures for code change blocks and code change comments (including the determination of input vectors and feature vectors) are the same as those for code blocks and the current defect report, and will not be elaborated upon here.

[0044] Specifically, there is a correspondence between code blocks and code change comments (defect reports). Code blocks and code change comments with correct correspondences constitute positive training samples, while code blocks and code change comments with incorrect correspondences constitute negative training samples. Assume that code block 1 corresponds to code change comment 1 (i.e., the change in code block 1 corresponds to code change comment 1), code block 2 corresponds to code change comment 2, code block 3 corresponds to code change comment 3, code block 4 corresponds to code change comment 4, and code block 5 corresponds to code change comment 5. "Code block-code change comment" pairs with a corresponding relationship are positive training samples, such as "code block 1-code change comment 1", "code block 2-code change comment 2", "code block 3-code change comment 3", etc. "Code block-code change comment" pairs without a corresponding relationship are negative training samples, such as "code block 1-code change comment 2", "code block 1-code change comment 3", "code block 1-code change comment 4", "code block 1-code change comment 2", "code block 2-code change comment 1", "code block 2-code change comment 3", etc. Positive training samples can be identified using the problem-solving identifier for defect tracking. It's important to note that during model training, positive and negative training samples are configured in a 1:1 ratio; that is, each code block has one positive and one negative training sample. Generally, the negative training samples used for training are selected from all the negative training samples for that code block, focusing on those most easily confused or misclassified. Specifically, matching prediction is performed on all negative training samples, and the sample with the highest score is used as the negative training sample for that code block. This method dynamically selects the most difficult-to-distinguish negative training samples, improving their quality and preventing overfitting in the early stages of training. It's also worth noting that after each iteration during model training, negative training samples are reselected to increase model stability. Using the cross-entropy function as the loss function maximizes the prediction probability of the correct class, improving model performance. The model after training (e.g., the number of training iterations meets the required training number, the model's training parameters meet the model convergence condition, and the loss value of the loss function is less than the model convergence metric) serves as a backup defect localization model.

[0045] Furthermore, after obtaining the backup defect localization model, this invention will further fine-tune the backup defect localization model using at least one code block and at least one historical defect report from the current project to obtain the defect localization model. The specific model fine-tuning process includes: performing word embedding encoding on at least one code block and at least one historical defect report from the current project to obtain at least one code block input vector and at least one historical defect report input vector; determining positive and negative fine-tuning samples of the defect localization model based on the correspondence between the at least one code block input vector and the at least one historical defect report input vector, wherein the number of positive and negative fine-tuning samples of the defect localization model is the same, the code block input vector and the historical defect report input vector are corresponding in the positive fine-tuning samples of the defect localization model, and the code block input vector and the historical defect report input vector are not corresponding in the negative fine-tuning samples of the defect localization model; and fine-tuning the backup defect localization model using the positive and negative fine-tuning samples of the defect localization model with the cross-entropy function as the loss function to obtain the defect localization model.

[0046] Historical defect reports can be understood as all defect reports for the current project. This invention uses historical data from the current project (at least one code block and at least one historical defect report) to fine-tune and optimize the backup defect localization model, thereby improving the accuracy of defect code localization in the current project. The processing procedure for historical defect reports is the same as that for code blocks; that is, the process for determining the input vector of historical defect reports is the same as that for defect reports, and the process for determining the feature vector of historical defect reports is the same as that for defect reports. Furthermore, the methods for determining the positive and negative fine-tuning samples of the defect localization model are the same as those for determining the positive and negative training samples of the defect localization model. The optimization process of the defect localization model is also the same as the training process, the only difference being the different samples, which will not be elaborated upon here. The defect localization model of this invention is pre-trained and periodically optimized to ensure the efficiency and quality of defect code localization.

[0047] Optionally, before model training, the training samples are classified in a 6:2:2 ratio to form training samples, test samples, and evaluation samples, with the same number of positive and negative training samples in each sample. In one implementation, the training process of the defect localization model includes: 1) classifying the positive and negative training samples in the training samples according to training batches to obtain multiple training batches. For example, if a batch has 8 samples, then 4 positive training samples and 4 negative training samples are selected. If the total number of samples is 32, then there are 4 training batches; 2) inputting the training samples into the defect localization model in batch order to calculate the cross-entropy function; 3) adjusting the parameters of the defect localization model using the cross-entropy function until all 4 training batches have been traversed; 4) using test samples to test the model after parameter optimization and calculating the accuracy of the model. If the accuracy of the model is greater than the model training index (e.g., defect code localization accuracy is higher than 97%, defect code localization accuracy is higher than 95%, etc.), then the model training is determined to be complete, and 5) is executed. Otherwise, negative training samples are re-determined, training samples, test samples, and evaluation samples are divided, and 1) is executed; 5) using evaluation samples to evaluate the model and recording the evaluation data in order to trace back the life cycle of the model and adjust the model parameters. It is worth noting that evaluation metrics can also be set according to model training and optimization requirements. When the evaluation results do not meet the evaluation metrics, negative training samples are re-determined, training samples, test samples and evaluation samples are divided, and step 1) is executed to improve the model's defect localization performance.

[0048] S103. Using a pre-trained defect localization model, perform defect localization and classification processing on the defect report feature vector and at least one code block feature vector to obtain at least one association probability, and determine the defect code of the current defect report based on at least one association probability.

[0049] The defect localization and classification process aims to analyze the correlation probability between the defect report feature vector and the feature vectors of various code blocks. The correlation probability can be understood as the degree of matching between the defect report feature vector and the code block feature vector. There is a one-to-one correspondence between the correlation probability and the code block feature vector. The defect code in the current defect report can be understood as the code block that may have caused the current defect report, determined based on various correlation probabilities. For example, the code block corresponding to the feature vector of the code block with a correlation probability higher than a preset defect code correlation threshold (0.8, 0.75, 0.85, etc.), and the code block corresponding to the feature vector of the code block with the highest correlation probability among preset defect codes (3, 5, 8, 10, etc.). This invention automatically identifies multiple code blocks as defect codes, enabling developers to locate the cause of project anomalies. This avoids the low debugging efficiency caused by manually locating each defect code block, effectively ensuring project debugging efficiency and reducing the project debugging time for developers.

[0050] Figure 3 This is a schematic diagram of the information processing process of a defect localization model provided by this invention. It illustrates the architecture of the correlation probability determination module of the defect localization model, demonstrating the process of determining the correlation probability between the defect report feature vector and the code block feature vector. The correlation probability includes relevant probability and irrelevant probability, representing binary information. Combined with... Figure 3 For any code block feature vector, a pre-trained defect localization model is used to perform defect localization and classification processing on the defect report feature vector and the code block feature vector to obtain the association probability. This includes: concatenating the defect report feature vector and the code block feature vector to obtain the first candidate feature vector, which is (U, V, |UV|), where U is the defect report feature vector and V is the code block feature vector; using the feature extraction layer in the defect localization model to compress the dimension of the first candidate feature vector and extract the high-level features of the first candidate feature vector to obtain the second candidate feature vector; and using the probability analysis layer in the defect localization model to perform relevance analysis processing on the second candidate feature vector to obtain the relevance probability and irrelevance probability of the second candidate feature vector.

[0051] The second candidate feature vector can be understood as a feature vector that extracts the important decision features (i.e., high-level features) from the first candidate feature vector and discards noise and redundant information. The feature extraction layer uses a fully connected layer with an additional regularization layer. This design aims not only to compress the dimensionality of the first candidate feature vector to extract high-level features but also to prevent overfitting during model training and optimization, thus improving model performance. The probabilistic analysis layer performs the following tasks: 1) applying an activation function to the second candidate feature vector using a non-linear mapping; 2) using a regularization layer and a fully connected layer to transform the mapped second candidate feature vector into a binary vector; 3) using a probability distribution function to determine the correlation between the defect report feature vector and the code block feature vector within the second candidate feature vector. The probability distribution function compresses the correlation between the defect report feature vector and the code block feature vector into a probability distribution between 0 and 1, i.e., a correlation probability and an irrelevance probability, with the sum of the correlation probability and the irrelevance probability being 1.

[0052] After obtaining the feature vector of the defect report and the association probability of the feature vector of each code block, the defect code of the current defect report will be determined according to the association probability. Specifically, this includes: sorting at least one association probability in descending order of association degree, and using a preset number of defect codes or a preset defect code association threshold to select backup association probabilities from the association probability sorting results; and determining the code block corresponding to the backup association probability as the defect code.

[0053] The backup association probability can be understood as the association probability that meets the defect code block screening criteria. For example, the association probability that is higher than the preset defect code association threshold (0.8, 0.75, 0.85, etc.), the association probability of the top preset number of defect codes (3, 5, 8, 10, etc.) in terms of association probability, and the code blocks corresponding to all the association probabilities that meet the defect code block screening criteria constitute the defect code. This invention adheres to the principle of over-selecting defect code blocks, selecting at least two code blocks as defect code blocks at a time. This avoids the phenomenon of secondary defect code location due to anomalies in the optimal recommended defect code block (the code block with the highest association probability). Developers can check the relationship between multiple potentially abnormal code blocks and defect reports at once, saving resources and improving the efficiency of defect location work.

[0054] The technical solution of the above embodiment automatically analyzes each code block in the current defect report and code change information to obtain the defect report input vector and the code block input vector of each code block. A pre-trained defect localization model is then used to perform language encoding processing on the defect report input vector and the code block input vector of each code block to obtain the defect report feature vector and the code block feature vector of each code block. The feature vector is an optimized input vector that reduces the amount of data while retaining vector features. Defect code localization is performed based on the correlation probability between the defect report feature vector and the code block feature vector of each code block. This allows for a quantitative measurement of the matching degree between each code block and the defect report, thereby improving the accuracy and efficiency of defect code localization. It is a low-cost, high-efficiency, and high-accuracy defect code localization method that can save on defect code localization costs while ensuring the effectiveness of defect code localization. It solves the problems of manual defect location methods, which require developers to have a thorough understanding of the software project's architecture and business logic and repeatedly debug the code, resulting in high labor and time costs and low efficiency and speed in locating defective code. It also solves the problems of automatic defect location methods, which require repeated running of test cases, resulting in limited efficiency in locating defective code, high demand for computing resources, low economic efficiency, and high cost in locating defective code.

[0055] Figure 4 This is a flowchart illustrating another method for determining project defect codes provided by the present invention. Based on the above embodiments, this embodiment provides a preferred method for determining project defect codes, essentially a method with more complete process details. Specifically, as shown... Figure 4 As shown, the method includes:

[0056] S201. Obtain the current defect report and code change information for the current project.

[0057] S202. Extract the title, content, and reporter of the current defect report to obtain the defect report feature text. Then, use a byte pair encoding algorithm to perform word embedding encoding on the defect report feature text to obtain the defect report input vector.

[0058] S203. According to the code change record, the code change information is split to obtain at least one code block, and the at least one code block is word-embedding encoding is performed on the at least one code block using the byte pair encoding algorithm to obtain the at least one code block input vector.

[0059] S204. Using a pre-trained defect localization model, perform language encoding processing on the defect report input vector and at least one code block input vector to obtain the defect report feature vector and at least one code block feature vector.

[0060] S205. Using a pre-trained defect localization model, perform defect localization and classification processing on the defect report feature vector and at least one code block feature vector to obtain at least one association probability.

[0061] The correlation probability corresponds one-to-one with the feature vector of the code block.

[0062] S206. Sort at least one association probability in descending order of association degree, and select alternative association probabilities from the association probability sorting results using a preset number of defect codes or a preset defect code association threshold.

[0063] S207. Identify the code block corresponding to the backup association probability as the defect code.

[0064] Figure 5 This is a schematic diagram of the structure of a device for determining project defect codes provided by the present invention. Figure 5 As shown, the device includes: an information acquisition module 301, a feature vector determination module 302, and a defect code localization module 303.

[0065] The information acquisition module 301 is used to acquire the current defect report and code change information of the current project, and to perform word embedding encoding processing on the current defect report and code change information to obtain the defect report input vector and at least one code block input vector; wherein, the code change information includes at least one code block, and the code block and the code block input vector correspond one-to-one.

[0066] The feature vector determination module 302 is used to perform language encoding processing on the defect report input vector and at least one code block input vector using a pre-trained defect localization model to obtain the defect report feature vector and at least one code block feature vector; wherein, the code block feature vector and the code block input vector correspond one-to-one.

[0067] The defect code localization module 303 is used to perform defect localization and classification processing on the defect report feature vector and at least one code block feature vector using a pre-trained defect localization model, to obtain at least one association probability, and to determine the defect code of the current defect report based on the at least one association probability; wherein, the association probability corresponds one-to-one with the code block feature vector.

[0068] Optionally, the information acquisition module 301 is specifically used to: extract the title, report content and reporter of the current defect report to obtain the defect report feature text, and use the byte pair encoding algorithm to perform word embedding encoding on the defect report feature text to obtain the defect report input vector; according to the code change record, split the code change information to obtain at least one code block, and use the byte pair encoding algorithm to perform word embedding encoding on at least one code block to obtain at least one code block input vector.

[0069] Optionally, the device for determining project defect codes may also include a model training module, which is used to determine the defect localization model.

[0070] Optionally, the model training module is specifically used for: obtaining at least one code change block and at least one code change comment from similar projects in the current project, with a one-to-one correspondence between the code change block and the code change comment; performing word embedding encoding on each code change block and each code change comment to obtain at least one code change block input vector and at least one code change comment input vector; determining the positive training samples and negative training samples of the defect localization model based on the correspondence between the at least one code change block input vector and at least one code change comment input vector, wherein the number of positive training samples and negative training samples of the defect localization model is the same, the code change block input vector and the code change comment input vector are corresponding in the positive training samples of the defect localization model, and the code change block input vector and the code change comment input vector are not corresponding in the negative training samples of the defect localization model; training the initial defect localization model using the positive training samples and negative training samples of the defect localization model with the cross-entropy function as the loss function to obtain a backup defect localization model; and fine-tuning the backup defect localization model using at least one code block and at least one historical defect report from the current project to obtain the defect localization model.

[0071] Optionally, the model training module is specifically used for: performing word embedding encoding on at least one code block and at least one historical defect report of the current project to obtain at least one code block input vector and at least one historical defect report input vector; determining positive and negative fine-tuning samples of the defect localization model based on the correspondence between the at least one code block input vector and the at least one historical defect report input vector, wherein the number of samples in the positive and negative fine-tuning samples of the defect localization model is the same, the code block input vector and the historical defect report input vector are corresponding in the positive fine-tuning samples of the defect localization model, and the code block input vector and the historical defect report input vector are not corresponding in the negative fine-tuning samples of the defect localization model; and fine-tuning the standby defect localization model using the positive and negative fine-tuning samples of the defect localization model with the cross-entropy function as the loss function to obtain the defect localization model.

[0072] Optionally, for any code block feature vector, the defect code localization module 303 is specifically used to: concatenate the defect report feature vector and the code block feature vector to obtain a first candidate feature vector, wherein the first candidate feature vector is (U, V, |UV|), where U is the defect report feature vector and V is the code block feature vector; use the feature extraction layer in the defect localization model to compress the dimension of the first candidate feature vector and extract the high-level features of the first candidate feature vector to obtain a second candidate feature vector; perform relevance analysis on the second candidate feature vector to obtain the relevance probability and irrelevance probability of the second candidate feature vector, the sum of the relevance probability and the irrelevance probability being 1.

[0073] Optionally, the defect code location module 303 is specifically used to: sort at least one association probability in descending order of association degree, and use a preset number of defect codes or a preset defect code association threshold to filter out alternative association probabilities from the association probability sorting results; and determine the code block corresponding to the alternative association probability as the defect code.

[0074] The device for determining project defect codes provided in this embodiment can execute the method for determining project defect codes provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the method.

[0075] Figure 6This is a schematic diagram of the structure of an electronic device provided by the present invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0076] like Figure 6 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the read-only memory (ROM) 12 or loaded from storage unit 18 into the random access memory (RAM) 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0077] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0078] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as methods for identifying project defect codes.

[0079] In some embodiments, the method for determining project defect codes may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the method for determining project defect codes described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the method for determining project defect codes by any other suitable means (e.g., by means of firmware).

[0080] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0081] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0082] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0083] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0084] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0085] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0086] In one embodiment, the present invention further includes a computer program product comprising a computer program that, when executed by a processor, implements a method for determining project defect codes according to any embodiment of the present invention.

[0087] In the implementation of a computer program product, computer program code for performing the operations of this invention can be written in one or more programming languages ​​or a combination thereof. Programming languages ​​include object-oriented programming languages ​​as well as conventional procedural programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0088] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0089] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A method for determining project defect codes, characterized in that, include: Obtain the current defect report and code change information of the current project, and perform word embedding encoding on the current defect report and the code change information to obtain a defect report input vector and at least one code block input vector; wherein, the code change information includes at least one code block, and the code block and the code block input vector correspond one-to-one; Using a pre-trained defect localization model, language encoding is performed on the defect report input vector and the at least one code block input vector to obtain a defect report feature vector and at least one code block feature vector; wherein, the code block feature vector and the code block input vector correspond one-to-one. Using a pre-trained defect localization model, defect localization and classification processing is performed on the defect report feature vector and the at least one code block feature vector to obtain at least one association probability, and the defect code of the current defect report is determined based on the at least one association probability; wherein, the association probability corresponds one-to-one with the code block feature vector.

2. The method according to claim 1, characterized in that, The step of performing word embedding encoding on the current defect report and the code change information to obtain a defect report input vector and at least one code block input vector includes: Extract the title, report content, and reporter of the current defect report to obtain the defect report feature text, and use a byte pair encoding algorithm to perform word embedding encoding on the defect report feature text to obtain the defect report input vector; According to the code change record, the code change information is split to obtain at least one code block, and the at least one code block is processed by word embedding encoding using a byte pair encoding algorithm to obtain the input vector of the at least one code block.

3. The method according to claim 1, characterized in that, Determining the defect localization model includes: Obtain at least one code change block and at least one code change comment from similar projects of the current project, wherein the code change block and the code change comment correspond one-to-one; Each code change block and each code change comment are subjected to word embedding encoding to obtain at least one code change block input vector and at least one code change comment input vector. Based on the correspondence between the at least one code change block input vector and the at least one code change comment input vector, the positive training samples and negative training samples of the defect localization model are determined, wherein the number of positive training samples and negative training samples of the defect localization model is the same, the code change block input vector and the code change comment input vector are in a corresponding relationship in the positive training samples of the defect localization model, and the code change block input vector and the code change comment input vector are not in a corresponding relationship in the negative training samples of the defect localization model; Using the positive and negative training samples of the defect localization model, and with the cross-entropy function as the loss function, the initial defect localization model is trained to obtain the backup defect localization model. The backup defect location model is fine-tuned using at least one code block of the current project and at least one historical defect report to obtain the defect location model.

4. The method according to claim 3, characterized in that, The step of fine-tuning the backup defect location model using at least one code block and at least one historical defect report from the current project to obtain the defect location model includes: Word embedding encoding is performed on at least one code block and at least one historical defect report of the current project to obtain at least one code block input vector and at least one historical defect report input vector; Based on the correspondence between the at least one code block input vector and the at least one historical defect report input vector, positive fine-tuning samples and negative fine-tuning samples of the defect localization model are determined, wherein the number of positive fine-tuning samples and negative fine-tuning samples of the defect localization model is the same, the code block input vector and the historical defect report input vector are corresponding in the positive fine-tuning samples of the defect localization model, and the code block input vector and the historical defect report input vector are not corresponding in the negative fine-tuning samples of the defect localization model; Using the positive and negative fine-tuning samples of the defect localization model, and with the cross-entropy function as the loss function, the backup defect localization model is fine-tuned to obtain the defect localization model.

5. The method according to claim 1, characterized in that, For any code block feature vector, a pre-trained defect localization model is used to perform defect localization and classification processing on the defect report feature vector and the code block feature vector to obtain the association probability, including: The defect report feature vector and the code block feature vector are concatenated to obtain a first candidate feature vector, wherein the first candidate feature vector is (U, V, |UV|), where U is the defect report feature vector and V is the code block feature vector; The first candidate feature vector is compressed in dimension by using the feature extraction layer in the defect localization model, and the high-level features of the first candidate feature vector are extracted to obtain the second candidate feature vector. The second candidate feature vector is subjected to relevance analysis to obtain the relevance probability and the irrelevance probability of the second candidate feature vector, and the sum of the relevance probability and the irrelevance probability is 1.

6. The method according to claim 1, characterized in that, Determining the defect code of the current defect report based on the at least one association probability includes: The at least one association probability is sorted in descending order of association degree, and a backup association probability is selected from the association probability sorting results using a preset number of defect codes or a preset defect code association threshold. The code block corresponding to the backup association probability is identified as the defect code.

7. A device for determining project defect codes, characterized in that, The device for determining project defect codes, used to implement the method for determining project defect codes according to any one of claims 1 to 6, comprises: The information acquisition module is used to acquire the current defect report and code change information of the current project, and to perform word embedding encoding processing on the current defect report and the code change information to obtain a defect report input vector and at least one code block input vector; wherein, the code change information includes at least one code block, and the code block and the code block input vector correspond one-to-one; The feature vector determination module is used to perform language encoding processing on the defect report input vector and the at least one code block input vector using a pre-trained defect localization model to obtain a defect report feature vector and at least one code block feature vector; wherein, the code block feature vector and the code block input vector correspond one-to-one; The defect code localization module is used to perform defect localization and classification processing on the defect report feature vector and the at least one code block feature vector using a pre-trained defect localization model, to obtain at least one association probability, and to determine the defect code of the current defect report based on the at least one association probability; wherein, the association probability corresponds one-to-one with the code block feature vector.

8. An electronic device, characterized in that, The electronic device includes: At least one processor; and a memory communicatively connected to said at least one processor; The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform the method for determining project defect codes as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when executed by a processor, implement the method for determining project defect codes as described in any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method for determining project defect codes as described in any one of claims 1 to 6.