Software Defect Localization Method, Device, Equipment, Storage Medium and Program Product
By obtaining defect description information and code file information, using preset models to determine defect keywords and split code blocks, the problem of low efficiency in software defect positioning is solved and efficient positioning of complex logical errors is achieved.
Patent Information
- Application Number
- CN202510046129.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-01-13
AI Technical Summary
In the prior art, software defect positioning efficiency is low, especially for complex logic errors, and static or dynamic code analysis tools cannot effectively locate.
By obtaining defect description information and code file information, the defect keyword is determined using the preset model, and the associated file is determined from the code file, split it into multiple code blocks, and the target code block matching the defect keyword is retrieved.
It improves the efficiency of software defect positioning, narrows the index range, reduces manual intervention, and improves the accuracy and speed of positioning complex logical errors.
Smart Images

Figure CN119441012B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computers, and in particular, to a method, apparatus, device, storage medium, and program product for software defect localization. Background Art
[0002] With the continuous evolution of software functions and code, the scale of software is getting larger and the structure is getting more complex. To ensure the normal operation of software, it is necessary to perform defect localization, defect repair, defect test verification, etc. on the software. Among them, defect localization, as a prerequisite for defect repair, its accuracy and efficiency affect subsequent steps and is an important and indispensable step in the entire defect repair process. However, this step becomes increasingly complex and inefficient as the software scale increases.
[0003] In related technologies, static or dynamic code analysis tools are usually used to detect the corresponding code files of the software or during the software runtime to locate code errors. However, static or dynamic code analysis tools can usually only locate code syntax errors or errors that conform to specific templates (for example, memory errors, deadlocks, etc.), and cannot locate complex logical errors; for relatively complex logical errors, manual location is required.
[0004] As can be seen from the above, the efficiency of software defect localization in related technologies is low. Summary of the Invention
[0005] Multiple aspects of the present application provide a method, apparatus, device, storage medium, and program product for software defect localization to solve the problem of low efficiency in software defect localization.
[0006] In a first aspect, an embodiment of the present application provides a method for software defect localization, including:
[0007] Obtain defect description information of a target software and code file information of the target software, where the code file information includes information describing the code files corresponding to the target software;
[0008] Based on the defect description information, code file information, and a preset model, determine defect keywords and determine associated files from the code files corresponding to the target software;
[0009] Split the associated files into multiple code blocks;
[0010] In the multiple code blocks of the associated files, retrieve a target code block that matches the defect keywords.
[0011] In a possible implementation, the associated files include a first associated file and a second associated file; determining defect keywords based on the defect description information, code file information, and a preset model, and determining associated files from the code files corresponding to the target software includes:
[0012] Determining defect keywords based on the defect description information, code file information, and a preset model, and determining a first associated file from the code files corresponding to the target software;
[0013] Determining a second associated file from the code files corresponding to the target software based on the relevance between the defect keywords and the code files.
[0014] In a possible implementation, the associated files further include a third associated file; the method further includes:
[0015] Based on the file dependency graph of the target software, determining, from the code files corresponding to the target software, the code files on which the first associated file depends as the third associated file.
[0016] In a possible implementation, the preset model is further configured to preprocess the defect description information to obtain preprocessed defect description information;
[0017] Retrieving a target code block that matches the defect keywords from multiple code blocks of the associated files includes:
[0018] Using the code blocks obtained by splitting the associated files as the index range, calculating the matching degree of each code block based on the defect description information, the defect keywords, and the preprocessed defect description information;
[0019] Determining the target code block from the code blocks obtained by splitting the associated files based on the matching degree.
[0020] In a possible implementation, determining the target code block from the code blocks obtained by splitting the associated files based on the matching degree includes:
[0021] Sorting the code blocks obtained by splitting the associated files and their dependent files in descending order of the matching degree;
[0022] Generating code block text vectors for the first N code blocks in the sorting result based on the code segments of the code blocks, where N is a positive integer greater than 1;
[0023] Determining the defect text vector of the preprocessed defect description information;
[0024] Determine the target code block from the first N code blocks in the sorting result based on the similarity between the defective text vector and the code block text vector.
[0025] In a possible implementation manner, determining the target code block from the first N code blocks in the sorting result based on the similarity between the defective text vector and the code block text vector includes:
[0026] Determine the similarity threshold corresponding to the target software;
[0027] Based on the similarity between the defective text vector and the code block text vector, determine the code blocks whose similarity is greater than or equal to the similarity threshold as the target code blocks among the first N code blocks in the sorting result.
[0028] In a possible implementation manner, the code file information further includes external dependency description information of the target software, and the external dependency description information is used to describe external components on which the code file corresponding to the target software depends.
[0029] In a possible implementation manner, splitting the associated file into multiple code blocks includes:
[0030] Obtain the abstract syntax tree of the associated file;
[0031] Based on the abstract syntax tree, split the associated file into multiple code blocks.
[0032] In a possible implementation manner, the method further includes:
[0033] Generate a defect location report of the target software based on the code snippet and location information of the target code block, where the location information is information for locating the target code block.
[0034] In a second aspect, an embodiment of the present application provides a software defect location device, including: an acquisition module, a first determination module, a splitting module, and a second determination module, where,
[0035] The acquisition module is configured to acquire defect description information of a target software and code file information of the target software, where the code file information includes information describing the code file corresponding to the target software;
[0036] The first determination module is configured to determine defect keywords based on the defect description information, code file information, and a preset model, and determine an associated file from the code file corresponding to the target software;
[0037] The splitting module is configured to split the associated file into multiple code blocks;
[0038] The second determination module is configured to retrieve a target code block that matches the defect keyword from multiple code blocks of the associated file.
[0039] In a possible implementation manner, the associated file includes a first associated file and a second associated file; the first determination module is specifically configured to:
[0040] Determine a defect keyword based on the defect description information, code file information, and a preset model, and determine the first associated file from the code files corresponding to the target software;
[0041] Determine the second associated file from the code files corresponding to the target software based on the relevance between the defect keyword and the code files.
[0042] In a possible implementation manner, the associated file further includes a third associated file;
[0043] The first determination module is further configured to determine, based on the file dependency graph of the target software, the code files on which the first associated file depends in the code files corresponding to the target software as the third associated file.
[0044] In a possible implementation manner, the preset model is further configured to preprocess the defect description information to obtain preprocessed defect description information; the second determination module is specifically configured to:
[0045] Taking the code blocks obtained by splitting the associated file as the index range, calculate the matching degree of each code block based on the defect description information, the defect keyword, and the preprocessed defect description information;
[0046] Based on the matching degree, determine the target code block from the code blocks obtained by splitting the associated file.
[0047] In a possible implementation manner, the second determination module is specifically configured to:
[0048] Sort the code blocks obtained by splitting the associated file and its dependent files in descending order of the matching degree;
[0049] Generate a code block text vector for the first N code blocks in the sorting result based on the code segments of the code blocks, where N is a positive integer greater than 1;
[0050] Determine the defect text vector of the preprocessed defect description information;
[0051] Based on the similarity between the defect text vector and the code block text vector, determine the target code block from the first N code blocks in the sorting result.
[0052] In a possible implementation, the second determination module is specifically configured to:
[0053] Determine a similarity threshold corresponding to the target software;
[0054] Based on the similarity between the defect text vector and the code block text vector, determine, among the top N code blocks in the sorting result, the code blocks whose similarity is greater than or equal to the similarity threshold as the target code blocks.
[0055] In a possible implementation, the code file information further includes external dependency description information of the target software, and the external dependency description information is used to describe external components on which the code file corresponding to the target software depends.
[0056] In a possible implementation, the splitting module is specifically configured to:
[0057] Obtain the abstract syntax tree of the associated file;
[0058] Based on the abstract syntax tree, split the associated file into multiple code blocks.
[0059] In a possible implementation, the second determination module is further configured to:
[0060] Generate a defect location report of the target software based on the code snippet and location information of the target code block, where the location information is information for locating the target code block.
[0061] In a third aspect, an embodiment of the present application provides an electronic device, including: a memory and a processor;
[0062] The memory stores computer execution instructions;
[0063] The processor executes the computer execution instructions stored in the memory, so that the processor executes the method according to any one of the first aspect.
[0064] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which computer execution instructions are stored, and when the computer execution instructions are executed by a processor, they are used to implement the method according to any one of the first aspect.
[0065] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the method according to any one of the first aspect.
[0066] An embodiment of the present application provides a software defect localization method, apparatus, device, storage medium, and program product. An electronic device can obtain defect description information of a target software and code file information of the target software, and determine defect keywords based on the defect description information, the code file information, and a preset model, and determine associated files from the code files corresponding to the target software. The electronic device can split the associated files into multiple code blocks, and determine target code blocks that match the defect keywords in the code blocks of the associated files. Since the electronic device can determine defect keywords through a preset model and determine associated files with high relevance to the problem, and then can determine target code blocks that match the defect keywords in the code blocks of the associated files, without manual defect localization, and determines the target code blocks with the code blocks of the associated files as the index range, narrowing the index range, the efficiency of defect localization for the target software is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.
[0068] Figure 1 A schematic diagram of an application scenario provided for an exemplary embodiment of the present application;
[0069] Figure 2 A flowchart of a software defect localization method provided for an exemplary embodiment of the present application;
[0070] Figure 3 A schematic diagram of splitting code blocks based on an AST syntax tree provided for an exemplary embodiment of the present application;
[0071] Figure 4 A flowchart of another software defect localization method provided for an exemplary embodiment of the present application;
[0072] Figure 5 A schematic diagram of a file dependency graph provided for an exemplary embodiment of the present application;
[0073] Figure 6 A schematic diagram of the process of a software defect localization method provided for an exemplary embodiment of the present application;
[0074] Figure 7 A schematic diagram of the structure of a software defect localization apparatus provided for an exemplary embodiment of the present application;
[0075] Figure 8 A schematic diagram of the structure of an electronic device provided for an exemplary embodiment of the present application.
[0076] Through the above-mentioned accompanying drawings, specific embodiments of the present application have been shown, and more detailed descriptions will be provided hereinafter. These drawings and textual descriptions are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. Detailed Description of the Embodiments
[0077] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data that have been authorized by the user or fully authorized by all parties. And the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards, and corresponding operation entrances are provided for users to choose to authorize or refuse.
[0078] To make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with specific embodiments of the present application and the corresponding accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0079] Next, for the convenience of understanding the technical solutions of the present application, the concepts involved in the present application will be explained first.
[0080] Abstract Syntax Tree (AST): It is a tree-like representation of the abstract syntax structure of a code file, which shows the syntax structure of the code in the form of a tree.
[0081] File Dependency Graph: Due to mutual references between files, a dependency relationship will be generated. According to the dependency relationship between files, a file dependency graph can be formed, which is a structured representation method for describing the mutual dependency relationship between files.
[0082] Next, in combination with Figure 1 , the application scenarios of the present application will be described.
[0083] Figure 1 FIG. Figure 1 shows a schematic diagram of an application scenario provided for an exemplary embodiment of the present application. Please refer to
[0084] When the target software exhibits unexpected behaviors (i.e., behaviors caused by software defects), the electronic device can, based on the defect description information, locate defects in multiple code files in the code library corresponding to the target software, so as to determine the target code block with defects among the multiple code files.
[0085] For example, if the target software is a shopping software, when a user performs an order placement operation through the software and an abnormal situation of duplicate order placement occurs, this indicates that there is a defect of duplicate order placement in the target software. The defect description information is "When the user placed an order to purchase product A through the target software at 15:00 today, a problem of duplicate order placement occurred. It is not clear whether it is caused by a submission logic error on the front-end page or an implementation defect in the back-end order processing logic. It is hoped to find out the specific reason for the duplicate order placement in order to solve this problem." Then the electronic device can, based on this defect description information, locate defects in 100 code files corresponding to the target software, so as to determine the target code block with defects among the 100 code files. For example, the target code blocks can be code block 1 and code block 3.
[0086] In the related art, usually static or dynamic code analysis tools are used to detect the code files corresponding to the software or when the software is running to locate code errors. However, static or dynamic code analysis tools can usually only locate code syntax errors or errors that conform to specific templates (such as memory errors, deadlocks, etc.); for relatively complex logical errors, manual location is required. As can be seen from the above, the efficiency of defect location for software in the related art is low.
[0087] In the embodiments of the present application, the electronic device can, based on the defect description information, the code file information of the target software, and a preset model, determine defect keywords and determine associated files from the code files corresponding to the target software. Furthermore, the electronic device can determine the target code block that matches the defect keywords in the code blocks of the associated files. Since the electronic device can determine the defect keywords through the preset model and determine the associated files with high relevance to the problem, and then can determine the target code block that matches the defect keywords in the code blocks of the associated files, without the need for manual defect location, and determines the target code block with the code blocks of the associated files as the index range, narrowing the index range, thus improving the efficiency of defect location for the target software.
[0088] Next, the technical solutions shown in the present application will be described in detail through specific embodiments. It should be noted that the following several embodiments can exist independently or be combined with each other. For the same or similar content, it will not be repeated in different embodiments.
[0089] The execution subject of the embodiments of this application can be an electronic device or a software defect localization device provided in the electronic device. The software defect localization device can be implemented by software or by a combination of software and hardware. The software defect localization device can be a processor in the electronic device. For the sake of easy understanding, in the following, the execution subject is taken as an electronic device as an example for description.
[0090] Figure 2 It is a schematic flowchart of a software defect localization method provided for an exemplary embodiment of this application. Please refer to Figure 2 , the method may include:
[0091] S201. Obtain defect description information of the target software and code file information of the target software.
[0092] Optionally, the defect description information of the target software can be input by the user or automatically generated by the electronic device according to the defects of the target software.
[0093] The code file information may include information describing the code files corresponding to the target software.
[0094] Optionally, the code files corresponding to the target software can be stored in the code library according to the directory tree corresponding to the target software. The electronic device can determine the directory tree in the code library and determine the directory tree as the information describing the code files corresponding to the target software, so the code file information may include the directory tree.
[0095] Optionally, the code library may include an external dependency description file. The electronic device can parse the external dependency description file to obtain external dependency description information. The external dependency description information can be used to describe the external components (such as dependency libraries, dependency frameworks, third-party services, etc.) on which the code files corresponding to the target software depend. For example, the external dependency description information may include the names, versions, sources, etc. of the external components on which the code files corresponding to the target software depend.
[0096] Optionally, the code file information may further include external dependency description information corresponding to the target software.
[0097] For example, if the target software is a shopping software, the defect description information of the target software input by the user is "When the user placed an order to purchase product A at 15:00 today through the target software, a problem of duplicate orders occurred. It is not clear whether it is caused by an error in the submission logic of the front-end page or a defect in the implementation of the back-end order processing logic. I hope to find out the specific reason for the duplicate order to solve this problem", and the code library corresponding to the shopping software is code library 1, then the electronic device obtains the defect description information and the code file information, and the code file information may include the external dependency description information and the directory tree 1 determined in the code library 1.
[0098] S202. Determine defect keywords based on defect description information, code file information, and a preset model, and determine associated files from the code files corresponding to the target software.
[0099] Since the defect description information can be described by users and may include redundant information, after the electronic device inputs the defect description information into the preset model, the preset model can preprocess the defect description information to obtain the preprocessed defect description information. Optionally, the preprocessing may include reference disambiguation, removal of redundant information, etc.
[0100] Optionally, in any embodiment of the present application, the preset model may be a pre-trained large language model (LLM). The preset model may contain more than one billion parameters, and the underlying transformer of the preset model may contain a series of neural networks. In one or more embodiments, the preset model may include an encoder and / or a decoder and has a self-attention function. Through the encoder and / or the decoder, the meaning can be extracted from the input text, and the structural relationship between contexts can be clarified. The preset model is pre-trained on a public dataset containing more than 1TB (terabyte) of text data to learn general language representations.
[0101] To adapt to the application scenario of software defect localization in the present application, we trained the large language model on a dedicated dataset containing thousands of defect description information and code file information to obtain the preset model. The preset model can receive defect description information and code file information as inputs, and can preprocess the defect description information, output the preprocessed defect description information and defect keywords; and can determine and output the first associated file in the code files corresponding to the target software according to the defect keywords and code file information.
[0102] Through the preset model, the semantics of the defect description information can be better captured, the defect description information can be preprocessed to generate a concise and accurate preprocessed defect description information, and then the defect keywords can be more accurately determined; the semantic relevance between the defect keywords and the code files can also be better captured to more accurately retrieve the first associated file, strengthening the objectivity and relevance of determining the first associated file, and thus improving the accuracy of retrieving the target code block.
[0103] For example, if the defect description information is as shown in the above example, the electronic device can preprocess the defect description information through a preset model, and the preprocessed defect description information is "When placing an order for product A at 15:00, there is a problem of duplicate orders. Please find out whether the problem is caused by an error in the submission logic of the front-end page or a defect in the implementation of the back-end order processing logic."
[0104] Optionally, the following method can be used to determine defect keywords based on the defect description information, code file information, and preset model, and to determine associated files from the code files corresponding to the target software: Determine defect keywords based on the defect description information, code file information, and preset model, and determine a first associated file from the code files corresponding to the target software; determine a second associated file from the code files corresponding to the target software based on the relevance between the defect keywords and the code files. The associated files can include the first associated file and the second associated file.
[0105] Optionally, the electronic device can extract defect keywords from the preprocessed defect description information through a preset model, and determine a first associated file from the code files corresponding to the target software according to the defect keywords and the code file information.
[0106] For example, if the preprocessed defect description information is as shown in the above example, the electronic device can extract 3 defect keywords from the preprocessed defect description information through a preset model, namely placing an order, submission logic, and order processing logic. If the code file information includes the directory tree corresponding to the shopping software and the external dependency description information, and there are 100 code files corresponding to the target software in the code library, the electronic device can, through the preset model, determine 4 code files related to order processing as the first associated files from the 100 code files based on "placing an order" and "order processing logic" and the code file information. Suppose the 4 code files can be Order / Dao.java, Order / Controller.java, Order / Handler.java, and Order / Transform.java respectively; it can determine 3 code files related to the submission logic as the first associated files from the 100 code files based on "submission logic" and the code file information. Suppose the 3 code files can be Submit / Handler.java, Submit / Fail.java, and Submit / Verify.java respectively. Then the electronic device can determine 7 first associated files, namely code file 1 (i.e., Order / Dao.java), code file 2 (i.e., Order / Controller.java),..., code file 7 (Submit / Verify.java).
[0107] Optionally, the electronic device can calculate the relevance between the defect keywords and the code files through a preset relevance algorithm, and then determine the code files with a relevance greater than or equal to the first relevance threshold as the second associated files in the code files corresponding to the target software.
[0108] Exemplarily, the preset relevance algorithm can be the BM25 algorithm, the Term Frequency-Inverse Document Frequency (TF-IDF) algorithm, etc.
[0109] Optionally, the first relevance threshold can be preset manually. For example, the first relevance threshold can be 0.8.
[0110] For example, if there are 100 code files corresponding to the target software in the code library, and there are 3 defect keywords, namely placing an order, submission logic, and order processing logic, and the preset relevance algorithm is the BM25 algorithm, the electronic device can calculate the relevance between these 3 defect keywords and the 100 code files through the BM25 algorithm to obtain 100 relevance degrees. If the first relevance threshold is 0.8, the code files with a relevance greater than 0.8 can be determined as the second associated files. Suppose the electronic device can determine 5 second associated files, namely code file 3, code file 10, code file 21, code file 22, and code file 23.
[0111] After the electronic device determines at least one first associated file and at least one second associated file, it can determine that the associated files include at least one first associated file and at least one second associated file.
[0112] For example, if at least one first associated file includes code file 1, code file 2, code file 3,..., code file 7, and at least one second associated file includes code file 3, code file 10, code file 21, code file 22, code file 23, the electronic device can determine that there are 11 associated files, namely code file 1, code file 2, code file 3, code file 4, code file 5, code file 6, code file 7, code file 10, code file 21, code file 22, and code file 23.
[0113] S203. Split the associated files into multiple code blocks.
[0114] Optionally, the associated files can be split into multiple code blocks in the following way: obtain the abstract syntax tree of the associated file; based on the abstract syntax tree, split the associated file into multiple code blocks.
[0115] For any associated file, the electronic device can parse the associated file through an AST parsing tool to obtain the AST syntax tree corresponding to the associated file, and then split the associated file into multiple code blocks (Chunks) according to the structure of the AST syntax tree.
[0116] A code block is a set of logically related code statements. For any code block, the code block may include code snippets and location information for defect keyword retrieval. For example, the location information may include the module name to which the code block belongs, the file path where the code block is located, and the line number range of the code block in the code file, etc. For example, the location information included in code block 1 may be: the module name to which code block 1 belongs is the order processing module, the file path where it is located is "shopping-software / order / file 1", and the line number range of code block 1 in code file 1 is 010 - 020.
[0117] Next, in conjunction with Figure 3 , the AST syntax tree will be described.
[0118] Figure 3 This is a schematic diagram of splitting code blocks based on the AST syntax tree provided for an exemplary embodiment of the present application. Please refer to Figure 3 , the AST syntax tree can be obtained by parsing code file 1. The AST syntax tree may include multiple code blocks, namely: code block 1, code block 2,..., code block 10. The electronic device can split code file 1 into 10 code blocks according to the structure of the AST syntax tree.
[0119] For example, if the electronic device determines that there are 11 associated files, namely code file 1, code file 2, code file 3, code file 4, code file 5, code file 6, code file 7, code file 10, code file 21, code file 22, and code file 23, then for code file 1, code file 1 can be split into 10 code blocks; code file 2 can be split into 8 code blocks;...; code file 23 can be split into 15 code blocks. Assume that there are a total of 200 code blocks corresponding to the 11 associated files, namely code block 1, code block 2,..., code block 200.
[0120] S204. Among the multiple code blocks of the associated file, retrieve the target code block that matches the defect keyword.
[0121] The electronic device can calculate the relevance between the defect keyword and the multiple code blocks corresponding to the associated file through a preset relevance algorithm, and thus determine the code block with a relevance greater than the second relevance threshold as the target code block that matches the defect keyword.
[0122] It should be noted that when calculating the relevance between multiple defect keywords and the code block through the preset relevance algorithm, the multiple defect keywords are regarded as a whole to calculate the relevance with the code block.
[0123] For example, if there are 200 code blocks, and there are 3 defect keywords which are placing an order, submission logic, and order processing logic respectively, and the preset relevance algorithm is the BM25 algorithm, the electronic device can calculate the relevance between these 3 defect keywords and each code block through the BM25 algorithm, obtaining 200 relevance degrees. If the second relevance threshold is 0.9, the electronic device can determine the code blocks with a relevance greater than 0.9 as the target code blocks. Suppose there are 8 target code blocks, which are code block 1, code block 3, code block 20, code block 65, code block 86, code block 126, code block 153, and code block 171.
[0124] In the embodiment of the present application, the electronic device can obtain the defect description information of the target software and the code file information of the target software, and determine the defect keywords and determine the associated files from the code files corresponding to the target software based on the defect description information, the code file information, and the preset model. The electronic device can split the associated files into multiple code blocks, and retrieve the target code blocks that match the defect keywords among the multiple code blocks of the associated files. Since the electronic device can determine the defect keywords through the preset model and determine the associated files with high relevance to the problem, and then can determine the target code blocks that match the defect keywords among the code blocks of the associated files, without manual defect location, and determining the target code blocks with the code blocks of the associated files as the index range, narrowing the index range, so the efficiency of defect location for the target software is improved.
[0125] Next, based on Figure 2 the embodiment shown, in combination with Figure 4 this, the above software defect location method will be described in detail.
[0126] Figure 4 It is a schematic flowchart of another software defect location method provided by an exemplary embodiment of the present application. Please refer to Figure 4 , the method may include:
[0127] S401. Obtain the defect description information of the target software and the code file information of the target software.
[0128] It should be noted that the execution process of step S401 can refer to step S201, and details will not be repeated here.
[0129] S402. Determine the defect keywords based on the defect description information, the code file information, and the preset model, and determine the first associated file from the code files corresponding to the target software.
[0130] The electronic device preprocesses the defect description information through the preset model to obtain the preprocessed defect description information, and extracts the defect keywords from the preprocessed defect description information.
[0131] After the electronic device determines the defect keywords, it can determine the first associated file in the code file corresponding to the target software according to the defect keywords and the code file information through a preset model.
[0132] Optionally, the first associated file can be determined in the following manner: The electronic device can determine the function of each code file according to the external dependency description information in the code file information through a preset model, and determine the semantic relevance between the defect keywords and each file path in the directory tree. Furthermore, the code file corresponding to the file path with a semantic relevance greater than or equal to the preset relevance threshold is determined as the first associated file.
[0133] It should be noted that the file path refers to the complete file path of each code file in the directory tree.
[0134] Optionally, the preset relevance threshold can be preset manually or determined according to the semantic relevance distribution.
[0135] For example, if the electronic device determines 3 defect keywords, namely placing an order, submission logic, and order processing logic, and there are 100 code files in the code library, the electronic device determines the functions of the 100 code files according to the external dependency description information through a preset model, and determines the semantic relevance between the 3 defect keywords and the file paths corresponding to the 100 files in the directory tree. For example, for code file 1, if the external dependency description information includes that code file 1 depends on the order database, it can be determined that the content in code file 1 is order-related content, and the semantic relevance between the defect keywords and the file path of code file 1 (Order / Dao.java) can be determined to be 0.9.
[0136] If the preset relevance threshold is 0.7, and assume that the file paths with semantic relevance greater than or equal to the preset relevance threshold are: file path 1 (Order / Dao.java), file path 2 (Order / Controller.java), file path 3 (Order / Handler.java), file path 4 (Order / Transform.java), file path 5 (Submit / Handler.java), file path 6 (Submit / Fail.java), and file path 7 (Submit / Verify.java), then the files corresponding to these 7 file paths can be determined as the first associated files, that is, the first associated files include: code file 1 corresponding to file path 1 (i.e., Order / Dao.java file), code file 2 corresponding to file path 2 (i.e., Order / Controller.java file),..., code file 7 (i.e., Submit / Verify.java file).
[0137] S403. Determine the second associated files from the code files corresponding to the target software based on the relevance between the defect keywords and the code files.
[0138] Optionally, the execution process of step S403 can refer to the relevant content in step S202, which will not be elaborated here.
[0139] S404. Based on the file dependency graph of the target software, determine the code files on which the first associated files depend as the third associated files among the code files corresponding to the target software.
[0140] The code library can include at least one code file, and there may be a dependency relationship between the at least one code file. The first associated files are determined among the at least one code file. Based on actual development experience, there may also be defects in the code files on which the first associated files depend. Therefore, when performing software localization, it is also necessary to consider calculating the association degree between the code files on which the first associated files depend and the defects.
[0141] The file dependency graph can include the dependency relationships between at least one code file corresponding to the target software.
[0142] Next, in combination with Figure 5 , the file dependency graph will be described.
[0143] Figure 5 It is a schematic diagram of a file dependency graph provided by an exemplary embodiment of the present application. Please refer to Figure 5 , if the target software corresponds to 100 code files, then the file dependency graph can include the dependency relationships between each of the 100 code files. For exampleFigure 5 As shown in Figure 5 , in the file dependency graph, code file 1 depends on code files 25 and 26, code file 2 depends on code file 24, code files 3 and 4 do not depend on any code files, code file 5 depends on code files 8 and 9, code file 6 does not depend on any code files, code file 7 depends on code files 14, 15 and 16, ……, code file 99 depends on code file 100.
[0144] Optionally, the electronic device may determine the file dependency graph of the target software, and based on the file dependency graph, determine the code files upon which the first associated file depends in the code library, and determine the code files upon which the first associated file depends as the third associated files.
[0145] For example, if there are 7 first associated files, namely code files 1, 2, 3, ……, 7, and the file dependency graph is as Figure 5 shown in Figure 5 , then the electronic device may, based on the file dependency graph, determine that the code files upon which code file 1 depends are code files 25 and 26, the code file upon which code file 2 depends is code file 24, ……, the code files upon which code file 7 depends are code files 14, 15 and 16. The electronic device may determine that there are a total of 8 code files upon which the 7 first associated files depend (i.e., code files 25, 26, 24, 8, 9, 14, 15, and 16), and determine these 8 code files as the third associated files.
[0146] S405. Determine that the associated files include the first associated files, the second associated files, and the third associated files.
[0147] For example, if the electronic device determines that there are 7 first associated files, 5 second associated files, and 8 third associated files as shown in Table 1 respectively:
[0148] Table 1
[0149]
[0150] Then the electronic device may determine that there are a total of 19 associated files, as shown in Table 1.
[0151] S406. Split the associated files into multiple code blocks.
[0152] It should be noted that the execution process of step S406 may refer to step S203, and will not be elaborated here.
[0153] S407: Using the code blocks obtained by splitting the associated files as the index range, calculate the matching degree of each code block based on the defect keywords and the preprocessed defect description information.
[0154] Optionally, after the electronic device obtains multiple code blocks obtained by splitting the associated files, it can perform text indexing on the multiple code blocks. Specifically, the electronic device can simplify the code blocks to generate labels corresponding to the code blocks, and generate a text index based on the labels of the code blocks. The text index includes the labels of each code block.
[0155] Optionally, for any one code block, the electronic device can remove the useless information in the code block (such as curly braces, semicolons, etc.) to simplify the code block, and use the simplified code block as the label of the code block. Then, the label of the code block includes the useful information in the code block.
[0156] For example, if there are 19 associated files, and these 19 associated files correspond to a total of 500 code blocks, the electronic device can simplify each of the 500 code blocks to obtain the label of each code block, and generate a text index based on the labels of the 500 code blocks. The text index can include the labels of the 500 code blocks.
[0157] Since the text index is established based on the labels of each code block, the electronic device using the text index as the index range is equivalent to using the code blocks obtained by splitting the associated files as the index range. Indexing based on the text index can improve the indexing efficiency.
[0158] The electronic device can calculate the matching degree of the labels of each code block in the text index based on the defect keywords and the preprocessed defect description information through a preset correlation algorithm.
[0159] Specifically, the electronic device can calculate the first correlation degree between the defect keywords and the labels of each code block in the text index through a preset correlation algorithm; it can calculate the second correlation degree between the preprocessed defect description information and the labels of each code block in the text index through a preset correlation algorithm.
[0160] For any one code block, the electronic device can determine the matching degree of the label of the code block according to the first correlation degree and the second correlation degree corresponding to the label of the code block, and use the matching degree of the label of the code block as the matching degree of the code block.
[0161] Optionally, the matching degree can be a statistical value calculated based on the first correlation degree and the second correlation degree. For example, the statistical value can be an average value or a weighted average value, etc.
[0162] For example, if the defect description information is "When the user placed an order to purchase product A through the target software at 15:00 today, a problem of duplicate orders occurred. It is not clear whether it was caused by a submission logic error on the front-end page or an implementation defect in the back-end order processing logic. It is hoped to identify the specific reason for the duplicate orders in order to solve this problem"; the preprocessed defect description information is "A problem of duplicate orders occurred when placing an order for product A at 15:00. Please find out whether the problem was caused by a submission logic error on the front-end page or an implementation defect in the back-end order processing logic"; the defect keywords are placing an order, submission logic, and order processing logic, and the text index includes tags of 500 code blocks, then the electronic device can calculate the first relevance between the preprocessed defect description information and the tags of these 500 code blocks respectively through the BM25 algorithm to obtain 500 first relevances; and can calculate the second relevance between the defect keywords and the tags of these 500 code blocks respectively through the BM25 algorithm to obtain 500 second relevances. Suppose the 500 first relevances and 500 second relevances are as shown in Table 2:
[0163] Table 2
[0164]
[0165] If the matching degree is the average of the first relevance and the second relevance, then the matching degree of the tag of each code block can be calculated as shown in Table 2, that is, the matching degree of each code block is as shown in Table 2.
[0166] S408. Determine the target code block from the code blocks split from the associated file based on the matching degree.
[0167] Optionally, the target code block can be determined from the code blocks split from the associated file based on the matching degree in the following way: sort the code blocks split from the associated file in descending order of the matching degree; generate the code block text vectors of the first N code blocks in the sorting result based on the code segments of the code blocks; determine the defect text vector of the preprocessed defect description information; and determine the target code block from the first N code blocks in the sorting result based on the relevance between the defect text vector and the code block text vector.
[0168] N can be a positive integer greater than 1, and N can be preset manually. For example, N can be 10.
[0169] Since the matching degree of each code block is calculated by a preset relevance algorithm without using the context information of the code snippets in the code block, after the top N code blocks in the sorting result are determined, the code block text vectors of the top N code blocks can be generated respectively based on the code snippets in the top N code blocks. The electronic device can determine the defect text vector of the preprocessed defect description information, calculate the similarity between the defect text vector and the code block text vectors of the N code blocks through a preset similarity algorithm, and determine the target code block from the N code blocks according to the similarity.
[0170] Exemplarily, the preset similarity algorithm can be a cosine similarity algorithm, an Euclidean distance algorithm, a Pearson correlation coefficient algorithm, etc.
[0171] Optionally, the target code block can be determined from the top N code blocks in the sorting result based on the similarity between the defect text vector and the code block text vector in the following way: determine the similarity threshold corresponding to the target software; based on the similarity between the defect text vector and the code block text vector, determine the code blocks with similarity greater than or equal to the similarity threshold as the target code blocks among the top N code blocks in the sorting result.
[0172] Through the code block text vector, the context information of the code snippets in the code block can be expressed. Therefore, based on the top N code blocks determined according to the matching degree, and based on the relevance between the defect text vector and the code block text vector, the target code block is determined among the top N code blocks, which further improves the accuracy of determining the target code block.
[0173] Optionally, the similarity threshold can be determined according to the relevance distribution between the defect text vector and the code block text vector. For example, the similarity threshold can be 0.8.
[0174] For example, if the matching degrees of the code blocks are as shown in Table 2, the electronic device can sort the 500 code blocks in descending order of the matching degrees to obtain a sorting result. If N is 10, the electronic device can determine the first 10 code blocks in the sorting result. Suppose the first 10 code blocks are Code Block 3, Code Block 1, Code Block 20, Code Block 126, Code Block 65, Code Block 309, Code Block 351, Code Block 201, Code Block 412, and Code Block 465. For Code Block 3, the electronic device can generate a code block text vector corresponding to Code Block 3 based on the code segments in Code Block 3; for Code Block 1, the electronic device can generate a code block text vector corresponding to Code Block 1 based on the code segments in Code Block 1;...; for Code Block 465, the electronic device can generate a code block text vector corresponding to Code Block 465 based on the code segments in Code Block 465. The electronic device can generate a defect text vector of the preprocessed defect description information. The electronic device can calculate the similarity between the defect text vector and each code block text vector through the cosine similarity algorithm. Suppose as shown in Table 3:
[0175] Table 3
[0176]
[0177] If the similarity threshold corresponding to the target software is 0.5, the electronic device can determine Code Block 3, Code Block 1, Code Block 20, Code Block 65, and Code Block 412 as the target code blocks among the 10 code blocks.
[0178] S409. Generate a defect location report for the target software based on the code segments and location information of the target code blocks.
[0179] The location information can be the information used to locate the target code block. For example, the location information may include the module name to which the code block belongs, the file path where the code block is located, and the line number range of the code block in the code file, etc. For example, the location information included in Code Block 1 can be: the module name to which Code Block 1 belongs is the order processing module, the file path where it is located is "shopping-software / order / file 1", and the line number range of Code Block 1 in Code File 1 is 010 - 020.
[0180] For example, if the target software is shopping software and the electronic device determines that there are 5 target code blocks, namely Code Block 3, Code Block 1, Code Block 20, Code Block 65, and Code Block 412, the electronic device can generate a defect location report for the shopping software based on the code segments and location information of the 5 target code blocks.
[0181] It should be noted that in Figure 4Each processing step (S401~S409) shown in the embodiment does not constitute a specific limitation on the software defect localization process. In some other embodiments of the present application, the software defect localization process may include more or fewer steps than Figure 4 the embodiment. For example, the software defect localization process may include Figure 4 some steps in the embodiment, or Figure 4 some steps in the embodiment may be replaced by steps with the same function, or Figure 4 some steps in the embodiment may be split into multiple steps, etc.
[0182] In the embodiment of the present application, the electronic device can obtain defect description information of the target software and code file information of the target software. The electronic device can determine defect keywords based on the defect description information, code file information, and a preset model, and determine a first associated file from the code files corresponding to the target software. Based on the relevance between the defect keywords and the code files, the electronic device can determine a second associated file from the code files corresponding to the target software. The electronic device can also determine, based on the file dependency graph of the target software, the code files on which the first associated file depends as the third associated file in the code files corresponding to the target software. Furthermore, the associated files can be determined to include the first associated file, the second associated file, and the third associated file. The electronic device can split the associated files into multiple code blocks, and calculate the matching degree of each code block based on the defect keywords and the preprocessed defect description information with the code blocks obtained by splitting the associated files as the index range. Furthermore, based on the matching degree, the target code block can be determined from the code blocks obtained by splitting the associated files. The electronic device can generate a defect localization report of the target software based on the code snippet and the localization information of the target code block. Since the electronic device can determine, according to the defect keywords, the associated files with high relevance to the problem from the code files corresponding to the target software, and then can determine the target code block with defects in the code blocks of the associated files, narrowing the index range and improving the speed of defect localization for the target software; and the top N code blocks with high relevance can be determined through a preset similarity algorithm, and then the target code block can be determined from these top N code blocks based on the code block text vectors, which not only considers the keyword matching relevance between the defect description information and the code blocks, but also considers the context information of the code snippets in the code blocks, improving the accuracy of defect localization for the target software. Therefore, the efficiency of defect localization for the target software is comprehensively improved.
[0183] Next, on the basis of any of the above embodiments, in combination with Figure 6 , the above software defect localization method will be further described.
[0184] Figure 6 It is a schematic diagram of the process of a software defect localization method provided by an exemplary embodiment of the present application. Please refer toFigure 6 , including steps ①②③④⑤⑥⑦⑧⑨.
[0185] In step ①, the electronic device can input defect description information and code file information into a preset model.
[0186] For example, the defect description information can be "When the user placed an order to purchase product A through the target software at 15:00 today, a problem of duplicate orders occurred. It is not clear whether it is caused by an error in the submission logic of the front-end page or a defect in the implementation of the back-end order processing logic. It is hoped to find out the specific reason for the duplicate order to solve this problem"; the code file information can include the directory tree determined in code library 1 and external dependency description information.
[0187] In step ②, the electronic device can preprocess the defect description information through the preset model to obtain the preprocessed defect description information. For example, the preprocessed defect description information can be: When placing an order for product A at 15:00, a problem of duplicate orders occurred. Please find out whether the problem is caused by an error in the submission logic of the front-end page or a defect in the implementation of the back-end order processing logic.
[0188] In step ③, the electronic device can determine defect keywords in the preprocessed defect description information through the preset model. For example, if the preprocessed defect description information is as shown in the above example, the defect keywords can be: placing an order, submission logic, and order processing logic.
[0189] In step ④, the electronic device can determine the first associated file in the code library corresponding to the target software according to the defect keywords and code file information. For example, the code library can include multiple code files corresponding to the target software, namely code file 1, code file 2, code file 3,..., code file 100.
[0190] Optionally, the specific execution process of step ④ can refer to the execution process of step S402, which will not be elaborated here.
[0191] For example, the electronic device can determine that there are 7 first associated files, as shown in Table 1.
[0192] In step ⑤, based on the relevance between the defect keywords and the code files, determine the second associated file from the code files corresponding to the target software.
[0193] Optionally, the specific execution process of step ⑤ can refer to the execution process of step S202, which will not be elaborated here.
[0194] For example, the electronic device can determine that there are 5 second associated files, as shown in Table 1.
[0195] In step ⑥, based on the file dependency graph of the target software, the code file with the first associated file dependency is determined as the third associated file in the code file corresponding to the target software.
[0196] Optionally, for the specific execution process of step ⑥, reference can be made to the execution process of step S404, which will not be elaborated here.
[0197] For example, the electronic device can determine that there are 8 third associated files, as shown in Table 1.
[0198] The electronic device can determine that the associated files include the first associated file, the second associated file, and the third associated file. For example, the associated files can be as shown in Table 1.
[0199] In step ⑦, the electronic device can perform splitting processing on the associated files to obtain multiple code blocks.
[0200] For example, if there are 19 associated files, the electronic device can use an AST parsing tool to parse each associated file to obtain the AST syntax tree corresponding to each associated file, and then perform splitting processing on the corresponding associated file according to the AST syntax tree to obtain multiple code blocks. Assuming that 19 associated files are split, a total of 500 code blocks can be obtained, as Figure 6 shown.
[0201] In step ⑧, the electronic device can determine the matching degree corresponding to each code block, and based on the matching degree, screen out the first N code blocks in the sorting result from multiple code blocks.
[0202] It should be noted that for the process of determining the matching degree corresponding to each code block, reference can be made to the relevant content in step S407, which will not be elaborated here.
[0203] For example, if the matching degrees corresponding to the 500 code blocks are as shown in Table 2 and N is 10, then the 500 code blocks are sorted in descending order of the matching degree to obtain the sorting result. The electronic device can determine the first 10 code blocks in the sorting result. Assuming the first 10 code blocks are code block 3, code block 1, code block 20, code block 126, code block 65, code block 309, code block 351, code block 201, code block 412, code block 465, as Figure 6 shown.
[0204] In step ⑨, the electronic device can generate a defect text vector of the preprocessed defect description information and a code block text vector of the first N code blocks, calculate the similarity between the defect text vector and each code block text vector, and then based on the defect text vector and the code block text vectors of each code block, determine the target code block among the first N code blocks in the sorting result.
[0205] For example, if the similarities between the defective text vector and the code block text vectors of the first 10 code blocks are as shown in Table 3, and if the similarity threshold is 0.5, the electronic device can determine code block 3, code block 1, code block 20, code block 65, and code block 412 as target code blocks among these 10 code blocks, as Figure 6 shown below.
[0206] In the embodiment of the present application, the electronic device can obtain the defect description information of the target software and the code file information of the target software. The electronic device can determine defect keywords based on the defect description information, the code file information, and a preset model, and determine a first associated file from the code files corresponding to the target software. Then, based on the relevance between the defect keywords and the code files, the electronic device can determine a second associated file from the code files corresponding to the target software. The electronic device can also determine, based on the file dependency graph of the target software, the code files on which the first associated file depends as third associated files among the code files corresponding to the target software. Furthermore, the associated files can be determined to include the first associated file, the second associated file, and the third associated file. The electronic device can split the associated files into multiple code blocks, calculate the matching degree of each code block based on the defect keywords and the preprocessed defect description information with the code blocks obtained by splitting the associated files as the index range, and determine the first N code blocks with the top sorting results from the multiple code blocks based on the matching degree. Then, based on the defective text vector and the code block text vectors of each code block, the target code blocks can be determined among the first N code blocks. Since the electronic device can determine the associated files with high relevance to the problem from the code files corresponding to the target software according to the defect keywords, and then can determine the target code blocks with defects among the code blocks of the associated files, narrowing the index range and improving the speed of defect location for the target software. Moreover, the first N code blocks with high relevance can be determined through a preset similarity algorithm, and then the target code blocks can be determined among the first N code blocks based on the code block text vectors, which not only considers the keyword matching relevance between the defect description information and the code blocks, but also considers the context information of the code segments in the code blocks, improving the accuracy of defect location for the target software. Therefore, the efficiency of defect location for the target software is comprehensively improved.
[0207] The technical solution of this application can be applied in the software development process to help developers quickly locate and fix software defects. For example, in automated testing, the technical solution of this application can be used to help testers quickly locate defects in test cases. In automated development and question-and-answer systems, it can provide code information related to questions as part of the Retrieval-Augmented Generation (RAG) process. In addition, the technical solution of this application can also be used in the fields of teaching and research to help students and researchers better understand the method of locating software defects.
[0208] Figure 7 The following is a schematic structural diagram of a software defect location device provided by an exemplary embodiment of this application. Please refer to Figure 7 The software defect location device 10 includes: an acquisition module 11, a first determination module 12, a splitting module 13, and a second determination module 14, where
[0209] The acquisition module 11 is configured to acquire defect description information of the target software and code file information of the target software, where the code file information includes information describing the code files corresponding to the target software;
[0210] The first determination module 12 is configured to determine defect keywords based on the defect description information, code file information, and a preset model, and determine associated files from the code files corresponding to the target software;
[0211] The splitting module 13 is configured to split the associated files into multiple code blocks;
[0212] The second determination module 14 is configured to retrieve a target code block that matches the defect keywords from the multiple code blocks of the associated files.
[0213] The software defect location device provided by the embodiments of this application can execute the technical solutions shown in the above method embodiments, and the implementation principles and beneficial effects are similar, so details are not described here again.
[0214] In a possible implementation manner, the associated files include a first associated file and a second associated file; the first determination module 12 is specifically configured to:
[0215] Determine defect keywords based on the defect description information, code file information, and a preset model, and determine the first associated file from the code files corresponding to the target software;
[0216] Determine the second associated file from the code files corresponding to the target software based on the relevance between the defect keywords and the code files.
[0217] In a possible implementation, the associated file further includes a third associated file;
[0218] The first determination module 12 is further configured to determine, based on the file dependency graph of the target software, the code file on which the first associated file depends in the code files corresponding to the target software as the third associated file.
[0219] In a possible implementation, the preset model is further configured to preprocess the defect description information to obtain preprocessed defect description information; the second determination module 14 is specifically configured to:
[0220] Using the code blocks obtained by splitting the associated file as the index range, calculate the matching degree of each code block based on the defect description information, the defect keywords, and the preprocessed defect description information;
[0221] Based on the matching degree, determine the target code block from the code blocks obtained by splitting the associated file.
[0222] In a possible implementation, the second determination module 14 is specifically configured to:
[0223] Sort the code blocks obtained by splitting the associated file in descending order of the matching degree;
[0224] Based on the code segments of the code blocks, generate code block text vectors for the first N code blocks in the sorting result, where N is a positive integer greater than 1;
[0225] Determine the defect text vector of the preprocessed defect description information;
[0226] Based on the similarity between the defect text vector and the code block text vectors, determine the target code block from the first N code blocks in the sorting result.
[0227] In a possible implementation, the second determination module 14 is specifically configured to:
[0228] Determine the similarity threshold corresponding to the target software;
[0229] Based on the similarity between the defect text vector and the code block text vectors, determine the code blocks with similarity greater than or equal to the similarity threshold as the target code blocks from the first N code blocks in the sorting result.
[0230] In a possible implementation, the code file information further includes external dependency description information of the target software, and the external dependency description information is used to describe the external components on which the code files corresponding to the target software depend.
[0231] In a possible implementation, the splitting module 13 is specifically configured to:
[0232] Obtain the abstract syntax tree of the associated file;
[0233] Based on the abstract syntax tree, split the associated file into multiple code blocks.
[0234] In a possible implementation, the second determination module 14 is further configured to:
[0235] Generate a defect location report for the target software based on the code snippet of the target code block and the location information, where the location information is information for locating the target code block.
[0236] The software defect location device provided by the embodiments of the present application can execute the technical solutions shown in the above method embodiments, and its implementation principle and beneficial effects are similar, which will not be elaborated here.
[0237] Figure 8 The figure is a schematic structural diagram of an electronic device provided by an exemplary embodiment of the present application. Please refer to Figure 8 , the electronic device 20 may include a processor 21 and a memory 22. Exemplarily, the processor 21 and the memory 22 are interconnected with each other through a bus 23.
[0238] The memory 22 stores computer-executable instructions;
[0239] The processor 21 executes the computer-executable instructions stored in the memory 22, so that the processor 21 executes the method shown in the above method embodiments.
[0240] Correspondingly, the embodiments of the present application provide a computer-readable storage medium, in which computer-executable instructions are stored, and when the computer-executable instructions are executed by a processor, they are used to implement the method described in the above method embodiments.
[0241] Correspondingly, the embodiments of the present application can also provide a computer program product, including a computer program, which when executed by a processor, can implement the method shown in the above method embodiments.
[0242] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0243] The present invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each flow and / or block in the flowchart illustrations and / or block diagrams, and combinations of flows and / or blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions may be provided to a processor of a general purpose computer, special purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions executed by the processor of the computer or other programmable data processing apparatus create means for implementing the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 block or multiple blocks.
[0244] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instruction means that implement the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 block or multiple blocks.
[0245] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, such that the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 block or multiple blocks.
[0246] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0247] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0248] A computer-readable medium includes both permanent and non-permanent, removable and non-removable media that can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information that can be accessed by a computing device. As defined herein, a computer-readable medium does not include transitory computer-readable media such as modulated data signals and carrier waves.
[0249] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or apparatus comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or apparatus. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or apparatus comprising the element.
[0250] The above description is only for the embodiments of the present application and is not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. A software defect localization method, characterized in that Including: Obtain defect description information of the target software and code file information of the target software, where the code file information includes a directory tree and external dependency description information corresponding to the target software; the external dependency description information is used to describe external components on which the code files corresponding to the target software depend; Based on the defect description information, code file information, and a preset model, determine defect keywords; where the preset model is used to preprocess the defect description information to obtain preprocessed defect description information and defect keywords; the preset model is a pre-trained large language model; Determine associated files from the code files corresponding to the target software; split the associated files into multiple code blocks; the code blocks include code segments for defect keyword retrieval and location information; In the multiple code blocks of the associated files, perform a simplification process on each code block in the multiple code blocks to generate a label corresponding to each code block, and generate a text index based on the labels of the code blocks; calculate a first correlation degree between the defect keywords and the labels of each code block in the text index through a preset correlation algorithm; calculate a second correlation degree between the preprocessed defect description information and the labels of each code block in the text index through a preset correlation algorithm; determine the matching degree of the labels of each code block according to the first correlation degree and the second correlation degree corresponding to the label of the code block, and use the matching degree of the label of the code block as the matching degree of the code block; sort the code blocks split from the associated files in descending order of the matching degree; generate code block text vectors for the first N code blocks in the sorting result based on the code segments of the code blocks; determine a defect text vector of the preprocessed defect description information; determine a target code block from the first N code blocks in the sorting result based on the correlation degree between the defect text vector and the code block text vectors, where the text vector of the code block is used to express the context information of the code segment in the code block; The determining the associated files from the code files corresponding to the target software includes: Determine the function of each code file based on the external dependency description information, determine the semantic correlation degree between the defect keywords and the path of each code file in the directory tree, and determine the code file corresponding to the file path with a semantic correlation degree greater than or equal to a preset correlation degree threshold as the first associated file.
2. The method according to claim 1, characterized in that The associated files further include a second associated file and a third associated file; the determining the associated files from the code files corresponding to the target software further includes: Determine the second associated file from the code files corresponding to the target software based on the correlation degree between the defect keywords and the code files; Based on the file dependency graph of the target software, determine the code files on which the first associated file depends in the code files corresponding to the target software as the third associated file.
3. The method according to claim 1, characterized in that The determining the target code block from the first N code blocks in the sorting result based on the similarity between the defect text vector and the code block text vectors includes: Determine a similarity threshold corresponding to the target software; Based on the similarity between the defective text vector and the code block text vector, determine the target code blocks from the top N code blocks in the sorting result where the similarity is greater than or equal to the similarity threshold.
4. The method according to any one of claims 1 to 3, characterized in that, Split the associated file into multiple code blocks, including: Obtain the abstract syntax tree of the associated file; Based on the abstract syntax tree, split the associated file into multiple code blocks.
5. The method according to any one of claims 1 to 3, characterized in that, The method further includes: Generate a defect location report for the target software based on the code snippet and location information of the target code block, where the location information is the information used to locate the target code block.
6. A software defect localization device, characterized in that, Including: An acquisition module, a first determination module, a splitting module, and a second determination module, where The acquisition module is configured to obtain defect description information of a target software and code file information of the target software, where the code file information includes a directory tree and external dependency description information corresponding to the target software; the external dependency description information is used to describe external components on which the code files corresponding to the target software depend; The first determination module is configured to determine defect keywords based on the defect description information, the code file information, and a preset model; where the preset model is used to preprocess the defect description information to obtain the preprocessed defect description information and defect keywords; the preset model is a pre-trained large language model; The first determination module is further configured to determine an associated file from the code files corresponding to the target software; The splitting module is configured to split the associated file into multiple code blocks; the code blocks include code snippets and location information for defect keyword retrieval; The second determination module is configured to, in the multiple code blocks of the associated file, perform a simplification process on each code block in the multiple code blocks to generate a label corresponding to each code block, and generate a text index based on the labels of the code blocks; calculate a first relevance between the defect keywords and the labels of the code blocks in the text index through a preset relevance algorithm; calculate a second relevance between the preprocessed defect description information and the labels of the code blocks in the text index through a preset relevance algorithm; determine the matching degree of the labels of the code blocks according to the first relevance and the second relevance corresponding to the labels of the code blocks, and use the matching degree of the labels of the code blocks as the matching degree of the code blocks; sort the code blocks split from the associated file in descending order of the matching degree; generate a code block text vector for the top N code blocks in the sorting result based on the code snippets of the code blocks; determine a defective text vector of the preprocessed defect description information; determine the target code blocks from the top N code blocks in the sorting result based on the relevance between the defective text vector and the code block text vector, where the text vector of the code block is used to express the context information of the code snippet in the code block; The first determination module is specifically configured to determine the function of each code file based on external dependency description information, determine the semantic relevance between the defect keywords and the paths of each code file in the directory tree, and determine the code files corresponding to the file paths with a semantic relevance greater than or equal to a preset relevance threshold as the first associated files.
7. An electronic device, characterized in that, Comprising: At least one processor; And A memory communicatively connected to the at least one processor; Wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to cause the electronic device to execute the method according to any one of claims 1-5.
8. A computer-readable storage medium, characterized in that, Computer-executable instructions are stored in the computer-readable storage medium, and when the processor executes the computer-executable instructions, the method according to any one of claims 1-5 is implemented.
9. A computer program product comprising a computer program, characterized in that, When the computer program is executed by the processor, the method according to any one of claims 1-5 is implemented.
Citation Information
Patent Citations
Software defect code file locating method based on reverse index technology
CN105095091A
Defect positioning technology based on positioning keyword extraction
CN118114098A