Defect positioning method and system based on code semantic search guidance

CN122884809APending Publication Date: 2026-10-09NAT UNIV OF DEFENSE TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511932752.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-19
Publication Date
2026-10-09

AI Technical Summary

Technical Problem

现有缺陷定位技术仅能通过分析单独的程序元素来判断缺陷位置,无法获取软件的运行时行为和全局功能模块信息,这限制了其定位的准确性

Benefits of technology

本发明通过获取存在缺陷软件的运行时行为,并利用大语言模型对缺陷软件的运行时行为进行多粒度项目特定知识提取,构建知识库,从而实现更准确的缺陷定位。在多粒度查询生成与动态信息补充方面,根据失败测试用例和软件报错信息,生成涵盖模块、方法、代码块三个粒度的多粒度查询,在信息不足时通过语义搜索补充模块摘要,形成新提示信息以生成更准确的查询。在多粒度检索与语义相似度计算过程中,将多粒度查询输入检索模块,通过词嵌入模型向量化查询和代码知识,在向量空间中计算语义相似度,检索出可能存在缺陷的代码方法列表,全面覆盖不同粒度代码元素,提高准确性,避免关键词不匹配导致的漏检和误检。通过聚合多个查询和粒度的检索结果,计算方法可疑分数并排序,提供全面评估,帮助开发人员快速定位故障,提升效率。此外,本发明提供易于理解的自然语言解释,减少推理步骤,降低机器资源消耗,提高定位效率,增强实用性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122884809A_ABST
    Figure CN122884809A_ABST
Patent Text Reader

Abstract

The application provides a defect positioning method and system based on code semantic search guidance, which comprises the following steps: firstly, a large language model is used to extract multi-granularity project-specific knowledge of a defective software to construct a software knowledge base; secondly, a multi-granularity query describing suspicious functions is generated according to a failed test case and error information; then, the query is input into a multi-granularity retrieval module, a word embedding model is used to vectorize each granularity query and the code knowledge of the corresponding granularity in the knowledge base, and the semantic similarity is calculated; finally, a suspicious score is calculated for the code method based on the semantic similarity, and a sorted defect code method list is generated according to the score. Through multi-granularity semantic search and sorting, the accuracy and efficiency of software defect positioning are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of automatic software program defect repair technology, and in particular to a defect localization method and system based on code semantic search guidance. Background Technology

[0002] Software development is an iterative cycle of continuous feature implementation and debugging. During the implementation phase, programmers write code to meet specific requirements, but may introduce errors into the software system. During the debugging phase, developers use error signals to find the root cause and fix any unexpected program behaviors. This iterative process is crucial for continuously improving software quality and flexibly adapting to changing needs.

[0003] Code search and defect localization are two crucial tasks in software development. Code search aims to help developers reuse specific code snippets from open-source repositories during the implementation phase, thus avoiding "reinventing the wheel." Defect localization, on the other hand, identifies erroneous program elements throughout the software system during the debugging phase to expedite the error resolution process.

[0004] For years, code search and defect localization technologies have tended to leverage advanced deep learning techniques. In code search, thanks to advances in representation learning, the technology has evolved from traditional keyword matching to learning the semantic association between search queries and code snippets. This allows for a more accurate understanding of the semantics of natural language queries and code snippets, leading to more precise retrieval of code that matches the search requirements. In defect localization, the technology has also progressed from coverage-based program spectrum analysis to representation learning, and recently, it has been supported by large-scale language models, which analyze the structural and semantic information of the code to locate potential defects.

[0005] However, existing technologies still have some shortcomings. In the field of code search, although state-of-the-art technologies can accurately identify the required code snippets for about 80% of needs, they may still fail to fully meet developers' expectations when dealing with complex requirements or domain-specific code. For defect localization technologies, their effectiveness is significantly lower than that of code search technologies. In benchmark tests in specific domains, the latest defect localization technologies can only pinpoint the erroneous program entity in about 30% of cases. Existing defect localization technologies can only determine the defect location by analyzing individual program elements, and cannot obtain information about the software's runtime behavior and global functional modules, which limits their accuracy. Furthermore, while some defect localization technologies based on large language models have improved the localization effect to some extent, they suffer from problems such as numerous reasoning steps, high machine resource consumption, room for improvement in repair accuracy, and a lack of easily understandable natural language interpretation, affecting their efficiency and usability in practical applications. Summary of the Invention

[0006] The technical problem to be solved by this invention is: in view of the technical problems existing in the prior art, this invention provides a method and system for accurate defect localization based on code semantic search, which aims to effectively improve the accuracy and efficiency of software defect localization through multi-granular semantic search and sorting.

[0007] To solve the above-mentioned technical problems, the technical solution proposed by this invention is as follows: A defect localization method guided by code semantic search includes: The runtime behavior of defective software is obtained, and a software knowledge base is obtained by extracting project-specific knowledge of the runtime behavior of the defective software at different granularities using a large language model. Based on failed test cases and software error messages, prompt messages are set. The large language model generates multi-granularity queries based on the prompt messages to describe suspicious functions in the software at different granularities. The prompt messages are used to guide the large language model to identify potential erroneous functions in the test software, or to request additional information when the existing information is insufficient to determine that the test software has potential erroneous functions. The additional information is a module summary that best matches the function description and error message of the current failed test case found from the module index using semantic search, and the module summary is added to the prompt messages to form new prompt messages. The generated multi-granularity query is input into the multi-granularity retrieval module; each granularity query in the multi-granularity query and the corresponding granularity code knowledge in the software knowledge base are vectorized using a word embedding model; the semantic similarity between each granularity query vector and the corresponding code knowledge vector is calculated in the vector space; a suspicious score is calculated for the code method based on the semantic similarity, and a sorted list of potentially defective code methods is generated based on the suspicious score.

[0008] As a further improvement to the method of the present invention: the project-specific knowledge of different granularities includes module-level knowledge, and the extraction of the module-level knowledge includes: For each failed test case, program instrumentation is used to trace method call information and generate a method call graph. The method call graph records the method ID, call edge, and frequency of each call, and excludes methods not covered at runtime and utility methods in external libraries. The call graphs of each test case are merged into a global call graph as additional internal software information input for the large language model. The global call graph is decomposed into functional modules using the Leiden algorithm; The size of the functional modules is optimized by using heuristic strategies, making the module division more reasonable.

[0009] As a further improvement to the method of the present invention: the step of decomposing the global call graph into functional modules using the Leiden algorithm further includes optimizing the partitioning of the functional modules by maximizing the quality function value, wherein the function expression of the quality function is:

[0010] in, Let represent the mass function, and the range of values ​​for the mass function is []. 1,1], This indicates the edges in the global call graph ( , The weight of the method Calling methods , called Second-rate, This represents the sum of the weights of all edges in the global call graph. Represents vertices The weighting degree, Represents vertices The weighting degree, It is the vertex The module to which it belongs. Represents vertices The module to which it belongs. Indicates an indicator function, if Its value is 1 if it is true, otherwise it is 0.

[0011] As a further improvement to the method of the present invention: the project-specific knowledge of different granularities also includes method-level knowledge and code block-level knowledge. The extraction of method-level knowledge and code block-level knowledge is to summarize the detailed workflow of the code methods in the functional modules using a large language model, so as to decompose the overall function of the code methods into more fine-grained block-level units.

[0012] As a further improvement to the method of the present invention: the method for generating multi-granularity queries includes: Based on the source code and software error information of failed test cases, extract the functional description and error information of the test cases, construct prompt information containing functional description and error information, and integrate the prompt information with test case code, exception stack trace, test output, module details and response process to form a complete prompt information input large language model; Based on the prompts, the large language model generates multi-granularity queries that include module granularity, method granularity, and code block granularity, which are used to describe suspicious functions in the software. When the large language model determines that the existing prompt information is insufficient to accurately identify potential erroneous functions, it retrieves the module summary that best matches the function description and error information of the current failed test case from the module index through semantic search, adds the retrieved module summary to the original prompt information to form a new prompt information, and inputs it back into the large language model to help generate more accurate multi-granular queries.

[0013] As a further improvement to the method of the present invention: the generated multi-granularity query is input into the multi-granularity retrieval module, and a multi-granularity index is constructed by multi-granular vectorizing the multi-granularity query and the code knowledge in the software knowledge base through a word embedding model, including the following steps: Constructing a module-level retrieval system: Vectorize the module-level queries and the module-level code knowledge in the software knowledge base, and calculate the semantic similarity between the module-level query vector and the code knowledge vector in the vector space to obtain a list of suspicious modules; Constructing a method-level retrieval system: Vectorize method-level queries and method-level code knowledge in the software knowledge base, and calculate the semantic similarity between the method-level query vector and the code knowledge vector in the vector space to obtain a list of suspicious methods; Constructing code block-level retrieval: Vectorize code block-level queries and code block-level code knowledge in the software knowledge base, and calculate the semantic similarity between code block-level query vectors and code knowledge vectors in the vector space to obtain a list of suspicious code blocks; Aggregation results: The search results at the module level, method level, and code block level are aggregated to obtain the final list of suspicious code methods.

[0014] As a further improvement to the method of the present invention: the module-level retrieval Method-level retrieval and code block-level search The expression is as follows:

[0015] in, This represents the set of all retrieved program elements. The function represents a retrieval function based on a similarity metric. This represents the set of all retrieved methods. This represents the set of modules retrieved. This represents the set of retrieved code blocks. ( ) indicates that the method Mapped to its respective module, ( ) indicates that the method Mapped to the collection of code blocks inside it , For the generated multi-granularity query, where, , , Queries are provided at the module, method, and code block levels, respectively. It is the number of failed test cases. Indicates hyperparameters, Representation method index, Indicates the module index. Indicates the code block index.

[0016] As a further improvement to the method of the present invention: the method for calculating the suspicious score is as follows:

[0017] in, This represents the semantic similarity between a program element and its corresponding query during the retrieval process, with values ​​ranging from [-1, 1]. Representation method suspicious values, Representation method With modules Match scores between them Representation method With Method Match scores between them Representation method With code blocks Match scores between them Indicates the first The first test case Each module Indicates the first The first test case One method, Indicates the first The first test case A code block, Indicates the number of failed test cases. Indicates the sequence number of the failed test case. Indicates the index of an element in the set.

[0018] The present invention also provides a defect localization system guided by code semantic search, comprising a microprocessor and a memory interconnected thereto, wherein the microprocessor is programmed or configured to execute the defect localization method guided by code semantic search.

[0019] The present invention also provides a computer-readable storage medium storing a computer program / instructions programmed or configured to perform a defect localization method guided by code semantic search as described in the present invention.

[0020] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention achieves more accurate defect localization by acquiring the runtime behavior of defective software and using a large language model to extract multi-granularity project-specific knowledge from this behavior, constructing a knowledge base. Regarding multi-granularity query generation and dynamic information supplementation, multi-granularity queries covering modules, methods, and code blocks are generated based on failed test cases and software error messages. When information is insufficient, semantic search is used to supplement module summaries, forming new prompts to generate more accurate queries. In the multi-granularity retrieval and semantic similarity calculation process, the multi-granularity queries are input into the retrieval module. The query and code knowledge are vectorized using a word embedding model, and semantic similarity is calculated in the vector space to retrieve a list of potentially defective code methods. This comprehensively covers code elements at different granularities, improving accuracy and avoiding missed or false detections due to keyword mismatches. By aggregating the retrieval results of multiple queries and granularities, a method suspicion score is calculated and ranked, providing a comprehensive assessment to help developers quickly locate faults and improve efficiency. Furthermore, this invention provides easily understandable natural language explanations, reducing reasoning steps, lowering machine resource consumption, improving localization efficiency, and enhancing practicality. Attached Figure Description

[0021] Figure 1 This is a flowchart illustrating the defect localization method based on code semantic search in an embodiment of the present invention.

[0022] Figure 2 This is a schematic diagram of the source code and error message of a failed test case for defective software in an embodiment of the present invention.

[0023] Figure 3 This is a schematic diagram illustrating the extraction of project-specific knowledge at different granularities using a large language model in an embodiment of the present invention.

[0024] Figure 4 This is a schematic diagram illustrating three different granularity queries in an embodiment of the present invention.

[0025] Figure 5 This is a schematic diagram of a list of potentially defective methods in an embodiment of the present invention. Detailed Implementation

[0026] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0027] like Figure 1As shown, this embodiment of the defect localization method based on code semantic search includes: Step 1: Obtain the runtime behavior of the defective software, and use a large language model to extract project-specific knowledge at different granularities from the runtime behavior of the defective software to obtain a software knowledge base.

[0028] In this embodiment, project-specific knowledge at different granularities includes module-level knowledge. Module-level knowledge is extracted by monitoring methods called during program runtime through failed test cases, and organizing them into several functional modules based on the calling relationships between these methods. Detailed methods include: Step 101: For each failed test case, use program instrumentation to trace method call information and generate a method call graph. The method call graph records the method ID, call edge, and frequency of each call, and excludes methods not covered at runtime and utility methods in external libraries. The call graphs of each test case are merged into a global call graph as additional internal software information input for the large language model. Step 102: Use the Leiden algorithm to decompose the global call graph into functional modules; Step 103: Optimize the size of the functional modules using heuristic strategies to make the module division more reasonable.

[0029] In this embodiment, using the Leiden algorithm to decompose the global call graph into functional modules further includes optimizing the functional module partitioning by maximizing the quality function value. The expression of the quality function is: (1) in, This represents the mass function, and its range of values ​​is [ 1,1], This indicates a global call to the edges in the graph. , The weight of the method Calling methods , called Second-rate, This represents the sum of the weights of all edges in the global call graph. Represents vertices The weighting degree, Represents vertices The weighting degree, It is the vertex The module to which it belongs. Represents vertices The module to which it belongs. Indicates an indicator function, if Its value is 1 if it is 1, otherwise it is 0.

[0030] Specifically, a call graph is a graphical structure used to represent the call relationships between methods in a software system. Let's assume a call graph... =( , In the diagram, each node V represents a method, and edges E represent the calling relationships between methods. , Weights Representation method Calling methods To evaluate the quality of module partitioning in the call graph, a quality function Q is defined as shown in Formula 1. This function measures the quality of partitioning by comparing the tightness of connections within modules with the looseness of connections between modules. The value of Q ranges from [ [1,1], a higher Q value indicates a better partitioning, meaning that the detected functional modules have tighter internal connections and looser connections between modules. Based on By definition, the Leiden algorithm continuously searches for the optimal partition of a graph. This invention sets the initial partition of the algorithm to treat each node as a separate community. The termination condition is defined as a quality metric. The algorithm stops improving or reaches the maximum number of iterations (set to 5 in this embodiment), and treats the call subgraph of each cluster as a candidate functional module. In this way, the Leiden algorithm can effectively identify the community structure, i.e., functional modules, in the call graph, thus providing a basis for subsequent defect localization. This method not only improves the accuracy of module partitioning but also ensures the efficiency of the partitioning process.

[0031] In this embodiment, project-specific knowledge at different granularities also includes method-level knowledge and code block-level knowledge. The extraction of method-level and code block-level knowledge involves summarizing the detailed workflow of the code methods within the functional modules using a large language model, thus decomposing the overall functionality of the code methods into finer-grained block-level units. Summarizing refers to generating a natural language summary describing the method's input, output, and main logic using the large language model, serving as method-level knowledge. Decomposition refers to automatically dividing the method code into code segments with independent sub-functions based on the code's control flow structure (such as loops, conditional statements, and function call sequences), and generating natural language labels describing the function of each code segment, serving as code block-level knowledge.

[0032] Step 2: Set prompt messages based on failed test cases and software error messages. The large language model generates multi-granularity queries based on the prompt messages to describe suspicious functions in the software at different granularities. The prompt messages are used to guide the large language model to identify potential erroneous functions in the test software, or to request additional information when the existing information is insufficient to determine that the test software has potential erroneous functions. The additional information is to use semantic search to find the module summary that best matches the function description and error message of the current failed test case from the module index, and add the module summary to the prompt messages to form new prompt messages.

[0033] In this embodiment, the method for generating multi-granularity queries includes: Step 201: Based on the source code and software error information of the failed test cases, extract the functional description and error information of the test cases, construct a prompt message containing the functional description and error information, and integrate the prompt message with the test case code, exception stack trace, test output, module details and response process to form a complete prompt message input large language model; Step 202: Based on the prompt information, the large language model generates multi-granularity queries with three levels of granularity: module level, method level, and code block level, which are used to describe suspicious functions in the software; Step 203: When the large language model determines that the existing prompt information is insufficient to accurately identify potential erroneous functions, it retrieves the module summary that best matches the function description and error information of the current failed test case from the module index through semantic search, adds the retrieved module summary to the original prompt information to form a new prompt information, and inputs it back into the large language model to help generate more accurate multi-granularity queries.

[0034] Step 3: Input the generated multi-granularity query into the multi-granularity retrieval module; vectorize each granularity query in the multi-granularity query and the corresponding granularity code knowledge in the software knowledge base using a word embedding model; calculate the semantic similarity between each granularity query vector and the corresponding code knowledge vector in the vector space; since the retrieval will return suspicious elements at multiple granularities, in order to sort the methods involved in the final defect, it is necessary to calculate a comprehensive suspicious score for each method and generate a sorted list of potentially defective code methods based on the suspicious scores.

[0035] In this embodiment, a text embedding model is used to vectorize program elements in the constructed software knowledge base at multiple granularities, creating indexes at three granularities. A set of program elements is defined. ,in Represents a collection of functional modules. The set of representation methods Represents the set of all code blocks. Representation method A collection of code blocks in the code.

[0036] Specifically, the generated multi-granularity query is input into the multi-granularity retrieval module. A multi-granularity index is constructed by vectorizing the multi-granularity query and code knowledge in the software knowledge base using a word embedding model. This includes the following steps: Step 301: Construct module-level retrieval: Vectorize the module-level query and the module-level code knowledge in the software knowledge base, and calculate the semantic similarity between the module-level query vector and the code knowledge vector in the vector space to obtain a list of suspicious modules; Step 302: Construct method-level retrieval: Vectorize the method-level query and the method-level code knowledge in the software knowledge base, and calculate the semantic similarity between the method-level query vector and the code knowledge vector in the vector space to obtain a list of suspicious methods; Step 303: Construct code block-level retrieval: Vectorize the code block-level query and the code block-level code knowledge in the software knowledge base, and calculate the semantic similarity between the code block-level query vector and the code knowledge vector in the vector space to obtain a list of suspicious code blocks; Step 304: Aggregate Results: Aggregate the search results at the module, method, and code block levels to obtain the final list of suspicious code methods.

[0037] In a specific application embodiment, after extracting knowledge at the module, method, and code block levels, a corresponding vectorized index is constructed to support efficient semantic retrieval. The specific process is as follows: using a pre-trained word embedding model, each type of knowledge summary is converted into a high-dimensional vector. The set of vectorized knowledge summaries from all functional modules is then used to construct the module index. The set of vectorized knowledge summaries of all methods is used to construct a method index. The set of vectorized knowledge tags for all code blocks is used to construct a code block index. Module index, method index and code block index Together, they form the core retrieval foundation of the software knowledge base.

[0038] In this embodiment, the expressions for module-level retrieval, method-level retrieval, and code block-level retrieval are as follows: (2) (3) (4) in, (·) represents a function that maps each program element to a corresponding knowledge summary. This represents a text embedding model. In this embodiment, the module level Method level and code block level The aggregation of search results also includes calculating a suspicion score for each method and sorting the methods from highest to lowest suspicion level. The formula for calculating the suspicion score for each method is as follows: (5) (6) (7) (8) The Retrieve function measures the semantic similarity between the query and the program element summary using cosine similarity, and returns suspicious program elements in descending order of similarity. This represents the set of all retrieved methods. Representing a single method , This represents knowledge at three different granularities stored in the knowledge base. ( ) indicates that the method Mapped to its respective module, ( ) indicates that the method Mapped to the collection of code blocks inside it , For the generated multi-granularity query, where, , , Queries are provided at the module, method, and code block levels, respectively. It is the number of failed test cases. This represents a hyperparameter used to control the number of program elements retrieved. Representation method index, Indicates the module index. Indicates the code block index.

[0039] Specifically, for each failed test case, Formula (5) is used to retrieve suspicious methods, Formula (6) and Formula (7) are used to retrieve suspicious modules and code blocks respectively, and Formula (8) represents the retrieval results of all failed test cases.

[0040] In this embodiment, after obtaining all retrieved program elements, the search results from multiple queries and granularities are aggregated to obtain a list of methods sorted in descending order of suspicion. If a method (including its related modules and blocks) is semantically relevant to more queries, then this method is more likely to be the actual location of the fault. The collection of all retrieved methods... Calculate a single method using the following formula. Suspicious scores: (9) (10) (11) (12) in, This represents the semantic similarity between a program element and its corresponding query during the retrieval process, with values ​​ranging from [-1, 1]. Representation method suspicious values, Representation method With modules Match scores between them Representation method With Method Match scores between them Representation method With code blocks Match scores between them Indicates the first The first test case Each module Indicates the first The first test case One method, Indicates the first The first test case A code block, Indicates the number of failed test cases. Indicates the sequence number of the failed test case. Indicates the index of an element in the set.

[0041] The present invention will be further illustrated below by taking the above method for defect localization of defective software based on code semantic search as an example in a specific application embodiment.

[0042] like Figure 2 As shown, the specific application steps of defect localization technology guided by code semantic search are as follows: (The text then lists the source code of the failed test cases and the software error messages for the defective software.) Step 1): Software Knowledge Base Construction: For defective software, first collect the software's runtime behavior and use a large language model to extract project-specific knowledge at different granularities (including module level, method level, and code block level), such as... Figure 3As shown, when running failed test cases, method call information collected through program instrumentation is used to generate a method call graph. Nodes in the graph represent methods, edges represent the call relationships between methods, and the weight of the edges represents the call frequency. The Leiden algorithm is used to partition the call graph into modules. The Leiden algorithm is a community detection algorithm that optimizes a quality function... Q The process begins by identifying the community structure (i.e., functional modules) within the diagram. Each functional module contains a set of methods with related functionalities. A large language model is then used to summarize the detailed workflow of each method. This model understands the logical structure and functionality of the code, generating natural language descriptions to help developers quickly understand the method's purpose and implementation details. This detailed workflow summary for each method forms method-level knowledge. The overall functionality of each method is then broken down into finer-grained code blocks. Here, a "block" refers to a relatively independent code segment within a method, such as loops, conditional statements, or function calls. By breaking down methods into these fine-grained block-level units, the functional structure of the code can be described and understood more precisely. The code block-level knowledge for each method forms a code block-level knowledge base. Through this process, a software knowledge base containing module-level, method-level, and code block-level knowledge is constructed, providing a foundation for subsequent defect localization and remediation.

[0043] Step 2): Multi-granularity query generation: To more comprehensively describe the characteristics and context of suspicious software functions, the source code of failed test cases and software error messages are understood through a large language model to generate queries at three different granularities (i.e., module, method, and code block), such as... Figure 4 As shown.

[0044] Step 3): Defect Retrieval: Using the multi-granularity query generated in Step 2) as input, a word embedding model is used to search for a list of potentially defective methods in the vector space. Successfully identified defective methods are ranked first. Figure 5 As shown.

[0045] The present invention also provides a defect localization system guided by code semantic search, including a microprocessor and a memory interconnected, wherein the microprocessor is programmed or configured to guide the defect localization method based on code semantic search.

[0046] This embodiment also provides a computer-readable storage medium storing a computer program / instruction that is programmed or configured to guide a defect localization method based on processor-based code semantic search.

[0047] Those skilled in the art will understand that the above embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-readable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create a machine for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The functions specified in one or more boxes. These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable apparatus for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the functions specified in one or more boxes. The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Therefore, any simple modifications, equivalent changes, and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention should fall within the protection scope of the present invention.

[0048] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the invention. Therefore, any simple modifications, equivalent changes, and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention should fall within the protection scope of the present invention.

Claims

1. A defect localization method based on code semantic search guidance, characterized in that, include: The runtime behavior of defective software is obtained, and a software knowledge base is obtained by extracting project-specific knowledge of the runtime behavior of the defective software at different granularities using a large language model. Based on failed test cases and software error messages, prompt messages are set. The large language model generates multi-granularity queries based on the prompt messages to describe suspicious functions in the software at different granularities. The prompt messages are used to guide the large language model to identify potential erroneous functions in the test software, or to request additional information when the existing information is insufficient to determine that the test software has potential erroneous functions. The additional information is a module summary that best matches the function description and error message of the current failed test case found from the module index using semantic search, and the module summary is added to the prompt messages to form new prompt messages. The generated multi-granularity query is input into the multi-granularity retrieval module; The word embedding model is used to vectorize each granularity query in the multi-granularity query and the corresponding granularity code knowledge in the software knowledge base; the semantic similarity between each granularity query vector and the corresponding code knowledge vector is calculated in the vector space; a suspicious score is calculated for the code method based on the semantic similarity, and a sorted list of potentially defective code methods is generated based on the suspicious score.

2. The defect localization method based on code semantic search as described in claim 1, characterized in that, The project-specific knowledge at different granularities includes module-level knowledge, and the extraction of module-level knowledge includes: For each failed test case, program instrumentation is used to trace method call information and generate a method call graph. The method call graph records the method ID, call edge, and frequency of each call, and excludes methods not covered at runtime and utility methods in external libraries. The call graphs of each test case are merged into a global call graph as additional internal software information input for the large language model. The global call graph is decomposed into functional modules using the Leiden algorithm; The size of the functional modules is optimized by using heuristic strategies, making the module division more reasonable.

3. The defect localization method based on code semantic search guidance according to claim 2, characterized in that, Using the Leiden algorithm to decompose the global call graph into functional modules also includes optimizing the partitioning of the functional modules by maximizing the value of the quality function, wherein the expression of the quality function is: in, This represents the mass function, and the range of values ​​for the mass function is []. 1,1], This indicates the edges in the global call graph ( , The weight of the method Calling methods , called Second-rate, This represents the sum of the weights of all edges in the global call graph. Represents vertices The weighting degree, Represents vertices The weighting degree, It is the vertex The module to which it belongs. Represents vertices The module to which it belongs. Indicates an indicator function, if Its value is 1 if it is true, otherwise it is 0.

4. The defect localization method based on code semantic search guidance according to claim 2, characterized in that, The project-specific knowledge at different granularities also includes method-level knowledge and code block-level knowledge. The extraction of method-level knowledge and code block-level knowledge involves summarizing the detailed workflow of the code methods in the functional modules using a large language model, so as to decompose the overall function of the code methods into more fine-grained block-level units.

5. The defect localization method based on code semantic search guidance according to claim 1, characterized in that, The method for generating multi-granularity queries includes: Based on the source code and software error information of failed test cases, extract the functional description and error information of the test cases, construct prompt information containing functional description and error information, and integrate the prompt information with test case code, exception stack trace, test output, module details and response process to form a complete prompt information input large language model; Based on the prompts, the large language model generates multi-granularity queries that include module granularity, method granularity, and code block granularity, which are used to describe suspicious functions in the software. When the large language model determines that the existing prompt information is insufficient to accurately identify potential erroneous functions, it retrieves the module summary that best matches the function description and error information of the current failed test case from the module index through semantic search, adds the retrieved module summary to the original prompt information to form a new prompt information, and inputs it back into the large language model to help generate more accurate multi-granular queries.

6. The defect localization method based on code semantic search as described in claim 1, characterized in that, The generated multi-granularity query is input into the multi-granularity retrieval module. The word embedding model is used to vectorize each granularity query in the multi-granularity query and the corresponding granularity code knowledge in the software knowledge base, including the following steps: Constructing a module-level retrieval system: Vectorize the module-level queries and the module-level code knowledge in the software knowledge base, and calculate the semantic similarity between the module-level query vector and the code knowledge vector in the vector space to obtain a list of suspicious modules; Constructing a method-level retrieval system: Vectorize method-level queries and method-level code knowledge in the software knowledge base, and calculate the semantic similarity between the method-level query vector and the code knowledge vector in the vector space to obtain a list of suspicious methods; Constructing code block-level retrieval: Vectorize code block-level queries and code block-level code knowledge in the software knowledge base, and calculate the semantic similarity between code block-level query vectors and code knowledge vectors in the vector space to obtain a list of suspicious code blocks; Aggregation results: The search results at the module level, method level, and code block level are aggregated to obtain the final list of suspicious code methods.

7. The defect localization method based on code semantic search as described in claim 6, characterized in that, The module-level search Method-level retrieval and code block-level search The expression is as follows: in, This represents the set of all retrieved program elements. The function represents a retrieval function based on a similarity metric. This represents the set of all retrieved methods. This represents the set of modules retrieved. This represents the set of retrieved code blocks. ( ) indicates that the method Mapped to its respective module, ( ) indicates that the method Mapped to the collection of code blocks inside it , For the generated multi-granularity query, where, , , Queries are provided at the module, method, and code block levels, respectively. It is the number of failed test cases. Indicates hyperparameters, Representation method index, Indicates the module index. Indicates the code block index.

8. The defect localization method based on code semantic search guidance according to claim 1, characterized in that, The method for calculating the suspicious score is as follows: in, This represents the semantic similarity between a program element and its corresponding query during the retrieval process, with values ​​ranging from [-1, 1]. Representation method suspicious values, Representation method With modules Match scores between them Representation method With Method Match scores between them Representation method With code blocks Match scores between them Indicates the first The first test case Each module Indicates the first The first test case One method, Indicates the first The first test case A code block, Indicates the number of failed test cases. Indicates the sequence number of the failed test case. Indicates the index of an element in the set.

9. A defect localization system guided by code semantic search, comprising a microprocessor and a memory interconnected, characterized in that, The microprocessor is programmed or configured to execute the defect localization method based on code semantic search as described in any one of claims 1 to 8.

10. A computer-readable storage medium storing a computer program / instructions, characterized in that, The computer program / instructions are programmed or configured to execute the defect localization method based on code semantic search as described in any one of claims 1 to 8 via a processor.