Code defect detection method and device, storage medium and electronic device

By building the target knowledge graph in the code warehouse and determining the code defect detection context, the problem of inaccurate code defect detection in the existing technology is solved, and higher detection accuracy and flexibility are achieved.

CN120196537APending Publication Date: 2025-06-24ZHEJIANG DAHUA TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510293116.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

In the existing technology, code defect detection is inaccurate, static code analysis tools are relatively regular in scanning, lack of flexibility, and poor performance when dealing with complex logic and algorithms. Dynamic code analysis relies on test cases and running environment and cannot detect non-runtime problems.

Method used

By obtaining all code source files of the target code from the code repository when receiving the code detection instructions, determine the difference between each code source file and the target code, and build a target knowledge graph based on these code source files. This knowledge graph indicates the relationship between code block instances and file-level community instances to determine the code defect detection context and ultimately detect the object code to determine the defect.

Benefits of technology

Improves the accuracy of code defect detection, can handle complex logic and algorithms more efficiently, and detects non-runtime problems that cannot be captured by dynamic code analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120196537A_ABST
    Figure CN120196537A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a code defect detection method and device, a storage medium and an electronic device.The method comprises the steps that under the condition that a code detection instruction is received, all code source files of a target code indicated by the code detection instruction are obtained from a code warehouse, determining difference content between each code source file and the target code; constructing a target knowledge graph based on all the code source files; determining a code defect detection context in the code source file based on the difference content and the target knowledge graph; and detecting the target code based on the code defect detection context and the difference content to determine the defect of the target code. By means of the code defect detection method and device, the problem that code defect detection is inaccurate in the prior art is solved, and then the effect of improving the code defect detection accuracy is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to the field of communications, and more particularly, to a method, apparatus, storage medium, and electronic device for detecting code defects. Background Art

[0002] In related technologies, code defect detection generally includes static code analysis and dynamic code analysis. However, the scanning of static code analysis tools is relatively regular, lacking flexibility and performing poorly when dealing with complex logic and algorithms. Dynamic code analysis depends on test cases, runtime environments, etc., and cannot detect some non-runtime problems.

[0003] It can be seen that there is a technical problem of inaccurate code defect detection in related technologies.

[0004] Currently, no effective solution has been proposed for the above problems existing in related technologies. Summary of the Invention

[0005] Embodiments of the present invention provide a method, apparatus, storage medium, and electronic device for detecting code defects, so as to at least solve the problem of inaccurate code defect detection in related technologies.

[0006] According to an embodiment of the present invention, a method for detecting code defects is provided, including: when receiving a code detection instruction, obtaining all code source files of the target code indicated by the code detection instruction from a code repository, and determining the difference content between each code source file and the target code; constructing a target knowledge graph based on all the code source files, where the target knowledge graph is used to indicate the relationship between code block instances included in the same code source file, and to indicate the relationship between file-level community instances in different code source files, and the file-level community instance in the code source file represents the role of each instance in the instance list included in the code source file; determining a code defect detection context in the code source file based on the difference content and the target knowledge graph; and detecting the target code based on the code defect detection context and the difference content to determine the defects of the target code.

[0007] In an exemplary embodiment, constructing a target knowledge graph based on all the code source files includes: splitting each of the code source files to obtain a plurality of code blocks; determining instances of each of the code blocks to obtain a plurality of the code block instances; constructing a low-level knowledge graph based on the plurality of the code block instances; determining an instance list of each of the code source files to obtain a plurality of instance lists; determining the instance role of each instance included in the plurality of the instance lists; determining the instance role as a file-level community instance; constructing a middle-level knowledge graph based on the file-level community instance; and determining the low-level knowledge graph and the middle-level knowledge graph as the target knowledge graph.

[0008] In an exemplary embodiment, constructing a low-level knowledge graph based on the plurality of the code block instances includes: determining first code block instances that belong to the same code source file among the plurality of the code block instances; determining the instance relationships between the first code block instances; determining the relationships between second code block instances that belong to different code source files among the plurality of the code block instances to obtain the relationships between code block instances; and determining the knowledge graph including the instance relationships and the relationships between code block instances as the low-level knowledge graph.

[0009] In an exemplary embodiment, constructing a middle-level knowledge graph based on the file-level community instance includes: determining the relationships between second code block instances that belong to different code source files among the plurality of the code block instances to obtain the relationships between code block instances; determining the mapping relationship between the file path of the code source file and the file-level community instance; constructing a prompt word based on the relationships between the code block instances; analyzing the mapping relationship based on the prompt word to determine the file-level community instance relationships; and determining the knowledge graph including the file-level community instance relationships and the file-level community instance as the middle-level knowledge graph.

[0010] In an exemplary embodiment, splitting each of the code source files to obtain a plurality of code blocks includes: determining the start position and the end position corresponding to each target structure included in the code source file; and splitting the code source file according to the start position and the end position to obtain a plurality of the code blocks.

[0011] In an exemplary embodiment, determining the code defect detection context based on the difference content and the target knowledge graph in the code source file includes: determining an initial region corresponding to the difference content in the code source file; expanding the initial region to obtain a target region, where the range of the target region is larger than that of the initial region, and the target region includes the initial region; determining target instances of target content included in the target region; determining first instances associated with the target instances in the underlying knowledge graph included in the target knowledge graph, where the underlying knowledge graph is used to indicate the relationships between code block instances included in the same code source file; determining second instances associated with the first instances in the middle-level knowledge graph included in the target knowledge graph, where the middle-level knowledge graph is used to indicate the relationships between file-level community instances in different code source files; obtaining the target file-level community instance relationships of the second instances and the instance descriptions of the second instances; and determining the target file-level community instance relationships and the instance descriptions as the code defect detection context.

[0012] In an exemplary embodiment, determining first instances associated with the target instances in the underlying knowledge graph included in the target knowledge graph includes: determining the most instances connected to the target instances included in the underlying knowledge graph; and determining the most instances as the first instances.

[0013] In an exemplary embodiment, detecting the target code based on the code defect detection context and the difference content to determine the defects of the target code includes: constructing a difference detection prompt word using the code defect detection context and the difference content; and detecting the target code based on the prompt word to determine the defects of the target code.

[0014] According to another embodiment of the present invention, a code defect detection device is provided, including: an acquisition module, configured to, when receiving a code detection instruction, acquire all code source files of the target code indicated by the code detection instruction from a code repository, and determine the difference content between each of the code source files and the target code; a construction module, configured to construct a target knowledge graph based on all the code source files, where the target knowledge graph is used to indicate the inter-code-block instance relationships between code block instances included in the same code source file, and to indicate the file-level community instance relationships between file-level community instances in different code source files, and the file-level community instance in the code source file represents the function of each instance in the instance list included in the code source file; a determination module, configured to determine a code defect detection context in the code source file based on the difference content and the target knowledge graph; and a detection module, configured to detect the target code based on the code defect detection context and the difference content to determine the defects of the target code.

[0015] According to still another embodiment of the present invention, a computer-readable storage medium is further provided. A computer program is stored in the computer-readable storage medium, where the computer program is configured to execute the steps in any one of the above method embodiments when running.

[0016] According to still another embodiment of the present invention, an electronic device is further provided, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.

[0017] According to still another embodiment of the present invention, a computer program product is further provided, including a computer program, where the steps of the methods described in various embodiments of the present application are implemented when the computer program is executed by a processor.

[0018] Through the present invention, in the case of receiving a code detection instruction, all code source files of the target code to be detected indicated by the code instruction can be obtained from the code repository, and the difference content between each code source file and the target code can be determined. All the code source files can be used to construct a target knowledge graph, which can indicate the relationships between code block instances included in the same code source file, and the relationships between file-level community instances in different code source files. Among them, the file-level community instance in the code source file can represent the function of each instance in the instance list included in the code source file. Through the above-mentioned difference content and the constructed target knowledge graph, the code defect detection context in the code source file can be determined, and then the target code can be detected through the code defect detection context and the difference content, so as to determine the defects of the target code. Therefore, the problem of inaccurate code defect detection in the related art can be solved, and the effect of improving the accuracy of code defect detection can be achieved. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 is a hardware structure block diagram of a mobile terminal for a method of detecting code defects according to an embodiment of the present invention;

[0020] Figure 2 is a flowchart of a method of detecting code defects according to an embodiment of the present invention;

[0021] Figure 3 is a flowchart of a method of detecting code defects according to a specific embodiment of the present invention;

[0022] Figure 4 is a structure block diagram of a device for detecting code defects according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0023] The embodiments of the present invention will be described in detail below with reference to the drawings and in conjunction with the embodiments.

[0024] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence.

[0025] The method embodiments provided in the embodiments of the present application can be executed on a mobile terminal, a computer terminal or a similar computing device. Taking running on a mobile terminal as an example, Figure 1 is a hardware structure block diagram of a mobile terminal for a method of detecting code defects according to an embodiment of the present invention. As Figure 1 shown, the mobile terminal may include one or more ( Figure 1Only one processor 102 (the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a field-programmable gate array FPGA) and a memory 104 for storing data are shown. Among them, the above mobile terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that Figure 1 The structure shown is only schematic and does not limit the structure of the above mobile terminal. For example, the mobile terminal may further include more or fewer components than Figure 1 shown therein, or have a different configuration from Figure 1 shown.

[0026] The memory 104 can be used to store computer programs. For example, software programs and modules of application software, such as the computer program corresponding to the code defect detection method in the embodiment of the present invention. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implements the above method. The memory 104 may include a high-speed random access memory and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely provided with respect to the processor 102, and these remote memories can be connected to the mobile terminal through a network. Examples of the above network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0027] The transmission device 106 is used to receive or send data via a network. Specific examples of the above network may include a wireless network provided by a communication provider of the mobile terminal. In one instance, the transmission device 106 includes a network adapter (abbreviated as NIC), which can be connected to other network devices through a base station and thus can communicate with the Internet. In one instance, the transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0028] In this embodiment, a method for detecting code defects is provided. Figure 2 is a flowchart of the method for detecting code defects according to the embodiment of the present invention, as Figure 2 shown, and the process includes the following steps:

[0029] Step S202, in the case of receiving a code detection instruction, obtain all code source files of the target code indicated by the code detection instruction from the code repository, and determine the difference content between each of the code source files and the target code;

[0030] Step S204: Construct a target knowledge graph based on all the code source files, where the target knowledge graph is used to indicate the relationships between code block instances included in the same code source file and the relationships between file-level community instances in different code source files, and the file-level community instance in the code source file represents the function of each instance in the instance list included in the code source file.

[0031] Step S206: Determine the code defect detection context in the code source file based on the difference content and the target knowledge graph.

[0032] Step S208: Detect the target code based on the code defect detection context and the difference content to determine the defects of the target code.

[0033] In the above embodiment, the code repository can be understood as a source code repository or a version control repository, which can be a central repository for storing and managing project source code, and can also allow the creation of multiple branches for development and repair, etc. When a branch is completed with development or repair, it can be merged back into the main branch, that is, the code repository can record the version of each code modification and can view and restore historical versions, such as a Git repository. Therefore, in the case of receiving a code submission request, the code detection instruction will specify the target code and a certain branch where it is located, and all the code source files of the target code can be obtained from the code repository, that is, the latest version of the source code of the current branch can be pulled from the code repository, and the difference content Data between the target code and each code source file can be determined. diff ; In the case of receiving a code branch merge request, the latest version of the source code of the target branch (i.e., the above target code) to be merged can be pulled from the code repository, and the difference content Data between the target branch and the source code of the target branch can be determined. diff Among them, the code detection instruction can include the detection instruction when submitting code and the detection instruction when merging code.

[0034] In the above embodiment, the target knowledge graph can be understood as a network structure that represents entities and relationships by taking entities (such as people, places, events, concepts, etc.) as nodes and the relationships between entities as edges in the form of a graph. Therefore, a target knowledge graph can be constructed through all the code source files to represent the relationships between code block instances in the same code source file and the file-level community instance relationships between file-level community instances in different code source files, where the file-level community instance can be understood as the function of each instance in the set of all instances (i.e., the above instance list).

[0035] In the above embodiments, during the process of determining the context of the code defect, the difference content Data between the obtained target code and the code source file can be used diff as input, extract relevant instances, and then search for relevant instances in the target knowledge graph to obtain the corresponding source code text, which can be used as the context for code defect detection, construct prompt words for code defect detection, and determine the defects of the target code to obtain content such as code defect detection results, scores, and modification suggestions.

[0036] Through the present invention, when a code detection instruction is received, all code source files of the target code to be detected indicated by the code instruction can be obtained from the code repository, and the difference content between each code source file and the target code can be determined. All the code source files can be used to construct a target knowledge graph, which can indicate the relationships between code block instances included in the same code source file and the relationships between file-level community instances in different code source files. Among them, the file-level community instance in the code source file can represent the role of each instance in the instance list included in the code source file. Through the above difference content and the constructed target knowledge graph, the code defect detection context in the code source file can be determined, and then the target code can be detected through the code defect detection context and the difference content to determine the defects of the target code. Therefore, the problem of inaccurate code defect detection in the related art can be solved, and the effect of improving the accuracy of code defect detection can be achieved.

[0037] Optionally, the execution subject of the above steps may be a client, a terminal, a server, or a virtual machine, etc., but not limited thereto.

[0038] In an exemplary embodiment, constructing a target knowledge graph based on all the code source files includes: splitting each of the code source files to obtain a plurality of code blocks; determining the instances of each of the code blocks to obtain a plurality of the code block instances; constructing a bottom-layer knowledge graph based on the plurality of code block instances; determining the instance list of each of the code source files to obtain a plurality of instance lists; determining the instance role of each instance included in the plurality of instance lists; determining the instance role as a file-level community instance; constructing a middle-layer knowledge graph based on the file-level community instances; and determining the bottom-layer knowledge graph and the middle-layer knowledge graph as the target knowledge graph.

[0039] In the above embodiments, each source code file can be segmented first, and large files or complex code structures can be decomposed into smaller, more understandable and processable code blocks. Based on the segmented code blocks, instances in each code block can be determined, and then a low-level knowledge graph can be constructed according to multiple code block instances. Among them, an instance can be understood as various elements in the code, such as variables, classes, function definitions, etc. In this embodiment, instances in all source code files can be constructed into multiple instance lists, the instance roles of all instances in each instance list can be analyzed, and the instance roles can be summarized to form an overview of the function of the entire file and a description of the community instances (i.e., the above-mentioned file-level community instances). For example, a function may be used for data processing, and a variable may be used to store intermediate results, etc. Among them, the instance role can be understood as the attribute of the instance.

[0040] In the above embodiments, after determining the file-level community instances, a middle-level knowledge graph can be constructed through the file-level community instances, that is, the dependency relationship, call relationship or data flow relationship between file levels, which can reflect the interaction and collaboration between different files or modules in the source code. Finally, the low-level knowledge graph (reflecting the internal structure of the code and the relationship between elements) and the middle-level knowledge graph (reflecting the interaction between files or modules) can be integrated to form a target knowledge graph, which can comprehensively reflect the local and global structure of the code, as well as the functions and relationships of code elements and file communities.

[0041] In an exemplary embodiment, constructing a low-level knowledge graph based on multiple code block instances includes: determining first code block instances that belong to the same source code file among the multiple code block instances; determining the instance relationship between the first code block instances; determining the relationship between second code block instances that belong to different source code files among the multiple code block instances to obtain the relationship between code block instances; and determining the knowledge graph including the instance relationship and the relationship between code block instances as the low-level knowledge graph.

[0042] In the above embodiments, the large language model can be used to determine, from multiple code block instances, those first code block instances that belong to the same code source file. These can include functions, class definitions, objects, structs, enums, global variables, local variables, formal parameters, actual parameters, templates, template instances, etc. The extracted first code block instances can contain information such as instance names, instance types, and descriptions of the instances, and are formatted and output as a triple containing the instance name, instance type, and instance description. These first code block instances share the same file environment, and the relationships between them are usually direct and close. For example, a function may use a variable defined in the same file, or a class may inherit from another class within the file. Among them, the large language model (LLM) can be understood as a deep learning model with a large number of parameters (usually ranging from hundreds of millions to hundreds of billions), which can generate high-quality natural language text through training on large-scale text data. In this embodiment, specific prompt words can be designed to determine the code block instances and their relationships. Among them, the prompt information used is shown in Table 1, and the code block Code id is used as input to construct a Prompt Node , for extracting the first code block instances; the code block Code id is used as input to construct a Prompt subNode , for extracting the instances of each component between the first code block instances (i.e., the above-mentioned instance relationships), and obtaining the corresponding formatted output. Among them, the prompt words can be understood as the key instructions for guiding the large language model to perform knowledge reasoning and generation, which can be used to trigger the model to identify and summarize the code relationships and functions across files, and can include descriptions of issues such as function calls, data flows, control flows, etc., as well as instructions on how to map these relationships to file-level community instances.

[0043] In the above embodiments, the relationships between the second code block instances can also be determined by the large language model from multiple different code source files, that is, by constructing a Prompt Node_RExtract the relationships between the above code block instances. The extracted relationships between instances include information such as the source instance name, target instance name, relationship description, and relationship strength, and are formatted and output as a triple containing the source instance name, target instance name, and relationship description between instances. Among them, the relationships between code block instances can include a function using a variable, a function calling another function, a variable being an actual parameter of a function, a variable being the return value of a function, etc. In this embodiment, finally, the determined instance relationships and the relationships between code block instances can be integrated together to form a bottom-layer knowledge graph, which not only includes the elements within a single code source file and the relationships between them, but also includes the dependencies and interactions between different code source files, laying a foundation for understanding the code structure and logic of the entire project.

[0044] Table 1

[0045]

[0046] In an exemplary embodiment, constructing a middle-layer knowledge graph based on the file-level community instances includes: determining the relationships between the second code block instances belonging to different code source files among the multiple code block instances to obtain the relationships between code block instances; determining the mapping relationship between the file paths of the code source files and the file-level community instances; constructing a prompt word based on the relationships between the code block instances; analyzing the mapping relationship based on the prompt word to determine the file-level community instance relationships; and determining the knowledge graph including the file-level community instance relationships and the file-level community instances as the middle-layer knowledge graph.

[0047] In the above embodiment, the relationships between the second code block instances can be determined from multiple different code source files through a large language model. As shown in Table 1, that is, a Prompt can be constructed Node_R Extract the relationships between the above code block instances. The extracted relationships between instances include information such as the source instance name, target instance name, relationship description, and relationship strength, and are formatted and output as a triple containing the source instance name, target instance name, and relationship description between instances. Among them, the relationships between code block instances can include a function using a variable, a function calling another function, a variable being an actual parameter of a function, a variable being the return value of a function, etc.

[0048] In the above embodiment, the code block instances in different code source files can be used as inputs to construct a Prompt fileNode , as shown in Table 2. By using the large language model to summarize the instance list, file-level community instances can be obtained and the mapping relationship between the file-level community instances and the code source file paths can be recorded. The relationships between code block instances between different code source files can be used as inputs to construct a prompt word Prompt file_R, among which, by analyzing and summarizing the mapping relationship through prompt words, the relationships between file-level community instances can be obtained. For example, the model may summarize a cross-file dependency mapping relationship such as "the functions in File A depend on the data structures in File B".

[0049] In the above embodiments, finally, the file-level community instance relationships can be integrated with the file-level community instances to form a middle-level knowledge graph. Among them, the middle-level knowledge graph connects the underlying code block instances and the high-level project functions and architectures, which helps to locate code defects, optimize the code structure, and improve the overall code quality. It can not only include the function descriptions of individual files, but also reflect the collaboration and dependencies between files, providing a deeper perspective for understanding the overall project architecture and code logic.

[0050] Table 2

[0051]

[0052]

[0053] In an exemplary embodiment, each of the code source files is segmented to obtain a plurality of code blocks, including: determining the start position and the end position corresponding to each target structure included in the code source file; segmenting the code source file according to the start position and the end position to obtain a plurality of the code blocks.

[0054] In the above embodiments, during the process of parsing the code source file, a concrete syntax tree parser library can be used to generate a concrete syntax tree (CST). By traversing the concrete syntax tree, the start lines and ends (i.e., the above start positions and end positions) corresponding to elements such as functions, classes, structs, enums, global variables, templates, etc. (i.e., the above target structures) in the code source file can be identified. Among them, the concrete syntax tree can be understood as a tree structure representing the source code, which can directly reflect the syntax structure of the source code, including all syntax details such as parentheses, spaces, comments, etc., and can accurately represent and operate on the source code, reflecting the hierarchical relationship of the source code.

[0055] In the above embodiments, after determining the start and end positions of the target structure, the code source file can be segmented according to these position information, and each structure or element is used as an independent code block, and a series of code blocks Code can be obtained id , where id = 1, 2, …, N. The embedding representations Embedding of these code blocks can also be obtained through a text embedding model id, and record the associations between code block embeddings and code blocks, and the source code block content can be obtained. For example, if a function definition is recognized, all the internal code lines included from the start position of the function declaration (such as def functionName(...): or functionName(...)) to the end position of the function body will be segmented into a code block. Through the above segmentation process, multiple code blocks will eventually be obtained, and each code block contains the complete definition and implementation of a target structure or element. Among them, the code blocks may include but are not limited to functions, classes, templates, conditional statements, loop statements, etc. By identifying the start and end positions of the target structure, it can be ensured that the segmented code blocks contain complete structure definitions, avoiding code logic breaks or redundancies, enabling modular processing of the code, facilitating parallel analysis and understanding, and improving the efficiency of code review and defect detection. In addition, the segmentation based on the target structure can also make the code analysis more structured, facilitating further instance and relationship extraction, and constructing an underlying knowledge graph.

[0056] In an exemplary embodiment, determining the code defect detection context in the code source file based on the difference content and the target knowledge graph includes: determining the initial region corresponding to the difference content in the code source file; expanding the initial region to obtain a target region, where the range of the target region is larger than that of the initial region, and the target region includes the initial region; determining the target instances of the target content included in the target region; determining a first instance associated with the target instance in the underlying knowledge graph included in the target knowledge graph, where the underlying knowledge graph is used to indicate the code block instance relationships between code block instances included in the same code source file; determining a second instance associated with the first instance in the middle-level knowledge graph included in the target knowledge graph, where the middle-level knowledge graph is used to indicate the file-level community instance relationships between file-level community instances in different code source files; obtaining the target file-level community instance relationship of the second instance and the instance description of the second instance; and determining the target file-level community instance relationship and the instance description as the code defect detection context.

[0057] In the above embodiment, the difference content Data diff can be used as input to parse the line number range [L1, L2] corresponding to the modified location of the code source file (i.e., the above initial region). However, in order to obtain more comprehensive context information, since the code around the initial region is all related and can help understand and capture the complete logic of the difference code context, the line number range [L1, L2] can be expanded upward and downward to the target region [L1 - s, L2 + s].

[0058] In the above embodiments, target instances can be extracted from the target content included in the extended target area, that is, the target instances related to the differential content Data can be determined in the underlying knowledge graph through a graph matching algorithm. diff Among them, the underlying knowledge graph can reflect the relationships between code block instances within the same code source file. Through the underlying knowledge graph, the direct connections between target instances and other code elements (such as the functions that call it or the variables it accesses) can be identified.

[0059] In the above embodiments, the target instance can be used as the entry point of the underlying knowledge graph to obtain the first instance related to the target instance, such as the instance connected to the target instance and the corresponding instance relationship. The first instance determined in the underlying graph can be used as the entry point of the middle-level knowledge graph to obtain the second instance related to the first instance, and determine the target file-level and target directory-level community instance relationships and instance descriptions of the second instance, as well as the mapping relationship between the second instance and the code block, which can be used as the code defect detection context. Among them, the middle-level knowledge graph can cover the file-level community instances between different code source files and the relationships between file-level community instances. The acquisition of the directory-level community can be understood as processing the code block instances corresponding to all source code files under a certain directory. The process of obtaining the target directory-level community instance relationships and instance descriptions can be seen in Table 3, which is the same as the operation of obtaining the target file-level community instance relationships and instance descriptions.

[0060] Table 3

[0061]

[0062]

[0063] In an exemplary embodiment, determining the first instance associated with the target instance in the underlying knowledge graph included in the target knowledge graph includes: determining the instance with the most connections to the target instance included in the underlying knowledge graph; and determining the most-connected instance as the first instance.

[0064] In the above embodiments, when obtaining the first instance related to the target instance in the underlying knowledge graph, the code block instance with the most connections to the target instance can be used as the first instance, that is, the first instance with the highest connection frequency and the closest relationship is identified, which can be the most frequently used or most dependent code instance. For example, in a function instance that frequently calls other functions, the function with the most call times will be selected as the first instance. The selection of the first instance is based on the strength of the direct relationship between it and the target instance, and the most representative and influential instance can be selected as the focus of analysis, thereby improving the accuracy and efficiency of code defect detection.

[0065] In an exemplary embodiment, the target code is detected based on the code defect detection context and the difference content to determine the defects in the target code, including: constructing a difference detection prompt word using the code defect detection context and the difference content; detecting the target code based on the prompt word to determine the defects in the target code.

[0066] In the above embodiment, the difference content Data diff and the code defect detection context can be used to construct a difference detection prompt word prompt. Among them, the difference detection prompt word prompt may include background information of the target code, descriptions of associated code instances and file-level community relationships, as well as specific difference content, etc. By integrating this information, the difference detection prompt word can provide a comprehensive perspective to the large language model, helping the model understand the position and role of the target code in the project, as well as the intention and possible impact of the modified or newly added code.

[0067] In the above embodiment, the large language model can analyze the syntax and semantics of the target code based on the information in the difference detection prompt word prompt, and identify possible defects, which may include syntax errors, logical loopholes, performance issues, security risks, specific locations, types, scopes of influence, and repair suggestions of the defects, etc. For example: "In the loop on line 10, the variable i is not properly incremented before the end of the loop, which may cause an infinite loop" or "The function readFile does not check whether the file is successfully opened, there is a potential risk of null pointer exception".

[0068] The following describes the code defect detection method in combination with specific embodiments:

[0069] Figure 3 It is a flowchart of the code defect detection method according to a specific embodiment of the present invention, including the following steps:

[0070] Step S302, obtain the diff file;

[0071] Step S304, parse all project source files;

[0072] Step S306, obtain code blocks through CST;

[0073] Step S308, extract instance and instance relationships;

[0074] Step S310, construct an underlying knowledge graph;

[0075] Step S312, construct a middle-level knowledge graph;

[0076] Step S314, construct a file-level community;

[0077] Step S316, construct a multi-level directory-level community;

[0078] Step S318, obtain the context;

[0079] Step S320, code defect detection.

[0080] In the above embodiments, when a code submission request is received, the latest version of the code in the branch where the target code is located can be pulled from the code repository as the source file to be processed (i.e., the above-mentioned code source file), and the difference content Data between the latest version of the code and the code is obtained. diff ; when a code branch merge request is received, the latest version of the code in the branch where the target code is located is pulled from the code repository as the source file to be processed (i.e., the above-mentioned code source file), and the difference content Data between the code and the source code of the target branch is obtained. diff ;

[0081] In the above embodiments, all source files to be processed (i.e., the above-mentioned code source files) can be parsed, a specific syntax tree is generated using a specific syntax tree parsing library, and the specific syntax tree is traversed to identify the start and end lines (i.e., the above-mentioned start position and end position) corresponding to elements such as functions, classes, structs, enums, global variables, templates, etc. in the source code, which are used to split the corresponding source files to obtain a series of code blocks Code id , where id = 1, 2,..., N. The embedding representation Embedding of these code blocks is obtained using a text embedding model. id , where id = 1, 2,..., N. And the association between the code block embeddings and the code blocks is recorded for obtaining the code block content through the embedding representation in subsequent steps;

[0082] In the above embodiments, instance extraction can be performed using a large language model, mainly including functions, class definitions, objects, structs, enums, global variables, local variables, formal parameters, actual parameters, templates, template instances, etc. The extracted instances include information such as instance names, instance types, and descriptions of the instances, and are formatted and output as a triple (instance name, instance type, instance description) containing the following fields. Among them, the prompt information used is shown in Table 1, and the code block Code id is used as the input to construct a Prompt Node , for extracting code block instances; further, the code block Code id is used as the input to construct a Prompt subNode , for extracting instances of each component inside the code block to obtain the corresponding formatted output; construct a Prompt Node_RUsing a large language model to extract the relationships between the above-mentioned instances, the relationships between the extracted instances include information such as the source instance name, target instance name, relationship description, and relationship strength, and are formatted and output as a triple (source instance name, target instance name, relationship description between instances). The relationships between instances include function using variables, function calling functions, variables being function arguments, variables being function return values, etc. All of the above instances and the relationships between instances together constitute the underlying knowledge graph required by this proposal;

[0083] In the above embodiment, the obtained code block instances and the relationships between code block instances can be used to construct a middle-level knowledge graph, and the mapping relationship between the instances and the source code block embeddings can be recorded; the code block instances in each source file can be used as input to construct a Prompt fileNode , using a large language model to summarize the instance list, obtain file-level community instances, and record the mapping relationship between the file-level community instances and the source file paths. Using the relationships between code block instances between different source files as input to construct a Prompt file_R , using a large language model to summarize the relationship list, obtain the relationships between file-level community instances, and obtain a file-level community, that is, the relationships between file-level community instances and file-level community instances; perform similar operations to obtain multi-level directory-level communities;

[0084] In the above embodiment, during the context acquisition process, the difference content Data between the obtained source code of the branch where the target code is located can be parsed diff As input, parse the line number range corresponding to the modified location of the source file (i.e., the above-mentioned initial area) [L1, L2], expand this range upward and downward to (i.e., the above-mentioned target area) [L1 - s, L2 + s], extract instances from the content of the target area, and use a graph matching algorithm to identify the target instances related to the difference content in the underlying knowledge graph. Use these target instances as the entry point of the underlying knowledge graph to obtain more instance information related to the domain target instances (i.e., the above-mentioned first instances), such as the instances connected to them and the corresponding relationships between instances. Among them, what is obtained after using the difference content for instance extraction processing is a subgraph in the underlying knowledge graph. The graph matching algorithm is an algorithm used to find similar or corresponding relationships in a graph structure, that is, to find the position of the subgraph in the complete underlying knowledge graph.

[0085] In the above embodiments, the code block instance with the most connections to the first instance (i.e., the second instance above) can be searched in the underlying knowledge graph as the entry of the middle-level knowledge graph, and relevant middle-level knowledge graph nodes can be obtained. The corresponding source code text content can be obtained by using the mapping between the nodes and the code blocks, and it is used as the context for code defect detection. According to the nodes of the middle-level knowledge graph used, the community instance descriptions and relationships of the file-level and directory-level communities can be obtained and used as the context for code defect detection;

[0086] In the above embodiments, a difference detection prompt "prompt" can be constructed by using the difference content between the source code of the branch where the target code is located and the code defect detection context, which can be used to obtain code defect detection results, scores, modification suggestions, etc.

[0087] Through the description of the above implementation manners, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation manner. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc) and includes several instructions for causing a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in various embodiments of the present invention.

[0088] In this embodiment, a code defect detection device is also provided. This device is used to implement the above embodiments and preferred implementation manners, and those that have been described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0089] Figure 4 is a structural block diagram of a code defect detection device according to an embodiment of the present invention, as Figure 4 shown, the device includes:

[0090] An acquisition module 42, configured to, when receiving a code detection instruction, acquire all code source files of the target code indicated by the code detection instruction from a code repository, and determine the difference content between each code source file and the target code;

[0091] A construction module 44 for constructing a target knowledge graph based on all the code source files, wherein the target knowledge graph is used to indicate the relationships between code block instances included in the same code source file, and to indicate the relationships between file-level community instances in different code source files, and the file-level community instances in the code source file represent the functions of each instance in the instance list included in the code source file;

[0092] A determination module 46 for determining a code defect detection context in the code source file based on the difference content and the target knowledge graph;

[0093] A detection module 48 for detecting the target code based on the code defect detection context and the difference content to determine the defects of the target code.

[0094] In an exemplary embodiment, the construction module 44 can construct the target knowledge graph based on all the code source files in the following manner: segment each code source file to obtain a plurality of code blocks; determine the instances of each code block to obtain a plurality of code block instances; construct a bottom-layer knowledge graph based on the plurality of code block instances; determine the instance list of each code source file to obtain a plurality of instance lists; determine the function of each instance included in the plurality of instance lists; determine the function of the instance as a file-level community instance; construct a middle-layer knowledge graph based on the file-level community instance; and determine the bottom-layer knowledge graph and the middle-layer knowledge graph as the target knowledge graph.

[0095] In an exemplary embodiment, the construction module 44 can construct the bottom-layer knowledge graph based on the plurality of code block instances in the following manner: determine the first code block instances that belong to the same code source file among the plurality of code block instances; determine the relationships between the first code block instances; determine the relationships between the second code block instances that belong to different code source files among the plurality of code block instances to obtain the relationships between code block instances; and determine the knowledge graph including the relationships between the instances and the relationships between the code block instances as the bottom-layer knowledge graph.

[0096] In an exemplary embodiment, the building module 44 may implement building a middle-level knowledge graph based on the file-level community instance in the following manner: determining the relationships between second code block instances included in multiple code block instances that belong to different code source files to obtain the relationships between code block instances; determining the mapping relationship between the file paths of the code source files and the file-level community instance; building prompts based on the relationships between the code block instances; analyzing the mapping relationship based on the prompts to determine the file-level community instance relationships; and determining the knowledge graph including the file-level community instance relationships and the file-level community instance as the middle-level knowledge graph.

[0097] In an exemplary embodiment, the building module 44 may implement splitting each code source file to obtain multiple code blocks in the following manner: determining the start position and end position corresponding to each target structure included in the code source file; and splitting the code source file according to the start position and the end position to obtain multiple code blocks.

[0098] In an exemplary embodiment, the determining module 46 may implement determining code defect detection context in the code source file based on the difference content and the target knowledge graph in the following manner: determining the initial region corresponding to the difference content in the code source file; expanding the initial region to obtain a target region, where the range of the target region is larger than that of the initial region and the target region includes the initial region; determining the target instances of the target content included in the target region; determining a first instance associated with the target instance in the underlying knowledge graph included in the target knowledge graph, where the underlying knowledge graph is used to indicate the relationships between code block instances included in the same code source file; determining a second instance associated with the first instance in the middle-level knowledge graph included in the target knowledge graph, where the middle-level knowledge graph is used to indicate the file-level community instance relationships between file-level community instances in different code source files; obtaining the target file-level community instance relationship and the instance description of the second instance; and determining the target file-level community instance relationship and the instance description as the code defect detection context.

[0099] In an exemplary embodiment, the determining module 46 may implement determining a first instance associated with the target instance in the underlying knowledge graph included in the target knowledge graph in the following manner: determining the instance with the most connections to the target instance included in the underlying knowledge graph; and determining the instance with the most connections as the first instance.

[0100] In an exemplary embodiment, the detection module 48 may detect the target code based on the code defect detection context and the difference content in the following manner to determine the defects of the target code, including: constructing a difference detection prompt word by using the code defect detection context and the difference content; and detecting the target code based on the prompt word to determine the defects of the target code.

[0101] It should be noted that the above-mentioned modules can be implemented by software or hardware. For the latter, it can be implemented in the following ways, but not limited to: all the above modules are located in the same processor; or, the above-mentioned modules are respectively located in different processors in any combination form.

[0102] An embodiment of the present invention also provides a computer-readable storage medium, in which a computer program is stored. Wherein, the computer program is configured to execute the steps in any one of the above method embodiments when running.

[0103] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media such as USB flash drives, read-only memories (ROM for short), random access memories (RAM for short), mobile hard disks, magnetic disks or optical discs that can store computer programs.

[0104] An embodiment of the present invention also provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.

[0105] In an exemplary embodiment, the above electronic device may further include a transmission device and an input / output device. Wherein, the transmission device is connected to the above processor, and the input / output device is connected to the above processor.

[0106] An embodiment of the present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the methods in various embodiments of the present application are implemented.

[0107] The specific examples in this embodiment may refer to the examples described in the above embodiments and exemplary embodiments, and will not be repeated here.

[0108] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the present invention can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. They can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order from here, or they can be separately made into individual integrated circuit modules, or multiple modules or steps among them can be made into a single integrated circuit module to implement. In this way, the present invention is not limited to any specific combination of hardware and software.

[0109] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for detecting code defects, characterized in that: include: When receiving the code detection instruction, obtaining all code source files of the target code indicated by the code detection instruction from the code repository, and determining the difference between each of the code source files and the target code; Constructing a target knowledge graph based on all the code source files, wherein the target knowledge graph is used to indicate the relationship between code block instances included in the same code source file, and to indicate the relationship between file-level community instances in different code source files, wherein the file-level community instances in the code source files represent the role of each instance in the instance list included in the code source files; Determine a code defect detection context in the code source file based on the difference content and the target knowledge graph; The target code is detected based on the code defect detection context and the difference content to determine the defect of the target code.

2. The method according to claim 1, characterized in that Building a target knowledge graph based on all the code source files includes: Splitting each of the code source files to obtain multiple code blocks; Determine an instance of each of the code blocks to obtain a plurality of the code block instances; Building an underlying knowledge graph based on multiple code block instances; Determine an instance list of each of the code source files to obtain multiple instance lists; determining an instance role of each instance included in a plurality of said lists of instances; Determining the instance role as a file-level community instance; Building a mid-level knowledge graph based on the file-level community instance; The underlying knowledge graph and the middle-level knowledge graph are determined as the target knowledge graph.

3. The method according to claim 2, characterized in that Building an underlying knowledge graph based on multiple code block instances includes: Determine a first code block instance included in the plurality of code block instances and belonging to the same code source file; Determining instance relationships between the first code block instances; Determine the relationship between the second code block instances belonging to different code source files included in the plurality of code block instances to obtain the relationship between the code block instances; The knowledge graph including the instance relationship and the relationship between the code block instances is determined as the underlying knowledge graph.

4. The method according to claim 2, characterized in that: Building a mid-level knowledge graph based on the file-level community instance includes: Determine the relationship between the second code block instances belonging to different code source files included in the plurality of code block instances to obtain the relationship between the code block instances; Determine a mapping relationship between a file path of the code source file and the file-level community instance; Constructing prompt words based on the relationship between the code block instances; Analyzing the mapping relationship based on the prompt word to determine the file-level community instance relationship; The knowledge graph including the file-level community instance relationship and the file-level community instance is determined as the middle-level knowledge graph.

5. The method according to claim 2, characterized in that: Each of the code source files is segmented to obtain multiple code blocks including: Determine the starting position and the ending position corresponding to each target structure included in the code source file; The code source file is segmented according to the starting position and the ending position to obtain a plurality of the code blocks.

6. The method according to claim 1, characterized in that Determining a code defect detection context in the code source file based on the difference content and the target knowledge graph includes: Determine an initial region corresponding to the difference content in the code source file; Expanding the initial area to obtain a target area, wherein the target area is larger than the initial area and includes the initial area; determining a target instance of target content included in the target area; Determine a first instance associated with the target instance in an underlying knowledge graph included in the target knowledge graph, wherein the underlying knowledge graph is used to indicate a relationship between code block instances included in the same code source file; Determining a second instance associated with the first instance in a middle-level knowledge graph included in the target knowledge graph, wherein the middle-level knowledge graph is used to indicate a file-level community instance relationship between file-level community instances in different code source files; Acquire a target file-level community instance relationship of the second instance and an instance description of the second instance; The target file-level community instance relationship and the instance description are determined as the code defect detection context.

7. The method according to claim 6, characterized in that Determining a first instance associated with the target instance in an underlying knowledge graph included in the target knowledge graph includes: Determine the largest number of instances in the underlying knowledge graph that are most connected to the target instance; The most instances are determined as the first instances.

8. The method according to claim 1, characterized in that Detecting the target code based on the code defect detection context and the difference content to determine that the defects of the target code include: Constructing a difference detection prompt word using the code defect detection context and the difference content; The target code is detected based on the prompt word to determine the defects of the target code.

9. A device for detecting code defects, characterized in that: include: An acquisition module is used to, upon receiving a code detection instruction, acquire all code source files of the target code indicated by the code detection instruction from a code repository, and determine the difference between each of the code source files and the target code; A construction module, configured to construct a target knowledge graph based on all the code source files, wherein the target knowledge graph is used to indicate the relationship between code block instances included in the same code source file, and to indicate the relationship between file-level community instances in different code source files, wherein the file-level community instance in the code source file represents the role of each instance in the instance list included in the code source file; A determination module, configured to determine a code defect detection context in the code source file based on the difference content and the target knowledge graph; A detection module is used to detect the target code based on the code defect detection context and the difference content to determine the defects of the target code.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program is configured to execute the method according to any one of claims 1 to 8 when executed.

11. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to run the computer program to perform the method according to any one of claims 1 to 8.

12. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method described in any one of claims 1 to 8 are implemented.

Citation Information

Cited By

  • Code scanning analysis method and system based on large model

    CN120893046A

  • Code retrieval method and device and related equipment

    CN121277885A