Code repository question and answer method based on hierarchical code knowledge graph

By building a hierarchical code knowledge graph and Monte Carlo tree search optimization Q&A path, the problem of insufficient structural and semantic correlation of code warehouse Q&A systems in large-scale code libraries in the existing technology is solved, and an efficient and accurate Q&A system is realized.

CN120523911APending Publication Date: 2025-08-22ZHEJIANG UNIV

Patent Information

Application Number
CN202510617262.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-08-22

AI Technical Summary

Technical Problem

When handling large-scale code repository Q&A systems, existing code repository Q&A systems cannot fully capture the hierarchical structure and semantic relationships between various parts of the code, and lack the ability to optimize Q&A paths.

Method used

The hierarchical code knowledge graph construction method is adopted to analyze the code structure through abstract syntax trees, combine community detection and hierarchical clustering to generate code knowledge graphs, and use large language models to generate community abstracts, and combine Monte Carlo tree search to optimize the Q&A path.

Benefits of technology

It improves the quality of code knowledge modeling and the accuracy and efficiency of inference of Q&A system, and significantly improves the adaptability of complex queries and the accuracy of Q&A results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120523911A_ABST
    Figure CN120523911A_ABST
Patent Text Reader

Abstract

The invention discloses a code repository question and answer method based on a hierarchical code knowledge graph, which comprises the following steps: collecting all files of a target code repository, and preprocessing to obtain code files and text files; generating an abstract syntax tree from the code file, and extracting related entities; carrying out block segmentation on the text file, and extracting to obtain text blocks; constructing a code knowledge graph representing a code structure based on the extracted entities and text blocks; generating a hierarchical code knowledge graph by adopting a community detection algorithm and a hierarchical clustering method, and generating a functional abstract for each community in combination with a large language model; a Monte Carlo tree search-based code warehouse question and answer agent is used for performing question and answer path optimization, an optimal query track is dynamically planned through the steps of node selection, expansion, simulation, reward evaluation and the like, and a path with the highest accumulated reward is selected to generate a final question and answer result. The code knowledge modeling quality and the reasoning accuracy and efficiency of the question answering system can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of code knowledge graph construction and question-answering system, and in particular to a code warehouse question-answering method based on a hierarchical code knowledge graph. Background Art

[0002] With the increasing abundance of code repositories and the continued expansion of software development, the efficient acquisition and utilization of structured information in code has become a focus of industry attention. Existing technologies, pre-trained code models (such as CodeT5 and CodeBERT) have achieved certain results in tasks such as code understanding, code generation, and defect detection.

[0003] However, these methods mainly focus on the analysis of single code snippets and usually use abstract syntax trees (ASTs) for code parsing, but they cannot fully capture the hierarchical structure and semantic associations between different parts of the code when dealing with large-scale code bases.

[0004] Chinese patent publication number CN118733737A discloses a question-and-answer processing method and apparatus, comprising: obtaining a natural language question text; matching the natural language question text against a code library to obtain candidate code snippets, the code library including code snippets obtained by parsing and extracting target code snippets; using intent categories obtained by performing intent recognition on the natural language question text to filter out candidate code snippets that do not match the intent categories; selecting a target code snippet from the filtered candidate code snippets, and using the target code snippet to generate a response to the natural language question text. This patent improves the accuracy of question-and-answer processing by performing intent recognition on the natural language question text and filtering candidate code snippets.

[0005] However, most existing code question-answering systems rely solely on a simple fusion of code and textual information, lacking in-depth exploration of the community structure and modularity hidden within the code. Traditional approaches often directly construct static knowledge graphs after collecting data, filtering files, and parsing code repositories. These approaches fail to utilize clustering algorithms to identify close connections between code entities, nor do they incorporate large language models (LLMs) to automatically generate community summaries. This limits the system's adaptability to complex queries and the accuracy of its Q&A results.

[0006] On the other hand, question-answering sampling methods based on Monte Carlo Tree Search (MCTS) have garnered significant attention in recent years in the fields of natural language processing and decision optimization. By sampling and evaluating the reward model multiple times on the question-answering trajectory, they can effectively mitigate the problems of target deviation and information forgetting during reasoning, improving the confidence of the final question-answering results. However, there is currently a lack of systematic solutions for combining Monte Carlo Tree Search (MCTS) with code knowledge graphs and further leveraging community clustering and large model (LLM) to automatically generate summaries based on the hierarchical structure to optimize query paths.

[0007] Therefore, how to design a code repository question-answering method that integrates code structured feature information to achieve accurate question-answering and efficient query optimization of code repositories has become an urgent problem to be solved. Summary of the Invention

[0008] To address the problems of insufficient structural information extraction, incomplete semantic representation, and insufficient question-answering path optimization capabilities in existing code repository question-answering methods, the present invention provides a code repository question-answering method based on a hierarchical code knowledge graph, which can effectively improve the quality of code knowledge modeling and the reasoning accuracy and efficiency of the question-answering system.

[0009] A code repository question-answering method based on a hierarchical code knowledge graph includes the following steps:

[0010] S1. Collect all files in the target code repository and preprocess them to obtain code files and text files.

[0011] S2. Generate an abstract syntax tree from the code file and extract related entities; segment the text file into blocks and extract text blocks;

[0012] S3. Build a code knowledge graph representing the code structure based on the extracted entities and text blocks;

[0013] S4. First, the constructed code knowledge graph is preliminarily clustered using a community detection algorithm to form preliminary communities. Then, these preliminary communities are refined and layered using a hierarchical clustering method to generate a hierarchical code knowledge graph. Clusters at different levels within the hierarchical code knowledge graph constitute the final communities. Finally, the text blocks within each final community are processed using a large model to form the corresponding community summary.

[0014] S5. Optimize the question-answering path through a code repository question-answering agent based on Monte Carlo tree search. The agent assists in node selection and expansion based on the community summaries contained in the community nodes of the hierarchical code knowledge graph (these nodes contain previously generated community summaries), and evaluates rewards based on the simulation expansion results. Subsequently, the agent dynamically updates the optimal query trajectory through backpropagation, and finally selects the path with the highest cumulative reward to generate the question-answering result.

[0015] In step S1, the preprocessing includes file format recognition, classification and filtering to distinguish code files from text files.

[0016] In step S2, the code file is generated into an abstract syntax tree in the following manner:

[0017] First, exclude the standard library, built-in functions, and classes of the code, then exclude third-party libraries introduced in the code file, and finally build an abstract syntax tree based on predefined grammar rules.

[0018] In step S2, the extracted related entities include classes, functions, methods and variables.

[0019] In step S3, the code knowledge graph uses a bidirectional graph to represent the bidirectional relationship between entities, as follows:

[0020] The nodes of the code file in the bidirectional graph include classes, functions, class member variables, and class member methods, and are divided into two types: definition def and reference ref;

[0021] The edges of the code files in the bidirectional graph include the connection between classes and methods, the connection between classes and member variables, the connection between class definitions and references, and the connection between function definitions and references;

[0022] The nodes of the text file in the bidirectional graph are the file name and the content after block segmentation;

[0023] The edges of text files in the bidirectional graph are the connections between the file names and the corresponding segmented text blocks.

[0024] The specific process of step S4 is:

[0025] S41. Use community detection algorithm to identify the connection between entities in the constructed code knowledge graph, and perform preliminary clustering on the code knowledge graph G = (V, E) based on the modularity index to form a preliminary community. Each node i∈V is assigned a community label c i , where the calculation formula of module Q is:

[0026]

[0027] Among them, k i =∑j∈v A ij is the degree of node i, is the total number of edges in the graph, δ(c i ,c j ) is the indicator function, when c i =c j Take 1 when it is, otherwise take 0;

[0028] S42, further use the hierarchical clustering method to refine the division and layering of the preliminary communities, and divide the preliminary community set {C1, C2, ..., C k According to the inter-community distance index d(C i ,C j ) gradually merge and layer until the predetermined clustering criteria are met to generate a code knowledge graph with a hierarchical structure;

[0029] S43. For each final community C, aggregate all text blocks and related information in the community to form a text set T C :T C ={t|t is a text block belonging to community C}, and calls the large language model summary generation function f LLM (·) for T C Generate community summary S C ; At the same time, the aggregated features and community summary information of each final community are recorded for subsequent query optimization.

[0030] In step S42, the inter-community distance index d (C i ,C j ).

[0031] The specific process of step S5 is:

[0032] S51. Node selection: Based on the community summary contained in the community node in the hierarchical code knowledge graph, the user's query is used as the root node in the hierarchical code knowledge graph, and the current node is selected according to a predefined selection strategy. The selection formula is:

[0033]

[0034] Where Q(s,a) represents the average reward for performing action a in state s, N(s) is the number of visits to state s, N(s,a) is the number of times action a is selected in state s, and c is the exploration constant.

[0035] S52, node expansion: Based on the selected node, several candidate actions are sampled at the same time, each candidate action is executed and a corresponding child node is generated, and the expanded node is added to the search tree;

[0036] S53, simulate expansion: randomly select a path from the currently expanded child node for simulation, that is, repeat the expansion and random selection operations until a leaf node is reached or a preset maximum search depth is reached;

[0037] S54, reward evaluation: For the leaf nodes reached during the simulation, the quality of the inference trajectory corresponding to the node is scored using the pre-trained reward model R(·) and its reward value r is calculated, where: r = R(x leaf ), x leaf Indicates the sampling trajectory results corresponding to the leaf node;

[0038] S55, back propagation update: propagate the reward value r obtained by the leaf node upward along the path, and update all the cumulative rewards on the path; the update formula is:

[0039]

[0040] Among them, r i is the reward value of the corresponding node in the i-th simulation;

[0041] S56. Repeat the simulation process of steps S51 to S55 above until the reward values ​​of all nodes tend to stabilize or reach the preset upper limit of the number of simulations; finally, based on the cumulative reward values ​​of each node, select the trajectory with the highest cumulative reward as the final question-and-answer result output.

[0042] Compared with the prior art, the present invention has the following beneficial effects:

[0043] 1. The present invention first collects and preprocesses the code repository comprehensively, uses the abstract syntax tree (AST) to perform structural analysis on the code, and divides the text into blocks, thereby achieving a full integration of code structure information and semantic information, and constructing a code knowledge graph with bidirectional association characteristics.

[0044] 2. During the knowledge graph construction process, this paper uses a community detection algorithm to identify close connections between entities in the graph and performs preliminary clustering based on metrics such as modularity. Subsequently, a hierarchical clustering method is used to refine and stratify these preliminary communities, ultimately generating a hierarchical code knowledge graph. This process fully explores the inherent connections between modules in the code base, providing rich structured information and semantic aggregation features for subsequent question-answering optimization.

[0045] 3. This invention further utilizes a large language model to automatically generate community summaries for the aggregated text blocks within each community, forming community-level semantic descriptions, effectively improving the accuracy of information representation and query matching efficiency in the knowledge graph. During the question-answering sampling phase, the Repo-QA Agent, based on Monte Carlo Tree Search (MCTS), performs global evaluation and backpropagation optimization of the question-answering trajectory through multiple simulation (rollout) processes, effectively alleviating the target deviation and information forgetting problems existing in traditional question-answering systems, and significantly improving the confidence and accuracy of the final question-answering results.

[0046] 4. The present invention not only makes full use of the structured information contained in the code itself, but also builds an end-to-end efficient question-answering system through community clustering, summary generation, and Monte Carlo Tree Search (MCTS) question-answering optimization. While reducing computing resources and training time consumption, the system demonstrates excellent retrieval accuracy and response speed in a variety of complex query scenarios, and has important theoretical significance and broad application prospects. Through multi-level information fusion and global question-answering optimization, the present invention significantly improves the performance and robustness of the code repository question-answering system, and has outstanding technical advantages and practical application value. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 This is a flow chart of a code repository question-answering method based on a hierarchical code knowledge graph in an embodiment of the present invention.

[0048] Figure 2 This is a framework diagram of a code repository question-answering method based on a hierarchical code knowledge graph in an embodiment of the present invention.

[0049] Figure 3 This is an example diagram of the knowledge graph in an embodiment of the present invention. DETAILED DESCRIPTION

[0050] The present invention will be described in further detail below with reference to the accompanying drawings and examples. It should be noted that the following examples are intended to facilitate understanding of the present invention and do not have any limiting effect on the present invention.

[0051] The present invention is implemented through the following technical solutions: First, source code files and document class files are extracted from the code repository and processed separately. For code files, an abstract syntax tree (AST) is used to extract entities such as functions, classes, member variables, and a bidirectional graph structure of definitions and references is constructed; for document files, they are divided into chunks, and then connecting edges are generated between file names and text chunks. Based on the above structure, a unified knowledge graph is constructed. Subsequently, community detection and hierarchical clustering methods are used to construct a hierarchical code knowledge graph (HCKG), and summary information is generated for each community to improve semantic aggregation capabilities. In the question-answering stage, a code repository question-answering agent (Repo-QAAgent) based on Monte Carlo tree search (MCTS) is introduced to plan the question-answering path, recall relevant code and text from the hierarchical code knowledge graph (HCKG), and generate accurate answers through a semantic aggregation module. This method integrates structured code information and semantic content, has high efficiency and strong generalization capabilities, and is suitable for scenarios such as code repository question-answering and code document generation.

[0052] like Figure 1 and Figure 2 As shown in the figure, a code repository question answering method based on a hierarchical code knowledge graph includes the following steps:

[0053] S1: Data collection and filtering: Collect code repositories and filter irrelevant files:

[0054] S1.1. Obtain all files in the target source code repository;

[0055] S1.2. Preprocess the files, including file format recognition, classification and filtering, and distinguishing code files from non-code files.

[0056] In this embodiment, code repositories are first collected on a code repository platform, such as Github (the world's largest code hosting platform), and the commits corresponding to the issues are selected as the dataset. Data acquisition methods include cloning the repository using the git command. After collecting and integrating the raw dataset, the raw dataset is processed. The data processing process includes file format recognition and filtering out some unnecessary format files, such as database files. In addition, due to the different processing methods for code files and text files, code files are also distinguished from non-code files.

[0057] S2. Code and text processing: Use the abstract syntax tree (AST) to parse code snippets and use chunk segmentation to split text:

[0058] S2.1. Load the code file and the above text files that meet the requirements;

[0059] S2.2. Generate an abstract syntax tree from the code file and split the text file into chunks;

[0060] S2.3. Extract related entities such as classes, functions, methods, and variables based on the abstract syntax tree.

[0061] In this embodiment, the code portion of the collected data set is converted into an abstract syntax tree using the tool tree-sitter (a grammar-based code parsing tool for generating abstract syntax trees). Tree-sitter is an open source generator tool that uses the grammatical rules of a programming language to generate an abstract syntax tree. The abstract syntax tree uses a tree structure to display the grammatical structure of the source code, where each node in the tree represents a part of the code. This tree-like representation allows the serialized code to display deep structural information, allowing development tools and systems to understand and operate the code more intelligently. In the present invention, the abstract syntax structure can help the model better understand the semantic information of the code.

[0062] In step S2.2, the system generates an abstract syntax tree from the code file. First, the system reads the target source code file from disk or a code repository in plain text format. Next, the system uses the programming language's corresponding parser to perform syntax analysis on the source code file. The parser breaks the code file into various grammatical units according to Python's grammatical rules and constructs a tree structure.

[0063] In step S2.2, the text file is split into chunks. First, a mapping from file extensions to files is created, and then different splitting methods are used depending on the corresponding files. First, different file formats are different, so different overlapping methods and splitters are used.

[0064] For json files, use the RecursiveJsonSplitter splitter with the overlap size set to 0;

[0065] For YAML files, use the RecursiveCharacterTextSplitter splitter with an overlap size of 100;

[0066] For markdown files, use the MarkdownTextSplitter splitter with an overlap size of 200;

[0067] For the other files, use the RecursiveCharacterTextSplitter splitter with an overlap size of 300.

[0068] S3. Knowledge graph construction: Build a knowledge graph representing the code structure based on the extracted entities and text blocks:

[0069] Use the networkx library (a Python graph theory analysis tool) to generate a bidirectional graph, including:

[0070] S3.1. Code file nodes in the bidirectional graph include classes, functions, class member variables, and member methods, and are divided into two types: definition (def) and reference (ref);

[0071] S3.2. Code edges in a bidirectional graph include connections between classes and methods, between classes and member variables, between class definitions and references, and between function definitions and references.

[0072] S3.3. The nodes of the text file in the bidirectional graph are the file name and the content after chunk segmentation;

[0073] S3.4. The edges of text files in the bidirectional graph are the connections between the file names and the corresponding chunk segments.

[0074] Combine Figure 3 , for code files ( Figure 3 Left side), nodes are classes (references or definitions), class member variables, class member methods and functions (references or definitions), and the connections between them are existence edges; for text files ( Figure 3 On the right side), the nodes are text file names and corresponding block segments, and the connection between them is the existence edge.

[0075] In the above steps, for code composition, through static analysis of the code, the node-based expression of code elements and the edge-based modeling of their relationships are achieved:

[0076] The node construction steps are as follows:

[0077] (1) Node Source: Through step S2, all key code elements (such as classes, functions, variables, references, etc.) are extracted and standardized into "Tag" objects. Each tag contains the following attributes:

[0078] name: a unique identifier for a code element, such as a function name or class name;

[0079] fname, rel_fname (path): the absolute path of the code element and the relative path within the project;

[0080] Category: such as class, function, etc., indicates the element type;

[0081] line (line number): the specific line number of the element in the source file;

[0082] kind: such as def (definition), reference (ref), etc., indicating the role of the element in the code;

[0083] info (supplementary information): such as class name, function signature, etc.

[0084] (2) Node uniqueness and naming conventions: For nodes with kind=def, the element name is used directly as the node unique identifier. For nodes with kind=ref, to avoid confusion between references with the same name, a unique node name is generated using the format of ref:{name}:{rel_fname}:{line}.

[0085] (3) Node Addition: A directed multigraph (MultiDiGraph) is used as the underlying data structure. Each node and its attributes are added to the graph sequentially by traversing all labels. Node attributes are stored as key-value pairs, ensuring that each node carries complete contextual information.

[0086] The steps for edge construction are as follows:

[0087] (1) Edge semantics: Edges in the graph are used to express various relationships between codes, including:

[0088] Calls (function call relationship): indicates that a function or method calls another function or method;

[0089] Contains (contains relationship): indicates that a class contains methods, member variables and other structural hierarchical relationships;

[0090] Ref-def (reference-definition relationship): indicates the correspondence between a reference to a variable, function, class, etc. in a certain code and its definition.

[0091] (2) Edge adding logic:

[0092] Traverse all definition and reference nodes, analyze function bodies, class bodies and other structures, and identify call, include, reference and other relationships;

[0093] For each pair of nodes that have a relationship, add an edge using the function add_edge(source, target, type = relationship type) and label the edge with an attribute (type) to distinguish different semantic relationships.

[0094] Before adding an edge, a check is performed to see if the same edge already exists to avoid redundancy.

[0095] In the above steps, for text file composition, first, based on the splitter and corresponding chunk size mentioned in S2, a file is divided into multiple chunk parts, and the nodes are divided into text name nodes and several corresponding text chunk nodes. All chunk nodes are connected to the corresponding text name nodes to form edges.

[0096] S4. Hierarchical code knowledge graph construction: Community clustering is performed using a clustering algorithm, and then a community summary is formed through a large model (LLM):

[0097] S4.1. Use community detection algorithm to identify the close connection between entities in the constructed code knowledge graph, and perform preliminary clustering on the graph G = (V, E) based on indicators such as modularity to form a preliminary community. Each node i∈V is assigned a community label c i , where the calculation formula of module Q is:

[0098]

[0099] Among them, k i =∑ j∈v A ij is the degree of node i, is the total number of edges in the graph, δ(c i ,c j ) is the indicator function, when c i =c j Take 1 when it is, otherwise take 0;

[0100] S4.2, further use the hierarchical clustering method to refine the division and stratification of the preliminary communities, and divide the preliminary community set {C1, C2, ..., C k According to the inter-community distance index d(C i ,C j ) (for example, based on similarity between nodes or edge weight calculation) to gradually merge and layer until a predetermined clustering criterion is met, thereby generating a code knowledge graph with a hierarchical structure;

[0101] S4.3. Aggregate all text blocks and related information in each final community C to form a text set T C :T C ={t|t is a text block belonging to community C}; and call the large language model summary generation function f LLM (·) for T C Generate community summary S C :

[0102] S C =f LLM (T C )

[0103] At the same time, the aggregated features and summary information of each community are recorded for subsequent query optimization.

[0104] In this example, if Figure 2 As shown, the community clustering (Cluster) in S4.2 is a structure-aware graph clustering method, that is, after completing the construction of the basic code knowledge graph, community-level clustering operations will be performed based on the graph. Unlike traditional document clustering, the present invention innovatively introduces a structured clustering mechanism driven by code semantic relationships. By analyzing the calls, dependencies, definitions-references and other edges between the nodes (classes, functions, file blocks, etc.) in the graph, based on the modularity optimization algorithm, nodes with similar semantics and tight structures are automatically divided into multiple semantic communities (Code Community). Each community represents a group of functionally related or highly coupled code entities, which can effectively map the distribution of modules, subsystems or functional blocks in the real code base. The figure uses a multi-layer "Cluster & Summary" structure for visualization, indicating that a multi-granularity clustering strategy is adopted to achieve hierarchical abstraction from function level, class level to module level.

[0105] In this example, if Figure 2 As shown, community summary generation (Summary) in S4.2: combined with the semantic enhancement mechanism of the large model, that is, in order to make each community not only structurally aggregated but also semantically readable and interpretable, the present invention further introduces an LLM (large language model) driven automatic summary module after each clustering. This module extracts key code segments, comments and document descriptions from each community, combines their call chain context, and inputs them into a large language model (such as Codex, GPT-4) to automatically generate a human-readable functional summary. The summary content covers the functional role represented by the community, the modules or interfaces involved, and the description of usage. The generated summary is not only used as a community semantic label for subsequent search and question-and-answer, but can also be provided to developers as a document supplement to reduce the cost of understanding.

[0106] S5. Repository Question-Answering Agent (Repo-QA Agent) Search: Monte Carlo Tree Search (MCTS) is used to search in the hierarchical code knowledge graph. The relevance of nodes is evaluated through semantic matching. The Upper Confidence Bound (UCB) formula is used to balance the exploration of unknown paths and the use of known high-value paths. The path with a higher reward value (such as code correctness, efficiency, etc.) is selected to generate the final answer.

[0107] The Repo-QA Agent search process includes:

[0108] S5.1. Node selection (Select): From the code knowledge graph, the user's query is used as the root node, and the current node is selected according to the predefined selection strategy. The selection formula is:

[0109]

[0110] Where Q(s,a) represents the average reward for performing action a in state s, N(s) is the number of visits to state s, N(s,a) is the number of times action a is selected in state s, and c is the exploration constant.

[0111] S5.2, node expansion (Expand), based on the selected node, simultaneously sample several candidate actions, execute each candidate action and generate corresponding child nodes, and add the expanded nodes to the search tree;

[0112] S5.3, Simulate expansion: Randomly select a path from the currently expanded child nodes for simulation (rollout). Repeat the "expand + random selection" operation until a leaf node is reached or the preset maximum search depth is reached.

[0113] S5.4, Reward Evaluation: For each leaf node reached during the simulation (rollout), the quality of the inference trajectory corresponding to the node is scored using the pre-trained reward model R(·) and its reward value r is calculated, where: r = R(x leaf )

[0114] Here x leaf Indicates the sampling trajectory results corresponding to the leaf node;

[0115] S5.5, Backpropagation update, propagates the reward value r obtained by the leaf node upward along the path, and updates all the cumulative rewards on the path. The update formula is:

[0116]

[0117] Among them, r i is the reward value of the corresponding node in the i-th simulation (rollout);

[0118] S5.6. Repeat the simulation (rollout) process of steps S5.1 to S5.5 above until the reward values ​​of all nodes stabilize or reach the preset simulation (rollout) limit. Finally, based on the cumulative reward values ​​of each node, the trajectory with the highest cumulative reward is selected as the final question-answering result output.

[0119] In this embodiment, the code repository question-answering agent (Repo-QA Agent) plans the search path through the Monte Carlo Tree Search (MCTS) algorithm, which includes the action (Action) stage and the result (Result) stage, and finally enters the code and text recall processing stage.

[0120] (1) Action phase: After completing community selection and community expansion, the Repo-QA Agent initiates a specific action from the target node based on the semantic connection between nodes in the hierarchical code knowledge graph, simulating an access path in the graph. This path can span different community levels (e.g., from function to class, from class to module or file block), accessing potentially relevant information units layer by layer. The core of this phase is to guide the agent to select the path branch that is most likely to contain the answer based on the semantic clues of the current question, thereby performing efficient information sampling.

[0121] (2) Result phase: After the simulation path is completed, the Repo-QA Agent will evaluate the content of the nodes visited in the path and calculate their relevance scores to the question. The evaluation results will be passed back to each decision node on the original path as a feedback signal and used as a basis for strategy optimization. This process is called Check & Backpropagation, which can dynamically adjust the selection weights of each path in the search tree, thereby improving the accuracy of subsequent sampling. Finally, the Repo-QA Agent selects the path with the highest score as the source of the answer candidate set.

[0122] (3) Recall relative code & text: After obtaining the optimal path, the recall phase begins. This phase extracts the corresponding code entities (such as class definitions, function definitions, member variables, call relationships, etc.) and document content (such as configuration instructions, comment blocks, etc.) from the path nodes. Since the text files have been segmented into chunks in the preprocessing phase, the text paragraphs and code snippets related to the question can be quickly located based on the structure in the graph to achieve accurate recall. This recall module not only supports the extraction of structured code snippets, but also supports cross-file and cross-module content fusion to ensure the completeness of the context. The recall results will eventually be passed to the aggregation and answer generation module (Aggregation & GenerateAnswer), which performs semantic understanding, summary, and organization on all relevant content, and generates answers to developer-oriented questions in natural language.

[0123] The embodiments described above provide a detailed description of the technical solutions and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not intended to limit the present invention. Any modifications, supplements and equivalent substitutions made within the scope of the principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A code repository question-answering method based on a hierarchical code knowledge graph, characterized in that: The following steps are involved: S1. Collect all files in the target code repository and preprocess them to obtain code files and text files. S2. Generate an abstract syntax tree from the code file and extract relevant entities; Divide the text file into blocks and extract the text blocks; S3. Build a code knowledge graph representing the code structure based on the extracted entities and text blocks; S4. First, the constructed code knowledge graph is clustered using a community detection algorithm to form a preliminary community; Subsequently, a hierarchical clustering method is used to refine and stratify these preliminary communities, thereby generating a hierarchical code knowledge graph. Clusters at different levels within the hierarchical code knowledge graph constitute the final communities. Finally, the text blocks within each final community are processed through a large model to form the corresponding community summary. S5. Optimize the question-answering path using a code repository question-answering agent based on Monte Carlo tree search. This agent assists in node selection and expansion based on the community summaries contained in the community nodes in the hierarchical code knowledge graph, and evaluates rewards based on the simulated expansion results. Subsequently, the optimal query trajectory is dynamically planned through backpropagation, and the path with the highest cumulative reward is finally selected to generate the question-answering result.

2. The code warehouse question-answering method based on hierarchical code knowledge graph according to claim 1 is characterized in that: In step S1, the preprocessing includes file format recognition, classification and filtering to distinguish code files from text files.

3. The code warehouse question-answering method based on hierarchical code knowledge graph according to claim 1 is characterized in that: In step S2, the code file is generated into an abstract syntax tree in the following manner: First, exclude the standard library, built-in functions, and classes of the code, then exclude third-party libraries introduced in the code file, and finally build an abstract syntax tree based on predefined grammar rules.

4. The code warehouse question-answering method based on hierarchical code knowledge graph according to claim 1 is characterized in that: In step S2, the extracted related entities include classes, functions, methods and variables.

5. The code repository question-answering method based on hierarchical code knowledge graph according to claim 1 is characterized in that: In step S3, the code knowledge graph uses a bidirectional graph to represent the bidirectional relationship between entities, as follows: The nodes of the code file in the bidirectional graph include classes, functions, class member variables, and class member methods, and are divided into two types: definition def and reference ref; The edges of the code files in the bidirectional graph include the connection between classes and methods, the connection between classes and member variables, the connection between class definitions and references, and the connection between function definitions and references; The nodes of the text file in the bidirectional graph are the file name and the content after block segmentation; The edges of text files in the bidirectional graph are the connections between the text blocks corresponding to the file names.

6. The code repository question-answering method based on hierarchical code knowledge graph according to claim 1 is characterized in that: The specific process of step S4 is: S41. Use community detection algorithm to identify the connection between entities in the constructed code knowledge graph, and perform preliminary clustering on the code knowledge graph G = (V, E) based on the modularity index to form a preliminary community. Each node i∈V is assigned a community label c i , where the calculation formula of module Q is: Among them, k i =∑ j∈v A ij is the degree of node i, is the total number of edges in the graph, δ(c i ,c j ) is the indicator function, when c i =c j Take 1 when it is, otherwise take 0; S42, further use the hierarchical clustering method to refine the division and stratification of the preliminary communities, and divide the preliminary community set {C1, C2, ..., C k According to the inter-community distance index d(C i ,C j ) gradually merge and layer until the predetermined clustering criteria are met to generate a code knowledge graph with a hierarchical structure; S43. For each final community C, aggregate all text blocks and related information in the community to form a text set T C :T C ={t|t is a text block belonging to community C}, and calls the large language model summary generation function f LLM (·) for T C Generate community summary S C ; At the same time, the aggregated features and community summary information of each final community are recorded for subsequent query optimization.

7. The code warehouse question-answering method based on hierarchical code knowledge graph according to claim 6 is characterized in that: In step S42, the inter-community distance index d (C i ,C j ).

8. The code warehouse question-answering method based on hierarchical code knowledge graph according to claim 1 is characterized in that: The specific process of step S5 is: S51. Node selection: Based on the community summary contained in the community node in the hierarchical code knowledge graph, the user's query is used as the root node in the hierarchical code knowledge graph, and the current node is selected according to a predefined selection strategy. The selection formula is: Where Q(s,a) represents the average reward for performing action a in state s, N(s) is the number of visits to state s, N(s,a) is the number of times action a is selected in state s, and c is the exploration constant. S52, node expansion: Based on the selected node, several candidate actions are sampled at the same time, each candidate action is executed and a corresponding child node is generated, and the expanded node is added to the search tree; S53, simulate expansion: randomly select a path from the currently expanded child node for simulation, that is, repeat the expansion and random selection operations until a leaf node is reached or a preset maximum search depth is reached; S54, reward evaluation: For the leaf nodes reached during the simulation, the quality of the inference trajectory corresponding to the node is scored using the pre-trained reward model R(·) and its reward value r is calculated, where: r = R(x leaf ), x leaf Indicates the sampling trajectory results corresponding to the leaf node; S55, back propagation update: propagate the reward value r obtained by the leaf node upward along the path, and update all the cumulative rewards on the path; the update formula is: Among them, r i is the reward value of the corresponding node in the i-th simulation; S56. Repeat the simulation process of steps S51 to S55 above until the reward values ​​of all nodes tend to stabilize or reach the preset upper limit of the number of simulations; finally, based on the cumulative reward values ​​of each node, select the trajectory with the highest cumulative reward as the final question-and-answer result output.

Citation Information

Patent Citations

  • Question and answer processing method and device

    CN118733737A

Cited By

  • Code retrieval method and device and related equipment

    CN121277885A

  • Code abstract generation method and system based on hierarchical context awareness

    CN121957613A

  • Intelligent understanding methods, devices, equipment, and storage media for code repositories

    CN122569948A