Code library question and answer method and device based on large model

By segmenting the code repository into words, initializing directory depth, dividing it into segments, and converting it into semantic vectors, the problem of low accuracy and efficiency in code repository retrieval and question answering is solved, and efficient and accurate code interpretation generation and querying under a large language model is achieved.

CN121303374AActive Publication Date: 2026-01-09SHAANXI TUDOU DATA TECH CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202511850984.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-10
Publication Date
2026-01-09
Estimated Expiration
2045-12-10

AI Technical Summary

Technical Problem

Existing code repository retrieval and question answering methods suffer from low accuracy and efficiency in large and complex code repositories. Traditional methods lack contextual understanding and cross-file semantic relationship analysis, and large language models have different understandings of the code, which leads to the need for manual intervention when generating answers.

Method used

By segmenting code files into word units, initializing directory depth, generating a set of code files and inputting it into a large language model, natural language description is performed, segments are converted into semantic vectors and stored in a vector database, query vectors are used for matching and searching, and the results are integrated and fed back.

Benefits of technology

It improves the accuracy and efficiency of code retrieval, reduces manual intervention, and achieves efficient understanding and accurate answers to code files under a large language model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121303374A_ABST
    Figure CN121303374A_ABST
Patent Text Reader

Abstract

The invention discloses a code library question answering method and device based on a large model, and relates to the technical field of software engineering. The method comprises the steps of performing word segmentation on code files to obtain word units, performing statistics on the number of the word units in each code file, initializing the directory depth of the code files, further determining the hierarchical relationship of the code files to generate a code file set, and inputting the code file set into a large language model to generate a natural language description set; inputting the code file set and the natural language description set into a large language model to obtain code interpretation; segmenting the code explanation and converting the code explanation into semantic vectors; converting a query statement of a user into a query vector, and performing matching search on the query vector and the semantic vector to obtain a matching result; matching results are integrated and fed back to the user interface. The problem that an existing code library retrieval method is low in accuracy and efficiency is solved, a more intelligent code retrieval method is achieved, and the accuracy of large language model questions and answers is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of software engineering technology, and in particular to a question-answering method and apparatus for a codebase based on a large model. Background Technology

[0002] With the continuous development of artificial intelligence, more and more question-answering retrieval systems in various fields have incorporated large language models to achieve more efficient and intelligent retrieval, in order to meet the diverse needs of customers.

[0003] However, limitations still exist when performing retrieval and question answering in large and complex codebases. Traditional codebase retrieval often relies on methods such as keyword matching, pattern matching, and manual location. These methods lack contextual understanding and cross-file semantic relationship analysis, resulting in inaccurate retrieval results.

[0004] Furthermore, large language models have limited input / output lengths. Current technologies typically truncate or complete large codebases, leading to data loss or redundancy. Moreover, large language models differ in their understanding of natural language and code, making it difficult to accurately generate answers based solely on existing knowledge bases, requiring manual intervention. Therefore, existing codebase retrieval and question-answering methods cannot meet users' needs for efficient question answering within codebases. Summary of the Invention

[0005] This application provides a codebase question-answering method and apparatus based on a large model, which solves the problem of low accuracy and efficiency in existing codebase retrieval methods. It utilizes natural language processing methods to describe the function and technology stack of each code set, rationally segments the code content according to a segmentation strategy, and designs an intelligent agent to further retrieve questions not matched by the large language model, thereby improving the efficiency and accuracy of code retrieval.

[0006] In a first aspect, embodiments of this application provide a code library question answering method based on a large model, comprising: segmenting code files in the code library to obtain word units, and counting the number of word units in each code file; initializing the directory depth of the code files according to the number of word units, determining the hierarchical relationship of the code files according to the directory depth and generating a code file set, inputting the code file set into a large language model to generate a natural language description set; inputting the code file set and the natural language description set into the large language model to obtain code explanations; segmenting the code explanations and converting them into semantic vectors, storing them in a vector database; converting the user's query statement into a query vector, matching the query vector with the semantic vector in the vector database to obtain matching results; integrating the matching results and feeding them back to the user interface.

[0007] In one possible implementation, the step of segmenting code files in the code library to obtain word units includes: traversing the code files; performing word segmentation on the Chinese and English text in the code files respectively to obtain segmentation units, and assigning IDs to the segmentation units; establishing a vocabulary based on the IDs of the segmentation units, and obtaining the word units using a data compression algorithm; and obtaining the IDs of the word units based on the vocabulary and mapping rules.

[0008] In one possible implementation, initializing the directory depth of the code file based on the number of word units includes: calculating the directory depth of each file in the code file based on the file path, and arranging them to determine the processing order; wherein, if the directory depths are the same, the files are processed in ascending order based on the number of word units in each code file.

[0009] In one possible implementation, after initializing the directory depth of the code file according to the number of word units, the step includes: truncating the word units to a predefined maximum word unit limit according to the number of word units.

[0010] In one possible implementation, segmenting the code interpretation includes: setting a segment length; segmenting the code interpretation according to the segment length to obtain code segments, and storing the content of the current code segment of a preset length into the next code segment.

[0011] In one possible implementation, the conversion to semantic vectors and storage in a vector database includes: vectorizing code segments to obtain the semantic vectors; and creating an index based on the semantic vectors and storing it in the vector database.

[0012] In one possible implementation, the step of matching the query vector with the semantic vector in the vector database to obtain matching results includes: the matching results include direct matching results and interactive matching results; the direct matching results are obtained by directly deriving the code explanations of all query statements in the vector database based on the large language model; the interactive matching results are obtained by using an agent to search in the vector database and matching based on the vector similarity between the searched semantic vector and the query vector.

[0013] Secondly, embodiments of this application provide a code library question-answering device based on a large model, comprising: a data processing module, used to segment code files in the code library to obtain word units, and count the number of word units in each code file; initialize the directory depth of the code files according to the number of word units, determine the hierarchical relationship of the code files according to the directory depth and generate a code file set, input the code file set into a large language model to generate a natural language description set; input the code file set and the natural language description set into the large language model to obtain code explanations; a vectorization processing module, used to segment the code explanations and convert them into semantic vectors, and store them in a vector database; a matching module, used to convert the user's query statement into a query vector, and perform a matching search between the query vector and the semantic vector in the vector database to obtain matching results; and a feedback module, used to integrate the matching results, and A codebase question-answering device based on a large model, comprising: a data processing module, used to segment code files in the codebase to obtain word units and count the number of word units in each code file; initialize the directory depth of the code files according to the number of word units, determine the hierarchical relationship of the code files according to the directory depth and generate a set of code files, input the set of code files into a large language model to generate a set of natural language descriptions; input the set of code files and the set of natural language descriptions into the large language model to obtain code explanations; a vectorization processing module, used to segment the code explanations and convert them into semantic vectors, and store them in a vector database; a matching module, used to convert the user's query statement into a query vector, and perform a matching search between the query vector and the semantic vector in the vector database to obtain matching results; and a feedback module, used to integrate the matching results and feed them back to the user interface.

[0014] Thirdly, embodiments of this application provide an apparatus for a codebase question-answering method based on a large model, the apparatus comprising: a processor; a memory for storing processor-executable instructions; wherein, when the processor executes the executable instructions, it implements the method as described in the first aspect or any possible implementation of the first aspect.

[0015] Fourthly, embodiments of this application provide a non-volatile computer-readable storage medium, the non-volatile computer-readable storage medium including storage for storing a computer program or instructions that, when executed, cause the method described in the first aspect or any possible implementation of the first aspect to be implemented.

[0016] One or more technical solutions provided in the embodiments of this application have at least the following technical effects or advantages: This application's embodiments, by breaking down code files into smaller word units, help capture key elements in the code, thus more accurately reflecting its function and semantics. Initializing the directory depth of code files clearly demonstrates the logical connections and hierarchical structure between code segments. By inputting the code file set and the natural language description set into a large language model, the original information of the code and the auxiliary explanations of natural language are fully utilized, enabling the large language model to more comprehensively and accurately understand the function and intent of the code, thereby generating more accurate and detailed code explanations. Through a vector database, efficient queries are possible, leading to rapid return of matching results and improved user experience. This addresses the low accuracy and efficiency issues of existing code repository retrieval methods, thereby improving the large language model's understanding of code files without human intervention, minimizing conceptual comprehension errors in the code repository, and achieving accurate and efficient responses to user queries. Users can obtain the required code explanation information without long waiting times. Attached Figure Description

[0017] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 A flowchart illustrating a codebase question-answering method based on a large model, provided in this application embodiment; Figure 2 A flowchart illustrating the method for obtaining search matching results provided in this application embodiment; Figure 3 This is a schematic diagram of the structure of a codebase question-answering device based on a large model, provided in an embodiment of this application. Detailed Implementation

[0019] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0020] The following description of some technologies involved in the embodiments of this application is provided to aid understanding and should be considered merely exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this application. Similarly, for clarity and brevity, some descriptions of well-known functions and structures are omitted in the following description.

[0021] Figure 1 This is a flowchart of a codebase question-answering method based on a large model provided in an embodiment of this application, including steps 101 to 106. Figure 1 This is merely one execution order shown in the embodiments of this application and does not represent the only execution order for question-answering methods based on large model codebases. Where the final result can be achieved, Figure 1 The steps shown can be performed in parallel or in reverse order, as detailed below.

[0022] Step 101: Segment the code files in the code library to obtain word units, and count the number of word units in each code file. In this embodiment, the code files are traversed. The Chinese and English text in the code files are segmented to obtain segmentation units, and IDs are assigned to the segmentation units. A vocabulary is built based on the segmentation unit IDs, and word units are obtained using a data compression algorithm. The IDs of the word units are obtained based on the vocabulary and mapping rules.

[0023] Specifically, the word segmentation process uses regularization to uniformly process the English text, converting it to lowercase and standardizing spaces and punctuation marks. Based on the standardized spaces, text segmentation is performed to obtain word-level segmentation units.

[0024] Using Chinese word segmentation algorithms, for example, HanLP (Natural Language Processing Toolkit) is used to segment Chinese text to identify basic segmentation units. HanLP provides interfaces for multiple programming languages ​​such as Python and C++, making it widely applicable. For new words (words not in the vocabulary), an HMM (Hidden Markov Model) with the ability to form words from Chinese characters is used for segmentation. The obtained segmentation units are then mapped to pre-defined unique IDs according to mapping rules.

[0025] Based on the vocabulary, word units are mapped to unique corresponding IDs. A data compression algorithm is used to iteratively merge the segmented units until a preset merging limit (set to 10 times) is reached and / or no word units can be merged.

[0026] For example, the data compression algorithm uses BPE (Byte-Pair Encoding), which, in each iteration, identifies and merges the most frequent and adjacent segmentation unit pairs in the current word unit sequence, converting them into a new word unit. This transforms words or character sequences in the text into a more compact representation while retaining sufficient information for subsequent training and processing of large language models.

[0027] By utilizing data compression algorithms, a vocabulary is constructed by merging frequently occurring consecutive word unit sequences. Based on the mapping rules of the vocabulary, the corresponding IDs of the word units are obtained, which are then used to convert user-input text into numerical input that the large language model can understand. The number of word units is also counted for subsequent processing during the training of the large language model. This step enhances the large language model's ability to handle rare or foreign words, improving its ability to process complex text.

[0028] Step 102: Initialize the directory depth of the code files based on the number of word units, determine the hierarchical relationship of the code files based on the directory depth and generate a set of code files, and input the set of code files into the large language model to generate a set of natural language descriptions.

[0029] In this embodiment, the directory depth of each file in the code file is calculated based on the file path, and they are arranged to determine the processing order. If the directory depths are the same, the files are processed in ascending order based on the number of word units in each code file.

[0030] Specifically, code file processing is performed based on a count of the number of word units in each code file. The directory depth of each code file in the code repository is calculated according to the file path, and the files are sorted in descending order of directory depth to ensure that the deepest code files are processed first.

[0031] In addition, it is also possible to: truncate word units to a predefined maximum word unit limit based on the number of word units.

[0032] Specifically, if multiple code files with the same directory depth exist during the sorting process, they are sorted in ascending order based on the number of word units in each code file, resulting in a list of files at that directory depth. Here, directory depth can be understood as the number of directory separators appearing in the file path.

[0033] Specifically, the sorted list of files is traversed, and if the number of word units exceeds the predefined max_token (maximum word unit limit), they are truncated to ensure that each code file contains at most max_token word units.

[0034] The hierarchy of the parent file set to which each code file to be processed belongs is determined based on directory depth. The parent file set represents a set of files above the current code file in the file hierarchy with a smaller directory depth. If a parent file set that meets the conditions is found, and the sum of the word counts of the parent file set and the current code file (denoted as the subset file) does not exceed the limit of `max_token`, then the word counts of the subset file are merged and input into the parent file set. A code file set is generated according to the merging rules, which include ensuring that the code set obtained after each merging of the parent file set and the subset file adheres to the `max_token` condition, thus simplifying the data volume and providing a good foundation for subsequent data processing or storage.

[0035] For example, large language models such as ChatGPT and Qwen are used to intelligently summarize and analyze a collection of code files, generating a set of natural language descriptions. During the processing of the code file collection, all code files are traversed, and natural language processing methods are used to perform semantic analysis on the code within each file to obtain the core themes and functions of the code files. The large language model reads the contextual content of the code files based on the natural language description set to extract the code content. This ensures that the large language model can understand the key classes, functions, or related code components within the code files, as well as how the code files interact with each other. It identifies the specific content of the technology stack elements used in each code file, such as routing logic, business logic, and data access. Simultaneously, the filename clarifies the contextual relationships and purpose of the code files, and allows locating the current code file relative to its position within the entire codebase for subsequent path retrieval. The functionality implemented by the collection of code files is described using natural language, generating a set of natural language descriptions. This set of natural language descriptions includes how the specific functions of the code within the collection are implemented and its impact on the entire project or codebase.

[0036] Step 103: Input the code file set and the natural language description set into the large language model to obtain code interpretation.

[0037] Specifically, by using natural language descriptions of the collection, the large language model can better understand the specific function or role of each code file. Furthermore, it identifies the technology stack elements applied to each code file collection, such as routing logic, business logic, and data access.

[0038] The code file set and the natural language description set are input into a large language model to obtain the processed code interpretation. The code interpretation can fully express the function or implementation of each function or method. For example, the function of int_to_string is described as "converting int type data to string type", instead of using vague terms such as "this function".

[0039] Among these features, code interpretation enables large language models to efficiently and accurately understand the overall architecture of a project, thereby enabling better use and management of the codebase.

[0040] Step 104: Segment the code explanation and convert it into semantic vectors, then store them in a vector database. In this embodiment, a segment length is set. The code explanation is segmented according to the segment length to obtain code segments, and the content of the current code segment with a preset length is stored in the next code segment. The code segments are vectorized to obtain semantic vectors. An index is built based on the semantic vectors and stored in the vector database.

[0041] Specifically, the segmentation process includes setting the segment length to 200 characters. In this embodiment, each code segment contains 10% redundant content from the end of the previous code segment, which can effectively promote the understanding of context by the large language model without excessively increasing the data burden.

[0042] Based on BCEmbedding (a library of bilingual and cross-lingual semantic representation algorithm models), code segments are vectorized to obtain semantic vectors. By using an Embedding Model, the text content of the code interpretation can be converted into accurate, high-dimensional semantic vectors, as follows: .

[0043] In the formula, Indicates the current number The semantic vector of the text content of each code segment. This represents the encoding function of the Embedding Model. Indicates the first The text of a code segment.

[0044] A semantic index is created, and the vectorized semantic vectors are saved to the Milvus vector database (an open-source vector database) for faster subsequent search and retrieval tasks.

[0045] Specifically, the index includes a primary key index and a vector index. The primary key index is a unique ID identifier assigned to each semantic vector. The vector index is built using IVF_FLAT (a vector similarity algorithm based on divide and conquer), such as 768-dimensional semantic vector values ​​[0.1, 0.2, ..., 0.768].

[0046] The vector database also includes attributes of semantic vectors, such as file name, file path, and the original string (i.e., code segment text) corresponding to the semantic vector.

[0047] Step 105: Convert the user's query statement into a query vector, and match the query vector with the semantic vector in the vector database to obtain matching results. In this embodiment, the matching results include direct matching results and interactive matching results. Direct matching results are obtained by directly deriving the code interpretations of all query statements in the vector database based on the large language model. Interactive matching results are obtained by using an agent to search the vector database and matching the semantic vectors with the query vectors based on the vector similarity.

[0048] Specifically, if the large language model yields answers to all query statements, a direct matching result is obtained. If the large language model cannot yield a direct matching result based on the query statement, a matching search is performed in a vector database. This includes: converting the user's query statement into a query vector, calling an agent to perform a vectorized search in the vector database, matching based on a similarity calculation method, and outputting the agent's answer. Based on the agent's answer, the code file under the corresponding path of the semantic vector is determined, resulting in an interactive matching result.

[0049] The similarity calculation method is as follows: .

[0050] In the formula, Represents cosine similarity. The Euclidean norm of a vector. Represents the query vector. This represents the i-th semantic vector in the vector database.

[0051] The cosine similarity value ranges from -1 to 1. If two query vectors are identical to their semantic vectors, the cosine similarity is 1. If they are completely opposite, the cosine similarity is -1. If they are perpendicular, indicating no correlation, the cosine similarity is 0.

[0052] For example, when a user's query is sent to the large language model, the model processes it into a query vector, parses the question, and understands its meaning. For instance, if the user's query is: "How to convert a string to an integer in code?", and the large language model can find the answer directly, then "Answer" is used as a prefix instruction, and the direct matching result is output: "Answer: In a codebase, the int() function is typically used to convert a string to an integer."

[0053] like Figure 2 As shown, if the large language model cannot directly answer all the user's input questions and needs to obtain more information to form an answer, it will call the agent to perform a matching search in the code base file to obtain the interactive matching result, including: performing a matching search on the semantic vector until the interactive matching result is output. If the matching search reaches a preset number of loops (e.g., 10 times) and no corresponding result is found within the preset number of loops, it will return "No corresponding result found".

[0054] For example, a large language model might use "Action" as a prefix to output the function and parameters invoked by the agent. For instance, "Action:retrieve_code('How to convert a string to an integer in code?')". Upon receiving the instruction, the agent will perform a search in the codebase and return the most relevant semantic vectors, i.e., files and functions, based on similarity.

[0055] Specifically, it returns a file path and function definition corresponding to the semantic vector. The `get_code_by_path` method (a function method that returns the code path) can retrieve and return the contents of the corresponding file. Then, it reads the file content based on the absolute path. This allows users to directly obtain the contents of a single file without traversing an entire directory, such as " / path / to / code.py".

[0056] Furthermore, retrieving the entire vector database for each query could lead to a rapid increase in the number of word units. The `get_function_define` method (a function retrieval scheme) allows the agent to obtain the function definition based on the filename and function name, reducing additional computation and avoiding increased search time.

[0057] Step 106: Integrate the matching results and feed them back to the user interface. Specifically, after receiving the information returned by the agent, the large language model integrates it into a more specific and easier-to-understand answer, i.e., the interactive matching result, such as "Answer: In the file / path / to / code.py, an int() function is defined that converts a string to an integer."

[0058] While this application provides the method operation steps as described in the embodiments or flowcharts, more or fewer operation steps may be included based on conventional or non-inventive labor. The order of steps listed in this embodiment is merely one possible execution order among many and does not represent the only execution order. In actual device or client product execution, the methods shown in this embodiment or the accompanying drawings can be executed sequentially or in parallel (e.g., in a parallel processor or multi-threaded processing environment).

[0059] like Figure 3 As shown in the figure, this application embodiment also provides a code library question answering device 300 based on a large model. The device includes: a data processing module 301, a vectorization processing module 302, a matching module 303, and a feedback module 304, as detailed below.

[0060] The data processing module 301 is used to segment code files in the code library to obtain word units and count the number of word units in each code file. It initializes the directory depth of the code files based on the number of word units, determines the hierarchical relationship of the code files based on the directory depth, and generates a set of code files. This set of code files is then input into a large language model to generate a set of natural language descriptions. Finally, the set of code files and the set of natural language descriptions are input into the large language model to obtain code explanations.

[0061] The vectorization processing module 302 is used to interpret and segment the code and convert it into semantic vectors, which are then stored in the vector database.

[0062] The matching module 303 is used to convert the user's query statement into a query vector, and then match the query vector with the semantic vector in the vector database to obtain the matching result.

[0063] Feedback module 304 is used to integrate the matching results and provide feedback to the user interface.

[0064] Some modules in the apparatus described in this application can be described in the general context of computer-executable instructions that are executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, classes, etc., that perform a specific task or implement a specific abstract data type. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0065] The apparatus or module described in the above embodiments can be implemented by a computer chip or physical entity, or by a product with a certain function. For ease of description, the above apparatus is described by dividing it into various modules according to their functions. When implementing the embodiments of this application, the functions of each module can be implemented in one or more software and / or hardware. Of course, a module that implements a certain function can also be implemented by combining multiple sub-modules or sub-units.

[0066] The methods, apparatus, or modules described in this application can be implemented in a computer-readable program code manner. The controller can be implemented in any suitable manner, such as a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of a memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code manner, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included within it for implementing various functions can also be considered as structures within the hardware component. Alternatively, the device used to implement various functions can be viewed as either a software module that implements the method or a structure within a hardware component.

[0067] This application also provides an apparatus for executing a codebase question-answering method based on a large model. The apparatus includes: a processor; a memory for storing processor-executable instructions; and when the processor executes the executable instructions, it implements the method described in this application.

[0068] This application also provides a non-volatile computer-readable storage medium storing a computer program or instructions thereon, which, when executed, enables the method described in this application embodiment to be implemented.

[0069] Furthermore, in the various embodiments of this application, each functional module can be integrated into one processing module, or each module can exist independently, or two or more modules can be integrated into one module.

[0070] The aforementioned storage media include, but are not limited to, Random Access Memory (RAM), Read-Only Memory (ROM), Cache, Hard Disk Drive (HDD), or Memory Card. The memory can be used to store computer program instructions.

[0071] As can be seen from the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary hardware. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product, or it can be embodied in the process of data migration. The computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, mobile terminal, server, or network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of this application.

[0072] The various embodiments described in this specification are presented in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on its differences from other embodiments. All or part of this application can be used in numerous general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, mobile communication terminals, multiprocessor systems, microprocessor-based systems, programmable electronic devices, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices, etc.

[0073] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of this application.

Claims

1. A codebase question-answering method based on a large model, characterized in that, include: The code files in the code library are segmented into word units, and the number of word units in each code file is counted. The directory depth of the code file is initialized based on the number of word units, the hierarchical relationship of the code file is determined based on the directory depth, and a set of code files is generated. The set of code files is then input into a large language model to generate a set of natural language descriptions. The code file set and the natural language description set are input into the large language model to obtain code interpretation; The code is interpreted, segmented, and converted into semantic vectors, which are then stored in a vector database. The user's query statement is converted into a query vector, and the query vector is matched and searched with the semantic vector in the vector database to obtain the matching result; The matching results are integrated and then fed back to the user interface.

2. The method according to claim 1, characterized in that, The step of segmenting code files in the code library to obtain word units includes: Traverse the code file; The Chinese and English texts in the code file are segmented into word segments to obtain segmentation units, and IDs are assigned to the segmentation units. A vocabulary is built based on the ID of the segmentation unit, and the word unit is obtained using a data compression algorithm; The ID of the word unit is obtained based on the vocabulary and mapping rules.

3. The method according to claim 1, characterized in that, The process of initializing the directory depth of the code file based on the number of word units includes: The directory depth of each file in the code file is calculated based on the file path, and they are arranged to determine the processing order; wherein, if the directory depths are the same, they are processed in ascending order based on the number of word units in each code file.

4. The method according to claim 1, characterized in that, After initializing the directory depth of the code file according to the number of word units, the process includes: The word units are truncated to a predefined maximum word unit limit based on the number of word units.

5. The method according to claim 1, characterized in that, The segmentation of the code interpretation includes: Set the segment length; The code interpretation is segmented according to the segment length to obtain code segments, and the content of the current code segment with a preset length is stored in the next code segment.

6. The method according to claim 1, characterized in that, The conversion into semantic vectors, stored in a vector database, includes: The code segment is vectorized to obtain the semantic vector; An index is created based on the semantic vector and stored in a vector database.

7. The method according to claim 1, characterized in that, The step of matching the query vector with the semantic vector in the vector database to obtain the matching result includes: The matching results include direct matching results and interactive matching results; The direct matching result is derived directly from the large language model, showing the corresponding code explanations of all query statements in the vector database. The interactive matching result is obtained by using an agent to search the vector database and matching the semantic vectors found with the query vectors based on the vector similarity.

8. A codebase question-answering device based on a large model, characterized in that, include: The data processing module is used to segment code files in the code library to obtain word units and count the number of word units in each code file; The directory depth of the code file is initialized based on the number of word units, the hierarchical relationship of the code file is determined based on the directory depth, and a set of code files is generated. The set of code files is then input into a large language model to generate a set of natural language descriptions. The code file set and the natural language description set are input into the large language model to obtain code interpretation; A vectorization processing module is used to interpret and segment the code and convert it into semantic vectors, which are then stored in a vector database. The matching module is used to convert the user's query statement into a query vector, and then match the query vector with the semantic vector in the vector database to obtain the matching result. The feedback module is used to integrate the matching results and feed them back to the user interface.

9. A device for executing a codebase question-answering method based on a large model, characterized in that, include: processor; Memory used to store processor-executable instructions; When the processor executes the executable instructions, it implements the method as described in any one of claims 1 to 7.

10. A non-volatile computer-readable storage medium, characterized in that, Includes storage of computer programs or instructions that, when executed, cause the method as described in any one of claims 1 to 7 to be implemented.

Citation Information

Patent Citations

  • Method, device and equipment for improving large model code capability and storage medium

    CN118349715A

  • Code retrieval method and device based on large language model

    CN118643120A

  • Super-long text retrieval question and answer method, device and equipment based on large language model and medium

    CN118820424A

  • Database reasoning method and device based on large model

    CN118861084A

  • Code storage method and device, code retrieval method and device and electronic equipment

    CN119201195A