Corpus retrieval method and device, cluster, storage medium and program product
By using file path information to search target corpus in the corpus, the problem of degradation of search accuracy in the prior art is solved, and the reasoning effect of LLM is significantly improved.
Patent Information
- Application Number
- CN202510040819.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-07
- Publication Date
- 2025-06-06
AI Technical Summary
With the update and enrichment of corpus, existing search methods are difficult to distinguish between high-similar corpus, resulting in a decrease in retrieval accuracy, which in turn affects the inference effect of the large language model LLM.
By using the path information of the file in the database to determine the target corpus in the corpus, the computing device obtains the similarity between the path information and the path information in the corpus, and determines the target corpus corresponding to the path information that meets the threshold range, thereby improving the search accuracy.
It significantly improves the search accuracy of computing devices in the corpus, ensures that the target corpus has a higher correlation with the user input propt, thereby improving the reasoning effect of LLM.
Smart Images

Figure CN120104726A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a corpus retrieval method, device, cluster, storage medium and program product. Background Art
[0002] At present, in order to improve the reasoning effect of large language models (LLM) and reduce the hallucination of LLM, more and more manufacturers choose to use retrieval-augmented generation (RAG) technology.
[0003] LLM usually performs reasoning based on the prompt word entered by the user, and the purpose of RAG technology is to retrieve corpus related to prompt from the corpus, providing LLM with richer and more accurate technical background, professional knowledge, etc., so as to improve the reasoning effect of LLM. In other words, the higher the correlation between the corpus retrieved by RAG and the prompt entered by the user, the more accurate the reasoning effect of LLM will be. The retrieval technology involved is a key link in determining whether the reasoning effect of LLM can be improved in the end.
[0004] In the scenario of LLM-assisted code generation, the current retrieval method used by RAG is usually to directly retrieve similar corpora from the corpus through the prompt entered by the user. However, as the corpus is continuously updated and enriched, the number of corpora in the corpus becomes larger and larger, and it is difficult to avoid the high similarity between different corpora, resulting in the retrieval accuracy becoming worse and worse, which further leads to the deterioration of LLM's reasoning effect. Summary of the invention
[0005] The present application provides a corpus retrieval method, apparatus, cluster, storage medium and program product. When a computing device retrieves a target corpus in a corpus, it makes full use of the path information of the file located in the database to determine the target corpus in the corpus, so that the correlation between the target corpus and the business functions of the code that the user needs to generate through the large language model LLM is higher, that is, the accuracy of the computing device in corpus retrieval is significantly improved, which helps to improve the reasoning effect of the LLM.
[0006] In a first aspect, the present application provides a corpus retrieval method, the method comprising: a computing device obtains first path information of a first file located in a first database, the first file including a target text, the target text being used to characterize the business functions of the code that the user needs to generate through a large language model (LLM); the computing device obtains a first similarity between the first path information and each path information in the corpus; the computing device determines a target corpus corresponding to second path information whose first similarity meets a first threshold range, the target corpus including text whose similarity to the target text meets a second threshold range, and the target corpus is used to assist the LLM in generating code with business functions.
[0007] It can be understood that there is a high degree of correlation between the files and corpora corresponding to the same or similar path information, especially from the perspective of business functions. The computing device can significantly improve the accuracy of corpus retrieval by retrieving the target corpus that meets the similarity requirements in the corpus through path information. Using target corpora with higher relevance to assist LLM can improve the accuracy of LLM in generating codes with business functions and improve the reasoning effect of LLM.
[0008] In one possible implementation, a computing device determines a target corpus corresponding to second path information whose first similarity meets a first threshold range, including: if the second path information does not exist, the computing device obtains third path information, wherein the directory corresponding to the third path information is the parent directory of the directory corresponding to the first path information; the computing device obtains the second similarity between the third path information and each path information in the corpus; the computing device determines a target corpus corresponding to fourth path information whose second similarity meets a second threshold range.
[0009] It can be understood that when there is no path information in the corpus that is highly similar to the first path information, the computing device searches again through the third path information corresponding to the parent directory of the directory corresponding to the first path information. By gradually expanding the search scope, it can still ensure that the final target corpus and the first file meet the corresponding similarity requirements in path information, thereby improving the accuracy of corpus retrieval.
[0010] In one possible implementation, obtaining a first similarity between first path information and each path information in a corpus includes: a computing device obtaining a first vector corresponding to the first path information; a computing device obtaining a first vector distance between the first vector and vectors corresponding to each path information in the corpus; and a computing device determining a first similarity based on the first vector distance.
[0011] It is understandable that the computing device can accurately determine the similarity between different path information by obtaining the first vector distance, such as cosine distance, between the first vector corresponding to the first path information and the vectors corresponding to each path information in the corpus.
[0012] In one possible implementation, obtaining the second similarity between the third path information and each path information in the corpus includes: a computing device obtaining a second vector corresponding to the third path information; a computing device obtaining a second vector distance between the second vector and a vector corresponding to each path information in the corpus; and a computing device determining the second similarity based on the second vector distance.
[0013] It is understandable that the similarity between path information at different levels can also be measured by calculating the vector distance. In this way, when the computing device gradually expands the search scope, the accuracy of determining the similarity between different path information can be guaranteed.
[0014] In one possible implementation, a computing device obtains a first similarity between the first path information and each path information in a corpus, including: the computing device obtains a first degree of overlap between a character string corresponding to the first path information and a character string of each path information in the corpus; and the computing device determines a first similarity based on the first degree of overlap.
[0015] It is understandable that the computing device can accurately determine the similarity between different path information by calculating the degree of overlap between character strings of different path information, for example, by matching characters through regular expressions, thereby improving the accuracy of retrieval.
[0016] In one possible implementation, the computing device obtains the second similarity between the third path information and each path information in the corpus, including: the computing device obtains the second degree of overlap between the character string corresponding to the third path information and the character string of each path information in the corpus; the computing device determines the second similarity based on the second degree of overlap.
[0017] It is understandable that when the computing device gradually expands the search scope, the similarity between different path information can be accurately determined by calculating the degree of overlap between character strings of different path information, such as matching characters through regular expressions, thereby improving the accuracy of the search.
[0018] In one possible implementation, a computing device determines a target corpus corresponding to second path information whose first similarity meets a first threshold range, including: the computing device obtains a third vector corresponding to the target text; the computing device obtains a third vector distance between the third vector and a vector corresponding to a character string of a target field in the corpus corresponding to the second path information; the character string of the target field is used to characterize the business function of the code in the corresponding corpus; the computing device determines a third similarity based on the third vector distance; the computing device determines that the corpus corresponding to the second path information whose third similarity meets a third threshold range and whose first similarity meets the first threshold range is the target corpus.
[0019] It can be understood that the target text in the first file is used to represent the business functions of the code that the user needs to generate through the large language model LLM, and the character string of the target field is used to represent the business functions of the code in the corresponding corpus. Therefore, by calculating the vector distance to determine the similarity between the target text and the character string of the target field, it can be ensured that the code in the target corpus has a high correlation with the target text in terms of business functions, thereby improving the accuracy of retrieval.
[0020] In one possible implementation, a computing device determines a target corpus corresponding to second path information whose first similarity meets a first threshold range, including: the computing device obtains a degree of overlap between a character string of a target field and a character string of a target text in the corpus corresponding to the second path information, the character string of the target field representing a business function of a code in the same corpus; the computing device determines a fourth similarity based on the degree of overlap; the computing device determines that the corpus corresponding to the second path information whose fourth similarity meets a fourth threshold range and whose first similarity meets the first threshold range is the target corpus.
[0021] It is understandable that when determining the similarity between the character string of the target field and the character string of the target text, the computing device may also improve the accuracy of the retrieval by calculating the degree of overlap between the character strings, for example, by substring matching.
[0022] In one possible implementation, a computing device determines a target corpus corresponding to second path information whose first similarity meets a first threshold range, including: the computing device obtains a sorting result of the corpus corresponding to the second path information; the sorting result is determined based on the degree of correlation between the corpus corresponding to the second path information and domain knowledge related to the target business function; the computing device determines the target corpus based on the sorting result.
[0023] It is understandable that sorting the corpus corresponding to the second path information by domain knowledge related to the target business function can ensure that the target corpus and the target text have a higher relevance in business function through more professional domain knowledge, thereby further improving the accuracy of retrieval.
[0024] In a second aspect, the present application provides a corpus retrieval device, which is used to execute any one of the corpus retrieval methods provided in the first aspect.
[0025] In a possible implementation, the present application may divide the corpus retrieval device into functional modules according to the method provided in the first aspect above. For example, each functional module may be divided according to each function, or two or more functions may be integrated into one module. Exemplarily, the present application may divide the corpus retrieval device into a first acquisition module, a second acquisition module, and a determination module, etc. according to the function. The description of the possible technical solutions and beneficial effects executed by each of the functional modules divided above can refer to the technical solutions provided by the first aspect above or its corresponding possible implementation, and will not be repeated here.
[0026] In a third aspect, an embodiment of the present application provides a computing device, which includes a processor and a memory, wherein the processor is coupled to the memory; the memory is used to store computer instructions, which are loaded and executed by the processor so that the computing device implements the corpus retrieval method described in the above aspects.
[0027] In a fourth aspect, an embodiment of the present application provides a computing device cluster, which includes at least one computing device, each computing device includes a processor and a memory, and the processor is coupled to the memory; the processor of at least one computing device is used to execute computer instructions stored in the memory of at least one computing device, so that the computing device cluster executes the corpus retrieval method provided in various optional implementations of the above-mentioned first aspect.
[0028] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium, in which at least one computer program instruction is stored, and the computer program instruction is loaded and executed by a processor to implement the corpus retrieval method as described in the above aspects.
[0029] In a sixth aspect, an embodiment of the present application provides a computer program product, the computer program product including computer instructions, the computer instructions stored in a computer-readable storage medium. A processor of a computing device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computing device executes the corpus retrieval method provided in various optional implementations of the first aspect above.
[0030] For the specific description of the second to sixth aspects and their various implementations in the present application, reference may be made to the detailed description in the first aspect and its various implementations; and for the beneficial effects of the second to sixth aspects and their various implementations, reference may be made to the beneficial effects analysis in the first aspect and its various implementations, which will not be repeated here.
[0031] These and other aspects of the present application will become more apparent from the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 A schematic diagram of a RAG application provided in an embodiment of the present application;
[0033] Figure 2 A schematic diagram of a RAG system provided in an embodiment of the present application;
[0034] Figure 3 A schematic diagram of the structure of a computing device provided in an embodiment of the present application;
[0035] Figure 4 A flowchart of corpus extraction provided in an embodiment of the present application;
[0036] Figure 5 for Figure 4 A schematic diagram of a corpus processing method involved in the illustrated embodiment;
[0037] Figure 6 A flowchart of a corpus retrieval method provided in an embodiment of the present application;
[0038] Figure 7 for Figure 6 A schematic diagram of triggering reasoning involved in the illustrated embodiment;
[0039] Figure 8 for Figure 6 A schematic diagram of a reasoning result involved in the illustrated embodiment;
[0040] Fig. 9 A schematic diagram of the structure of a corpus retrieval device provided in an embodiment of the present application;
[0041] Fig.10 A schematic diagram of a computing device provided in an embodiment of the present application;
[0042] Fig.11 A schematic diagram of a computing device cluster provided in an embodiment of the present application;
[0043] Fig.12 A schematic diagram of a connection method between computing device clusters provided in an embodiment of the present application. DETAILED DESCRIPTION
[0044] In order to make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the implementation methods of the present application will be further described in detail below in conjunction with the accompanying drawings.
[0045] The term "multiple" as used herein refers to two or more than two. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the related objects are in an "or" relationship.
[0046] Furthermore, in the description of the embodiments of the present application, unless otherwise specified, "multiple" refers to two or more than two. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.
[0047] In addition, in order to facilitate the clear description of the technical solutions of the embodiments of the present application, in the embodiments of the present application, the words "first", "second" and the like are used to distinguish the same items or similar items with substantially the same functions and effects. Those skilled in the art will understand that the words "first", "second" and the like do not limit the quantity and execution order, and the words "first", "second" and the like do not necessarily limit the differences. At the same time, in the embodiments of the present application, the words "exemplary" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or design solutions. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a concrete manner for ease of understanding.
[0048] First, the application scenarios of the embodiments of the present application are exemplarily introduced.
[0049] Currently, when developers write programs, for example, using programming languages such as Python and Java to write code to test the performance of a certain interface or device, they can use some automation tools to improve work efficiency, such as summarizing some commonly used test cases and forming corresponding code templates. However, these methods are usually not flexible enough and have great limitations.
[0050] With the emergence of large language models (LLMs), this problem is expected to be solved. Users can enter a prompt, where the prompt describes the business function that the user hopes the code generated by the LLM to have, for example, "Please write a function in Python to calculate the sum of two numbers." In some simple scenarios, the efficiency of code generation can be improved. However, with the development of technology, problems such as LLM's hallucination problems, insufficient reasoning effect, and difficulty in timely self-updating have gradually emerged. The code generated by LLM is difficult to use directly, but requires detailed manual review.
[0051] In order to improve the reasoning effect of LLM and provide users with more accurate results, the industry has proposed retrieval-augmented generation (RAG) technology. Figure 1 As shown, Figure 1 A RAG application diagram provided in an embodiment of the present application, according to the prompt word prompt input by the user, from Figure 1 More relevant corpora (such as professional knowledge, technical background, etc.) are retrieved from the corpus shown, and then these professional knowledge corpora and prompts are input into the LLM to assist the LLM in reasoning, and the reasoning results output by the LLM are obtained, thereby reducing or even completely solving the hallucination problem of the LLM, and greatly improving the reasoning effect of the LLM, so that the LLM can keep up with the rapid development of technology without the need to train the LLM again.
[0052] Among them, the more accurate the corpus retrieved in RAG is, the better the reasoning effect of LLM is usually. At present, the commonly used retrieval method is to perform a global search in the corpus for the prompt input by the user, and the similarity of the retrieval results is calculated using the best matching 25 (BM25) algorithm, edit distance and other algorithms to score, and finally the corpus with the required score is used as the retrieval result to improve the reasoning effect of LLM. However, when the user uses LLM to generate code, the prompt input by the user is usually only a fragmentary description of a certain function, for example, "Please use Python to write a function to test the forwarding efficiency of interface A". In this case, the above solution may cause the retrieval results to be out of business during retrieval, because it is difficult to know whether "interface A" is the interface of a router device or the interface of a smart phone, or the interface of any device. This reduction in retrieval accuracy will be further amplified by a corpus with richer content, because the corpus contains a large amount of corpus related to "testing the forwarding efficiency of an interface", which leads to a significant reduction in retrieval accuracy, and then leads to poor reasoning effect of LLM.
[0053] According to research, developers usually follow many naming conventions when writing code, such as variable naming conventions, function naming conventions, and directory naming conventions, etc. Therefore, code engineering files usually have a relatively clear directory hierarchy. For example, the path information of a file is: cpu-memory / test-task / io-test / main.py, where the names of the directories involved in the path information usually reflect the business functions of the code therein. Taking the aforementioned path information as an example, without knowing the specific code, developers can infer that the code named "main.py" can be used to test the input / output (IO) performance between the central processing unit (CPU) and the memory (memory) based on the path information. In other words, there is a correlation between the path information of the file and its business function.
[0054] Based on the above research, the embodiment of the present application proposes a corpus retrieval method, which makes full use of the path information of the file in the database. By comparing the similarities between different path information, the accuracy of the retrieval can be further improved, thereby improving the reasoning effect of LLM.
[0055] In some feasible embodiments, the method includes: a computing device obtains first path information of a first file located in a first database, the first file includes a target text, and the target text is used to represent the business function of the code that the user needs to generate through the large language model LLM; the computing device obtains the first similarity between the first path information and each path information in the corpus; the computing device determines the target corpus corresponding to the second path information whose first similarity meets the first threshold range, the target corpus includes text whose similarity with the target text meets the second threshold range, and the target corpus is used to assist LLM in generating code with business functions. Since code files under the same path often have a high business correlation, when the computing device retrieves the target corpus in the corpus through the first path information, it can significantly improve the accuracy of the retrieval and help improve the reasoning effect of LLM.
[0056] Secondly, the system architecture of the embodiment of the present application is exemplarily introduced.
[0057] Figure 2 A schematic diagram of a RAG system provided in an embodiment of the present application, Figure 2 The RAG system 1000 shown can be applied to Figure 1 In the illustrated scenario, the RAG system 1000 includes: a corpus processing module 1100 , a corpus 1200 , a corpus retrieval module 1300 and an inference module 1400 .
[0058] Specifically, Figure 2 The corpus processing module 1100 shown can be used to process the corpus to be stored, such as performing key asset KIA inspection, corpus rewriting, corpus increment and other processing, and can also perform quality inspection on the processed corpus and eliminate unqualified low-quality corpus, so as to improve the quality of the stored corpus. The corpus processing module 1100 stores the processed corpus in the corpus 1200; the corpus retrieval module 1300 can be used to retrieve relevant corpus from the corpus 1200 according to the prompt words input by the user when the reasoning is triggered, and determine the final target corpus, etc.; the reasoning module 1400 can be used to generate reasoning results using the target corpus, prompt words and trained LLM, and can further verify the reasoning results, etc.
[0059] In the embodiments of the present application, Figure 2 The corpus retrieval module 1300 shown can also be used to obtain path information of a file in a database, and use the path information, prompt words input by a user and other information to accurately retrieve the target corpus.
[0060] Figure 3 A schematic diagram of a computing device provided in an embodiment of the present application. One or more components of the aforementioned RAG system 1000 can be operated on Figure 3 On the computing device shown, Figure 3 The computing device 3000 shown in the figure includes at least: a memory 3010, a processor 3020 and a bus 3030. The computing device 3000 can at least be used to run Figure 2 The corpus retrieval module 1300 is shown.
[0061] The processor 3020 can be used to obtain the path information of the file in the database, obtain the similarity between each path information in the corpus and the aforementioned path information, and determine the target corpus corresponding to the path information whose similarity meets the corresponding threshold range, etc. The memory 3010 can be used to store the logic code corresponding to the corpus retrieval method provided in the embodiment of the present application.
[0062] Obtain first path information of a first file located in a first database, wherein the first file includes a target text, and the target text is used to characterize the business function of the code that the user needs to generate through a large language model (LLM); a computing device obtains a first similarity between the first path information and each path information in a corpus; the computing device determines a target corpus corresponding to the second path information whose first similarity meets a first threshold range, wherein the target corpus includes text whose similarity to the target text meets a second threshold range, and the target corpus is used to assist the LLM in generating code with business functions.
[0063] Optionally, the computing device 3000 can be various types of server devices such as a rack server and a whole cabinet server. The computing device 3000 can also be a terminal computing device such as a computer, a mobile terminal, a tablet computer, a laptop computer, a desktop computer, an all-in-one computer, a personal digital assistant (PDA), an ultra-mobile personal computer (UMPC), etc.
[0064] Optionally, the memory 3010 may include a random access memory (RAM), a read-only memory (ROM), etc., wherein the memory 3010 can run a necessary operating system in its RAM, as well as a first acquisition module, a second acquisition module, a determination module, and other modules for executing the corpus retrieval method provided in the present application.
[0065] Optionally, the processor 3020 may be a central processing unit (CPU), a graphics processing unit (GPU) or other general-purpose processors, an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processor (DSP) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or any conventional processor, etc.
[0066] Optionally, bus 3030 may be a peripheral component interconnect (PCI) bus, etc., and the present application does not limit the type of bus. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 3 The bus 3030 is represented by only one line, but it does not mean that there is only one bus or one type of bus. The bus 3030 may include a path for transmitting information between various components of the computing device 3000 (for example, the memory 3010 and the processor 3020).
[0067] It should be noted that the system architecture and application scenarios described in the embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided in the embodiments of the present application. Ordinary technicians in this field can know that with the evolution of the system architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.
[0068] In the embodiment of the present application, in order to improve the accuracy of the computing device in retrieving the target corpus from the corpus, the corpus construction process is also improved accordingly. The corpus construction process is explained and illustrated in conjunction with the accompanying drawings. Figure 4 As shown, Figure 4 A flowchart of corpus extraction provided in an embodiment of the present application specifically includes:
[0069] S110, the computing device obtains a code file.
[0070] In this step, the computing device can obtain code files from public code repositories and other platforms, such as some open source code hosting platforms, which host a large number of code resources. The computing device can retrieve and download corresponding code files from these platforms according to the set keywords. When the computing device obtains the code file, it can obtain the project file where the code file is located, so as to maximize the retention of various directories and the code files under each directory.
[0071] Exemplarily, the computing device obtains corresponding code files from the code repository according to the set keywords "test project", "automation script", and "test script", and obtains the project files of each code file. Taking the path information: cpu-memory / test-task / io-test / main.py as an example, the computing device can obtain all subdirectories and files in the directory: cpu-memory.
[0072] It is worth noting that code files under the same path often have a high business relevance. For example, the code files under the path information: cpu-memory / test-task / io-test are likely to be related to the IO performance test between the CPU and the memory. Another example is the code files under the path information: cpu-memory / test-task are likely to be related to the test between the CPU and the memory. Furthermore, the code files corresponding to similar paths often have a high business relevance. For example, the code files under the path information: cpu-memory / auto-test / io-test and the code files under the aforementioned path information: cpu-memory / test-task / io-test are likely to have a high business relevance, and are both related to the IO performance test between the CPU and the memory.
[0073] S120, the computing device obtains path information of the code file.
[0074] In this step, the computing device can specifically obtain the relative path of the code file in the code repository through instructions related to obtaining path information in the code version management tool (git tool). Developers can also manage the code by using this tool to implement code version management and other operations.
[0075] Exemplarily, the relative path of the code file obtained by the computing device through the git tool and located in the code repository is: cpu-memory / test-task / io-test / main.py.
[0076] S130, the computing device processes the code file to obtain the corpus to be stored.
[0077] In this step, the computing device can construct the code file into a standardized code corpus. For example, the computing device can extract the content through a set regular expression, extract the target content from the code file and fill it into a preset corpus template. The computing device can also provide a regular extractor that the user can configure by himself, or provide a unified standard interface to support pluggable custom extractors, etc. Furthermore, the computing device can also remove meaningless characters in the target content, such as: digital numbers, # and other characters.
[0078] For example, see Figure 5 , Figure 5 for Figure 4A schematic diagram of a corpus processing involved in the illustrated embodiment, with two code files obtained by a computing device on the left, wherein, based on information such as the file name suffix, the computing device can determine that the programming language used in code file 1 is python, and the programming language used in code file 2 is java, and extract the target content from them through a regular extractor and fill it into a preset corpus template. Figure 5 The corpus template involved may include the following four parts: path information (path), test step description (test_step), code (code) and programming language (language). After the computing device constructs the corpus slice, it obtains Figure 5 The corpus slice 1 and corpus slice 2 shown, specifically, the path information in corpus slice 1 is: "autotest / script" "path": "autotest / script", the test step description is: "1. Add phone 1 (10001)", the code is: "ret_code = add_phone (1, 10001)", and the programming language is "language": "python". The path information in corpus slice 2 is: "autotest / script" "path": "autotest / script", the test step description is: "1. Add phone 1 (10001)", the code is: "this.addPhone (1, 10001)", and the programming language is "language": "java". Through further inspection and processing, the computing device removes the numerical numbers and punctuation in the test step description fields in corpus slice 1 and corpus slice 2, and finally obtains Figure 5 The two corpora to be stored (corpus 1 to be stored and corpus 2 to be stored) are shown on the right side, wherein the specific content corresponding to the test step description is a description of the business function of the specific content corresponding to the code in the corpus.
[0079] S140, the computing device performs a quality check on the corpus to be stored.
[0080] In this step, the computing device can perform quality checks through built-in multiple corpus checking specifications, or the computing device can also support user-defined corpus checking specifications, and then filter low-quality corpora according to the specifications set by the user. For example, check whether there are code errors in the corpus to be stored, whether there are problems such as incoherence and unclear semantics in the test step description, or check whether there are sensitive characters set by the user in the corpus to be stored, etc. Through this step, the computing device can improve the quality of the corpus, and a high-quality corpus helps to improve the reasoning effect of LLM.
[0081] S150, the computing device stores the corpus that meets the requirements in the corpus.
[0082] In this step, the computing device can obtain the corpus to be stored that meets the requirements through the aforementioned steps. When storing these corpora to be stored in the corpus, in order to improve the efficiency and accuracy of subsequent retrieval, the computing device can index and vectorize the key information such as path information and test step description in the corpus to be stored. That is to say, the computing device can obtain the corresponding vector of one or more of the key information such as path information and test step description of a corpus to be stored by inputting it into an embedding model or other means, and store it in the corpus as the index value of this corpus to facilitate subsequent RAG reasoning.
[0083] Through the above steps S110-S150, the computing device can build a high-quality corpus, and each corpus has a vector corresponding to its own path information, test step description and other key information as an index, which helps to improve the efficiency and accuracy of subsequent retrieval.
[0084] Based on the corpus built above, the computing device can implement the corpus retrieval method provided by this application. It should be noted here that the computing device for building the corpus and the computing device for executing the corpus retrieval method provided by this application may be the same computing device or may not be the same computing device, and this application does not impose any restrictions on this.
[0085] For ease of understanding, the following is an exemplary introduction to the corpus retrieval method provided in the embodiment of the present application in conjunction with the accompanying drawings. The corpus retrieval method is applicable to Figure 3 The computing device shown. Figure 6 As shown, Figure 6 A flowchart of a corpus retrieval method provided in an embodiment of the present application, the method specifically comprises the following steps:
[0086] S210: The computing device obtains first path information of a first file located in a first database.
[0087] The first file includes a target text, and the target text is used to represent the business functions of the code that the user needs to generate through the large language model LLM.
[0088] In the embodiment of the present application, the first database includes but is not limited to a code hosting platform (or code warehouse).
[0089] In a possible implementation, in addition to obtaining the first path information, the computing device may also obtain key information such as the target text in the first file and the identifier of the first database.
[0090] Furthermore, in some feasible embodiments, the computing device may also determine the first path information of the first file located in the first database based on the obtained script path (script path) of the first file located in the first database and the working path (working path).
[0091] For example, Figure 7 As shown, Figure 7 for Figure 6 A schematic diagram of triggering reasoning involved in the illustrated embodiment, wherein the user enters the following prompt words in the code file 1: "#1. Add phone 1 (10001), phone 2 (20002), phone 3 (30003)", which is also the target text that represents the business function of the code that the user needs to generate through the large language model LLM. The user can Figure 7 The "code generation" control shown or other shortcut keys are used to send inference trigger instructions to the computing device.
[0092] Furthermore, after obtaining the aforementioned trigger instruction, the computing device obtains the script path of the code file 1 as D:\document\Python program\phone\test_task\call_test\main.py, and the working path is D:\document\Python program. The computing device determines that the first path information of the file is: phone\test_task\call_test\main.py, or the file is a file hosted in the code repository, and the computing device obtains the first path information of the file through the git tool as: phone\test_task\call_test\main.py, and the computing device also obtains the target text of the code file 1. The computing device can also determine that the programming language used by the file is python based on the name of the file main.py. This information will be used for subsequent retrieval.
[0093] S220: The computing device obtains a first similarity between the first path information and each piece of path information in the corpus.
[0094] When the computing device executes step S220, at least the following two possible implementations may be adopted:
[0095] 1) The computing device obtains a first vector corresponding to the first path information, and then obtains a first vector distance between the first vector and vectors corresponding to each path information in the corpus, and the computing device determines a first similarity based on the first vector distance.
[0096] Specifically, when calculating the distance between two vectors, the computing device can calculate the cosine distance (or cosine similarity), Manhattan distance, Euclidean distance or Hamming distance, etc. Generally speaking, the closer the distance between the two vectors, the higher the similarity between the two path information. Therefore, the computing device can determine the similarity between the path information based on the vector distance.
[0097] 2) The computing device obtains a first degree of overlap between the character string corresponding to the first path information and the character string of each path information in the corpus; the computing device determines a first similarity based on the first degree of overlap.
[0098] Specifically, the computing device can match the character string corresponding to the first path information with the character string of each path information in the corpus according to methods such as regular matching. During the specific matching, a certain degree of fuzzy matching can be allowed (for example, not distinguishing between uppercase and lowercase letters, etc.), or strict matching can be performed. These can be determined based on actual needs, and this application does not impose any restrictions on this.
[0099] In some feasible embodiments, taking into account the ambiguity of measuring similarity based on vector distance, the computing device may also combine the above two possible implementation methods. For example, the computing device first obtains the degree of overlap between the character string of the first path information and the character string of each path information in the corpus through a regular matching method, and searches the corpus for path information that is highly consistent with or even completely consistent with the first path information. If not, the computing device obtains the first similarity by calculating the vector distance, and then searches the corpus for the path information.
[0100] Exemplarily, the first path information is: phone\test_task\call_test\main.py. The computing device can first use the regular matching method to search whether there is completely consistent path information in the corpus. If so, the corresponding first similarity is 100%; if not, the first similarity can be determined according to the degree of character matching, or the first similarity can be determined by calculating the vector distance. The computing device converts the calculated vector distance to the interval [0,1]. If the vector distance is 0, the corresponding first similarity is 100%; if the vector distance is 1, the corresponding first similarity is 0%; if the vector distance is 0.5, the corresponding first similarity is 50%. Similarly, the similarities corresponding to other distances in the interval are evenly distributed.
[0101] S230: The computing device determines the target corpus corresponding to the second path information whose first similarity meets the first threshold range.
[0102] Among them, the target corpus includes texts whose similarity with the target text meets the second threshold range, and the target corpus is used to assist LLM in generating code with business functions. Specifically, the target corpus and the target text can be filled in according to a preset prompt template and then input into the LLM, or the computing device can also fuse the target corpus and the target text and then input them into the LLM, etc. to assist LLM in generating code with business functions. This application does not impose any restrictions on this.
[0103] The computing device can obtain the first similarity between the first path information and each path information in the corpus through step S220, and then in step S230, the computing device can determine the second path information in the corpus (that is, the path information in the corpus whose first similarity meets the first threshold range) by comparing the first similarity and the first threshold range (which can be set by the user).
[0104] By comparison, the computing device can obtain the following two results:
[0105] 1) There is no second path information that meets the requirements in the corpus.
[0106] In this case, the computing device can further obtain third path information, the directory corresponding to the third path information is the parent directory of the directory corresponding to the first path information, and then, the computing device obtains the second similarity between the third path information and each path information in the corpus; and determines the target corpus corresponding to the fourth path information whose second similarity meets the second threshold range.
[0107] That is to say, in an embodiment of the present application, the computing device can expand the path information used for retrieval as needed. If there is no second path information that meets the requirements in the corpus, the third path information corresponding to the parent directory of the directory corresponding to the first path information can be used for retrieval. Similarly, the computing device can also search through the fifth path information corresponding to the parent directory of the directory corresponding to the third path information, and so on.
[0108] Among them, when the computing device obtains the similarity between path information of different levels and each path information in the corpus (including the aforementioned second similarity), it can adopt a method similar to that used to obtain the first similarity between the first path information and each path information in the corpus, including regular matching, vector distance calculation, etc., which will not be elaborated here.
[0109] Exemplarily, the first path information is: phone\test_task\call_test\main.py, the computing device obtains the first similarity between the first path information and each path information in the corpus, which is up to 90%, which does not meet the first threshold range: more than 95%, and determines that there is no second path information that meets the requirements in the corpus, and then the computing device obtains the third path information as: phone\test_task\call_test, and then the computing device obtains the second similarity between the third path information and each path information in the corpus, which is up to 98%, which meets the second threshold range: more than 95%. In this way, the computing device can then refer to the following 2) The implementation method involved in the existence of one or more second path information that meets the requirements in the corpus for further processing. If the aforementioned second similarity still does not meet the second threshold range, the computing device can also obtain the fifth path information as phone\test_task, and search again, and so on, which will not be repeated here.
[0110] 2) There is one or more second path information meeting the requirements in the corpus.
[0111] In a possible implementation, the computing device obtains a third vector corresponding to the target text, and then obtains a third vector distance between the third vector and a vector corresponding to a character string of a target field in the corpus corresponding to the second path information, wherein the character string of the target field is used to represent the business function of the code in the corresponding corpus, as described above. Figure 5 The string corresponding to the "test step description" field shown, and then the computing device determines the third similarity based on the third vector distance; the computing device determines that the corpus corresponding to the second path information whose third similarity meets the third threshold range and the first similarity meets the first threshold range is the target corpus.
[0112] Alternatively, in other feasible embodiments, before step S230, the computing device has screened out a portion of candidate corpora that meet the requirements from the corpus through the aforementioned implementation method, and then the computing device executes step S230 again to further determine, from these candidate corpora, the target corpus corresponding to the second path information whose first similarity meets the first threshold range.
[0113] That is to say, when the computing device retrieves the target corpus from the corpus, it can comprehensively compare the key information such as the path information and the target field of each corpus in the corpus and the similarity between the corresponding information of the first file, thereby improving the accuracy of the retrieval and improving the reasoning effect of the LLM.
[0114] In addition to the above two implementations, the computing device may also adopt the following two possible implementations to further improve the accuracy of the search, including:
[0115] a. The computing device obtains the degree of overlap between the character string of the target field and the character string of the target text in the corpus corresponding to the second path information, the character string of the target field representing the business function of the code in the same corpus; the computing device determines the fourth similarity based on the degree of overlap; the computing device determines that the corpus corresponding to the second path information whose fourth similarity meets the fourth threshold range and whose first similarity meets the first threshold range is the target corpus.
[0116] Specifically, the computing device can use methods such as substring matching to determine whether the character string of the target field in the corpus corresponding to the second path information includes the character string of the target text of one or more corpora in the corpus. Since the character string of the target field and the target text are both used to describe business functions, that is, if the character string of the target field in the corpus corresponding to the second path information includes the character string of the target text of a corpus in the corpus, it can be said that the code generated by the user requiring LLM is similar to the code in the corpus. Inputting such corpus together with the prompt input by the user into LLM can significantly improve the reasoning effect of LLM.
[0117] b. The computing device obtains the sorting result of the corpus corresponding to the second path information; the sorting result is determined based on the degree of correlation between the corpus corresponding to the second path information and the domain knowledge related to the target business function; the computing device determines the target corpus based on the sorting result.
[0118] Specifically, some specific domain knowledge can be set in the computing device, and the corpus corresponding to the second path information can be sorted by comparing the semantic similarity between the corpus corresponding to the second path information and these domain knowledge, and the corpus with a higher degree of relevance can be determined as the target corpus.
[0119] In some feasible embodiments, the computing device may also send the corpus corresponding to the second path information to an external domain expertise screening component, and then obtain the sorting result returned by the component, and determine the corpus with a higher degree of relevance indicated by the sorting result as the target corpus.
[0120] Through the above steps S210-S230 and various possible implementations, the computing device can use the key information such as the path information, the character string of the target field (test step description), etc. to accurately determine the target corpus that is highly relevant to the target text from the many corpora in the corpus, and then input it into Figure 2 The reasoning module 1400 shown is used to assist LLM in generating corresponding codes, for example, Figure 8 As shown, Figure 8 for Figure 6 A schematic diagram of a reasoning result involved in the illustrated embodiment, Figure 8 The code generated by LLM can be used directly without modification, which significantly improves the user's work efficiency.
[0121] The above mainly introduces the scheme of the embodiment of the present application from the perspective of the method. It is understandable that in order to realize the above functions, the corpus retrieval device includes at least one of the hardware structure and software modules corresponding to the execution of each function. Those skilled in the art should easily realize that, in combination with the units and algorithm steps of each example described in the embodiment disclosed herein, the embodiment of the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the embodiment of the present application.
[0122] The embodiment of the present application can divide the corpus retrieval device into functional units according to the above method example. For example, each functional unit can be divided according to each function, or two or more functions can be integrated into one processing unit. The above integrated unit can be implemented in the form of hardware or in the form of software functional units. It should be noted that the division of units in the embodiment of the present application is schematic and is only a logical functional division. There may be other division methods in actual implementation.
[0123] For example, Fig. 9 A structural diagram of a corpus retrieval device is provided for an embodiment of the present application. The corpus retrieval device 800 is applied to a computing device, or the corpus retrieval device 800 may be a computing device. The corpus retrieval device 800 includes:
[0124] The first acquisition module 810 is used to obtain first path information of a first file located in a first database, where the first file includes a target text, and the target text is used to represent the business function of the code that the user needs to generate through the large language model LLM.
[0125] The second acquisition module 820 is used to acquire a first similarity between the first path information and each path information in the corpus.
[0126] The determination module 830 is used to determine the target corpus corresponding to the second path information whose first similarity meets the first threshold range, wherein the target corpus includes text whose similarity with the target text meets the second threshold range, and the target corpus is used to assist the LLM in generating code with the business function.
[0127] For example, combined with Figure 6, the first acquisition module 810 can be used to perform the following steps: Figure 6 As shown in S210, the second acquisition module 820 can be used to perform the following steps: Figure 6 As shown in S220, the determination module 830 can be used to perform the following steps: Figure 6 S230 shown.
[0128] In a possible implementation, the determination module 830 is also used to obtain third path information if the second path information does not exist; the directory corresponding to the third path information is the parent directory of the directory corresponding to the first path information; obtain the second similarity between the third path information and each path information in the corpus; and determine the target corpus corresponding to the fourth path information whose second similarity meets the second threshold range.
[0129] In a possible implementation, the determination module 830 is further used to obtain a first vector corresponding to the first path information; obtain a first vector distance between the first vector and vectors corresponding to each path information in the corpus; and determine the first similarity based on the first vector distance.
[0130] In a possible implementation, the determination module 830 is further used to obtain a second vector corresponding to the third path information; obtain a second vector distance between the second vector and the vectors corresponding to each path information in the corpus; and determine the second similarity based on the second vector distance.
[0131] In a possible implementation, the determination module 830 is further configured to obtain a first degree of overlap between the character string corresponding to the first path information and the character string of each path information in the corpus; and determine the first similarity according to the first degree of overlap.
[0132] In a possible implementation, the determination module 830 is further configured to obtain a second degree of overlap between the character string corresponding to the third path information and the character string of each path information in the corpus; and determine the second similarity according to the second degree of overlap.
[0133] In a possible implementation, the determination module 830 is also used to obtain a third vector corresponding to the target text; obtain a third vector distance between the third vector and a vector corresponding to a character string of a target field in the corpus corresponding to the second path information; the character string of the target field is used to characterize the business function of the code in the corresponding corpus; determine a third similarity based on the third vector distance; determine that the corpus corresponding to the second path information whose third similarity meets a third threshold range and whose first similarity meets a first threshold range is the target corpus.
[0134] In a possible implementation, the determination module 830 is also used to obtain the degree of overlap between the character string of the target field in the corpus corresponding to the second path information and the character string of the target text, the character string of the target field representing the business function of the code in the same corpus; determine the fourth similarity based on the degree of overlap; determine that the corpus corresponding to the second path information whose fourth similarity meets the fourth threshold range and whose first similarity meets the first threshold range is the target corpus.
[0135] In a possible implementation, the determination module 830 is also used to obtain a sorting result of the corpus corresponding to the second path information; the sorting result is determined based on the degree of correlation between the corpus corresponding to the second path information and the domain knowledge related to the target business function; and the target corpus is determined based on the sorting result.
[0136] As a feasible example, the corpus retrieval device 800 provided in the present application is implemented through a software module. For example, the software module can be provided to users through a cloud service subscription model, and users can choose different subscription levels according to their needs. For example, the software module can also provide enterprise-level customized services with professional domain customization, interface personalization and extended functions according to the needs of users or enterprises.
[0137] In addition, the corpus retrieval device 800 provided by the present application can also be provided to users as a value-added service, which is not limited by the present application. When the corpus retrieval device 800 is implemented by a software module, the corpus retrieval device 800 can also be embedded in a retrieval software, a search engine or a large language model reasoning system.
[0138] The present application embodiment also provides a computing device 100. Fig.10 As shown, Fig.10 A schematic diagram of a computing device provided in an embodiment of the present application, Fig.10 The computing device 100 shown includes: a bus 102, a processor 104, a memory 106, and a communication interface 108. The processor 104, the memory 106, and the communication interface 108 communicate through the bus 102. The computing device 100 can be a server or a terminal device. It should be understood that the embodiment of the present application does not limit the number of processors and memories in the computing device 100.
[0139] The bus 102 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Fig.10 The bus 102 is represented by only one line, but it does not mean that there is only one bus or one type of bus. The bus 102 may include a path for transmitting information between various components of the computing device 100 (eg, the memory 106, the processor 104, and the communication interface 108).
[0140] The processor 104 may include any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0141] The memory 106 may include a volatile memory, such as a random access memory (RAM). The processor 104 may also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid state drive (SSD).
[0142] The memory 106 stores executable program codes, and the processor 104 executes the executable program codes to respectively implement the functions of the first acquisition module, the second acquisition module, and the determination module, thereby implementing the corpus retrieval method. That is, the memory 106 stores instructions for executing the corpus retrieval method.
[0143] The embodiment of the present application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smart phone.
[0144] like Fig.11 As shown, Fig.11 A schematic diagram of a computing device cluster provided in an embodiment of the present application, Fig.11 The computing device cluster shown includes at least one computing device 100. The memory 106 in one or more computing devices 100 in the computing device cluster may store the same instructions for executing the corpus retrieval method.
[0145] In some possible implementations, the memory 106 of one or more computing devices 100 in the computing device cluster may also store partial instructions for executing the corpus retrieval method. In other words, the combination of one or more computing devices 100 may jointly execute the instructions for executing the corpus retrieval method.
[0146] It should be noted that the memory 106 in different computing devices 100 in the computing device cluster can store different instructions, which are respectively used to execute part of the functions of the corpus retrieval device. That is, the instructions stored in the memory 106 in different computing devices 100 can implement the functions of one or more modules among the first acquisition module, the second acquisition module and the determination module.
[0147] In some possible implementations, one or more computing devices in the computing device cluster may be connected via a network, which may be a wide area network or a local area network. Fig.12 A possible implementation is shown. Fig.12 A schematic diagram of a connection method between computing device clusters provided in an embodiment of the present application, such as Fig.12 As shown, two computing devices 100A and 100B are connected via a network. Specifically, the network is connected via a communication interface in each computing device. In this type of possible implementation, the memory 106 in the computing device 100A stores instructions for executing the functions of the first acquisition module. At the same time, the memory 106 in the computing device 100B stores instructions for executing the functions of the second acquisition module and the determination module.
[0148] It should be understood that Fig.12 The functions of the computing device 100A shown in FIG. 1 may also be completed by multiple computing devices 100. Similarly, the functions of the computing device 100B may also be completed by multiple computing devices 100.
[0149] The embodiment of the present application also provides a computer program product including instructions. The computer program product may be software or a program product including instructions that can be run on a computing device or stored in any available medium. When the computer program product is run on at least one computing device, the at least one computing device executes the corpus retrieval method.
[0150] The embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state hard disk). The computer-readable storage medium includes instructions that instruct a computing device to execute a corpus retrieval method, or instructs a computing device to execute a corpus retrieval method.
[0151] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present invention.
Claims
1. A corpus retrieval method, characterized in that: The method comprises: Acquire first path information of a first file located in a first database, where the first file includes a target text, and the target text is used to represent a business function of a code that a user needs to generate through a large language model (LLM); Obtaining a first similarity between the first path information and each path information in the corpus; Determine a target corpus corresponding to the second path information whose first similarity meets a first threshold range, wherein the target corpus includes text whose similarity with the target text meets a second threshold range, and the target corpus is used to assist the LLM in generating a code having the business function.
2. The method according to claim 1, characterized in that The step of determining the target corpus corresponding to the second path information whose first similarity meets the first threshold range includes: If the second path information does not exist, obtain third path information; the directory corresponding to the third path information is the parent directory of the directory corresponding to the first path information; Obtaining a second similarity between the third path information and each path information in the corpus; Determine the target corpus corresponding to the fourth path information whose second similarity meets the second threshold range.
3. The method according to claim 1, characterized in that The obtaining of the first similarity between the first path information and each path information in the corpus includes: Obtaining a first vector corresponding to the first path information; Obtaining a first vector distance between the first vector and a vector corresponding to each path information in the corpus; The first similarity is determined according to the first vector distance.
4. The method according to claim 2, characterized in that: The obtaining of the second similarity between the third path information and each path information in the corpus includes: Obtaining a second vector corresponding to the third path information; Obtaining a second vector distance between the second vector and a vector corresponding to each path information in the corpus; The second similarity is determined according to the second vector distance.
5. The method according to claim 1 or 3, characterized in that: The obtaining of the first similarity between the first path information and each path information in the corpus includes: Obtaining a first degree of overlap between a character string corresponding to the first path information and a character string of each path information in the corpus; The first similarity is determined according to the first overlap degree.
6. The method according to claim 2 or 4, characterized in that: The obtaining of the second similarity between the third path information and each path information in the corpus includes: Obtaining a second degree of overlap between the character string corresponding to the third path information and the character strings of each path information in the corpus; The second similarity is determined according to the second overlap degree.
7. The method according to any one of claims 1 to 6, characterized in that: The step of determining the target corpus corresponding to the second path information whose first similarity meets the first threshold range includes: Obtaining a third vector corresponding to the target text; Obtaining a third vector distance between the third vector and a vector corresponding to a character string of a target field in the corpus corresponding to the second path information; the character string of the target field is used to represent a business function of a code in the corresponding corpus; Determining a third similarity according to the third vector distance; It is determined that the corpus corresponding to the second path information whose third similarity meets the third threshold range and whose first similarity meets the first threshold range is the target corpus.
8. The method according to any one of claims 1 to 6, characterized in that: The step of determining the target corpus corresponding to the second path information whose first similarity meets the first threshold range includes: Obtaining a degree of overlap between a character string of a target field in the corpus corresponding to the second path information and a character string of the target text, wherein the character string of the target field represents a business function of a code in the same corpus; Determining a fourth similarity according to the overlap degree; It is determined that the corpus corresponding to the second path information whose fourth similarity meets a fourth threshold range and whose first similarity meets a first threshold range is the target corpus.
9. The method according to any one of claims 1 to 8, characterized in that: The step of determining the target corpus corresponding to the second path information whose first similarity meets the first threshold range includes: Obtaining a sorting result of the corpus corresponding to the second path information; the sorting result is determined according to a degree of relevance between the corpus corresponding to the second path information and the domain knowledge related to the target business function; The target corpus is determined according to the sorting result.
10. A corpus retrieval device, characterized in that: The device comprises: A first acquisition module is used to acquire first path information of a first file located in a first database, wherein the first file includes a target text, and the target text is used to represent a business function of a code that a user needs to generate through a large language model LLM; A second acquisition module, used for acquiring a first similarity between the first path information and each path information in the corpus; A determination module is used to determine a target corpus corresponding to the second path information whose first similarity meets a first threshold range, wherein the target corpus includes text whose similarity with the target text meets a second threshold range, and the target corpus is used to assist the LLM in generating code with the business function.
11. The device according to claim 10, characterized in that The determination module is further used to: If the second path information does not exist, obtain third path information; the directory corresponding to the third path information is the parent directory of the directory corresponding to the first path information; Obtaining a second similarity between the third path information and each path information in the corpus; Determine the target corpus corresponding to the fourth path information whose second similarity meets the second threshold range.
12. The device according to claim 10, characterized in that The determination module is further used to: Obtaining a first vector corresponding to the first path information; Obtaining a first vector distance between the first vector and a vector corresponding to each path information in the corpus; The first similarity is determined according to the first vector distance.
13. The device according to claim 11, characterized in that The determination module is further used to: Obtaining a second vector corresponding to the third path information; Obtaining a second vector distance between the second vector and a vector corresponding to each path information in the corpus; The second similarity is determined according to the second vector distance.
14. The device according to claim 10 or 12, characterized in that The determination module is further used to: Obtaining a first degree of overlap between a character string corresponding to the first path information and a character string of each path information in the corpus; The first similarity is determined according to the first overlap degree.
15. The device according to claim 11 or 13, characterized in that The determination module is further used to: Obtaining a second degree of overlap between the character string corresponding to the third path information and the character strings of each path information in the corpus; The second similarity is determined according to the second overlap degree.
16. The device according to any one of claims 10 to 15, characterized in that: The determination module is further used to: Obtaining a third vector corresponding to the target text; Obtaining a third vector distance between the third vector and a vector corresponding to a character string of a target field in the corpus corresponding to the second path information; the character string of the target field is used to represent a business function of a code in the corresponding corpus; Determining a third similarity according to the third vector distance; It is determined that the corpus corresponding to the second path information whose third similarity meets the third threshold range and whose first similarity meets the first threshold range is the target corpus.
17. The device according to any one of claims 10 to 15, characterized in that The determination module is further used to: Obtaining a degree of overlap between a character string of a target field in the corpus corresponding to the second path information and a character string of the target text, wherein the character string of the target field represents a business function of a code in the same corpus; Determining a fourth similarity according to the overlap degree; It is determined that the corpus corresponding to the second path information whose fourth similarity meets a fourth threshold range and whose first similarity meets a first threshold range is the target corpus.
18. The device according to any one of claims 10 to 17, characterized in that: The determination module is further used to: Obtaining a sorting result of the corpus corresponding to the second path information; the sorting result is determined according to a degree of relevance between the corpus corresponding to the second path information and the domain knowledge related to the target business function; The target corpus is determined according to the sorting result.
19. A computing device cluster, characterized in that: comprising at least one computing device, each computing device comprising a processor and a memory; The processor of the at least one computing device is used to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the corpus retrieval method according to any one of claims 1 to 9.
20. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes computer instructions; when the computer instructions are executed in a computing device, the computing device executes the corpus retrieval method according to any one of claims 1 to 9.
21. A computer program product, characterized in that When the computer program product is run in a computing device, the computing device executes the corpus retrieval method according to any one of claims 1 to 9.
Citation Information
Cited By
Project-level code generation method and device based on file sorting, computer equipment and storage medium
CN121116264A