Document retrieval enhancement method, device and equipment for large language model

By adopting heuristic rules, text segmentation method of sliding windows and generating document directory structure in the large language model, combined with the secondary recall mechanism, the problem of illusion and insufficient recall in text retrieval enhancement of large language model is solved, and more accurate text similarity calculations and higher text recall rates are achieved.

CN119938884APending Publication Date: 2025-05-06BEIJING KEDONG ELECTRIC POWER CONTROL SYST CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510055368.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing large language models have hallucinations, mathematical numerical calculation problems and lack of interpretability in text retrieval enhancement, and the RAG method has shortcomings in text block granularity and recall.

Method used

By adopting text segmentation method based on heuristic rules and sliding windows in a large language model, a document directory structure is generated and vectorized text fragments are stored in a vector database, combining the secondary recall mechanism to improve the focus and recall rate of text blocks.

Benefits of technology

Improves the accuracy and text recall of text similarity calculations, reduces illusions and irrelevant problems during the generation process, improves user experience, and provides flexibility and scalability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938884A_ABST
    Figure CN119938884A_ABST
Patent Text Reader

Abstract

The invention provides a document retrieval enhancement method, device and equipment for a large language model. The method comprises the steps that for each original text, the starting position of the original text serves as a starting point, a text segmentation symbol is searched for according to the text sequence of the text, text segmentation is conducted on the basis of a preset heuristic rule and a sliding window according to the starting point and the text segmentation symbol, and a plurality of text blocks meeting the text length range are obtained; obtaining all word sizes in the original text, and generating a document directory structure corresponding to the original text; storing the document directory structure corresponding to the original text and the vectorized text fragments of the corresponding text blocks into a vector database; recalling topK recalled text fragments most related to user input; and for each recalled text fragment, according to the corresponding document directory structure, obtaining a vectorized text fragment corresponding to the same level or higher level of the recalled text fragment, and obtaining a recalled long text meeting the length range of the recalled text.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application provides a document retrieval enhancement method, apparatus, and device for a large language model. Background Art

[0002] Large language models (LLMs), due to their large parameter size and structural complexity, possess rich world knowledge, strong general task-solving capabilities, excellent complex task reasoning capabilities, and strong ability to follow human instructions. They demonstrate outstanding performance across a wide range of tasks in fields such as text, images, speech, and vision. However, large language models also have limitations, such as hallucinations, numerical computational challenges in mathematics, and lack of interpretability. As a solution, due to their versatile task-solving capabilities, their capabilities can be enhanced through fine-tuning, external knowledge bases, agents, and other methods that effectively interact with the external world.

[0003] Retrieval Augmented Generation (RAG) is one of the key approaches to addressing the hallucination problem in large language models. RAG typically consists of two phases: retrieving contextually relevant information and using the retrieved knowledge to guide the generation process. During the retrieval phase, data preprocessing, text segmentation, text block vectorization, and index construction are typically performed offline. Online, user input is vectorized, the similarity between the input text vector and the text block vector is calculated, and the k text blocks with the highest similarity are selected as augmented context information for the user input and fed into the large language model. Summary of the Invention

[0004] In order to better achieve text retrieval enhancement effects and improve text recall rates, this application provides a document retrieval enhancement method, device, and equipment for a large language model. The technical solutions proposed in this application are as follows:

[0005] In a first aspect, the present application provides a document retrieval enhancement method for a large language model, comprising:

[0006] For each original text, the starting position of the original text is used as the starting point, and the text segmentation symbol is searched according to the order of the text. According to the starting point and the text segmentation symbol, the text is segmented based on the preset heuristic rules and the sliding window to obtain multiple text blocks that meet the text length range;

[0007] Obtain all font sizes in the original text and generate a document directory structure corresponding to the original text;

[0008] Storing the document directory structure corresponding to the original text and the vectorized text segments of each corresponding text block in a vector database;

[0009] After vectorizing the user input, similarity is calculated with the vectorized text segments in the vector database, and the top K recalled text segments that are most relevant to the user input are retrieved;

[0010] For each recalled text segment, according to the corresponding document directory structure, a vectorized text segment corresponding to the same level or a higher level of the recalled text segment is obtained to obtain a recalled long text that meets the recalled text length range.

[0011] In combination with the first aspect above, in a possible implementation, for each original text, the starting position of the original text is used as the starting point, a text segmentation symbol is searched according to the order of the text, and the text is segmented based on the starting point and the text segmentation symbol based on a preset heuristic rule and a sliding window to obtain multiple text blocks that meet the text length range, including:

[0012] For each original text, the starting position of the original text is used as the starting point, and the text segmentation symbol is searched according to the order of the text. Each of the text segmentation symbols is used as a cutting point in turn, and it is determined whether the text length between the starting point and the cutting point is greater than the minimum length for document segmentation and less than the maximum length for document segmentation;

[0013] If so, the text between the starting point and the cutting point is taken as the first text block;

[0014] If the length of the text between the starting point and the cutting point is greater than the maximum length of the document segmentation, the position where the starting point reaches the maximum length of the document segmentation is used as the new cutting point;

[0015] Determine whether the position of the new cutting point is within the range of the original text, and whether the length of the text between the new cutting point and the starting point is greater than or equal to the minimum length of the document segmentation; if so, take the text between the starting point and the new cutting point as the first text block;

[0016] If the length of the text between the starting point and the new cutting point is less than or equal to the minimum length for document segmentation, the starting point is moved toward the end of the text, and the above judgment process is repeated until a cutting point that meets the length requirement is found to obtain the first text block;

[0017] A sliding window is moved from the cutting point corresponding to the first text block toward the beginning of the text to obtain a new starting point. Based on the new starting point, the above text segmentation process is repeated until the new starting point reaches the end of the original text, thereby obtaining multiple text blocks.

[0018] In conjunction with the first aspect above, in a possible implementation, obtaining all font sizes in the original text and generating a document directory structure corresponding to the original text includes:

[0019] Obtain all font sizes in the original text and establish a mapping relationship between font sizes and directory levels;

[0020] For each text block, generating a single-layer directory hierarchy structure corresponding to the text block according to the mapping relationship between the font size and the directory hierarchy;

[0021] Each single-layer directory hierarchy is traversed to determine the hierarchical relationship of all text blocks, and a document directory structure corresponding to the original text is generated.

[0022] In conjunction with the first aspect above, in a possible implementation, for each text block, generating a single-layer directory hierarchy structure corresponding to the text block according to the mapping relationship between font size and directory hierarchy includes:

[0023] For each text block, based on the mapping relationship between the font size and the directory level, Beautiful Image Soup is used to generate a single-layer directory level structure corresponding to the text block.

[0024] In conjunction with the first aspect above, in one possible implementation, traversing each single-layer directory hierarchy to determine the hierarchical relationship of all text blocks and generating a document directory structure corresponding to the original text includes:

[0025] Create an empty stack to store HTML tags;

[0026] For each text block, if the single-layer directory hierarchy structure corresponding to the text block is a title, determine the level of the current title, compare the level of the current title with the elements in the stack, pop out the tag to be closed, add the start tag of the current title, and push the current title into the stack;

[0027] If the single-layer directory hierarchy structure corresponding to the text block is text, add a text tag;

[0028] After all text blocks are processed, corresponding end tags are popped up and added in sequence to obtain a document directory structure of the HTML structure corresponding to the original text.

[0029] In conjunction with the first aspect above, in one possible implementation, for each recalled text segment, obtaining, from the vectorized text segments corresponding to the original text according to the corresponding document directory structure, vectorized text segments corresponding to the recalled text segment at the same level or higher than the recalled text segment, to obtain a recalled long text that meets the recalled text length range, includes:

[0030] For each recalled text segment, obtaining a vectorized text segment at the same level as the recalled text segment according to the corresponding document directory structure;

[0031] Determining whether the length of the obtained vectorized text segment at the same level of the recalled text segment meets the recalled text length range;

[0032] If yes, the vectorized text segment at the same level as the recalled text segment is used as the recalled long text;

[0033] If the length of the obtained vectorized text segments at the same level of the recalled text segments is greater than the length of the longest recalled text, the vectorized text segments at the same level of the recalled text segments are truncated in the order of arrangement to obtain the recalled long text having a length equal to the length of the longest recalled text;

[0034] If the length of the vectorized text segment at the same level of the recalled text segment is less than the shortest recalled text length, then obtain the vectorized text segment at the previous level of the recalled text segment, and re-judge whether the length of the vectorized text segment at the previous level of the recalled text segment is greater than the shortest recalled text length and less than the longest recalled text length, until the recalled long text that meets the recall length range is obtained, or the top level of the document directory structure is reached, and all vectorized text segments corresponding to the original text are used as the recalled long text.

[0035] In combination with the first aspect above, in a possible implementation, after receiving the user input, the method further includes:

[0036] Vectorize the user input to obtain the corresponding vectorized user input.

[0037] In a second aspect, the present application provides a document retrieval enhancement device for a large language model, comprising:

[0038] A text segmentation module is configured to, for each original text, use the starting position of the original text as a starting point, search for a text segmentation symbol according to the order of the text, and perform text segmentation based on the starting point and the text segmentation symbol based on a preset heuristic rule and a sliding window to obtain multiple text blocks that meet a text length range;

[0039] A directory structure building module is used to obtain all font sizes in the original text and generate a document directory structure corresponding to the original text;

[0040] A storage module, configured to store the document directory structure corresponding to the original text and the vectorized text segments of each corresponding text block in a vector database;

[0041] A first recall module is configured to perform similarity calculation on the vectorized user input and the vectorized text segments in the vector database, and recall the top K recalled text segments that are most relevant to the user input;

[0042] The second recall module is used to obtain, for each recalled text segment, a vectorized text segment corresponding to the same level or a higher level of the recalled text segment according to the corresponding document directory structure, and obtain a recalled long text that meets the recalled text length range.

[0043] In a third aspect, the present application provides a computer-readable storage medium having instructions stored therein. When the instructions are executed on a terminal, the terminal executes the document retrieval enhancement method for a large language model as described in the first aspect.

[0044] In a fourth aspect, the present application provides a computer program product comprising instructions, which, when executed on a computer device, enables the computer device to execute the document retrieval enhancement method for a large language model as described in the first aspect.

[0045] In a fifth aspect, the present application provides a computer device, characterized in that it includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other via the communication bus;

[0046] Memory for storing computer programs;

[0047] The processor is configured to implement the document retrieval enhancement method for a large language model as described in the first aspect when executing the program stored in the memory.

[0048] The descriptions of the second to fifth aspects of this application can refer to the detailed description of the first aspect; and the beneficial effects of the descriptions of the second to fifth aspects can refer to the analysis of the beneficial effects of the first aspect, which will not be repeated here.

[0049] Based on the above technical solution, the beneficial effects of this application compared with the prior art are as follows:

[0050] The document retrieval enhancement method for a large language model provided in an embodiment of the present application can divide text blocks more finely by integrating heuristic rules and a sliding window segmentation method, so that the content of each text block is more focused and the interference of irrelevant information is reduced. By generating a document directory structure of the original text and storing the corresponding vectorized text fragments in a vector database, it helps to more accurately match user input in the retrieval stage, and can quickly recall the most relevant text fragments when the user inputs, thereby improving the accuracy of text similarity calculation and making the retrieval results more in line with expectations. Moreover, by combining the document structured segmentation and the secondary recall segmentation and recall method, for each recalled text fragment, the corresponding vectorized text fragment at the same level or higher level is further obtained to form a recalled long text, which effectively improves the relevance of the recalled text to the question, thereby improving the recall rate of the text.

[0051] The document retrieval enhancement method for a large language model provided in the embodiments of this application improves the accuracy of text similarity calculations and the text recall rate. These improvements work together to improve the generation process of the large language model, enabling the model to generate high-quality text based on more accurate and rich contextual information. This helps reduce problems such as hallucinations and irrelevance during the generation process, improves the accuracy and fluency of the generated text, and thus enhances the user experience.

[0052] The document retrieval enhancement method for a large language model provided in the embodiment of the present application can conveniently adjust parameters such as heuristic rules, the size of the sliding window, and the method of generating the document directory structure to adapt to different application scenarios and needs. It combines the document structured segmentation and secondary recall segmentation and recall methods, and has strong flexibility and scalability. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 A flowchart of a document retrieval enhancement method for a large language model provided in an embodiment of the present application;

[0054] Figure 2 A schematic diagram of an operation and maintenance document sample provided for an embodiment of the present application;

[0055] Figure 3 A schematic diagram of the execution process of the document retrieval enhancement method for a large language model provided in an embodiment of the present application;

[0056] Figure 4 Provide the embodiment of this application for Figure 2 The following is a schematic diagram of the recall text content of the operation and maintenance document sample shown;

[0057] Figure 5 A schematic diagram of the structure of a document retrieval enhancement device for a large language model provided in an embodiment of the present application;

[0058] Figure 6 This is a structural block diagram of the computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0059] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0060] The term "and / or" in this article is merely a description of the association relationship between associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone.

[0061] The terms "first" and "second" and the like in the specification and drawings of this application are used to distinguish different objects, or to distinguish different processing of the same object, rather than to describe a specific order of objects.

[0062] Furthermore, the terms "including," "having," and any variations thereof, as used in the description of this application are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or units is not limited to the listed steps or units, but may optionally include other steps or units not listed, or may optionally include other steps or units inherent to the process, method, product, or apparatus.

[0063] It should be noted that in the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be interpreted as being more preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.

[0064] In the description of the present application, unless otherwise specified, “plurality” means two or more.

[0065] The inventors discovered that the prior art RAG-based text retrieval enhancement method has at least the following problems: First, during offline document segmentation, the large granularity of text blocks reduces the accuracy of document similarity calculations, and not all blocks in the retrieval set are relevant to the user input. Second, directly using text blocks from similarity calculations as enhanced context information results in low text recall, affecting the quality of generated results, such as hallucinations and irrelevance. Based on this, the present embodiment provides a document retrieval enhancement method for a large language model. The following describes the implementation of the present embodiment in detail in conjunction with the accompanying drawings.

[0066] First, the symbols and formulas used in the embodiments of this application are explained as follows:

[0067]

[0068]

[0069] Reference Figure 1 As shown, the document retrieval enhancement method for a large language model provided by this application includes the following steps:

[0070] S101: For each original text, taking the starting position of the original text as the starting point, searching for a text segmentation symbol according to the order of the text, and performing text segmentation based on the starting point and the text segmentation symbol based on a preset heuristic rule and a sliding window to obtain multiple text blocks that meet a text length range;

[0071] S102: Obtain all font sizes in the original text and generate a document directory structure corresponding to the original text;

[0072] S103: storing the document directory structure corresponding to the original text and the vectorized text segments of each corresponding text block in a vector database;

[0073] S104: Calculate similarity between the vectorized user input and the vectorized text segments in the vector database, and recall the top K recalled text segments that are most relevant to the user input;

[0074] S105: For each recalled text segment, according to the corresponding document directory structure, obtain a vectorized text segment corresponding to the same level or a higher level of the recalled text segment, and obtain a recalled long text that meets the recalled text length range.

[0075] In the embodiment of the present application, to achieve efficient and accurate RAG, the data needs to be clean, consistent, and context-rich. Therefore, before executing step S101, all acquired raw text can be preprocessed to improve the quality of data input. Preprocessing methods may include document format standardization and conversion, removal of special characters and irrelevant information, deletion of duplicate documents or redundant information, and use of expiration mechanisms to handle time-sensitive data.

[0076] Exemplarily, the process of document preprocessing may specifically include:

[0077] 1. Document format standardization and conversion

[0078] Unified format: First, you need to ensure that all documents follow a unified format standard. This includes text encoding, file type (such as PDF, DOCX converted to plain text or a specific format), paragraph and sentence separators, etc.

[0079] Structural processing: For documents with specific structures (such as tables, lists, etc.), appropriate structural processing is required so that subsequent steps can extract and parse information more efficiently.

[0080] 2. Remove special characters and irrelevant information

[0081] Special characters: Special characters in the document need to be removed, such as HTML tags, garbled characters, mathematical symbols (unless they are useful to the system), etc.

[0082] Irrelevant information: Remove information that is not related to the topic, such as advertisements, headers and footers, copyright notices, etc. This helps reduce noise and improve the retrieval and generation efficiency of the system.

[0083] 3. Delete duplicate documents or redundant information

[0084] Duplicate Documents: By comparing document content, we can identify and remove duplicate or highly similar documents. This can reduce storage space usage and avoid duplicate results during retrieval.

[0085] Redundant information: Remove repeated or redundant sentences and paragraphs in a single document. This helps make the document more concise and clear, and improves the system's processing speed.

[0086] 4. Use invalidation mechanisms to handle time-sensitive data

[0087] Timestamp: Add a timestamp to a document to identify when it was created or updated.

[0088] Expiration mechanism: Set an expiration mechanism based on the time sensitivity of the document. For example, for time-sensitive documents such as news and market reports, you can set an expiration date. Once a document expires, the system can automatically remove it from the search index or reduce its weight.

[0089] 5. Text cleaning: In addition to removing irrelevant information, the text also needs to be cleaned, including correcting spelling errors, standardizing word usage, etc. This helps improve the system's retrieval and generation accuracy.

[0090] 6. Metadata Extraction: Extract useful metadata (such as author, title, date, etc.) from documents and store them in the database. These metadata can be used as additional conditions during retrieval to improve the accuracy of retrieval.

[0091] 7. Word segmentation and indexing: For Chinese documents, word segmentation is required and the segmentation results are stored in the index. This helps speed up retrieval and improve the system's responsiveness.

[0092] In the embodiment of the present application, in the above step S101, the text length range of the text block is that the text length is within the range of the minimum length for document segmentation and the maximum length for document segmentation. For each original text, the starting position of the original text is used as the starting point, and the text segmentation symbol is searched according to the order of the text. According to the starting point and the text segmentation symbol, the text is segmented based on the preset heuristic rules and the sliding window to obtain multiple text blocks that meet the text length range, specifically including:

[0093] For each original text, the starting position of the original text is used as the starting point, and the text segmentation symbol is searched according to the order of the text. Each of the text segmentation symbols is used as a cutting point in turn, and it is determined whether the text length between the starting point and the cutting point is greater than the minimum length for document segmentation and less than the maximum length for document segmentation;

[0094] If so, the text between the starting point and the cutting point is taken as the first text block;

[0095] If the length of the text between the starting point and the cutting point is greater than the maximum length of the document segmentation, the position where the starting point reaches the maximum length of the document segmentation is used as the new cutting point;

[0096] Determine whether the position of the new cutting point is within the range of the original text, and whether the length of the text between the new cutting point and the starting point is greater than or equal to the minimum length of the document segmentation; if so, take the text between the starting point and the new cutting point as the first text block;

[0097] If the length of the text between the starting point and the new cutting point is less than or equal to the minimum length for document segmentation, the starting point is moved toward the end of the text, and the above judgment process is repeated until a cutting point that meets the length requirement is found, thereby obtaining the first text block;

[0098] A sliding window is moved from the cutting point corresponding to the first text block toward the beginning of the text to obtain a new starting point. Based on the new starting point, the above text segmentation process is repeated until the new starting point reaches the end of the original text, thereby obtaining multiple text blocks.

[0099] In an embodiment of the present application, the above-mentioned text segmentation symbol seperator can be, for example, ".", ":", "!" and "?", etc. The text is recursively segmented by customizing the text segmentation symbol, while considering heuristic rules such as the minimum length of document segmentation min_split_len and the minimum length of document segmentation max_split_len, so as to take into account the length of the segmented text and the integrity of the text semantics, ensure that the text block segmentation size is close, while maintaining complete semantic information, and at the same time, use a sliding window to maintain some overlap between text blocks to ensure that the semantic context is not lost between blocks. In an embodiment of the present application, the size of the minimum length of document segmentation and the minimum length of document segmentation can be set according to actual conditions. For example, min_split_len=10 and max_split_len=50 can be set. At the same time, the size of the sliding window can also be set according to actual conditions. Preferably, the size of the sliding window is 0≤window_size≤min_split_len.

[0100] In order to facilitate those skilled in the art to implement the process of step S101 above, the pseudo code for the specific implementation of step S101 is provided below:

[0101] enter:

[0102] text: the original text to be segmented

[0103] separator: text segmentation symbol

[0104] min_split_len: minimum length of document segmentation

[0105] max_split_len: Maximum length of document split

[0106] window_size: sliding window

[0107] Output:

[0108] segments: list of sentences after segmentation

[0109] algorithm:

[0110] 1. Initialize an empty list segments to store the results;

[0111] 2. Set the current position cursor to 0;

[0112] 3. When the cursor is smaller than the length of the text, perform the following steps:

[0113] a. Starting from the current cursor position, find the position of the next separator pos;

[0114] b. If pos-cursor>max_spl it_len, the position of cursor+max_spl it_len is used as the cutting point;

[0115] c. Check whether the cutting point is less than the length of text and whether the length after cutting is greater than min_spl i t_len;

[0116] d. If yes, add the string between the cursor and the cut point to segments;

[0117] e. Update the cursor to the cutting point - window_size;

[0118] f. If the conditions are not met, try to move backward until a suitable cutting point is found;

[0119] 4. Return segments.

[0120] In the embodiment of the present application, in the above step S102, obtaining all font sizes in the original text and generating a document directory structure corresponding to the original text specifically includes:

[0121] Obtain all font sizes in the original text and establish a mapping relationship between font sizes and directory levels;

[0122] For each text block, generating a single-layer directory hierarchy structure corresponding to the text block according to the mapping relationship between the font size and the directory hierarchy;

[0123] Each single-layer directory hierarchy is traversed to determine the hierarchical relationship of all text blocks, and a document directory structure corresponding to the original text is generated.

[0124] In order to improve the relevance of the recalled text to the user input, the inventors considered constructing a document directory structure based on the hierarchical and progressive relationships of the document content. Documents usually contain directory structures such as first-level titles, second-level titles, third-level titles, and text to help humans understand the content of the document. Taking docx documents as an example, open source docx parsing tools such as pydocx can extract docx documents and content, including document structure, paragraphs, fonts, colors, font sizes, tables, pictures, and other content. However, the parsing of open source tools is completely based on the directory structure in the original docx document, and this type of directory structure is usually of poor quality. To this end, the inventors proposed the above-mentioned method for constructing a document directory structure based on text font size. The specific implementation process may include the following steps:

[0125] (1) Use open source tools such as pydocx to obtain all font sizes of the original text (i.e., document D):

[0126] {SIZE} D ={s1,s2,s3,s4,s5};

[0127] (2) Construct a set of mapping relationships between font sizes and directory levels based on font size and quantity:

[0128] {SIZE_MAP} D ={h1:s1,h2:s2,h3:s3,h4:s4,p:s5};

[0129] (3) For each text block, based on the mapping relationship between font size and directory level, use Beautiful Image Soup to generate a single-layer directory hierarchy structure corresponding to the text block, as shown below:

[0130]

[0131] (4) Traverse each single-layer directory hierarchy to determine the hierarchical relationship of all text blocks and generate the document directory structure corresponding to the original text:

[0132] Create an empty stack to store HTML tags;

[0133] For each text block, if the single-layer directory hierarchy structure corresponding to the text block is a title, determine the level of the current title, compare the level of the current title with the elements in the stack, pop out the tag to be closed, add the start tag of the current title, and push the current title into the stack;

[0134] If the single-layer directory hierarchy structure corresponding to the text block is text, add a text tag;

[0135] After all text blocks are processed, the corresponding end tags are popped up and added in sequence to obtain the document directory structure HTML structure corresponding to the original text D .

[0136] In the embodiment of the present application, the inventors consider that the HTML structure can efficiently represent the hierarchical relationship, and thus construct the HTML structure of the document D. Traversing the single-layer HTML in the document D i , confirm the hierarchical relationship of different text blocks according to their directory levels, and thus generate the HTML structure, placing the content belonging to the same directory level under the same to facilitate subsequent text recall.

[0137] In order to make it easier for those skilled in the art to implement the document directory structure HTML in step S102 D Build, the document directory structure HTML is provided below D The pseudo code for the specific implementation of the construction is as follows:

[0138] Define an empty stack to store HTML tags

[0139] Define an empty string html for splicing the final HTML code

[0140] For each document fragment:

[0141] If fragment is a title:

[0142] Determine the level of the title (e.g., Title 1 is 1, Title 2 is 2, etc.);

[0143] When the stack is not empty and the level of the top element of the stack is greater than or equal to the level of the current title:

[0144] Pop the top element of the stack;

[0145] Add the corresponding end tag to html (such as < / h1>, etc.);

[0146] Push the current title onto the stack;

[0147] html add the corresponding opening tag (such as 〈h1> <h2>wait);

[0148] If fragment is body:

[0149] HTML adds the body content surrounded by tags;

[0150] When all fragments are processed:

[0151] When the stack is not empty:

[0152] Pop the top element of the stack;

[0153] Add the corresponding end tag to html;

[0154] Returns html.

[0155] The above-mentioned document directory structure construction method based on text font size provided by the embodiment of the present application has the following significant advantages compared to directly relying on the directory structure in the original docx document or only using open source tools for parsing:

[0156] Improve the quality of the directory structure: The directory structure in the original docx document may be poor due to various reasons (such as improper editing, confusing formatting, etc.), resulting in it not accurately reflecting the hierarchy and progression of the document content. By using a method based on text size, the directory structure can be constructed more flexibly to make it more consistent with the actual organization of the document content.

[0157] Improved relevance of text recall: An optimized document directory structure helps search engines or information retrieval systems more accurately understand and index document content. When a user enters a query, the system can more effectively recall relevant text snippets, improving retrieval accuracy and efficiency.

[0158] Improve user experience: A clear, well-organized document directory structure helps you quickly locate required information and reduces search time. This is especially true when working with long documents or complex reports.

[0159] Strong adaptability: It is not dependent on a specific document format or tool, but is built based on the universal feature of text font size. Therefore, it is suitable for various types of docx documents and has wide applicability and flexibility.

[0160] Easy to implement and maintain: The overall process is relatively clear and easy to understand and implement. At the same time, since it does not rely on the specific functions or format requirements of external tools, it is also more convenient to maintain and update.

[0161] In the embodiment of the present application, the vectorized text segments of each text block in step S103 can be obtained by vectorizing the text block in the following manner:

[0162] Use semantic vectorization model to vectorize text and convert text blocks into i Convert to numerical vector representation

[0163] In the embodiments of this application, text vectorization is the foundation for large language models to process and parse text. On the one hand, numerical vector methods are used for computer representation and processing. On the other hand, pre-trained models such as BERT and GPT consider the contextual information of the text when training the vectorization model, so that the text vector contains the contextual and semantic information of the text, which facilitates subsequent similarity calculations.

[0164] In the embodiments of the present application, the semantic vectorization model may be, for example, a BGE model. The BGE model uses the RetroMAE algorithm to pre-train large text pairs of a large language model. It first encodes low-masking rate input into semantic vectors, and then concatenates high-masking rate input with the semantic vectors to reconstruct the original input.

[0165] In a specific embodiment, before executing step S104, after receiving the user input, the method may further include: performing text vectorization on the user input to obtain corresponding vectorized user input. Specifically, the user input may be vectorized using the same semantic vectorization model as that used for the text vectorization of the text block.

[0166] In the embodiment of the present application, in the above step S104, the method of calculating the similarity between the vectorized user input obtained after vectorizing the user input and the vectorized text fragments in the vector database can adopt various representation methods, such as cosine similarity, Pearson correlation coefficient, Euclidean distance, etc. For example, the cosine similarity is used to calculate the similarity between the two vectors, and the calculation method is as follows:

[0167]

[0168] Where n is the vector dimension.

[0169] In the embodiment of the present application, in the above step S105, for each recalled text segment, according to the corresponding document directory structure, obtaining a vectorized text segment corresponding to the same level or higher level of the recalled text segment from each vectorized text segment corresponding to the original text, and obtaining a recalled long text that meets the recalled text length range specifically includes:

[0170] For each recalled text segment, obtaining a vectorized text segment at the same level as the recalled text segment according to the corresponding document directory structure;

[0171] Determining whether the length of the obtained vectorized text segment at the same level of the recalled text segment meets the recalled text length range;

[0172] If yes, the vectorized text segment at the same level as the recalled text segment is used as the recalled long text;

[0173] If the length of the obtained vectorized text segments at the same level of the recalled text segments is greater than the length of the longest recalled text, the vectorized text segments at the same level of the recalled text segments are truncated in the order of arrangement to obtain the recalled long text having a length equal to the length of the longest recalled text;

[0174] If the length of the vectorized text segment at the same level of the recalled text segment is less than the shortest recalled text length, then obtain the vectorized text segment at the previous level of the recalled text segment, and re-judge whether the length of the vectorized text segment at the previous level of the recalled text segment is greater than the shortest recalled text length and less than the longest recalled text length, until the recalled long text that meets the recall length range is obtained, or the top level of the document directory structure is reached, and all vectorized text segments corresponding to the original text are used as the recalled long text.

[0175] In the embodiment of the present application, in order to facilitate those skilled in the art to realize the long text recall in the above step S105, it is assumed that a recalled text segment is the i-th text block t in document D i , the text block t during the second recall i In the document directory structure HTML D Recall Get Recall Long Text RT i The pseudo code of the algorithm flow is as follows:

[0176] enter:

[0177] t i : The i-th text block of document D

[0178] HTML D : Document directory structure of document D

[0179] min_recall_len: Minimum length of document recall

[0180] max_recall_len: maximum length of document recall

[0181] Output:

[0182] RT i : Recall long text

[0183] Algorithm:

[0184] 1. Obtain t i The corresponding text content text at the same level. If min_recall_len <= len(text) <= max_recall_len, then return RT i ;

[0185] 2. If len(text) < min_recall_len, then obtain t i The corresponding text content text at the previous level, and repeat the judgment process until min_recall_len <= len(text) <= max_recall_len, then return RT i ;

[0186] 3. If len(text) > max_recall_len, then truncate the obtained text and return RT i .

[0187] In the embodiments of the present application, due to the use of fine-grained text segmentation based on the document structure in the above step S101, after obtaining the top K texts most relevant to the input through similarity calculation, a secondary recall is performed in the original document to obtain a longer context related to the user input. Through fine-grained text segmentation based on the document structure, information units in the document can be captured more precisely. Combining the secondary recall mechanism can ensure that after obtaining the most relevant text fragments through preliminary similarity calculation, a context more relevant and longer to the user input can be further recalled from the original document. This method reduces the interference of irrelevant information and improves the relevance between the recalled text and the user query. Moreover, more precise text recall means that the large language model can receive more relevant and richer context information. This helps the model generate more accurate and coherent responses or answers, thus improving the overall performance of the model. By setting the minimum length and maximum length of document recall, users can adjust the length range of the recalled text according to actual needs, and can be flexibly applied to different scenarios and different types of query requirements. And users can receive more relevant and accurate responses, which will significantly improve user satisfaction. At the same time, due to the controllability of the length of the recalled text, a response of an appropriate length can be selected according to the user's reading habit or needs, further enhancing the user experience.

[0188] To more clearly illustrate the specific implementation process of the embodiments of the present application, taking the operation and maintenance documents for internal desensitization of an enterprise as an example, the specific implementation process of text retrieval enhancement of RAG is described in detail as follows:

[0189] Assume that there are 500 anonymized operation and maintenance documents in the enterprise. The example of the anonymized operation and maintenance documents in the enterprise is as follows: Figure 2 shown.

[0190] 1.Reference Figure 3 As shown, when executing this method, the 500 operation and maintenance documents can be preprocessed first. The specific implementation process can refer to the description of step S101 above, delete duplicate documents based on the similarity of document titles and contents, remove special characters and irrelevant information in the documents, and finally 485 documents remain.

[0191] 2. Next, refer to Figure 3 As shown, through the above step S101, based on the text segmentation method of integrating heuristic rules and sliding windows, each document D in the remaining 485 documents is segmented to obtain corresponding multiple text blocks.

[0192] Take the text to be split "When starting a Ceph storage cluster, each daemon searches for its own configuration from the same configuration file (the default ceph.conf). There are many configurable parameters in ceph.conf. Sometimes we need to modify some parameters according to the actual environment. There are two ways to modify them: directly modify the parameter values ​​in the ceph.conf configuration file, and then restart the Ceph process for the modification to take effect. Or dynamically adjust the parameters during operation without restarting the process." as an example. Set min_split_len = 10, max_split_len = 50, window_size = 0, and seporator = "." "?" "?", and the split result is:

[0193] When you start a Ceph Storage Cluster, each daemon looks for its own configuration in the same configuration file (the default ceph.conf).

[0194] There are many configurable parameters in ceph.conf, and sometimes we need to modify some parameters according to the actual environment.

[0195] There are two ways to modify the parameters: directly modify the parameter value in the ceph.conf configuration file. After the modification, you need to restart the Ceph process for the modification to take effect.

[0196] Or dynamically adjust parameters during operation without restarting the process."

[0197] It should be noted that in the above example, the sliding window window_size is set to 0 only to show the segmentation result. In actual execution, the sliding window window_size can be set to a positive integer less than or equal to min_split_len.

[0198] 3. Then, through step S102, the document directory structure of each operation and maintenance document is constructed:

[0199] With the above Figure 2 Take the example of the operation and maintenance document shown above as an example, and use open source tools such as pydocx to obtain all font sizes of document D.

[0200] {SIZE} D =[1,2,3,4];

[0201] Then, a mapping set between font size and directory level is constructed based on the font size and quantity:

[0202] {SIZE_MAP} D ={h1:1,h2:2,h3:3,p:4};

[0203] Then, based on the font size and quantity, a mapping set between font size and directory level is constructed, and BeautifulSoup is used to generate each single layer of HTML of the operation and maintenance document sample. i ;

[0204] Finally, traverse the single-layer html in document D i , confirm the hierarchical relationship of different text blocks according to their directory levels, thereby generating the HTML structure, placing the content belonging to the same directory level under the same , and finally obtaining the HTML document directory structure.

[0205] 4. Next, the BGE model is used to vectorize the text blocks of each operation and maintenance document using the semantic vectorization model and store them in the vector database.

[0206] 5. Then, after receiving the user input, based on the above step S104, the user input is vectorized using the same semantic vectorization model, and then the top K text fragments most relevant to the user input are recalled from the vector database through similarity calculation between vectors.

[0207] For example, assuming K is 3, the user input is "What issues should be paid attention to when installing MySQL?", and the similarity text obtained by recall is: "Issues that need to be paid attention to when installing" (such as Figure 2 ), "MySQL installation failed" and "MySQL installation process".

[0208] Finally, a second recall is performed on each recalled fine-grained similarity text through the above step S105.

[0209] Or Figure 2 Take the document shown as an example, set min_recall_len = 300, max_recall_len = 1000, and recall long text content such as Figure 4 shown.

[0210] The document retrieval enhancement method for a large language model provided in an embodiment of the present application can divide text blocks more finely by integrating heuristic rules and a sliding window segmentation method, so that the content of each text block is more focused and the interference of irrelevant information is reduced. By generating a document directory structure of the original text and storing the corresponding vectorized text fragments in a vector database, it helps to more accurately match user input in the retrieval stage, and can quickly recall the most relevant text fragments when the user inputs, thereby improving the accuracy of text similarity calculation and making the retrieval results more in line with expectations. Moreover, by combining the document structured segmentation and the secondary recall segmentation and recall method, for each recalled text fragment, the corresponding vectorized text fragment at the same level or higher level is further obtained to form a recalled long text, which effectively improves the relevance of the recalled text to the question, thereby improving the recall rate of the text.

[0211] The document retrieval enhancement method for a large language model provided in the embodiments of this application improves the accuracy of text similarity calculations and the text recall rate. These improvements work together to improve the generation process of the large language model, enabling the model to generate high-quality text based on more accurate and rich contextual information. This helps reduce problems such as hallucinations and irrelevance during the generation process, improves the accuracy and fluency of the generated text, and thus enhances the user experience.

[0212] The document retrieval enhancement method for a large language model provided in the embodiment of the present application can conveniently adjust parameters such as heuristic rules, the size of the sliding window, and the method of generating the document directory structure to adapt to different application scenarios and needs. It combines the document structured segmentation and secondary recall segmentation and recall methods, and has strong flexibility and scalability.

[0213] Example 2

[0214] Based on the same inventive concept, the present application embodiment also provides a document retrieval enhancement device for a large language model, referring to Figure 5 As shown, the device includes:

[0215] The text segmentation module 101 is configured to, for each original text, use the starting position of the original text as a starting point, search for a text segmentation symbol according to the order of the text, and perform text segmentation based on the starting point and the text segmentation symbol based on a preset heuristic rule and a sliding window to obtain multiple text blocks that meet a text length range;

[0216] A directory structure building module 102 is configured to obtain all font sizes in the original text and generate a document directory structure corresponding to the original text;

[0217] A storage module 103 is configured to store the document directory structure corresponding to the original text and the vectorized text segments of each corresponding text block in a vector database;

[0218] A first recall module 104 is configured to perform similarity calculation between the vectorized user input and the vectorized text segments in the vector database, and recall the top K recalled text segments that are most relevant to the user input;

[0219] The second recall module 105 is used to obtain, for each recalled text segment, a vectorized text segment corresponding to the same level or a higher level of the recalled text segment according to the corresponding document directory structure, and obtain a recalled long text that meets the recalled text length range.

[0220] Example 3

[0221] Based on the same inventive concept, an embodiment of the present application also provides a computer-readable storage medium, which stores instructions. When the instructions are executed on a terminal, the terminal executes the document retrieval enhancement method for the large language model as described above.

[0222] Among them, computer-readable storage media, for example, can be but are not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices or components, or any combination thereof. More specific examples (non-exhaustive list) of computer-readable storage media include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a register, a hard disk, an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above, or any other form of computer-readable storage medium well known in the art. An exemplary storage medium is coupled to a processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be an integral part of the processor. The processor and the storage medium can be located in an application-specific integrated circuit (ASIC). In the embodiments of the present application, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0223] Example 4

[0224] Based on the same inventive concept, an embodiment of the present application further provides a computer program product comprising instructions. When the computer program product is run on a computer, the computer is enabled to execute the document retrieval enhancement method for a large language model as described above.

[0225] Example 5

[0226] Based on the same inventive concept, an embodiment of the present application further provides a computer device, characterized in that it includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other via the communication bus;

[0227] Memory for storing computer programs;

[0228] The processor is configured to implement the above-mentioned document retrieval enhancement method for a large language model when executing the program stored in the memory.

[0229] Figure 6 A possible structural diagram of the computer device involved in the above embodiment is shown. The computer device includes: a processor 1002 and a communication interface 1003. The processor 1002 is used to control and manage the actions of the computer device, for example, to execute the document retrieval enhancement method for the large language model described above, and / or to execute other processes of the technology described herein. The communication interface 1003 is used to support communication between the computer device and other network entities, for example, to execute the steps performed by the above communication unit 902. The computer device may also include a memory 1001 and a bus 1004. The memory 1001 is used to store program code and data of the computer device.

[0230] Among them, the memory 1001 can be a memory in a computer device, etc., and the memory can include a volatile memory, such as a random access memory; the memory can also include a non-volatile memory, such as a read-only memory, a flash memory, a hard disk or a solid-state drive; the memory can also include a combination of the above types of memory.

[0231] The processor 1002 may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. The processor may be a central processing unit, a general-purpose processor, a digital signal processor, an application-specific integrated circuit, a field-programmable gate array, or other programmable logic device, a transistor logic device, a hardware component, or any combination thereof. The processor may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. The processor may also be a combination that implements computing functions, such as a combination of one or more microprocessors, or a combination of a DSP and a microprocessor.

[0232] The bus 1004 may be an Extended Industry Standard Architecture (EISA) bus or the like. The bus 1004 may be divided into an address bus, a data bus, a control bus, or the like. For ease of representation, Figure 6 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.

[0233] Through the description of the above embodiments, those skilled in the art will clearly understand that for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. The specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0234] Since the computer-readable storage medium, computer program product, and computer device in the embodiments of the present application can be applied to the above method, the technical effects that can be obtained can also refer to the above method embodiments, and the embodiments of the present application will not be repeated here.

[0235] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0236] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0237] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0238] The above are only specific embodiments of the present application, but the scope of protection of the present application is not limited thereto. Any changes or replacements within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.< / h2>

Claims

1. A document retrieval enhancement method for a large language model, characterized in that: include: For each original text, the starting position of the original text is taken as the starting point, and the text segmentation symbol is searched according to the text sequence, and the text is segmented based on the starting point and the text segmentation symbol based on the preset heuristic rules and the sliding window to obtain multiple text blocks that meet the text length range; Obtain all font sizes in the original text, and generate a document directory structure corresponding to the original text; The document directory structure corresponding to the original text and the vectorized text segments of each corresponding text block are stored in a vector database; After the user input is vectorized, similarity calculation is performed with the vectorized text segments in the vector database, and the top K recalled text segments that are most relevant to the user input are recalled; For each recalled text segment, according to the corresponding document directory structure, a vectorized text segment corresponding to the same level or a higher level of the recalled text segment is obtained to obtain a recalled long text that meets the recalled text length range.

2. The method according to claim 1, characterized in that For each original text, the starting position of the original text is taken as the starting point, and the text segmentation symbol is searched according to the text sequence. According to the starting point and the text segmentation symbol, the text is segmented based on the preset heuristic rules and the sliding window to obtain multiple text blocks that meet the text length range, including: For each original text, the starting position of the original text is taken as the starting point, and the text segmentation symbol is searched according to the text sequence, and each of the text segmentation symbols is taken as the cutting point in turn, and it is determined whether the text length between the starting point and the cutting point is greater than the minimum length of the document segmentation and less than the maximum length of the document segmentation; If so, the text between the starting point and the cutting point is taken as the first text block; If the length of the text between the starting point and the cutting point is greater than the maximum length of the document segmentation, the position where the starting point reaches the maximum length of the document segmentation is used as the new cutting point; Determine whether the position of the new cutting point is within the range of the original text, and whether the length of the text between the new cutting point and the starting point is greater than or equal to the minimum length of the document segmentation; if so, take the text between the starting point and the new cutting point as the first text block; If the text length between the starting point and the new cutting point is less than or equal to the minimum length for document segmentation, the starting point is moved toward the end of the text, and the above judgment process is repeated until a cutting point that meets the length requirement is found to obtain the first text block; A sliding window is moved from the cutting point corresponding to the first text block toward the beginning of the text to obtain a new starting point. According to the new starting point, the above text segmentation process is repeated until the new starting point reaches the end of the original text to obtain multiple text blocks.

3. The method according to claim 1, characterized in that The obtaining of all font sizes in the original text and generating a document directory structure corresponding to the original text includes: Obtaining all font sizes in the original text and establishing a mapping relationship between font sizes and directory levels; For each text block, generating a single-layer directory hierarchy structure corresponding to the text block according to the mapping relationship between the font size and the directory hierarchy; Each single-layer directory hierarchy structure is traversed to determine the hierarchical relationship of all text blocks, and a document directory structure corresponding to the original text is generated.

4. The method according to claim 1, characterized in that For each text block, generating a single-layer directory hierarchy structure corresponding to the text block according to the mapping relationship between the font size and the directory hierarchy includes: For each text block, BeautifulSoup is used to generate a single-layer directory hierarchy structure corresponding to the text block according to the mapping relationship between the font size and the directory hierarchy.

5. The method according to claim 1, characterized in that Traverse each single-layer directory hierarchy structure, determine the hierarchical relationship of all text blocks, and generate a document directory structure corresponding to the original text, including: Create an empty stack to store HTML tags; For each text block, if the single-layer directory hierarchy structure corresponding to the text block is a title, determine the level of the current title, compare the level of the current title with the elements in the stack, pop out the tag to be closed, add the start tag of the current title, and push the current title into the stack; If the single-layer directory hierarchy structure corresponding to the text block is text, then add a text tag; After all text blocks are processed, corresponding end tags are popped up and added in sequence to obtain a document directory structure of the HTML structure corresponding to the original text.

6. The method according to claim 1, characterized in that For each recalled text segment, according to the corresponding document directory structure, from each vectorized text segment corresponding to the original text, a vectorized text segment corresponding to the same level or a higher level of the recalled text segment is obtained to obtain a recalled long text that meets the recalled text length range, including: For each recalled text segment, according to the corresponding document directory structure, obtain a vectorized text segment at the same level of the recalled text segment; Determine whether the length of the obtained vectorized text segment at the same level of the recalled text segment meets the recalled text length range; If yes, the vectorized text segment at the same level as the recalled text segment is used as the recalled long text; If the length of the obtained vectorized text segment at the same level of the recalled text segment is greater than the longest recalled text length, the vectorized text segment at the same level of the recalled text segment is truncated according to the arrangement order to obtain the recalled long text having a length equal to the longest recalled text length; If the length of the vectorized text segment at the same level of the recalled text segment is less than the shortest recalled text length, then the vectorized text segment at the previous level of the recalled text segment is obtained, and it is re-determined whether the length of the vectorized text segment at the previous level of the recalled text segment is greater than the shortest recalled text length and less than the longest recalled text length, until the recalled long text that meets the recall length range is obtained, or the top level of the document directory structure is reached, and all the vectorized text segments corresponding to the original text are used as the recalled long text.

7. A document retrieval enhancement device for a large language model, characterized in that: include: A text segmentation module is used for, for each original text, taking the starting position of the original text as the starting point, searching for a text segmentation symbol according to the order of the text, and performing text segmentation based on the starting point and the text segmentation symbol and a preset heuristic rule and a sliding window to obtain a plurality of text blocks that meet the text length range; A directory structure building module, used to obtain all font sizes in the original text and generate a document directory structure corresponding to the original text; A storage module, used for storing the document directory structure corresponding to the original text and the vectorized text segments of each corresponding text block into a vector database; A first recall module, configured to perform similarity calculation between vectorized user input and vectorized text segments in the vector database, and recall top K recalled text segments that are most relevant to the user input; The second recall module is used to obtain, for each recalled text segment, a vectorized text segment corresponding to the same level or a higher level of the recalled text segment according to the corresponding document directory structure, and obtain a recalled long text that meets the recalled text length range.

8. A computer-readable storage medium storing instructions, which, when executed on a terminal, enable the terminal to execute the document retrieval enhancement method for a large language model as described in any one of claims 1 to 6.

9. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instruction is executed by a processor, the document retrieval enhancement method for a large language model described in any one of claims 1 to 6 is implemented.

10. A computer device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; Memory, used to store computer programs; The processor is used to implement the document retrieval enhancement method for a large language model as described in any one of claims 1 to 6 when executing the program stored in the memory.

Citation Information

Cited By

  • High-capacity text clue retrieval method and system based on enhanced retrieval, product and medium

    CN120336494A

  • Large-capacity text clue retrieval method, system, product and medium based on enhanced retrieval

    CN120336494B