Retrieval response method and device, storage medium, electronic equipment and program product
By generating parent-child text chunking based on the sequence termination probability segmentation and vectorization processing of data sub-blocks in the document, the problem of semantic breakage and low retrieval accuracy in traditional chunking methods is solved, and more efficient and accurate document retrieval and response generation is achieved.
Patent Information
- Application Number
- CN202510728904.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-07-04
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The traditional text chunking method leads to semantic rupture and low retrieval accuracy, and cannot effectively capture the internal logic and semantic structure of the document, resulting in the inability to efficiently and accurately locate relevant text fragments during retrieval.
By segmenting according to the sequence termination probability of different data sub-blocks in the document, N parent text blocks are generated, and paragraph splitting is performed, N×M sub-text blocks are generated, and then vectorized to form query reference vectors, and search responses are performed based on these vectors.
Improves the accuracy and efficiency of retrieval, avoids semantic breaks, ensures the consistency and completeness of generated answers, and adapts to different types of documents and scenarios.
Smart Images

Figure CN120256546A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of big data technology. Specifically, the embodiments of the present application relate to a retrieval response method, device, storage medium, electronic device, and program product. Background Art
[0002] Text chunking is the process of dividing a long document into smaller, semantically coherent text segments. This step is particularly crucial in a RAG (Retrieval-Augmented Generation) system because it directly affects the subsequent retrieval efficiency and the quality of the generated content. Traditional chunking methods mainly include: (1) Fixed-length chunking: This method divides the text according to a preset number of characters or sentences. Although it is simple to implement, it is prone to semantic breaks, especially when dealing with complex documents, which may result in incomplete information within the chunks and damaged logical relationships between chunks. (2) Rule-driven chunking: It relies on grammar and punctuation marks for chunking and can preserve semantics to a certain extent. However, for documents with loose structures or multi-topic intersections, it is still difficult to capture the deep semantic structure and is prone to introducing noise.
[0003] In addition, in related technologies, neither the fixed-length nor the rule-driven chunking method fully considers the internal logic and semantic structure of the document, resulting in the inability to efficiently and accurately locate the text segments most relevant to the query during retrieval. At the same time, during the chunking process, larger text chunks help maintain context coherence but reduce retrieval efficiency; while smaller chunks improve retrieval speed but may sacrifice semantic integrity within the chunks, making the generated answers lack context background and have illogical coherence. Moreover, they have insufficient adaptability to complex documents and cannot meet the retrieval requirements for documents in different scenarios.
[0004] In view of the problems of semantic breaks and low retrieval accuracy caused by traditional fixed chunking in related technologies, no effective solution has been proposed yet. Summary of the Invention
[0005] The embodiments of the present application provide a retrieval response method, device, storage medium, electronic device, and program product to at least solve the problems of semantic breaks and low retrieval accuracy caused by traditional fixed chunking in related technologies.
[0006] According to an embodiment of the present application, a retrieval response system is provided, including: when receiving a prompt instruction for processing an original document, segmenting the original document according to the sequence termination probabilities corresponding to different data sub-blocks in the original document to obtain N parent text blocks; performing paragraph splitting processing on each of the N parent text blocks to generate N×M sub-text blocks, where N and M are positive integers; vectorizing the N×M sub-text blocks to obtain N×M query reference vectors; and performing a retrieval response to a query question issued by a target object based on the N×M query reference vectors.
[0007] According to another embodiment of the present application, a retrieval response system is provided, including: a segmentation module, configured to segment the original document according to the sequence termination probabilities corresponding to different data sub-blocks in the original document when receiving a prompt instruction for processing the original document, to obtain N parent text blocks; a splitting module, configured to perform paragraph splitting processing on each of the N parent text blocks to generate N×M sub-text blocks, where N and M are positive integers; a vector module, configured to vectorize the N×M sub-text blocks to obtain N×M query reference vectors; and a retrieval module, configured to perform a retrieval response to a query question issued by a target object based on the N×M query reference vectors.
[0008] According to still another embodiment of the present application, a computer-readable storage medium is further provided, in which a computer program is stored, and the computer program is configured to execute the steps in any one of the above method embodiments when running.
[0009] According to still another embodiment of the present application, an electronic device is further provided, including a memory and a processor, a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0010] According to still another embodiment of the present application, a computer program product is further provided, including a computer program, and the computer program realizes the steps in any one of the above method embodiments when executed by a processor.
[0011] Through this application, after receiving a prompt instruction to process the original document, it is segmented according to the sequence termination probability of different data sub-blocks in the document to obtain N parent text chunks; then, each parent text chunk is split into paragraphs to generate N×M sub-text chunks; then, these sub-text chunks are converted into vector form to form query reference vectors; finally, a retrieval response is made to the query question sent by the target object based on these query reference vectors. Subsequently, the technical effect of improving the retrieval accuracy through the generation of sub-blocks and their association with parent blocks is achieved, and the technical problem of semantic fracture and low retrieval accuracy caused by traditional fixed chunking is solved by the above method. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] To more clearly illustrate the embodiments of the present application, the accompanying drawings required for use in the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0013] Figure 1 is a hardware structure block diagram of a server device for a retrieval response method according to an embodiment of the present application;
[0014] Figure 2 is a flowchart of a retrieval response method according to an embodiment of the present application;
[0015] Figure 3 is an overall architecture schematic diagram of a novel text chunking framework for RAG according to an embodiment of the present application;
[0016] Figure 4 is a flowchart of the principle of parent-child chunking according to an embodiment of the present application;
[0017] Figure 5 is a structure block diagram of a retrieval response device according to an embodiment of the present application;
[0018] Figure 6 is a computer system structure block diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0019] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.
[0020] It should be noted that in the description of this application, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in this application are used to distinguish similar objects, rather than to describe a specific order or sequence.
[0021] In order to enable those skilled in the art of this technology to better understand the solution of this application, the following further detailed description of this application will be given in conjunction with the accompanying drawings and specific embodiments.
[0022] As an optional implementation manner, the method embodiments provided in the embodiments of this application can be executed in a server device or a similar computing device. Taking running on a server device as an example, Figure 1 is a hardware structure block diagram of a server device for a retrieval response method according to an embodiment of this application. As Figure 1 shown, the server device may include one or more ( Figure 1 only one is shown in the figure) processors 102 (the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. Among them, the above-mentioned server device may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that, Figure 1 the structure shown is only schematic, and it does not limit the structure of the above-mentioned server device. For example, the server device may further include more or fewer components than Figure 1 shown in the figure, or have a different configuration from Figure 1 shown in the figure.
[0023] The memory 104 can be used to store computer programs. For example, software programs and modules of application software, such as the computer program corresponding to the retrieval response method in the embodiments of this application. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, the above-mentioned method is implemented. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely provided relative to the processor 102, and these remote memories can be connected to the server device through a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and their combinations.
[0024] The transmission device 106 is used to receive or send data via a network. Specific examples of the above network may include a wireless network provided by a communication provider of a server device. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 may be a Radio Frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0025] In this embodiment, a retrieval response method is provided. Figure 2 It is a flowchart of the retrieval response method according to the embodiment of the present application, as Figure 2 shown, and the process includes the following steps:
[0026] Step S202, when receiving a prompt instruction for processing the original document, segment the original document according to the sequence termination probability corresponding to different data sub-blocks in the original document to obtain N parent text chunks;
[0027] It can be understood that in the process of using a document retrieval system, the document retrieval system first receives a prompt instruction for processing the original document, and this instruction usually includes requirements for intelligent analysis and segmentation of the document. Based on this instruction, the document retrieval system inputs the text content of the original document into an LLM (Large Language Model, abbreviated as LLM). The LLM analyzes each data sub-block (such as words, phrases or sentences) through its pre-trained ability and calculates the probability of the end-of-sequence [EOS] (End-of-Sequence Probability, abbreviated as EOS) after this sub-block. This probability reflects the tendency of this sub-block to be an independent and complete semantic unit, that is, the possibility of the data sub-block to end by itself. The document retrieval system automatically adjusts a dynamic probability threshold according to the context complexity of the document, and the threshold range is set between 0.65 and 0.92. The setting of this threshold is to calculate the document complexity coefficient C through a preset formula and adjust the threshold of the EOS probability accordingly to adapt to different types of documents. When the calculated EOS probability of a certain data sub-block exceeds the dynamic threshold, it is determined as a semantic boundary node, that is, segmentation should be performed after this sub-block. This decision ensures that the selection of the segmentation point can be based on the internal semantic structure of the document rather than a fixed number of characters or sentences. Once the segmentation points are determined, the original document is segmented according to these points to generate a series of text chunks with appropriate lengths and relatively complete semantics, that is, "parent text chunks".
[0028] Through the above process, the original document is intelligently segmented into a series of parent text chunks, which not only maintain the basic semantic structure of the document but also provide a basis for subsequent refined processing (such as sub-text chunk generation) and retrieval. This method overcomes the problem of context fragmentation that may be caused by fixed-length or rule-driven chunking, improving the accuracy and efficiency of document retrieval and generating responses in the RAG system.
[0029] Step S204: Perform paragraph splitting on each of the N parent text chunks to generate N×M sub-text chunks, where N and M are positive integers.
[0030] It can be understood that the N parent text chunks are obtained by analyzing the semantic boundaries of the original document and combining dynamic probability thresholds. Each parent chunk has relatively complete semantic and context information. For each of these N parent text chunks, "paragraph splitting" is performed. This process further splits the parent text chunks into smaller semantic units, namely sub-text chunks, based on preset delimiter rules (such as common punctuation marks, line breaks, etc.) and the maximum allowed length of each sub-chunk. Each parent text chunk will be split into M sub-text chunks according to the semantic units and predefined rules within it. The specific value of M depends on the content richness, complexity, and chunking rules of the parent chunk. For example, if the parent chunk is a paragraph containing multiple complete sentences, then each sentence may form a sub-text chunk. Finally, the entire original document is segmented into N×M sub-text chunks. Here, N represents the number of parent text chunks, and M represents the average number of sub-text chunks into which each parent text chunk is further divided. Both N and M are positive integers, and the value of M depends on the content and structure of the parent chunk. By implementing the above "parent-child" chunking strategy, the retrieval efficiency and context integrity can be effectively balanced, which is a key step in improving the performance of the Retrieval-Augmented Generation (RAG) system.
[0031] Step S206: Vectorize the N×M sub-text chunks to obtain N×M query reference vectors.
[0032] Step S208: Based on the N×M query reference vectors, perform a retrieval response to the query question issued by the target object.
[0033] Optionally, convert the text in natural language into a mathematical vector representation that can be parsed and manipulated by a machine. This is a processing step in text processing and information retrieval, aiming to encode semantic information into a numerical form for subsequent calculations and matching. That is, through the use of a pre-trained large language model, vectorize the N×M sub-text chunks that have been generated, and complete the encoding of each sub-text chunk to generate a vector corresponding to its semantics. It should be noted that the above large language model is trained with a large amount of text data and can capture the deep semantic features of the text. Therefore, each sub-text chunk is encoded as a query reference vector, and these vectors represent the semantic information of the sub-chunks. After vectorization, each sub-chunk of the original document is converted into a point in a high-dimensional space, and the position of this point reflects the features and meaning of the text. When receiving a query question from a target object (such as a user), first parse and preprocess the query, and then use a similar vectorization technique to convert the query question into one or more vector representations, that is, question vectors. Then, use the N×M query reference vectors in the vector database to perform matching retrieval on the user's question vectors. For example, the above matching can be performed by calculating the similarity (such as cosine similarity) between the question vector and each vector in the database to find the most relevant sub-text chunk.
[0034] Furthermore, according to the matching results, select the most relevant sub-text chunks, and restore the context through the parent text chunk information they carry. When necessary, merge the parent chunk information corresponding to multiple relevant sub-chunks to ensure the coherence and integrity of the answer. Then, use this information as input and use a generation model such as an LLM to generate an accurate and detailed answer. Finally, present the generated response content to the user, realizing the full automation of the whole process from user query to document information extraction, processing, and return. Thus, significantly improving the retrieval efficiency and answer accuracy of the RAG system.
[0035] Through the above method, after receiving the prompt instruction to process the original document, split it according to the sequence termination probability of different data sub-chunks in the document to obtain N parent text chunks; then, perform paragraph splitting on each parent text chunk to generate N×M sub-text chunks; then, convert these sub-text chunks into vector form to form query reference vectors; finally, perform retrieval response on the query question sent by the target object based on these query reference vectors. Subsequently, the technical effect of improving the retrieval accuracy through the generation of sub-chunks and the association with the parent chunks is achieved, and the technical problem of semantic breakage and low retrieval accuracy caused by traditional fixed chunking is solved through the above method.
[0036] In an exemplary embodiment, the original document is segmented according to the sequence termination probabilities corresponding to different data sub - blocks in the original document to obtain N parent text chunks, including: obtaining the first sequence termination probabilities corresponding to different data sub - blocks in the original document to get P first sequence termination probabilities, where P is a positive integer greater than or equal to N; screening out N second sequence termination probabilities greater than the dynamic probability threshold from the P first sequence termination probabilities, where the dynamic probability threshold is determined by calculating the complexity of the original document through a preset first formula; obtaining the sub - block positions of the N data sub - blocks corresponding to the N second sequence termination probabilities, and segmenting the original document based on the sub - block positions to obtain N parent text chunks.
[0037] Briefly, first calculate the sequence termination probability (EOS probability) for each data sub - block in the original document, so as to analyze the tendency of each sub - block (such as a sentence or a paragraph) to end as an independent semantic unit. Through this process, P first sequence termination probabilities are obtained, where P is a positive integer, and the size of P is at least equal to the number N of the finally generated parent text chunks.
[0038] Optionally, in order to determine which data sub - blocks truly constitute complete semantic units, a dynamic probability threshold needs to be set. This threshold is not fixed, but is dynamically adjusted according to the complexity of the original document. The complexity coefficient of the document can be calculated through a preset first formula, which needs to consider various attributes of the document, such as length, topic diversity, density of professional vocabulary, etc., so as to more accurately identify semantic boundaries.
[0039] From the P first sequence termination probabilities, screen out N second sequence termination probabilities whose probability values are greater than the dynamic threshold. The N data sub - blocks corresponding to these N probability values are regarded as having relatively high semantic integrity and are potential segmentation points. Further obtain the specific positions of these N data sub - blocks in the document, and segment the original document based on these positions, so as to obtain N relatively complete semantic units, that is, parent text chunks.
[0040] In summary, through the above - mentioned segmentation logic, it is ensured that the parent text chunks can be generated according to the natural boundaries of the document content, rather than simply being mechanically segmented according to a preset length or rule. By dynamically adjusting the threshold of the EOS probability, the document structure can be more intelligently identified, avoiding problems such as semantic breakage or information redundancy, and improving the efficiency and accuracy of subsequent retrieval and generation tasks.
[0041] In an exemplary embodiment, before screening out N second sequence termination probabilities greater than the dynamic probability threshold from P first sequence termination probabilities, the above method further includes: performing text parsing on the original document; determining the context complexity coefficient between different data sub-blocks in the original document according to the parsing result; using a preset first formula to perform calculation processing on the context complexity coefficient to obtain the dynamic probability threshold, where the preset first formula is: θ = max(0.65, min(0.92 - 0.27C, 0.92)), C is the context complexity coefficient, and θ is the dynamic probability threshold.
[0042] Optionally, first perform text parsing on the original document, which is a preprocessing step aimed at understanding the structure and content of the document. This may include identifying the format of the document, parsing elements such as words, phrases, sentences, paragraphs, etc. in the document, and understanding the logical relationships between them. Based on the text parsing, further analyze the context relationships between different data sub-blocks (such as sentences, paragraphs) in the document, and calculate a coefficient C that reflects the overall complexity of the document. The level of the context complexity coefficient can reflect the tightness of the internal logic of the document and the richness of semantic information. For example, a scientific and technological paper containing many professional terms, multi-topic discussions, and complex logical inferences may have a higher complexity coefficient C than a simple daily report. According to the calculated context complexity coefficient C, use the preset first formula to calculate the dynamic probability threshold θ; when C approaches 0 (indicating relatively simple context), the formula tends to (0.92), so the value of θ is higher, which implies that the system will be more cautious during segmentation and tend to perform segmentation at the clear boundaries of the text to avoid unnecessary segmentation and maintain larger semantic units. When C approaches 1 (indicating very complex context), the formula tends to (0.65), so the value of θ is lower, which indicates that the system will be more sensitive to the EOS probability in complex contexts and tend to perform more frequent segmentation to ensure that each segmented sub-block can maintain semantic integrity as much as possible, even if this means generating more segmentation points.
[0043] After calculating the dynamic probability threshold θ, screen out those second sequence termination probabilities greater than θ from the previously calculated P first sequence termination probabilities. This screening process is crucial because it determines which data sub-blocks will be regarded as independent semantic units and will ultimately be segmented out as part of the parent text chunk. A higher value of θ means a stronger requirement for semantic integrity, and a lower value of θ allows for more flexible segmentation decisions in complex contexts.
[0044] In summary, through the above content, it is possible to dynamically adjust the segmentation strategy based on fully considering the complexity of the document context, so as to generate the most suitable series of parent text chunks. This can not only ensure the rationality and effectiveness of segmentation, but also avoid the problems of context fragmentation or information redundancy caused by fixed thresholds, providing a more accurate and efficient text processing method for the Retrieval-Augmented Generation (RAG) system.
[0045] In an exemplary embodiment, before performing segmentation processing on the original document based on the sub-block positions to obtain N parent text chunks, the above method further includes: performing segmentation detection on the original document to determine the cumulative number of characters and the number of titles corresponding to the original document; adding a forced segmentation flag to the original document when the cumulative number of characters is greater than or equal to a preset number of characters and the number of titles is greater than or equal to a preset number; generating a first prompt message indicating that there is no forced segmentation flag for the original document when the cumulative number of characters is less than the preset number of characters or the number of titles is less than the preset number.
[0046] That is to say, first perform segmentation detection on the original document, which includes counting the cumulative number of characters in the document and identifying the number of titles. This is to better understand the structure and content complexity of the document, so as to determine whether forced segmentation measures need to be taken.
[0047] Furthermore, by setting certain thresholds for the number of characters and the number of titles. If it is detected that the cumulative number of characters in the original document reaches the preset upper limit of the number of characters (for example, 2000 characters), or the number of titles in the document exceeds the preset number of titles (which usually means that the document is clearly divided into different chapters or sections), it is considered that the complexity of the document is sufficient to affect the subsequent segmentation and retrieval efficiency. Subsequently, in this case, a forced segmentation flag will be added to the original document. This flag guides the subsequent segmentation process to perform segmentation whenever a chapter title is encountered or the cumulative number of characters reaches the upper limit, regardless of the EOS probability, to generate new parent chunks. This can ensure that the length of each parent chunk is appropriate, and at the same time respects the original structure and segmentation intention of the document.
[0048] On the contrary, if the cumulative number of characters is less than the preset upper limit of the number of characters and the number of titles is also lower than the preset number, a first prompt message will be generated indicating that the current document does not require forced segmentation. This means that the segmentation decision will be completely based on the EOS probability, and will focus more on maintaining semantic coherence and context integrity, rather than mechanically following structural segmentation. That is, the above preprocessing steps reflect the sensitivity to the document complexity. By dynamically adjusting the forced segmentation strategy, it is possible to effectively control the size of the parent chunks while maintaining semantic coherence, improving the retrieval efficiency and the accuracy of the generated content. Especially for documents with a clear structure and titles, forced segmentation can ensure that each chunk does not cross the title boundary, thus avoiding information chaos.
[0049] In summary, by preprocessing the original document, it is possible to intelligently determine when and where to perform forced segmentation, and when to rely on the EOS probability based on the LLM for semantic boundary recognition. This hybrid strategy significantly improves the flexibility and accuracy of document segmentation.
[0050] In an exemplary embodiment, before performing segmentation processing on the original document according to the sequence termination probabilities corresponding to different data sub-blocks in the original document, the above method further includes: comparing a plurality of identification positions corresponding to the forced segmentation identifiers in the original document with N sub-block positions greater than the dynamic probability threshold; setting the target positions where the identification positions and the sub-block positions overlap as segmentation monitoring points according to the comparison results.
[0051] Optionally, check whether there are any forced segmentation identifiers in the original document. These identifiers may be manually added based on the structural characteristics of the document (such as chapter headings, specific format changes, etc.) or content complexity (such as text length exceeding a certain threshold), indicating that the system needs to perform segmentation here regardless of the EOS probability. Subsequently, determine the positional relationship between the specific positions corresponding to the forced segmentation identifiers in the document and the N sub-block positions determined by calculating the EOS probability. Here, the N sub-block positions refer to those positions where the EOS probability exceeds the dynamic probability threshold, and these positions are regarded as natural semantic boundaries suitable for segmentation. During the comparison process, if it is found that a certain position of the forced segmentation identifier overlaps with the sub-block position determined based on the EOS probability, then the overlapping position will be specially marked as the target position and set as the segmentation monitoring point. This means that segmentation will be preferentially considered here, which not only meets the need for forced segmentation but also takes into account the requirements of semantic integrity and context coherence.
[0052] That is to say, by setting up segmentation monitoring points, more flexible and reasonable segmentation decisions can be made when processing complex documents. For example, if a long paragraph contains multiple potential EOS segmentation points, but according to the content structure, a clear chapter heading is recognized, then the position of the chapter heading will become the preferred segmentation monitoring point to ensure that the chapter boundary will not be damaged by the segmentation algorithm.
[0053] In summary, by combining the natural segmentation points of the EOS probability and the monitoring points of the forced segmentation identifiers, comprehensive segmentation processing is performed on the original document. This ensures that the parent text chunks can not only follow the internal logic and semantics of the document but also follow the externally set rules when necessary, thereby improving the efficiency of subsequent retrieval and content generation while maintaining semantic integrity. Thus, the purpose of achieving efficient and accurate text segmentation in various document types and scenarios is achieved.
[0054] In an exemplary embodiment, paragraph splitting is performed on each of the N parent text chunks to generate N×M sub - text chunks, including: performing paragraph splitting on the N parent text chunks using a preset delimiter rule and the maximum block length allowed by the minimum semantic unit to obtain N×M sub - paragraphs; adding the block identifier corresponding to the data sub - block to which the parent text chunk belongs and the position offset corresponding to the data sub - block to which the parent text chunk belongs to each of the N×M sub - paragraphs to generate N×M sub - text chunks.
[0055] Optionally, for each parent text chunk, paragraph splitting is performed using a preset delimiter rule and the maximum block length limit of the sub - paragraphs. Here, the delimiter rule may include, but is not limited to, splitting based on common punctuation marks, line breaks, or specific text patterns. Each parent chunk is thus subdivided into multiple sub - paragraphs, and the lengths of these sub - paragraphs generally do not exceed the preset maximum block length, e.g., 400 characters. After completing the paragraph splitting, an additional information - adding operation is performed on each sub - paragraph. For each generated sub - text chunk, the block identifier of the data sub - block from its parent text chunk and the position offset relative to the original document are added. The block identifier helps track which specific parent text chunk the sub - chunk belongs to, which is crucial for subsequent retrieval and information recombination. The position offset provides the exact position of the sub - chunk in the original document, helping to restore the context. The finally obtained N×M sub - text chunks are not only independent text segments but also carry metadata information about their sources, including the parent block identifier they belong to and the starting character position relative to the original document. This property enables efficient localization of the specific context in the original document, even when only the sub - chunk is matched during the retrieval process.
[0056] It can be understood that by generating fine - grained sub - text chunks, the retrieval speed and accuracy during subsequent retrieval calls are guaranteed to more quickly find the specific information related to the question. Secondly, by retaining the information of the parent text chunk to which each sub - chunk belongs, it can also ensure that when generating a response, sufficient rich context is provided, avoiding information fragmentation and ensuring the integrity and coherence of the answer.
[0057] In summary, the process of performing paragraph splitting on the parent text chunks and generating sub - text chunks is an important pre - processing step in the entire RAG system, aiming to balance retrieval efficiency and context integrity to ensure that the system can not only respond quickly to queries but also provide accurate and coherent information.
[0058] In an exemplary embodiment, paragraph splitting of N parent text chunks is performed using a preset delimiter rule and the maximum block length allowed by the minimum semantic unit, including: parsing the preset delimiter rule, determining the default number of segment identifiers according to the parsing result, and the number of characters between adjacent segment identifiers; performing segment verification on the N parent text chunks respectively according to the default number, the number of characters, and the maximum block length; and determining paragraph splitting of the N parent text chunks based on the verification result.
[0059] Optionally, by parsing the preset delimiter rule, define the segment identifiers included in the rule (such as common full stops, line breaks, etc.) and the default number of corresponding identifiers. For example, it may be default to divide the text into multiple paragraphs according to the line breaks appearing in each paragraph of text. At the same time, the number of characters between adjacent segment identifiers is also set as a benchmark for measuring the paragraph length.
[0060] Based on the parsed delimiter rule, segment verification is performed on each parent text chunk. The verification process takes into account the number of segment identifiers, the number of characters between adjacent identifiers, and the maximum allowed length of each segment. Through this verification, it can be confirmed whether each parent text chunk meets the preset segment standard, that is, whether it can be reasonably split according to the set identifiers and lengths. After segment verification, the actual paragraph splitting of the parent text chunk is determined based on the verification result. If the text within the parent chunk can be reasonably split into multiple small paragraphs that meet the maximum block length limit according to the delimiter rule, the splitting operation will be performed to generate a set of sub-text chunks. Each sub-text chunk will be assigned a block identifier to indicate the parent text chunk it belongs to, and the position offset relative to the original document will be recorded to maintain context information.
[0061] Optionally, its splitting behavior is dynamically adjusted according to the preset maximum block length limit. This means that even if there are multiple segment identifiers in the parent text chunk, it is still ensured that each generated sub-chunk will not exceed the maximum allowed length, thereby controlling the size of the retrieval unit while maintaining the integrity of the semantic unit and avoiding the problem of decreased retrieval efficiency caused by overly large information chunks.
[0062] In summary, through the above paragraph splitting operation, the original document is effectively decomposed into a series of smaller and more manageable sub-text chunks. These chunks can not only retain sufficient context coherence but also meet the requirements of retrieval efficiency. The generation of sub-chunks not only simplifies the retrieval process but also ensures that when the user's query involves specific details, relevant and accurate content can be quickly located and provided.
[0063] In an exemplary embodiment, determining paragraph splitting for N parent text chunks based on the verification result includes: when the verification result indicates that the current parent text chunk corresponds to a paragraph-type text, determining to perform paragraph splitting on the current parent text based on individual sentences in the paragraph, where each individual sentence corresponds to a sub-text chunk; when the verification result indicates that the current parent text chunk corresponds to a full-text-type text, determining to perform paragraph splitting on the current parent text based on each individual sentence in the full text, where each individual sentence corresponds to a sub-text chunk.
[0064] Simply put, when the verification result indicates that the current parent text chunk is a paragraph-type text, a finer-grained splitting strategy will be adopted. In this case, each independent sentence in the current parent text chunk is regarded as the basis for a sub-text chunk. This means that a parent chunk will be split into multiple sub-chunks, and each sub-chunk contains only one complete sentence. This strategy is applicable to paragraphs containing multiple independent and complete sentences, which can ensure the semantic integrity of each sentence and provide a more precise retrieval unit for subsequent retrieval operations. On the contrary, if the verification result indicates that the current parent text chunk is a full-text-type text, that is, the entire chunk contains a continuous text paragraph without obvious paragraph divisions. Paragraph splitting will instead be performed based on each individual sentence in the full text. Similarly, each individual sentence will form a sub-text chunk. This strategy is particularly applicable to texts with a relatively loose structure or without clear section divisions, such as free-form conversation records, continuous descriptive paragraphs, etc. By splitting each sentence into a sub-chunk separately, it can better meet the retrieval needs of such documents and improve the accuracy of matching.
[0065] It should be noted that regardless of whether it is processing paragraph-type text or full-text-type text, sentence boundary detection is the basis for generating sub-text chunks. Use a large language model to parse the text and identify the start and end points of each sentence. After detecting the sentence boundary, splitting will be performed at this boundary to generate sub-text chunks, and the position information of each sub-chunk in the original document will be recorded, such as the identifier of the parent text chunk it belongs to and the start offset.
[0066] In summary, through a flexible paragraph splitting strategy, it is ensured that both clearly structured paragraphs and continuous full-text-type texts can be effectively subdivided into a series of sub-text chunks. These sub-text chunks not only carry complete sentence information but also retain their context associations in the original document, thereby improving retrieval efficiency while also ensuring the accuracy and integrity of the generated content.
[0067] In an exemplary embodiment, before splitting the original document according to the sequence termination probabilities corresponding to different data sub - blocks in the original document, the above - mentioned method further includes: determining whether there is delimiter custom information in the prompt instruction; in the case where there is delimiter custom information, summarizing the initial delimiters and the newly added delimiters corresponding to the delimiter custom information, and updating the delimiter set in the preset delimiter rules according to the summarization result; in the case where there is no delimiter custom information, generating a second prompt message for the delimiter set in the preset delimiter rules that is not updated.
[0068] Briefly speaking, first, parse the prompt instruction for guiding text chunking to determine whether it contains custom information about delimiters. Delimiters here refer to a set of special symbols or patterns used to identify the boundaries of text paragraphs or sentences, such as common full stops (.), question marks (?), exclamation marks (!), or line breaks ( ) etc. If delimiter custom information is found in the prompt instruction, these user - specified delimiters will be combined with the predefined initial delimiter set to create an updated delimiter set. This set contains the default delimiters and the newly added delimiters specified by the user according to the specific document characteristics or personal preferences, providing more refined control for subsequent splitting operations. With the updated delimiter set, adjust the preset delimiter rules according to this result. To ensure that when splitting the original document, in addition to following the semantic boundary detection based on EOS probability, the user - defined delimiter criteria will also be considered, so that the splitting result is closer to the true structure of the document and the user's needs.
[0069] Optionally, in the case where no delimiter custom information is detected, no update will be made to the preset delimiter rules. At this time, a second prompt message will be generated, indicating that the document will be split according to the default delimiter set and rules. This ensures that when specific custom information is lacking, the splitting process can still be carried out based on widely applicable criteria without falling into the dilemma of missing parameter configuration.
[0070] In summary, the above - mentioned process reflects the flexibility and robustness of the system in the splitting strategy. By allowing users to customize delimiters, it can better adapt to different types of documents and specific application scenarios, improving the accuracy of splitting. When there is a lack of custom information, relying on the preset rules built into the system ensures the basic efficiency and rationality of document splitting, enhancing the versatility of the system.
[0071] In an exemplary embodiment, after performing paragraph splitting processing on each of the N parent text chunks to generate N×M sub - text chunks, the above - mentioned method further includes: establishing a data mapping relationship among the original document, the N parent text chunks, and the N×M sub - text chunks; storing the data mapping relationship in a preset database.
[0072] It is understandable that a multi - layer index or mapping table is created to map the relationship between the original document and N parent text chunks, as well as between N parent text chunks and N×M sub - text chunks. The detailed information of each sub - chunk is recorded, such as the ID (identity document, unique code, simply referred to as ID) of the parent chunk it belongs to, the starting position, ending position in the original document, and the text content within the chunk. This mapping relationship ensures that even when only sub - chunks are involved in retrieval or content generation, it is possible to quickly trace back to the corresponding parent chunk and even the original document to obtain complete context information. In addition, to ensure the smoothness of the overall retrieval process, the established mapping relationship will subsequently be stored in a preset database. The preset database here can be a database specifically designed to store document structure information, such as a relational database or a document storage system. Storing these mapping relationships is for the convenience of subsequent retrieval operations. When a user query is received, it can quickly locate the relevant sub - text chunks and trace back to the complete parent text chunks and the original document through the mapping relationships in the database, so as to provide a comprehensive and coherent information response.
[0073] It should be noted that the stored data mapping relationship plays a bridging role in the retrieval process. When a user asks a question related to a certain sub - chunk, first, based on the embedding vector of the sub - chunk, a match is made to determine the most relevant sub - chunk. Subsequently, by querying the mapping relationship in the database, the information of the parent text chunk to which the sub - chunk belongs is obtained, including its exact position in the original document. This enables not only the provision of directly relevant sub - chunk information but also the supplementation of additional context when generating an answer, ensuring the coherence and integrity of the answer.
[0074] Optionally, by storing these mapping relationships in the database, the hierarchical structure of the document is effectively maintained. Even when the document is chunked and reconstructed during processing, its content's logical chain and context association can be maintained. This is crucial for processing complex documents such as technical manuals, research reports, or legal documents, as these documents often rely on their internal structure and context information.
[0075] Optionally, the establishment and storage of the data mapping relationship also provide a basis for further optimization and functional expansion. For example, by analyzing these mapping relationships, important paragraphs or topics in the document can be identified, so that these parts can be given priority in retrieval, improving the retrieval efficiency. In addition, this structured storage method also facilitates maintaining the internal chunk relationship when the document is updated or modified subsequently, ensuring the coherence of the document processing flow.
[0076] In summary, by establishing and storing data mapping relationships between the original document, the parent text chunks, and the sub - text chunks, it is possible to ensure that even when dealing with a large number of fine - grained chunks during document processing, the integrity of information and the coherence of context can be maintained, thus providing a high - quality retrieval - enhanced generation service.
[0077] In one exemplary embodiment, retrieving a response to a query problem issued to a target object based on N×M query reference vectors includes: performing text parsing on the query problem to obtain Q query sub - problem texts, where Q is a positive integer; vectorizing the Q query sub - problem texts to obtain Q problem vectors; performing vector matching on the Q problem vectors and the N×M query reference vectors, and retrieving a response to the query problem issued to the target object based on the configuration result.
[0078] In summary, by converting the user query into a vector representation and matching it with the pre - processed vectors in the document library, it is possible to quickly locate the most relevant information in a large number of documents. At the same time, through context restoration and combination, the generated answer is ensured to be both accurate and have good coherence, meeting the high requirements for retrieval quality and content generation under the RAG framework.
[0079] In one exemplary embodiment, performing vector matching on the Q problem vectors and the N×M query reference vectors, and retrieving a response to the query problem issued to the target object based on the configuration result includes: determining Q target query reference vectors whose vector similarity to the Q problem vectors is greater than a preset similarity according to the vector matching result; determining corresponding target parent text chunks according to the chunk identifiers of the parent text chunks corresponding to the Q target query reference vectors and the position offsets corresponding to the parent text chunks; sending the target parent text chunks and the query problem to a large - language model for background information supplementation; determining the supplementation result as the response content for retrieving a response to the query problem issued to the target object.
[0080] Optionally, compare the Q query sub-questions (each query sub-question has been transformed into a vector representation) proposed by the user with the N×M query reference vectors stored in the database. This is achieved by calculating the similarity scores between the two sets of vectors, and common similarity metrics include cosine similarity, Euclidean distance, etc. Use the reference vectors whose similarity to each question vector exceeds a preset threshold. These reference vectors are regarded as potential sources of relevant information. After determining the query reference vectors with relatively high relevance, the next step is to locate the sub-text chunks from which these vectors are derived and find the corresponding parent text chunks. This is because each query reference vector not only contains the semantic information of the sub-text chunk but also stores the chunk identifier of its parent text chunk and the position offset of the sub-chunk within the parent chunk. Through this metadata, the system can accurately trace the exact location of the relevant paragraphs in the original document. Once the relevant target parent text chunks are located, send these parent text chunks, along with the user's original query question, to a large language model for further processing. The large language model utilizes its rich language understanding and generation capabilities, not only to supplement the complete background information to make the answer more coherent and complete but also to generate a more accurate and detailed answer based on the context of the query question.
[0081] Finally, the system determines the complete information supplemented and generated by the large language model as the response content and returns it to the user. This process may also include sorting and filtering the generated response content to ensure that the information received by the user is the most relevant and valuable. Through this series of steps, the system under the RAG framework can not only provide accurate information retrieval based on the user's query but also, through the intelligent generation of the large language model, provide coherent and detailed background information, greatly enhancing the user experience and retrieval efficiency.
[0082] In summary, by implementing the process of vector matching, locating the parent chunk, supplementing the background by the large language model, and generating and sending the response content, it is possible to accurately locate the required content in the vast amount of information, while ensuring that the answer is both accurate and coherent, perfectly balancing the requirements of efficiency and quality.
[0083] In an exemplary embodiment, after determining the corresponding target parent text chunks according to the chunk identifiers of the parent text chunks corresponding to the Q target query reference vectors and the position offsets corresponding to the parent text chunks, the above method further includes: when the Q target query reference vectors correspond to the same original document in the preset database, merging the multiple target parent text chunks corresponding to the Q target query reference vectors to obtain a partial document; performing a retrieval response to the query question based on the partial document.
[0084] In short, when processing multiple relevant retrieval results from the same original document, generating partial documents through a merging strategy can significantly enhance the coherence and integrity of the answers, providing users with a retrieval response that is closer to the real context and rich in content. This process not only improves the efficiency of retrieval but also greatly enhances user satisfaction because the information received by users is not just isolated facts but a comprehensive answer that can reflect the overall context and deep semantics of the document.
[0085] In an exemplary embodiment, after the original document is segmented according to the sequence termination probabilities corresponding to different data sub - blocks in the original document to obtain N parent text chunks, the method further includes: performing text verification on the original partial document content corresponding to the parent text chunks, where the text verification is used to determine the integrity of the original partial document content relative to the original document and the coherence degree values of different statements in the original partial document; determining whether to re - segment the original document based on the verification result.
[0086] Text verification is an inspection work carried out after the original document is segmented into N parent text chunks. Its main purpose is to evaluate whether the content of the part of the original document represented by each parent text chunk is complete and the coherence degree between the statements within the chunk. Integrity refers to whether the chunk contains sufficient information to completely express a semantic unit, while the coherence degree value measures the logical relationship and fluency between the statements within the chunk. Through text verification, it can be ensured that the segmented document chunks neither split key information nor lose good internal coherence, providing high - quality input for subsequent retrieval and generation processes.
[0087] Optionally, the specific implementation of text verification may include but is not limited to the following methods:
[0088] Semantic integrity analysis: Check whether each parent text chunk constitutes an independent and complete semantic unit through a large - language model (LLM) or a pre - trained text coherence evaluation model.
[0089] Statement coherence evaluation: Use a language model to evaluate the coherence between statements within the chunk, which may involve analyzing connecting words, logical relationships, etc. between sentences to ensure the fluency and coherence of the chunk content.
[0090] Information density detection: Check whether the chunk content is dense or redundant to ensure that the segmented chunks contain key information while avoiding unnecessary redundancy and maintaining the refinement of information.
[0091] Processing of verification results: Once the text verification is completed, it is determined whether the original document needs to be re-segmented based on the verification results. If the verification shows insufficient integrity or coherence of the chunks, it may indicate that the current segmentation strategy fails to effectively capture the structural features or semantic unit boundaries of the document. In this case, the segmentation parameters, such as the EOS probability threshold, may be adjusted, or more detailed segmentation rules may be introduced to re-segment the original document.
[0092] For example, if it is found that certain types of documents perform poorly under specific segmentation parameters, these parameters can be adaptively adjusted, or different segmentation rules can be adopted for different types of documents to improve the overall segmentation quality. This dynamic optimization based on verification results can help the system better adapt to various document structures and corpus characteristics, enhancing the performance and user experience of the entire RAG framework.
[0093] In summary, by performing text verification and adjusting the document segmentation strategy based on the verification results, the logical structure and information units of the document can be captured more accurately, thus providing more accurate and coherent results in the retrieval and information generation processes. It ensures that user queries can be answered based on accurate and complete information, improving the overall retrieval performance and user satisfaction.
[0094] In an exemplary embodiment, determining whether to re-segment the original document based on the verification results includes: determining to re-segment the original document when the verification results indicate that the integrity is less than a preset integrity threshold or the coherence degree value is less than a preset coherence degree threshold; determining not to re-segment the original document when the verification results indicate that the integrity is greater than or equal to the preset integrity threshold or the coherence degree value is greater than or equal to the preset coherence degree threshold.
[0095] Optionally, the above decision-making process is essentially a dynamic optimization process, which allows the segmentation strategy to be adjusted according to the actual segmentation effect until the integrity and coherence degree values of all parent text chunks meet the preset standards. Through continuous verification and necessary re-segmentation, the quality of document segmentation can be gradually improved, ensuring that the final chunks can not only meet the granularity requirements of retrieval but also maintain the coherence and integrity of information, thus enhancing the overall performance and user satisfaction of the RAG framework.
[0096] In summary, based on the verification results of integrity and coherence degree, it is intelligently determined whether the original document needs to be re-segmented. This mechanism ensures that the segmented parent text chunks can serve as high-quality basic units for retrieval and information generation, effectively supporting efficient and coherent content retrieval and service generation under the RAG framework.
[0097] In an exemplary embodiment, after retrieving and responding to a query question sent to a target object based on N×M query reference vectors, the method further includes: obtaining evaluation information made by the target object on the retrieval response, and the overall response duration of the query question; analyzing an effectiveness parameter of the retrieval response according to the evaluation information and the overall response duration, where the effectiveness parameter at least includes: retrieval delay duration, retrieval hit rate.
[0098] By actively obtaining the immediate feedback of the target object on the retrieval response. Such feedback can be ratings, comments or satisfaction surveys provided by the user through the interface, or can be indirect, such as whether the user will further submit queries, or whether they stay on a certain retrieval result page for a long time. Evaluation information is a key indicator for measuring whether the retrieval response meets the user's needs, and whether the content is accurate and useful. Optionally, the overall response duration refers to the time span from when the user submits a query until the retrieval result is returned. This includes all processing times from text parsing, vectorization, matching to finally generating a response. The response duration reflects the reaction speed and processing efficiency, which has a direct impact on the user experience, especially for application scenarios that require immediate feedback.
[0099] Based on the collected evaluation information and the overall response duration, analyze a series of parameters reflecting the effectiveness of the retrieval response to evaluate and optimize the performance.
[0100] Optionally, the retrieval delay duration refers to the time interval from when the user submits a query to when a preliminary response is returned. A low delay duration is one of the key goals pursued by the RAG framework because it directly affects the user's waiting experience. By analyzing user feedback and response time, it is possible to identify which operations or strategies increase the delay, and thus take measures to reduce the processing time. Retrieval hit rate: The hit rate of the retrieval response reflects the matching degree between the returned document chunks or information fragments and the user's query. A high hit rate means that the user's needs can be accurately understood and responded to. Conversely, the retrieval algorithm or chunking strategy needs to be adjusted to improve the relevance and satisfaction of the retrieval results.
[0101] In summary, by regularly analyzing and monitoring effectiveness parameters, performance bottlenecks such as too long processing time and low hit rate can be identified, and corresponding optimization measures can be taken. For example, if it is found through analysis that the retrieval delay duration exceeds the expectation, it may be necessary to optimize the vectorization process or upgrade the hardware resources; if the hit rate is not high, then it may be necessary to improve the chunking strategy or enhance the intelligence of the retrieval algorithm. In addition, the relevance threshold, retrieval strategy priority, etc. can also be adjusted according to user feedback to improve the overall service effect and user experience.
[0102] Among them, the execution subject of the above steps can be a server, a terminal, etc., but is not limited thereto.
[0103] To facilitate the understanding of the implementation of this application, the relevant scenarios are now explained, but they do not limit this application.
[0104] To better understand the technical solution of this application, the relevant technical terms are now explained, but they do not limit this application.
[0105] RAG (Retrieval-Augmented Generation): Retrieval-Augmented Generation technology, which combines external knowledge retrieval with a generation model to improve the accuracy and relevance of the generated content.
[0106] Semantic Chunking: Based on the semantic features of the text (such as sentence embedding similarity), the document is segmented into semantically coherent fragments, rather than relying on fixed rules.
[0107] Parent-Child Chunking: A hierarchical chunking structure where the parent chunk is a larger semantic unit (such as a paragraph or a topic module), and the child chunk is a fine-grained fragment (such as a sentence or a phrase), supporting multi-level retrieval and generation.
[0108] Dynamic Chunking Algorithm: An algorithm that dynamically adjusts the chunk boundaries according to semantic relevance, context coherence, and chunk length.
[0109] LLM: Large Language Model, a large language pre-training model (such as GPT, etc.) used to predict the probability of text boundaries.
[0110] EOS: End-of-Sequence Probability, the probability value of the end of a sentence predicted by the language model, used to determine the chunk boundaries.
[0111] Tokens: In the context of natural language processing (NLP) and large language models (LLMs), Token (mark) is a basic concept used to indicate the smallest semantic unit for text processing. Tokens represent the number corresponding to the smallest semantic unit. For example, for the English sentence: "ChatGPT-4 is powerful!", it is tokenized as: ["Chat","G","PT","-4","is","powerful","!"], a total of 7 Tokens.
[0112] As an optional implementation, the optional embodiment of this application provides a hybrid chunking method that combines LLM semantic chunking and parent-child chunking. This hybrid chunking method can be divided into two major parts. The first is the semantic chunking module, and the second is the parent-child chunking module.
[0113] Optionally, the semantic chunking module relies on the deep parsing ability of large-scale pre-trained language models for text semantics and achieves intelligent chunking by analyzing the sequence termination prediction signals output by the model. Specifically, the model's judgment on text continuity is stimulated through specific instruction templates to obtain the conditional probability distribution of the post-terminators ([EOS]) of each language unit. This probability value objectively reflects the model's cognitive tendency regarding whether the current text constitutes a complete semantic unit. When the probability value exceeds the dynamic threshold, it is determined as a semantic boundary node.
[0114] Compared with traditional semantic segmentation algorithms, the semantic chunking method based on large language models (LLMs) shows a fundamental breakthrough in technical principles and application effects. Traditional methods usually rely on manually set rules or local semantic features. For example, the segmentation point is detected by the similarity fluctuation of adjacent sentence embedding vectors (such as the Text Tiling algorithm), or mechanical segmentation is performed based on surface features such as punctuation marks and paragraph indents. Such methods perform well in structured texts (such as the chapter division of technical manuals), but in scenarios that require in-depth understanding of context associations, such as complex logical derivations and literary descriptions, they often result in fragmentation due to the inability to capture long-distance semantic dependencies.
[0115] Optionally, the parent-child chunking module improves the accuracy of the retrieval system by dividing the document into multi-level text chunks and using vector embedding technology. It is a multi-granularity chunking strategy aimed at solving the contradiction between "retrieval accuracy" and "context integrity" in traditional chunking methods through hierarchical text segmentation and collaborative retrieval mechanisms. The core principle is to balance semantic integrity and retrieval efficiency through a hierarchical segmentation strategy. Specifically, the document is first divided into parent text chunks Parent Chunks (about 2000 characters) containing multiple paragraphs or chapters to retain context coherence, and then further divided into smaller child text chunks Child Chunks (about 400 characters) to improve retrieval accuracy. The child text chunks generate semantic vectors through a pre-trained model (such as BERT, "Bidirectional Encoder Representations from Transformers", a pre-trained model in the field of natural language processing). After matching with the query vector, the relevant parent text chunks are located to ensure that the retrieval results are both accurate and have complete context. By merging the parent text chunks Parent Chunks corresponding to the retrieved child text chunks Child Chunks and filtering based on relevance ranking, coherent information that meets the user's needs is finally output, effectively solving the semantic fragmentation and information redundancy problems caused by traditional chunking.
[0116] Optionally, Figure 3It is a schematic diagram of the overall architecture of a new text chunking framework for RAG according to an embodiment of the present application. The core of the above overall framework includes: an LLM semantic chunking module 32 and a parent-child chunking module 34. Among them, the above LLM semantic chunking module 32 is used to determine the end-of-sequence probability EOS identification of different paragraph sequences for the original document according to the input prompt words; the above parent-child chunking module 34 is used to match the database reference vector stored after the sub-chunk is vectorized previously with the user's vectorized query vector when vectorizing the user's question and retrieving it in the database, and obtain the parent chunk corresponding to the sub-chunk (equivalent to the sub-text chunk in the above embodiment) corresponding to the matched reference vector based on the matching result, so as to complete the complete background information of the user's question based on the rich content corresponding to the parent chunk, and thus give a more accurate response answer.
[0117] Further, when the above new text chunking framework for RAG runs, the following steps are executed:
[0118] Step 1: For the original document in multiple formats such as pdf, word, and ppt, use the LLM to calculate the sentence end probability by inputting the prompt word "You are an expert in semantic boundary detection. You need to perform semantic integrity analysis on the input text sequence and identify the boundaries of complete semantic units by calculating the probability of the [EOS] marker appearing at each position. The output format is a JSON array containing the EOS probability values for each token position." When the probability value of the end-of-sequence probability EOS obtained is greater than a set threshold, sentence segmentation is performed. Here, the EOS probability threshold needs to be automatically adjusted according to the context complexity, and the threshold range is set to 0.65 - 0.92; at the same time, if the number of characters in a single block exceeds 2000 or a chapter title is detected, forced segmentation is performed. The mathematical expression for the above content can be expressed as:
[0119] 1. Dynamic threshold calculation: Let the context complexity coefficient be C ∈ [0, 1], then the dynamic segmentation threshold: θ = max(0.65, min(0.92 - 0.27C, 0.92)); where: when C = 1, the threshold θ = 0.65 (extremely complex context), and when C = 0, the threshold θ = 0.92 (simple context).
[0120] 2. Forced segmentation conditions: Define the forced segmentation flag: ; where, L represents the cumulative number of characters, (H tide = 1) represents the title detection flag.
[0121] 3. Final segmentation decision, for each Token position i: ; where: P eos (i) is the EOS probability of each token position calculated by the LLM.
[0122] Step 2: Define the chunks generated in Step 1 as parent chunks, and split the text into paragraphs according to the preset delimiter rules and the maximum chunk length. Each paragraph is applicable to documents with a large amount of text, clear content, and relatively independent paragraphs. The following splitting rules are mainly supported for parent chunks:
[0123] Segment identifier, with a default value of , that is, split according to text paragraphs. You can customize the chunking rules by following regular expressions, and the system will automatically perform segmentation when the segment identifier appears in the text.
[0124] Maximum segment length, which specifies the maximum upper limit of the number of text characters within a segment. When the length is exceeded, forced segmentation will occur. For example, the default value is 500 Tokens, and the maximum upper limit of the segment length is 4000 Tokens.
[0125] Step 3: The sub-segmented text is split based on the parent text segments by the delimiter rules, and is used to find and match the most relevant and direct information to the problem keywords. If the default sub-segmentation rules are used, the presented segmentation effect is:
[0126] When the parent segment is a paragraph, the sub-segments correspond to individual sentences in each paragraph; when the parent segment is the full text, the sub-segments correspond to individual sentences in the full text. Add the parent chunk ID and position offset to the sub-chunks to facilitate positioning the corresponding position and content of the parent chunk. Then vectorize the sub-chunk text of the chunks and store it in the vector database. Here, the vector model can use pre-trained models (such as: BERT, GPT (Generative Pre-trained Transformer)) etc.) to convert each sub-chunk into a vector. These vectors represent the semantic information of the text chunks and can be compared and matched through vector retrieval algorithms. The corresponding expression can be expressed as:
[0127] 1. Definition of parent segment and sub-segment:
[0128] Parent segment: ; where P is used to indicate the original document; P1, P2...P N are multiple parent segments corresponding to the original document respectively;
[0129] Sub-segment: ; where S is used to indicate the sub-segment corresponding under the parent segment, and S j are the characters corresponding under the sub-segment.
[0130] 2. Default sub-segment generation rules, ; The discriminant above means that when the parent segment is a paragraph, the child segment corresponds to a single sentence in each paragraph; when the parent segment is the whole text, the child segment corresponds to each individual sentence in the whole text.
[0131] 3. If it is necessary to support custom delimiters (such as question marks, semicolons), the splitting function can be extended: ; where D is a set of user-defined delimiters. Optionally, the default D is a set of characters consisting of a period (。), a full stop (.), an exclamation mark (!), a question mark (?), etc. Delimiters are a set of special symbols or patterns used to identify the boundaries of text paragraphs or sentences, such as common full stops (.), question marks (?), exclamation marks (!), or line breaks ( ) and so on.
[0132] 4. Hierarchical retrieval; ; where, o jk =offset(s jk ) is the starting character position of the k-th child chunk in the j-th parent chunk. |s jk |=len(s jk ) is the length of the child chunk.
[0133] Step 4: For the question input by the user, first vectorize the user's question using a vector model, then query the vector database to match the corresponding child chunk (sentence), and then map it to the parent chunk (paragraph) through the parent chunk ID and position offset, and send it to the LLM together to complete the full background information of the question and give a more accurate answer. When multiple relevant child chunks are returned in the retrieval result, the system will automatically merge the parent chunks corresponding to these child chunks to form a complete document part. To ensure that the user obtains the most relevant results, the parent chunks corresponding to the child chunks can be sorted and filtered to ensure that the content the user sees is the most in line with their query requirements.
[0134] Optionally, Figure 4 is a flowchart of the parent-child chunk principle according to an embodiment of the present application, including: determining the text after semantic segmentation by the LLM, generating parent chunks, further dividing the parent chunks into smaller text chunks to achieve the purpose of child chunk segmentation. Further, use a pre-trained model to convert each child chunk into a vector form to realize child chunk embedding. After that, record the association relationship between the child chunks and the parent chunks. When executing the user query process, first perform child chunk search, then locate the parent chunk. On this basis, when there are multiple parent chunks, merge the parent chunks and perform sorting and filtering, and then perform precise retrieval through the child chunks to ensure relevance, and then obtain the corresponding parent chunk to supplement the context information, and send it to the LLM together to give a more accurate answer. Thus, both accuracy and complete answer information can be ensured when generating a response.
[0135] It should be noted that through the EOS probability prediction and dynamic threshold adjustment of the LLM, the block error rate is reduced to 3.2%, effectively avoiding the problem of context fragmentation. Secondly, the retrieval accuracy is improved. By dividing into finer-grained sub-blocks, the retrieval precision is enhanced. Then, by associating with the parent block, the context content is enriched, the loss of sentence background is reduced, and a more accurate answer is given.
[0136] In summary, the large language model (LLM) is used to calculate the sequence termination probability (EOS probability) at each position in the text sequence in real time. Parent blocks are generated through semantic boundary detection, and the parent blocks are recursively cut into multi-level sub-blocks. The sub-blocks carry the metadata of the parent blocks to achieve dynamic association. The semantic boundary detection uses the LLM to predict the sentence end probability and dynamically combines adjacent semantic units. Here, the EOS probability threshold is automatically adjusted according to the context complexity, and the threshold range is set to 0.65 - 0.92; when the EOS probability exceeds the dynamic threshold, the boundary point determination is triggered and the parent block is generated. Subsequently, by introducing the EOS (End-of-Sequence Probability) prediction and dynamic threshold adjustment of the LLM (Large Language Model), the semantic integrity of text chunking is significantly improved. Through the vectorization and retrieval of sub-blocks, combined with the background information supplement of the parent blocks, the retrieval efficiency and information accuracy are effectively balanced, and the overall performance of the RAG system is improved. Moreover, the technical solution of this application supports the processing of multiple file formats, such as PPT, PDF, DOC, TXT, emails, etc., and can dynamically adjust the EOS probability threshold and chunking rules according to the characteristics of different documents, enhancing the adaptability of the system and the ability to process complex document structures.
[0137] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc), including several instructions to enable a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of this application.
[0138] In this embodiment, a retrieval response system is also provided. This system is used to implement the above embodiments and preferred implementation manners, and those that have been described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0139] Figure 5 is a structural block diagram of a retrieval response device according to an embodiment of the present application. As Figure 5 shown, the system includes:
[0140] A splitting module 52, configured to, when receiving a prompt instruction for processing an original document, split the original document according to the sequence termination probabilities corresponding to different data sub-blocks in the original document, to obtain N parent text chunks;
[0141] A splitting module 54, configured to perform paragraph splitting processing on each of the N parent text chunks to generate N×M sub-text chunks, where N and M are positive integers;
[0142] A vector module 56, configured to vectorize the N×M sub-text chunks to obtain N×M query reference vectors;
[0143] A retrieval module 58, configured to perform a retrieval response to a query problem issued by a target object based on the N×M query reference vectors.
[0144] Through the above device, after receiving a prompt instruction for processing an original document, splitting is performed according to the sequence termination probabilities of different data sub-blocks in the document to obtain N parent text chunks; then, paragraph splitting processing is performed on each parent text chunk to generate N×M sub-text chunks; then, these sub-text chunks are converted into a vector form to form query reference vectors; finally, a retrieval response is performed on the query problem issued by the target object based on these query reference vectors. Subsequently, the technical effect of improving the retrieval accuracy through the generation of sub-blocks and the association with parent blocks is achieved, and the technical problem of semantic breakage and low retrieval accuracy caused by traditional fixed chunking is solved through the above method.
[0145] In an exemplary embodiment, the above splitting module is further configured to obtain the first sequence termination probabilities corresponding to different data sub-blocks in the original document to obtain P first sequence termination probabilities, where P is a positive integer greater than or equal to N; screen out N second sequence termination probabilities greater than a dynamic probability threshold from the P first sequence termination probabilities, where the dynamic probability threshold is determined by calculating the complexity of the original document through a preset first formula; obtain the sub-block positions of the N data sub-blocks corresponding to the N second sequence termination probabilities, and perform splitting processing on the original document based on the sub-block positions to obtain N parent text chunks.
[0146] In an exemplary embodiment, the above-mentioned device further includes: a processing module, which is further configured to perform text parsing on the original document before screening out N second sequence termination probabilities greater than the dynamic probability threshold from P first sequence termination probabilities; determine the context complexity coefficient between different data sub-blocks in the original document according to the parsing result; use a preset first formula to perform calculation processing on the context complexity coefficient to obtain the dynamic probability threshold, where the preset first formula is: θ = max(0.65, min(0.92 - 0.27C, 0.92)), C is the context complexity coefficient, and θ is the dynamic probability threshold.
[0147] In an exemplary embodiment, the above-mentioned processing module further includes: an adding unit, which is configured to perform splitting processing on the original document based on the sub-block position to obtain N parent text chunks. Before that, the above method further includes: performing splitting detection on the original document to determine the cumulative number of characters and the number of titles corresponding to the original document; adding a forced splitting identifier to the original document when the cumulative number of characters is greater than or equal to the preset number of characters and the number of titles is greater than or equal to the preset number; generating a first prompt message indicating that there is no forced splitting identifier for the original document when the cumulative number of characters is less than the preset number of characters or the number of titles is less than the preset number.
[0148] In an exemplary embodiment, the above-mentioned processing module further includes: a comparing unit, which is configured to compare multiple identifier positions corresponding to the forced splitting identifier in the original document with N sub-block positions greater than the dynamic probability threshold before performing splitting processing on the original document according to the sequence termination probabilities corresponding to different data sub-blocks in the original document; set the target positions where the identifier positions and the sub-block positions are repeated as splitting monitoring points according to the comparison result.
[0149] In an exemplary embodiment, the above-mentioned splitting module is further configured to perform paragraph splitting on the N parent text chunks using a preset separation rule and the maximum block length allowed by the minimum semantic unit to obtain N×M sub-paragraphs; add the block identifier corresponding to the data sub-block to which the parent text chunk belongs and the position offset corresponding to the data sub-block to which the parent text chunk belongs to each of the N×M sub-paragraphs to generate N×M sub-text chunks.
[0150] In an exemplary embodiment, the above-mentioned splitting module is further configured to parse the preset separation rule, determine the default setting quantity of the paragraph identifier and the number of characters between adjacent paragraph identifiers according to the parsing result; perform paragraph verification on the N parent text chunks respectively according to the default setting quantity, the number of characters, and the maximum block length; determine to perform paragraph splitting on the N parent text chunks based on the verification result.
[0151] In an exemplary embodiment, the above splitting module is further configured to, when the verification result indicates that the current parent text block corresponds to a paragraph type text, determine to perform paragraph splitting on the current parent text based on individual sentences in the paragraph, where each individual sentence corresponds to a sub-text block; when the verification result indicates that the current parent text block corresponds to a full-text type text, determine to perform paragraph splitting on the current parent text based on each individual sentence in the full text, where each individual sentence corresponds to a sub-text block.
[0152] In an exemplary embodiment, the above device further includes an updating module, configured to determine whether there is separator custom information in the prompt instruction before performing segmentation processing on the original document according to the sequence termination probability corresponding to different data sub-blocks in the original document; in the case where there is separator custom information, summarize the initial separator and the newly added separators corresponding to the separator custom information, and update the separator set in the preset separation rule according to the summary result; in the case where there is no separator custom information, generate a second prompt information indicating that the preset separation rule does not update the separator set.
[0153] In an exemplary embodiment, the above device further includes a mapping module, configured to perform paragraph splitting processing on each of the N parent text blocks to generate N×M sub-text blocks. After that, the method further includes: establishing a data mapping relationship among the original document, the N parent text blocks, and the N×M sub-text blocks; storing the data mapping relationship in a preset database.
[0154] In an exemplary embodiment, the above retrieval module is further configured to perform text parsing on the query question to obtain Q query sub-question texts, where Q is a positive integer; vectorize the Q query sub-question texts to obtain Q question vectors; perform vector matching on the Q question vectors and the N×M query reference vectors, and perform a retrieval response to the query question sent to the target object based on the configuration result.
[0155] In an exemplary embodiment, the above retrieval module is further configured to determine Q target query reference vectors whose vector similarity to the Q question vectors is greater than a preset similarity according to the vector matching result; determine the corresponding target parent text blocks according to the block identifiers of the parent text blocks corresponding to the Q target query reference vectors and the position offsets corresponding to the parent text blocks; send the target parent text blocks and the query question to a large language model for background information supplementation; determine the supplementation result as the response content for performing a retrieval response to the query question sent to the target object.
[0156] In an exemplary embodiment, the above-mentioned retrieval module further includes: a merging unit, which is configured to, after determining corresponding target parent text chunks according to the chunk identifiers of the parent text chunks corresponding to the Q target query reference vectors and the position offsets corresponding to the parent text chunks, when the Q target query reference vectors correspond to the same original document in the preset database, merge the multiple target parent text chunks corresponding to the Q target query reference vectors to obtain a partial document; and perform a retrieval response to the query question based on the partial document.
[0157] In an exemplary embodiment, the above-mentioned device further includes: a verification module, which is configured to, after performing a segmentation process on the original document according to the sequence termination probabilities corresponding to different data sub-chunks in the original document to obtain N parent text chunks, perform text verification on the original partial document content corresponding to the parent text chunks, where the text verification is used to determine the integrity of the original partial document content relative to the original document and the coherence degree values of different sentences in the original partial document; and determine whether to re-segment the original document based on the verification result.
[0158] In an exemplary embodiment, the above-mentioned verification module is further configured to determine to re-segment the original document when the verification result indicates that the integrity is less than the preset integrity threshold or the coherence degree value is less than the preset coherence degree threshold; and determine not to re-segment the original document when the verification result indicates that the integrity is greater than or equal to the preset integrity threshold or the coherence degree value is greater than or equal to the preset coherence degree threshold.
[0159] In an exemplary embodiment, the above-mentioned device further includes: an evaluation module, which is configured to, after performing a retrieval response to the query question sent by the target object based on N×M query reference vectors, obtain the evaluation information made by the target object in response to the retrieval response, and the overall response duration of the query question; and analyze the effectiveness parameters of the retrieval response according to the evaluation information and the overall response duration, where the effectiveness parameters at least include: retrieval delay duration, retrieval hit rate.
[0160] It should be noted that the above-mentioned various modules can be implemented by software or hardware. For the latter, it can be implemented in the following ways, but not limited thereto: the above-mentioned modules are all located in the same processor; or, the above-mentioned various modules are respectively located in different processors in any combination form.
[0161] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, where the computer program is configured to execute the steps in any one of the above method embodiments when running.
[0162] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media such as USB flash drives, read-only memory (ROM), random access memory (RAM), external hard drives, magnetic disks, or optical discs that can store computer programs.
[0163] An embodiment of the present application also provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0164] Optionally, Figure 6 is a block diagram of a computer system structure of an electronic device according to an embodiment of the present application. As Figure 6 shown, the computer system 800 includes a central processing unit 801 (CPU), which can perform various appropriate actions and processes according to a program stored in a read-only memory 802 (ROM) or a program loaded from a storage section 808 into a random access memory 803 (RAM). In the random access memory 803, various programs and data required for system operations are also stored. The central processing unit 801, the read-only memory 802, and the random access memory 803 are connected to each other via a bus 804. An input / output interface 805 (Input / Output interface, i.e., I / O interface) is also connected to the bus 804.
[0165] The following components are connected to the input / output interface 805: an input section 806 including a keyboard, a mouse, etc.; an output section 807 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker; a storage section 808 including a hard disk, etc.; and a communication section 809 including a network interface card such as a local area network card, a modem, etc. The communication section 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to the input / output interface 805 as needed. A removable medium 811, such as a magnetic disk, an optical disc, a magneto-optical disc, a semiconductor memory, etc., is installed on the drive 810 as needed so that a computer program read from it can be installed into the storage section 808 as needed.
[0166] In an exemplary embodiment, the above electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the above processor, and the input / output device is connected to the above processor.
[0167] An embodiment of the present application further provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the steps in any one of the above method embodiments are implemented.
[0168] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any one of the above method embodiments are implemented.
[0169] An embodiment of the present application further provides a computer program. The computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium; a processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps in any one of the above method embodiments.
[0170] Specific examples in this embodiment may refer to the examples described in the above embodiments and exemplary embodiments, and will not be repeated here.
[0171] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0172] Obviously, those skilled in the art should understand that the above modules or steps of the present application can be implemented by a general-purpose computing device. They can be concentrated on a single computing device, or distributed on a network composed of multiple computing devices. They can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order than here, or they can be separately made into individual integrated circuit modules, or multiple modules or steps among them can be made into a single integrated circuit module to implement. In this way, the present application is not limited to any specific combination of hardware and software.
[0173] The above has introduced in detail a retrieval response system, method, medium, electronic device and program product provided by the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.
Claims
1. A retrieval response method, characterized in that, including: In the case of receiving a prompt instruction to process the original document, segment the original document according to the sequence termination probabilities corresponding to different data sub-blocks in the original document to obtain N parent text blocks; Perform paragraph splitting on each of the N parent text blocks to generate N×M sub-text blocks, where N and M are positive integers; Vectorize the N×M sub-text blocks to obtain N×M query reference vectors; Based on the N×M query reference vectors, perform a retrieval response to the query questions issued by the target object.
2. The retrieval response method according to claim 1, wherein Segment the original document according to the sequence termination probabilities corresponding to different data sub-blocks in the original document to obtain N parent text blocks, including: Obtain the first sequence termination probabilities corresponding to different data sub-blocks in the original document to obtain P first sequence termination probabilities, where P is a positive integer greater than or equal to N; Select N second sequence termination probabilities greater than the dynamic probability threshold from the P first sequence termination probabilities, where the dynamic probability threshold is determined by calculating the complexity of the original document using a preset first formula; Obtain the sub-block positions of the N data sub-blocks corresponding to the N second sequence termination probabilities, and segment the original document based on the sub-block positions to obtain N parent text blocks.
3. The retrieval response method according to claim 2, wherein Before selecting N second sequence termination probabilities greater than the dynamic probability threshold from the P first sequence termination probabilities, the method further includes: Perform text parsing on the original document; Determine the context complexity coefficients between different data sub-blocks in the original document according to the parsing results; Perform calculation processing on the context complexity coefficients using the preset first formula to obtain the dynamic probability threshold, where the preset first formula is: θ = max(0.65, min(0.92 - 0.27C, 0.92)), C is the context complexity coefficient, and θ is the dynamic probability threshold.
4. The retrieval response method according to claim 2, wherein Before segmenting the original document based on the sub-block positions to obtain N parent text blocks, the method further includes: Perform segmentation detection on the original document to determine the cumulative number of characters and the number of titles corresponding to the original document; In the case where the cumulative number of characters is greater than or equal to the preset number of characters and the number of titles is greater than or equal to the preset number, add a forced segmentation identifier to the original document; In the case where the cumulative number of characters is less than the preset number of characters or the number of titles is less than the preset number, generate a first prompt message indicating that there is no forced segmentation identifier for the original document.
5. The retrieval response method according to claim 1, wherein Before segmenting the original document according to the sequence termination probabilities corresponding to different data sub-blocks in the original document, the method further includes: Compare the multiple identifier positions corresponding to the forced segmentation identifier in the original document with the N sub-block positions greater than the dynamic probability threshold; Set the target positions where the identifier positions and the sub-block positions are repeated as segmentation monitoring points according to the comparison results.
6. The retrieval response method according to claim 1, wherein Perform paragraph splitting on each of the N parent text blocks to generate N×M sub-text blocks, including: Perform paragraph splitting on the N parent texts in chunks using a preset separation rule and the maximum chunk length allowed by the minimum semantic unit, to obtain N×M sub-paragraphs; Add the chunk identifier corresponding to the data sub-chunk to which the parent text chunk belongs and the position offset corresponding to the data sub-chunk to which the parent text chunk belongs to each of the N×M sub-paragraphs, to generate N×M sub-text chunks.
7. The retrieval response method according to claim 6, characterized in that Performing paragraph splitting on the N parent texts in chunks using a preset separation rule and the maximum chunk length allowed by the minimum semantic unit includes: Parse the preset separation rule, and determine the default setting quantity of the segment identifiers and the number of characters between adjacent segment identifiers according to the parsing result; Perform segment verification on the N parent text chunks respectively according to the default setting quantity, the number of characters, and the maximum chunk length; Determine paragraph splitting of the N parent text chunks based on the verification result.
8. The retrieval response method according to claim 7, wherein Determining paragraph splitting of the N parent text chunks based on the verification result includes: When the verification result indicates a paragraph type text corresponding to the current parent text chunk, determine to perform paragraph splitting on the current parent text based on individual sentences in the paragraph, where each individual sentence corresponds to a sub-text chunk; When the verification result indicates a full text type text corresponding to the current parent text chunk, determine to perform paragraph splitting on the current parent text based on each individual sentence in the full text, where each individual sentence corresponds to a sub-text chunk.
9. The retrieval response method according to claim 1, wherein Before performing splitting processing on the original document according to the sequence termination probabilities corresponding to different data sub-chunks in the original document, the method further includes: Determine whether there is separator custom information in the prompt instruction; When there is the separator custom information, summarize the initial separator and the new separators corresponding to the separator custom information, and update the separator set in the preset separation rule according to the summarization result; When there is no such separator custom information, generate a second prompt message for the separator set in the preset separation rule that is not updated.
10. The retrieval response method according to claim 1, wherein After performing paragraph splitting processing on each of the N parent text chunks to generate N×M sub-text chunks, the method further includes: Establish a data mapping relationship among the original document, the N parent text chunks, and the N×M sub-text chunks; Store the data mapping relationship in a preset database.
11. The retrieval response method according to claim 1, wherein Performing a retrieval response to a query question sent to a target object based on the N×M query reference vectors includes: Perform text parsing on the query question to obtain Q query sub-question texts, where Q is a positive integer; Vectorize the Q query sub-question texts to obtain Q question vectors; Perform vector matching on the Q question vectors and the N×M query reference vectors, and perform a retrieval response to the query question sent to the target object based on the configuration result.
12. The retrieval response method according to claim 11, wherein Performing vector matching on the Q question vectors and the N×M query reference vectors, and performing a retrieval response to the query question sent to the target object based on the configuration result includes: Determine Q target query reference vectors whose vector similarity to the Q question vectors is greater than a preset similarity according to the vector matching result; Determine corresponding target parent text chunks according to the chunk identifiers of the parent text chunks corresponding to the Q target query reference vectors and the position offsets corresponding to the parent text chunks; Send the target parent text chunks and the query question to a large language model for background information supplementation; Determine the supplementation result as the response content for retrieving and responding to the query question sent to the target object.
13. The retrieval response method according to claim 12, wherein After determining the corresponding target parent text chunks according to the chunk identifiers of the parent text chunks corresponding to the Q target query reference vectors and the position offsets corresponding to the parent text chunks, the method further includes: When the Q target query reference vectors correspond to the same original document in a preset database, merge the multiple target parent text chunks corresponding to the Q target query reference vectors to obtain a partial document; Retrieve and respond to the query question based on the partial document.
14. The retrieval response method according to claim 1, wherein After performing segmentation processing on the original document according to the sequence termination probabilities corresponding to different data sub-chunks in the original document to obtain N parent text chunks, the method further includes: Perform text verification on the original partial document content corresponding to the parent text chunks, where the text verification is used to determine the integrity of the original partial document content relative to the original document and the coherence degree values of different sentences in the original partial document; Determine whether to re-segment the original document based on the verification result.
15. The retrieval response method according to claim 14, wherein Determining whether to re-segment the original document based on the verification result includes: Determine to re-segment the original document when the verification result indicates that the integrity is less than a preset integrity threshold or the coherence degree value is less than a preset coherence degree threshold; Determine not to re-segment the original document when the verification result indicates that the integrity is greater than or equal to a preset integrity threshold or the coherence degree value is greater than or equal to a preset coherence degree threshold.
16. The retrieval response method according to claim 1, wherein After retrieving and responding to the query question sent by the target object based on the N×M query reference vectors, the method further includes: Obtain the evaluation information made by the target object for the retrieval response and the overall response duration of the query question; Analyze the effectiveness parameters of the retrieval response according to the evaluation information and the overall response duration, where the effectiveness parameters at least include: retrieval delay duration, retrieval hit rate.
17. A retrieval response device, characterized in that, Includes: A segmentation module, configured to, when receiving a prompt instruction to process the original document, perform segmentation processing on the original document according to the sequence termination probabilities corresponding to different data sub-chunks in the original document to obtain N parent text chunks; A splitting module, configured to perform paragraph splitting processing on each of the N parent text chunks to generate N×M sub-text chunks, where N and M are positive integers; A vector module, configured to vectorize the N×M sub-text chunks to obtain N×M query reference vectors; A retrieval module, configured to retrieve and respond to the query question sent by the target object based on the N×M query reference vectors.
18. An electronic device, characterized in that, Includes: A memory, configured to store a computer program; A processor, configured to implement the steps of the retrieval response method according to any one of claims 1 to 16 when executing the computer program.
19. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein the computer program, when executed by a processor, implements the steps of the retrieval response method according to any one of claims 1 to 16.
20. A computer program product, comprising a computer program, characterized in that, The computer program, when executed by a processor, implements the steps of the retrieval response method according to any one of claims 1 to 16.
Citation Information
Patent Citations
Document knowledge base-oriented multi-granularity structured retrieval enhancement generation method and device
CN118585615A
Method, system and equipment for retrieval enhancement generation based on large language model and medium
CN118964387A
Document retrieval method and device and storage medium
CN119357318A
Government affair large model semantic perception text segmentation technology
CN119862261A
Retrieval enhancement generation method based on dynamic document block segment optimization
CN120045696A
Cited By
Enhanced retrieval generation method based on father-child segmentation and multi-source recall
CN120723894A
An enhanced retrieval generation method based on parent-child segmentation and multi-source recall
CN120723894B
Test verification evaluation method based on AI intelligent agent and related device
CN120873522A
A method and related apparatus for experimental verification and evaluation based on AI intelligent agents
CN120873522B
Text information compression method and compression device for index distributed database
CN121144271A