Text block dynamic segmentation method and system based on RAG
Through the dynamic segmentation method, combined with directory structure and topic modeling, the semantic incompleteness and information segmentation problems caused by text segmentation in the RAG system are solved, and more efficient knowledge retrieval and generation effects are achieved.
Patent Information
- Application Number
- CN202510489841.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-04-18
AI Technical Summary
Prior Art In RAG systems, traditional text segmentation methods cannot adapt to the semantic structural characteristics of text in different fields, resulting in critical context breaks and information breaks, affecting the matching degree of search results and generation requirements, especially when dealing with long text and multimodal knowledge bases.
By analyzing the multi-level directory structure of the document, combining paragraph separators and end-of-sentence symbols as slicing boundaries, potential Dirichlet allocation theme modeling and sliding windows calculate topic similarity, dynamically adjust the slicing strategy, insert hard slicing marks and soft slicing marks to ensure the integrity and semantic coherence of the directory structure.
It realizes more accurate text semantic unit capture in the RAG system, improves the accuracy of knowledge retrieval and the consistency of generation, adapts to different text types, reduces memory usage, and supports streaming throughput.
Smart Images

Figure CN120353880A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of text data processing, and particularly to a method and system for dynamically segmenting text blocks based on RAG. Background Art
[0002] In a Retrieval-Augmented Generation (RAG) system, the segmentation quality of text blocks directly affects the accuracy of knowledge retrieval and the reliability of generated content. Traditional text segmentation methods use fixed-length windows or static segmentation strategies based on punctuation marks, making it difficult to adapt to the semantic structure characteristics of texts in different fields. Especially when dealing with long texts containing nested semantics such as technical documents and legal provisions, existing segmentation methods are prone to break critical contexts, resulting in semantic deviation during the retrieval stage. In addition, when facing unstructured data in a multi-modal knowledge base, static segmentation strategies cannot dynamically adjust the text granularity, leading to a decrease in the matching degree between retrieval results and generation requirements, severely restricting the practical application effect of the RAG system. Existing technologies include fixed-length chunking and rule-based chunking. Fixed-length chunking divides a long text into chunks of a certain fixed length, such as 100 words per chunk. Rule-based chunking mainly divides text into chunks according to the number of sentences, such as ten sentences per chunk. The main disadvantages of this technical solution are as follows: (1) It is easy to split a topic into two chunks, making each chunk semantically incomplete; (2) The content within a chunk is not coherent, and information is fragmented, which is not conducive to subsequent retrieval and generation.
[0003] Prior Art One, Application No.: CN202411426611.2 discloses a text segmentation method, device, storage medium, and electronic device. The method includes: obtaining the text to be segmented; determining an initial chunking method for segmenting the text to be segmented based on the text type of the text to be segmented; performing data segmentation processing on the text to be segmented using the initial chunking method to obtain each initial data chunk corresponding to the text to be segmented; and performing data segmentation processing on each initial data chunk at least by using a method of dynamically adjusting the chunk length to obtain each target data chunk corresponding to the text to be segmented. Although, for different text types, performing data segmentation processing on each initial data chunk by using a method of dynamically adjusting the chunk length or by using a method of dynamically adjusting the chunk length and dynamically adjusting the repeated chunk length can ensure that the content of each original initial data chunk is complete and make the chunking result less likely to be ambiguous; however, there is a lack of guarantee for semantic coherence: only relying on preset rules of text types and not considering semantic breaks caused by topic drift; poor adaptability to directory structures: not establishing a dynamic mapping between directory levels and chunk granularity, resulting in the loss of context association during cross-chapter retrieval; lack of a conflict handling mechanism: when preset rules conflict with semantic boundaries, the retrieval consistency cannot be maintained through a compensation mechanism.
[0004] Prior Art Two, Application No.: CN 202410571033.5 discloses a text segmentation method, device, and medium for large language models. The method includes: obtaining the text to be segmented; selecting a matching segmentation method according to the text type to segment the text so that each segmented text block does not exceed the input capacity limit of the large language model; storing each text block in a vector library for providing reference for the large language model to perform specific learning tasks. Although it considers the impact of large language models on text segmentation and provides an adapted text segmentation method for the application characteristics and requirements of large language models, there are deficiencies in multi-granularity demand response: only adjusting segmentation through a single dimension of text type cannot cope with the retrieval needs of different granularities within the same document; multi-modal structure processing is missing: the special segmentation characteristics of non-continuous texts such as table of contents levels, tables, and footnotes are not considered; the dynamic adjustment dimension is single: only the number of characters is used as the adjustment basis, and adaptive optimization is not combined with semantic density.
[0005] Prior Art Three, Application No.: CN202411805410.3 discloses a text segmentation method and device for retrieval-augmented generation, including: constructing a text tree for the original text, where the text tree uses the original text attributes as the root node, each level of headings as child nodes, and non-heading paragraphs as leaf nodes; segmenting text blocks based on the text tree so that each segmented text block starts with a table of contents structure, ends with a non-heading paragraph, and is constrained by a preset text block word limit. The table of contents structure included in each text block is: the table of contents structure composed of all the corresponding headings at each level and the original text attributes of all the parent nodes obtained step by step upward starting from the head leaf node included in each text block. Although it can segment the text in a more suitable tree structure manner, reducing the loss of text information caused during text segmentation, thereby improving the recall effect of the RAG system, there are static chunking constraints: relying on a preset word limit to force segmentation, which destroys natural semantic units; the lack of response to semantic drift: no topic continuity monitoring mechanism is established, resulting in the breakage of cross-block semantic associations during retrieval; the lack of a dynamic optimization closed loop: the segmentation strategy is fixed and cannot optimize subsequent chunking decisions based on retrieval feedback.
[0006] Currently, Prior Art One, Prior Art Two, and Prior Art Three have the problem of semantic incompleteness, content incoherence, and information fragmentation caused by traditional fixed chunking technical means for text chunking, and they cannot meet the complex long text and short text chunking requirements and achieve dynamic text chunking based on the combination of structure and semantics. Therefore, the present invention provides a method and system for dynamic text chunking based on RAG. Summary of the Invention
[0007] To solve the above technical problems, the present invention provides a method for dynamic text chunking based on RAG, including the following steps:
[0008] Parse the multi-level table of contents structure of the document, and split text blocks based on table of contents nodes; detect paragraph separators and end-of-sentence symbols as secondary split boundaries; for documents with a complete table of contents, generate a tree-like text block structure that strictly corresponds to the table of contents hierarchy. For documents without a table of contents or with an incomplete table of contents, switch to a mixed mode of rules and semantics for splitting, retain the existing table of contents nodes as the main split points, and enable semantic splitting in the areas where the table of contents is missing for supplementary splitting;
[0009] Perform Latent Dirichlet Allocation (LDA) topic modeling on the continuous text stream, calculate the topic distribution vector of each text segment in real time, calculate the topic similarity of adjacent windows through a sliding window. When the topic similarity exceeds the threshold, it is determined as a topic boundary. Perform adaptive splitting on the detected topic boundaries, insert hard split markers at topic mutation points, and use soft splitting for gradually changing topic regions;
[0010] The initial window size of the sliding window is set according to the document type, and the semantic density within the window is monitored in real time; hierarchical chunk reorganization; when there are conflicts between the results of rule-based splitting and semantic splitting: prioritize maintaining the integrity of the table of contents structure, insert bidirectional index anchors in the conflict areas, and maintain the context association during retrieval.
[0011] Optionally, among them, the multi-level table of contents structure includes chapter, section, or article numbers, paragraph separators include blank lines or indents, and end-of-sentence symbols include full stops, question marks, or exclamation marks.
[0012] Optionally, the process of parsing the multi-level table of contents structure of the document includes the following steps:
[0013] Identify the hierarchical marking features in the document, extract the explicit chapter, section, or article numbering system as the main basis for division; automatically generate independent text containers for each numbered node to form an initial tree-like skeleton structure; for each container unit in the skeleton structure, perform paragraph-level analysis: detect the visual isolation belt formed by blank lines as the paragraph separation identifier, and identify the text alignment boundary generated by indentation changes as the secondary division clue;
[0014] Within the already divided paragraph units, perform sentence-level parsing, and construct the minimum semantic unit boundary through the end-of-sentence punctuation triple; automatically activate the list item detection mode for paragraphs containing numbered sequences, and identify the sub-structures formed by consecutive numbers; when a directory missing area is detected, automatically start the compensation mechanism, retain the existing directory nodes as fixed anchors, and use semantic continuity analysis to construct virtual nodes in the vacant interval;
[0015] Calibrate the positions of the initial splitting scheme exported from the table of contents and the results of symbol analysis, mark abnormal flags for the areas that still have conflicts after calibration, and switch to the sliding window analysis process.
[0016] Optionally, when the deviation between the table of contents splitting boundary and the natural paragraph boundary exceeds a preset tolerance, select the following processing solutions according to the context density;
[0017] High-density technical term area: Keep the table of contents boundary as the priority;
[0018] Narrative text area: Adopt the natural paragraph boundary;
[0019] Transition area: Insert collapsible virtual segmentation markers.
[0020] Optionally, the process of constructing the minimum semantic unit boundary includes the following steps:
[0021] Use the paragraph units obtained by the paragraph separation identifier as the input basic structure of the minimum semantic unit boundary; the text stream within each container unit is used as the initial processing object; within the given paragraph unit, scan the text character sequence to locate the position coordinates of the end-of-sentence punctuation triple, and the existence verification of the end-of-sentence punctuation triple is based on the paragraph boundary constraint; the text interval between the triples generates primary semantic blocks, and the length of the primary semantic blocks is restricted by the paragraph granularity analysis result;
[0022] For the primary semantic blocks containing numbered prefixes, trigger the list item detection mode, and judge the similarity between the sub-structure formed by the consecutive numbers identified by the trigger condition and the established hierarchical marking features. The consecutive numbered items automatically form a sub-structure tree, and the root node of the sub-structure tree establishes a parent-child association with the paragraph container node;
[0023] When the sequence of primary semantic blocks is interrupted, start the vacant area compensation mechanism, use the parsed normal structure as the anchor reference system, and deduce the potential structural continuity by analyzing the verb tense pattern and entity reference relationship of the semantic blocks on both sides of the interrupted area. The generated virtual nodes inherit the same hierarchical attributes as the adjacent real nodes.
[0024] Optionally, the process of performing adaptive segmentation on the detected topic boundary includes the following steps:
[0025] Based on the text content within the sliding window, calculate the topic distribution vector of the current window. The topic distribution vector represents the probability weights of each topic within the window; if the text contains a table of contents structure, the window is initialized to align with the table of contents node boundary; if there is no table of contents structure, the window size is generated according to the initial setting;
[0026] Adopt a symmetric probability distribution difference metric to compare the topic distribution vectors of adjacent windows; if the similarity is lower than the preset threshold, it is determined as a topic mutation; if the similarity is higher than the threshold but shows a continuous downward trend, it is determined as a topic gradual change area; the window size is dynamically adjusted according to the detection result;
[0027] Insert segmentation markers at mutation points and expandable segmentation markers in the gradual change region, and cross - border associate contexts during retrieval; if the segmentation point conflicts with the directory node, give priority to preserving the directory integrity and insert bidirectional index anchors at the conflict position; structured output of the segmentation result: type marker (hard / soft segmentation), original directory position (if any), associated topic distribution vector, and hierarchical integration.
[0028] Optionally, the process of calculating the topic similarity includes the following steps:
[0029] Obtain the topic distribution vectors of adjacent windows as input, where each topic distribution vector contains the probability weight values of each topic within the window; if it is the first calculation, initialize a reference set containing all known topics; if not the first time, inherit the valid topic set confirmed in the previous calculation.
[0030] Normalize the topic distribution vectors of adjacent windows to ensure that the sum of all components is 1; calculate the symmetric difference value between the two distribution vectors, and the symmetric difference value satisfies the three basic properties of the distance function: non - negativity, symmetry, and triangle inequality.
[0031] Map the distribution difference value to a similarity score through a monotonically decreasing function, set an adaptively adjustable conversion threshold, and the default value range is based on historical statistical results. Output a normalized similarity score, and the value range is fixed at [0,1]; based on the characteristics of the current document type, load a preset determination threshold; output the final topic similarity score of adjacent windows, compare the final score with the determination threshold, and generate a determination result of mutation or gradual change.
[0032] Optionally, the process of outputting the normalized similarity score includes the following steps:
[0033] Receive the symmetric difference value, which represents the difference between two topic distributions; verify that the input value meets three basic mathematical constraints; record the current difference value in the historical statistical sequence; establish a conversion relationship from the difference value to the similarity, dynamically determine the initial form parameters of the conversion curve according to the characteristics of historical data, and set the boundary conditions of the conversion curve: the minimum difference corresponds to the highest similarity, and the theoretical maximum difference corresponds to zero similarity.
[0034] Substitute the current symmetric difference value into the mapping function to obtain the initial similarity, and calibrate the stability of the initial result according to the recent data fluctuation characteristics; determine the reference baseline based on the historical data distribution characteristics, adjust the calibrated topic similarity to the standard measurement interval, and truncate the result within the specified range through upper and lower limits.
[0035] Establish a dynamic baseline by integrating the historical similarity level and the current document characteristics, adjust the threshold considering the topic diversity index and context consistency factors, and apply a boundary protection mechanism to prevent the threshold from drifting beyond a reasonable range.
[0036] Optionally, the hierarchical block reorganization includes:
[0037] First-level block: Inherit the directory structure segmentation result;
[0038] Second-level block: The semantic segmentation boundary applied inside the directory node;
[0039] Third-level block: Implement dynamic granularity adjustment of the sliding window for complex paragraphs.
[0040] A text block dynamic segmentation system based on RAG provided by the present invention includes:
[0041] The segmentation mode module is responsible for parsing the multi-level directory structure of the document, segmenting text blocks based on directory nodes; detecting paragraph delimiters and end-of-sentence symbols as secondary segmentation boundaries; for documents with a complete directory, generating a tree-like text block structure that strictly corresponds to the directory hierarchy, and for documents without a directory or with an incomplete directory, switching to a mixed segmentation mode of rules and semantics, retaining the existing directory nodes as the main segmentation points, and enabling semantic segmentation in the directory missing area for supplementary segmentation;
[0042] The marking confirmation module is responsible for performing Latent Dirichlet Allocation (LDA) topic modeling on the continuous text stream, and calculating the topic distribution vector of each text segment in real time; calculating the topic similarity of adjacent windows through a sliding window, and determining a topic boundary when the topic similarity exceeds a threshold, performing adaptive segmentation on the detected topic boundary; inserting hard segmentation marks at topic mutation points, and using soft segmentation for gradually changing topic regions;
[0043] The block reorganization module is responsible for setting the initial window size of the sliding window according to the document type, and monitoring the semantic density within the window in real time; hierarchical block reorganization; when there is a conflict between the rule-based segmentation and the semantic segmentation results: prioritize maintaining the integrity of the directory structure, inserting bidirectional index anchors in the conflict area, and maintaining the context association during retrieval.
[0044] The present invention forms an orthogonal segmentation dimension by the explicit hierarchy of chapter headings and the implicit semantic boundaries of the LDA topic model, and realizes the adaptive switching of hard / soft segmentation strategies by combining the dynamic topic similarity calculation of a sliding window. The directory nodes serve as the main segmentation anchors to ensure the logical integrity of the document. Semantic segmentation supplements the boundary recognition of unstructured areas at a fine-grained level. The two establish a topological mapping relationship through bidirectional index anchors to ensure the compatibility of structural navigation and semantic retrieval during retrieval. The three-level segmentation hierarchy realizes progressive processing from macro directory inheritance (level 1), meso-semantic segmentation (level 2) to micro sliding window granularity regulation (level 3). After the window size is initially set according to the document type, dynamic adjustment is triggered by real-time semantic density monitoring: the window is enlarged for high-density professional term areas to maintain concept integrity, and the window is reduced for low-density narrative areas to improve segmentation accuracy. When boundary conflicts occur between regular segmentation (directory first) and semantic segmentation (topic mutation points), a hybrid storage strategy with a tree-like text block structure as the main trunk and semantic segmentation as the branches is adopted, and the context association path of the conflict area is recorded through bidirectional anchors. During the retrieval stage, structural traversal (directory tree jump) or semantic penetration (topic vector retrieval) can be selected as needed to achieve cross-level context reconstruction with an O(1) time complexity. The incremental calculation of the LDA topic distribution and the streaming cache mechanism of the sliding window enable the system to complete real-time segmentation without global document loading, which is particularly suitable for real-time log streams or progressive parsing of long documents; the linked adjustment of the topic similarity threshold and the window size provides the possibility of FPGA hardware offloading.
[0045] Other features and advantages of the present invention will be described in the following specification, and, in part, will be obvious from the specification, or will be understood by implementing the present invention. The objectives and other advantages of the present invention can be achieved and obtained by the structure specifically pointed out in the written specification and the drawings.
[0046] The technical solution of the present invention will be further described in detail below through the drawings and embodiments. Description of the Drawings
[0047] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention, and do not constitute a limitation to the present invention. In the drawings:
[0048] Figure 1 It is a flowchart of the dynamic text block segmentation method based on RAG in Embodiment 1 of the present invention;
[0049] Figure 2 It is a schematic diagram of the dynamic text block segmentation method based on RAG in Embodiment 1 of the present invention;
[0050] Figure 3 It is a schematic diagram of semantic segmentation in Embodiment 1 of the present invention;
[0051] Figure 4 It is a process diagram of the multi-level directory structure of the parsed document in Embodiment 2 of the present invention;
[0052] Figure 5 It is a process diagram of adaptively splitting the detected topic boundaries in Embodiment 4 of the present invention;
[0053] Figure 6 It is a system block diagram of text block dynamic splitting based on RAG in Embodiment 7 of the present invention. Detailed implementation manners
[0054] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.
[0055] The terms used in the embodiments of the present application are only for the purpose of describing specific embodiments, and are not intended to limit the embodiments of the present application. The singular forms "a", "the" and "said" used in the embodiments of the present application are also intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0056] When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application. In the description of the present application, it should be understood that the terms "first", "second", "third", etc. are only used to distinguish similar objects, and do not have to be used to describe a specific order or sequence, nor can they be understood as indicating or implying relative importance. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific circumstances.
[0057] Embodiment 1: As Figure 1 shown, the embodiment of the present invention provides a method for dynamically splitting text blocks based on RAG, including the following steps:
[0058] S100: Parse the multi-level directory structure of the document, and split the text blocks based on the directory nodes; detect the paragraph delimiter and the end-of-sentence symbol as the secondary split boundaries; for the document with a complete directory, generate a tree-like text block structure that strictly corresponds to the directory level, and for the document without a directory or with an incomplete directory, switch to the mixed splitting mode of rules and semantics, retain the existing directory nodes as the main split points, and enable semantic splitting in the directory missing area for supplementary splitting;
[0059] Among them, the multi-level directory structure includes chapter / section / article numbers, the paragraph separator includes blank lines and indentation, and the end-of-sentence symbols include full stops / question marks / exclamation marks;
[0060] S200: Perform Latent Dirichlet Allocation (LDA) topic modeling on the continuous text stream, calculate the topic distribution vector of each text segment in real time, calculate the topic similarity of adjacent windows through a sliding window. When the topic similarity exceeds the threshold, it is determined as a topic boundary. Perform adaptive segmentation on the detected topic boundary, insert hard segmentation marks at topic mutation points, and use soft segmentation for the gradual topic area;
[0061] S300: The initial window size of the sliding window is set according to the document type, and the semantic density within the window is monitored in real time; hierarchical block reorganization; when there is a conflict between the rule-based segmentation and the semantic segmentation results: prioritize maintaining the integrity of the directory structure, insert bidirectional index anchors in the conflict area, and maintain the context association during retrieval;
[0062] Among them, the hierarchical block reorganization includes:
[0063] First-level block: Inherit the segmentation result of the directory structure;
[0064] Second-level block: The semantic segmentation boundary applied within the directory node;
[0065] Third-level block: Perform dynamic granularity adjustment of the sliding window for complex paragraphs.
[0066] The working principle and beneficial effects of the above technical solution are as follows: In this embodiment, the multi-level directory structure of the document is first parsed, and the text blocks are segmented based on the directory nodes; the paragraph separators and end-of-sentence symbols are detected as secondary segmentation boundaries; for the document with a complete table of contents, a tree-like text block structure that strictly corresponds to the directory level is generated. For the document without a table of contents or with an incomplete table of contents, it is transferred to a mixed segmentation mode of rules and semantics, retaining the existing directory nodes as the main segmentation points, and enabling semantic segmentation in the missing area of the table of contents for supplementary segmentation; among them, the multi-level directory structure includes chapter / section / article numbers, the paragraph separators include blank lines and indents, and the end-of-sentence symbols include full stops / question marks / exclamation marks; secondly, latent Dirichlet allocation topic modeling is performed on the continuous text stream, the topic distribution vector of each paragraph of text is calculated in real time, and the topic similarity between adjacent windows is calculated through a sliding window. When the topic similarity exceeds the threshold, it is determined as a topic boundary, and adaptive segmentation is performed on the detected topic boundary, a hard segmentation mark is inserted at the topic mutation point, and soft segmentation is adopted for the gradual topic area; finally, the initial window size of the sliding window is set according to the document type, and the semantic density within the window is monitored in real time; hierarchical block reorganization; when there is a conflict between the rule segmentation and the semantic segmentation results: the integrity of the directory structure is preferentially retained, a two-way index anchor point is inserted in the conflict area, and the context association during retrieval is maintained; among them, the hierarchical block reorganization includes: primary block: inheriting the segmentation result of the directory structure; secondary block: the semantic segmentation boundary applied inside the directory node; tertiary block: dynamically adjusting the granularity of the sliding window for complex paragraphs (the principle is referred to in Appendix Figure 2 , and the principle of semantic segmentation is referred to in Appendix Figure 3). The above scheme forms orthogonal segmentation dimensions through the explicit hierarchy of the catalog and the implicit semantic boundary of the LDA topic model, and realizes the adaptive switching of hard / soft segmentation strategies by combining the dynamic topic similarity calculation of the sliding window. The catalog node is used as the main segmentation anchor to ensure the logical integrity of the document, and the semantic segmentation supplements the boundary identification of the unstructured area at the fine-grained level. The two establish a topological mapping relationship through the bidirectional index anchor to ensure the compatibility of structural navigation and semantic retrieval during retrieval. The three-level block hierarchy realizes the progressive processing from macro catalog inheritance (first level), meso semantic segmentation (second level) to micro sliding window granularity control (third level). After the window size is initially set according to the document type, it is dynamically adjusted through real-time semantic density monitoring: the window is expanded for high-density professional terminology areas to maintain conceptual integrity, and the window is reduced for low-density narrative areas to improve segmentation accuracy. When there is a boundary conflict between rule segmentation (catalog priority) and semantic segmentation (topic mutation point), a hybrid storage strategy with a tree-like text block structure as the trunk and semantic segmentation as branches is adopted, and the context-related path of the conflicting area is recorded through bidirectional anchors. In the retrieval phase, you can choose structure traversal (directory tree jump) or semantic penetration (topic vector retrieval) as needed to achieve cross-level context reconstruction with O(1) time complexity. The incremental calculation of LDA topic distribution and the streaming cache mechanism of sliding windows enable the system to complete real-time segmentation without global document loading, which is particularly suitable for real-time log streams or progressive parsing of long documents; the linkage adjustment of topic similarity threshold and window size provides the possibility of FPGA hardware offloading.
[0067] In summary, compared with the traditional fixed-granularity blocking method, this embodiment improves the accuracy of RAG retrieval and reduces memory usage on the standard test set, while supporting streaming throughput.
[0068] In this embodiment, the RAG system can capture text semantic units more accurately through dynamic segmentation technology, thereby improving the quality of knowledge retrieval and generation, which is a key step in optimizing complex document processing. It achieves accurate retrieval: avoiding the fragmentation of key information caused by fixed segmentation (such as half a sentence containing entity + relationship); generation coherence: ensuring that the context used by the model has complete semantic units when generating answers; domain adaptability: adapting to different text types (technical documents, social media conversations) by adjusting the segmentation strategy.
[0069] This embodiment is mainly a data processing strategy and method designed for segmenting text in word, pdf, and Markdown files when building a knowledge base in the RAG (Retrieval-Augmented Generation) project. The core idea is to adjust the size of the blocks according to the structure or semantics of the content, making each block more complete semantically, realizing dynamic segmentation of text blocks, and achieving better results during retrieval. At the same time, this technology is universal and can be flexibly extended to related fields of large language models (LLM: Large Language Model).
[0070] Embodiment 2: As Figure 4 shown, on the basis of Embodiment 1, the process of parsing the multi-level directory structure of the document provided by the embodiment of the present invention includes the following steps:
[0071] S101: Identify the hierarchical marking features in the document, extract the explicit chapter, section, or article numbering system as the main basis for division; automatically generate independent text containers for each numbered node to form an initial tree-like skeleton structure; for each container unit in the skeleton structure, perform paragraph granularity analysis; detect the visual isolation belt formed by blank lines as the paragraph separation identifier, and identify the text alignment boundary generated by the indentation change as the secondary division clue;
[0072] S102: Implement sentence-level parsing within the divided paragraph units, construct the minimum semantic unit boundary through the end-of-sentence punctuation triples (period / question mark / exclamation mark); automatically activate the list item detection mode for paragraphs containing numbered sequences, and identify the sub-structures formed by consecutive numbers; when a directory missing area is detected, automatically start the compensation mechanism, retain the existing directory nodes as fixed anchor points, and construct virtual nodes in the vacant interval using semantic continuity analysis;
[0073] S103: Calibrate the position of the initial segmentation scheme exported from the directory with the symbol analysis result, mark the abnormal flag for the areas that still have conflicts after calibration, and transfer to the sliding window analysis process; when the deviation between the directory segmentation boundary and the natural paragraph boundary exceeds the preset tolerance, select the following processing scheme according to the context density;
[0074] High-density technical term area: Keep the directory boundary as the priority;
[0075] Narrative text area: Adopt the natural paragraph boundary;
[0076] Transition area: Insert a collapsible virtual segmentation mark.
[0077] The working principle and beneficial effects of the above technical solution are as follows: In this embodiment, the hierarchical marking features in the document are first identified, and the explicit chapter, section, or article numbering system is extracted as the main basis for division; each numbered node automatically generates an independent text container to form an initial tree-like skeleton structure; for each container unit in the skeleton structure, paragraph-level analysis is performed: the visual isolation belt formed by detecting blank lines is used as the paragraph separation identifier, and the text alignment boundary generated by identifying the indentation change is used as the secondary division clue; secondly, within the divided paragraph units, sentence-level parsing is implemented, and the minimum semantic unit boundary is constructed through the end-of-sentence punctuation triple (period / question mark / exclamation mark); for paragraphs containing numbered sequences, the list item detection mode is automatically activated to identify the sub-structures formed by consecutive numbers; when a missing directory area is detected, the compensation mechanism is automatically started, and the existing directory nodes are retained as fixed anchor points, and virtual nodes are constructed using semantic continuity analysis in the vacant interval; finally, the initial segmentation scheme exported by the directory is positionally calibrated with the symbol analysis result, and abnormal flags are marked for the areas that still have conflicts after calibration, and the sliding window analysis process is entered. The hierarchical system formed by the processing flow of this embodiment has downward compatibility: the segmentation results of the upper-level directory nodes directly constrain the scope boundaries of the lower-level processing, and the correction signals obtained from the lower-level analysis adjust the division strategy of the upper-level containers through the feedback channel to form a closed-loop optimization system; the output text block structure maintains topological isomorphism with the original directory, and at the same time embeds the verified fine-grained division marks.
[0078] Embodiment 3: On the basis of Embodiment 2, the process of constructing the minimum semantic unit boundary provided by the embodiment of the present invention includes the following steps:
[0079] S1021: Using the paragraph unit obtained by the paragraph separation identifier as the input basic structure of the minimum semantic unit boundary; the text stream within each container unit is used as the initial processing object; within the given paragraph unit, scan the text character sequence to locate the position coordinates of the end-of-sentence punctuation triple, and the existence verification of the end-of-sentence punctuation triple is based on the paragraph boundary constraint; the text interval between the triples generates a primary semantic block, and the length of the primary semantic block is restricted by the paragraph-level analysis result;
[0080] S1022: For the primary semantic block containing a numbered prefix, trigger the list item detection mode, and the triggering condition is to identify the similarity between the sub-structure formed by consecutive numbers and the established hierarchical marking features. The consecutive numbered items automatically form a sub-structure tree, and the root node of the sub-structure tree is associated with the paragraph container node as a parent-child relationship;
[0081] S1023: When the sequence of primary semantic chunks is interrupted, start the vacant area compensation mechanism. Use the parsed normal structure as the anchor reference system. By analyzing the verb tense patterns and entity reference relationships of the semantic chunks on both sides of the interrupted area, deduce the potential structural continuity. The generated virtual nodes inherit the same hierarchical attributes as the adjacent real nodes.
[0082] The working principle and beneficial effects of the above technical solution are as follows: In this embodiment, first, use the paragraph units obtained by the paragraph separation identifier as the input basic structure of the boundary of the smallest semantic unit; the text flow within each container unit is used as the initial processing object; within the given paragraph unit, scan the text character sequence to locate the position coordinates of the end-of-sentence punctuation triple. The existence verification of the end-of-sentence punctuation triple is based on paragraph boundary constraints; the text interval between the triples generates primary semantic chunks, and the length of the primary semantic chunks is limited by the result of paragraph granularity analysis; secondly, for the primary semantic chunks containing number prefixes, trigger the list item detection mode, and judge the similarity between the sub-structure formed by the consecutive numbers identified by the trigger condition and the established hierarchical marking features. The consecutive numbered items automatically form a sub-structure tree, and the root node of the sub-structure tree establishes a parent-child relationship with the paragraph container node; finally, when the sequence of primary semantic chunks is interrupted, start the vacant area compensation mechanism. Use the parsed normal structure as the anchor reference system. By analyzing the verb tense patterns and entity reference relationships of the semantic chunks on both sides of the interrupted area, deduce the potential structural continuity. The generated virtual nodes inherit the same hierarchical attributes as the adjacent real nodes. Each step of the above solution automatically constructs a document object model with multi-layer semantic associations based on the structured input data stream; through the coupling of syntactic unit segmentation under paragraph granularity constraints and the dynamic compensation mechanism, a unified representation system of the surface layout features and deep semantic associations of the document is established. A complete mapping chain from physical layout to logical structure is formed. The space constraints provided by the paragraph container and the primary semantic chunks generated by the end-of-sentence punctuation detection constitute the basic segmentation framework. The list item pattern recognition establishes a fine-grained sub-tree structure within this framework. Finally, the closed-loop repair of the broken structure is achieved through vacancy compensation. The three processing stages jointly ensure the stable conversion of the document discrete symbol sequence to the continuous semantic topology. Realize the preservation of cross-hierarchical information integrity. The length constraint of the primary semantic chunks and the hierarchical inheritance of the list item sub-tree form a vertical protection mechanism, while the verb tense and entity reference analysis provide horizontal continuity guarantee, eliminating the semantic fragmentation problem caused by simply relying on layout features in traditional methods. The dual protection enables the output structure to meet both the layout boundary constraints and the semantic coherence requirements at the same time. Establish an adaptive error control system. The initial segmentation based on the paragraph container provides an error propagation boundary for processing. The similarity judgment of the list item detection realizes local structure verification. The generation of virtual nodes in the vacant area forms global structure compensation. The three-level fault tolerance mechanism constitutes a progressive error control network, controlling the local feature recognition error within the single sub-tree range and avoiding the risk of structure collapse.
[0083] Example 4: As Figure 5 shown, on the basis of Example 1, the process of adaptively segmenting the detected topic boundaries provided by the embodiments of the present invention includes the following steps:
[0084] S201: Based on the text content within the sliding window, calculate the topic distribution vector of the current window. The topic distribution vector represents the probability weights of each topic within the window; if the text contains a directory structure, the window is initially aligned with the directory node boundaries; if there is no directory structure, the window size is generated according to the initial setting;
[0085] S202: Adopt a symmetric probability distribution difference metric to compare the topic distribution vectors of adjacent windows; if the topic similarity is lower than the preset threshold, it is determined as a topic mutation; if the topic similarity is higher than the threshold but shows a continuous downward trend, it is determined as a topic gradual change region; the window size is dynamically adjusted according to the detection result;
[0086] S203: Insert a segmentation mark at the mutation point and an expandable segmentation mark in the gradual change region, and retrieve the cross-boundary associated context during retrieval; if the segmentation point conflicts with the directory node, the directory integrity is preferentially retained, and a two-way index anchor point is inserted at the conflict position; the structured output of the segmentation result: type mark (hard / soft segmentation), original directory position (if any), associated topic distribution vector, and hierarchical integration.
[0087] The working principle and beneficial effects of the above technical solution are as follows: In this embodiment, first, based on the text content within the sliding window, the topic distribution vector of the current window is calculated. The topic distribution vector represents the probability weights of each topic within the window. If the text contains a table of contents structure, the window is initialized to align with the boundaries of the table of contents nodes. If there is no table of contents structure, the window size is generated according to the initial setting. Secondly, a symmetric probability distribution difference metric is used to compare the topic distribution vectors of adjacent windows. If the similarity is lower than the preset threshold, it is determined as a topic mutation. If the similarity is higher than the threshold but shows a continuous downward trend, it is determined as a topic gradual change region. The window size is dynamically adjusted according to the detection results. Finally, a segmentation mark is inserted at the mutation point, and an expandable segmentation mark is inserted in the gradual change region to associate the context across boundaries during retrieval. If the segmentation point conflicts with the table of contents node, the integrity of the table of contents is preferentially retained, and a two-way index anchor point is inserted at the conflict position. The structured output of the segmentation result includes: type mark (hard / soft segmentation), original table of contents position (if any), associated topic distribution vector, and hierarchical integration. The above solution ensures the spatial alignment of regular segmentation and semantic segmentation through the joint constraint of the window boundary initialized by the table of contents node and the topic distribution vector. When a topic mutation is detected, the hard segmentation mark forcibly interrupts the current semantic unit, while the soft segmentation in the gradual change region maintains long-distance dependencies through expandable marks. The combination of the two achieves a balance between fine-grained cutting of local mutations and retention of global coherence. The symmetric probability difference metric feedback drives the adjustment of the window size, enabling automatic expansion of the window in high-semantic-density regions to reduce over-segmentation and contraction of the window in low-density regions to capture fine-grained topic transitions. This mechanism and the two-way index anchor point in case of a table of contents conflict constitute a redundant design to ensure that fragmented semantic segments can still be associated through the anchor point while giving priority to the integrity of the table of contents. The output structured metadata (type mark / table of contents position / topic vector) forms a three-level index system: the hard segmentation mark directly maps the physical storage boundary, the soft segmentation mark supports probabilistic boundary expansion during retrieval, and the topic distribution vector provides term-level semantic routing for dense retrieval. The hierarchical integration result enables the document set to meet the dual requirements of exact retrieval (based on the table of contents hierarchy) and semantic retrieval (based on the topic distribution) simultaneously. The initial alignment of the window with the table of contents reduces the ineffective calculation range, the dynamic adjustment strategy reduces the repeated modeling overhead, and the conflict anchor point mechanism avoids data redundancy at the storage level. The final output solution achieves Pareto optimality among the computational complexity (O(n) sliding window detection), storage efficiency (hard segmentation compresses the original text), and retrieval efficiency (soft segmentation supports approximate nearest neighbor).
[0088] Embodiment 5: On the basis of Embodiment 4, the calculation process of the topic similarity provided by the embodiment of the present invention includes the following steps:
[0089] S2021: Obtain the topic distribution vectors of adjacent windows as input, where each topic distribution vector contains the probability weight values of each topic within the window; if it is the first calculation, initialize the reference set containing all known topics; if not the first time, inherit the valid topic set confirmed in the previous calculation.
[0090] S2022: Normalize the topic distribution vectors of adjacent windows to ensure that the sum of all components is 1; calculate the symmetric difference value between the two distribution vectors, and the symmetric difference value satisfies the three basic properties of non-negativity, symmetry, and triangle inequality of the distance function.
[0091] S2023: Map the distribution difference value to a similarity score through a monotonically decreasing function, set an adaptively adjustable conversion threshold, and the default value range is based on historical statistical results. Output the normalized similarity score, and the value range is fixed at [0,1]; based on the characteristics of the current document type, load the preset decision threshold; output the final topic similarity score of adjacent windows, compare the final score with the decision threshold, and generate a mutation or gradual change decision result.
[0092] The working principle and beneficial effects of the above technical solution are as follows: In this embodiment, the topic distribution vectors of adjacent windows are first obtained as inputs, where each topic distribution vector contains the probability weight values of each topic within the window; if it is the first calculation, a reference set containing all known topics is initialized; if not, the valid topic set confirmed in the previous calculation is inherited; secondly, the topic distribution vectors of adjacent windows are normalized to ensure that the sum of all components is 1; the symmetric difference value between the two distribution vectors is calculated, and the symmetric difference value satisfies the three basic properties of the distance function: non-negativity, symmetry, and triangle inequality; finally, the distribution difference value is mapped to a similarity score through a monotonically decreasing function, and an adaptively adjustable conversion threshold is set. The default value range is based on historical statistical results, and the normalized similarity score is output, with the value range fixed at [0,1]; based on the characteristics of the current document type, a preset determination threshold is loaded; the final topic similarity score of adjacent windows is output, and the final score is compared with the determination threshold to generate a determination result of mutation or gradual change. The above solution eliminates the probability deviation caused by the window size difference through normalization, ensuring the comparability of topic vectors of text units with different granularities; the mathematical properties of the symmetric difference value guarantee the measurement consistency of cross-document comparison, solving the problem that traditional similarity calculations do not converge in the probability distribution space. The conversion threshold based on historical statistics and the document type characteristics constitute a two-layer adjustment mechanism: the global default value maintains the basic determination standard, and the local document feature compensation ensures the sensitivity of specific fields; the non-linear distribution difference is transformed into a directly comparable similarity scalar through a monotonic function mapping. The initialization of the reference set in the first calculation and the inheritance of the valid topic set in non-first calculations achieve the progressive expansion of the topic space; avoid the explosion of topic dimensions, while supporting the dynamic absorption of out-of-vocabulary topics, forming a stable incremental learning framework. The comparison result between the final score and the determination threshold not only outputs a discrete determination (mutation / gradual change), but also feeds back and updates the historical statistics database, enabling the determination threshold parameter to have continuous adaptability and effectively coping with the dynamic characteristics of different corpora and topic evolution patterns.
[0093] Embodiment 6: On the basis of Embodiment 5, the process of outputting the normalized similarity score provided in this embodiment includes the following steps:
[0094] S2031: Receive the symmetric difference value, which represents the difference between two topic distributions; verify that the input value meets the three basic mathematical constraints: non-negativity, symmetry, and triangle inequality; record the current difference value in the historical statistical sequence; establish a conversion relationship from the difference value to the similarity, dynamically determine the initial form parameters of the conversion curve according to the characteristics of historical data, and set the boundary conditions of the conversion curve: the minimum difference corresponds to the highest similarity, and the theoretical maximum difference corresponds to zero similarity;
[0095] S2032: Substitute the current symmetric difference value into the mapping function to obtain the initial similarity, and calibrate the stability of the initial result according to the recent data fluctuation characteristics; determine the reference baseline based on the historical data distribution characteristics, adjust the calibrated topic similarity into the standard measurement interval, and truncate the result within the specified range through the upper and lower limits.
[0096] S2033: Establish a dynamic baseline by integrating the historical similarity level and the current document characteristics, adjust the threshold considering the topic diversity index and the context consistency factor, and apply the boundary protection mechanism to prevent the threshold drift from exceeding the reasonable range.
[0097] The working principle and beneficial effects of the above technical solution are as follows: In this embodiment, first, receive the symmetric difference value, which characterizes the difference between two topic distributions; verify that the input value meets three basic mathematical constraints: non-negativity, symmetry, and triangle inequality; record the current difference value in the historical statistical sequence; establish the conversion relationship from the difference value to the similarity, dynamically determine the initial form parameters of the conversion curve according to the historical data characteristics, and set the boundary conditions of the conversion curve: the minimum difference corresponds to the highest similarity, and the theoretical maximum difference corresponds to zero similarity. Secondly, substitute the current symmetric difference value into the mapping function to obtain the initial similarity, and calibrate the stability of the initial result according to the recent data fluctuation characteristics; determine the reference baseline based on the historical data distribution characteristics, adjust the calibrated topic similarity into the standard measurement interval, and truncate the result within the specified range through the upper and lower limits. Finally, establish a dynamic baseline by integrating the historical similarity level and the current document characteristics, adjust the threshold considering the topic diversity index and the context consistency factor, and apply the boundary protection mechanism to prevent the threshold drift from exceeding the reasonable range. The initial difference value of the above solution is verified by non-negativity, symmetry, and triangle inequality to ensure that the mathematical basis of subsequent calculations meets the requirements of the metric space. At the same time, the historical sequence record provides data support for dynamic parameter adjustment; the preset boundary conditions ensure the theoretical completeness of the output range. The dynamic parameter determination mechanism based on historical data characteristics makes the similarity conversion curve adaptive, maintaining both the monotonically decreasing mathematical property and the ability to adjust the sensitivity according to the data distribution changes; the fluctuation calibration link suppresses the interference of outliers and enhances the computational robustness. The scale normalization is achieved through a two-layer adjustment mechanism: first, calibrate based on the baseline of the historical distribution, and then constrain through interval truncation to ensure that the output is strictly limited within the range of [0,1]. This processing maintains the comparability between different documents. The dynamic baseline integrates the historical statistical law and the current document characteristics, and combines the topic diversity and context coherence for weighted adjustment, making the decision threshold have the ability of context awareness; the boundary protection mechanism maintains the effectiveness and stability of the threshold.
[0098] Example 7: As Figure 6 shown, based on Embodiments 1 - 6, the RAG-based dynamic text block segmentation system provided by the embodiment of the present invention includes:
[0099] The segmentation mode module is responsible for parsing the multi-level directory structure of the document, segmenting text blocks based on directory nodes; detecting paragraph delimiters and end-of-sentence symbols as secondary segmentation boundaries; for documents with a complete table of contents, generating a tree-like text block structure that strictly corresponds to the directory hierarchy, and for documents without a table of contents or with an incomplete table of contents, switching to a mixed segmentation mode of rules and semantics, retaining the existing directory nodes as the main segmentation points, and enabling semantic segmentation in the areas where the table of contents is missing for supplementary segmentation;
[0100] Among them, the multi-level directory structure includes chapter / section / article numbers, paragraph delimiters include blank lines and indents, and end-of-sentence symbols include full stops / question marks / exclamation marks;
[0101] The tag confirmation module is responsible for performing Latent Dirichlet Allocation topic modeling on the continuous text stream, calculating the topic distribution vector of each paragraph of text in real time; calculating the topic similarity of adjacent windows through a sliding window, and when the topic similarity exceeds the threshold, determining it as a topic boundary and performing adaptive segmentation on the detected topic boundary; inserting hard segmentation tags at topic mutation points and using soft segmentation for gradually changing topic regions;
[0102] The block reorganization module is responsible for setting the initial window size of the sliding window according to the document type and monitoring the semantic density within the window in real time; hierarchical block reorganization; when there is a conflict between the rule-based segmentation and the semantic segmentation results: prioritize maintaining the integrity of the directory structure, inserting bidirectional index anchors in the conflict area, and maintaining the context association during retrieval;
[0103] Among them, the hierarchical block reorganization includes:
[0104] First-level block: Inheriting the segmentation result of the directory structure;
[0105] Second-level block: The semantic segmentation boundary applied within the directory node;
[0106] Third-level block: Dynamically adjusting the granularity of the sliding window for complex paragraphs.
[0107] The working principle and beneficial effects of the above technical solution are as follows: The segmentation mode module of this embodiment parses the multi-level directory structure of the document, and segments text blocks based on directory nodes; detects paragraph delimiters and end-of-sentence symbols as secondary segmentation boundaries; for a document with a complete directory, generates a tree-like text block structure that strictly corresponds to the directory level. For a document without a directory or with an incomplete directory, it switches to a hybrid segmentation mode of rules and semantics, retains the existing directory nodes as the main segmentation points, and enables semantic segmentation in the missing directory area for supplementary segmentation; the marking confirmation module performs Latent Dirichlet Allocation (LDA) topic modeling on the continuous text stream, and calculates the topic distribution vector of each paragraph of text in real time; calculates the topic similarity of adjacent windows through a sliding window. When the topic similarity exceeds the threshold, it is determined as a topic boundary, and adaptive segmentation is performed on the detected topic boundary; inserts a hard segmentation mark at the topic mutation point, and uses soft segmentation for the gradual topic area; the initial window size of the sliding window of the block recombination module is set according to the document type, and the semantic density within the window is monitored in real time; hierarchical block recombination; when there is a conflict between the rule-based segmentation and the semantic segmentation results: prioritize maintaining the integrity of the directory structure, insert bidirectional index anchors in the conflict area, and maintain the context association during retrieval. This embodiment realizes the precise alignment of the document directory structure and the text content through the segmentation mode module, ensuring that the inherent hierarchical relationships such as chapters / sections / articles are not lost during the block segmentation process; for a document with a complete directory, it can generate a tree-like text structure, and for a document without a directory, it maintains the existing structural framework through a hybrid segmentation of rules and semantics. Comprehensively uses format features (blank lines / indentation) and language features (end-of-sentence symbols) for initial division, combines the LDA topic modeling of the marking confirmation module, and dynamically detects topic continuity through a sliding window; the dual judgment mechanism (threshold determination and gradual change analysis) improves the accuracy of the segmentation boundary. The three-level block recombination system of the block recombination module realizes: the first level maintains the macro structure, the second level ensures topic coherence, and the third level processes micro semantic units. This mechanism supports multi-scale information organization from the chapter level to the inside of paragraphs. When the rule-based segmentation (directory-oriented) is inconsistent with the semantic segmentation (topic-oriented), the retrieval context is retained through the bidirectional index anchor technology, balancing the contradictory requirements between structural integrity and semantic rationality.
[0108] In summary, this embodiment realizes the coordinated unity of preserving the hierarchical structure, identifying topic boundaries, and meeting the multi-granularity block requirements during the document parsing process, providing text block inputs with reasonable structure and semantic coherence for the retrieval-augmented generation task.
[0109] The block processing method for text data in this embodiment is mainly applied to scenarios including: processing complex structure documents, multi-turn conversations and interactive systems, vertical domain specific scenarios, etc. A typical scenario example diagram is shown in Table 1:
[0110] Table 1
[0111]
[0112]
[0113] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of equivalent technologies of the present invention, the present invention is also intended to include these changes and modifications.
Claims
1. A method for dynamically segmenting text blocks based on RAG, characterized in that, The following steps are involved: Parse the multi-level directory structure of the document and segment the text blocks based on the directory nodes; detect paragraph separators and sentence end symbols as secondary segmentation boundaries; For documents with complete directories, a tree-like text block structure that strictly corresponds to the directory hierarchy is generated. For documents without directories or with incomplete directories, a mixed segmentation mode of rules and semantics is used, with the existing directory nodes retained as the main segmentation points, and semantic segmentation enabled for supplementary segmentation in the directory-missing areas. Perform latent Dirichlet allocation topic modeling on continuous text streams, calculate the topic distribution vector of each text segment in real time, calculate the topic similarity of adjacent windows through sliding windows, and determine the topic boundary when the topic similarity exceeds the threshold. Perform adaptive segmentation on the detected topic boundary, insert hard segmentation markers at the topic mutation point, and use soft segmentation for the gradual topic area; The initial window size of the sliding window is set according to the document type, and the semantic density in the window is monitored in real time; hierarchical block reorganization; when the rule segmentation conflicts with the semantic segmentation results: priority is given to preserving the integrity of the directory structure, inserting a bidirectional index anchor in the conflicting area, and maintaining the context association during retrieval.
2. The RAG-based dynamic text block segmentation method according to claim 1, wherein in, The multi-level directory structure includes chapter, section or article numbers, paragraph separators include blank lines or indents, and sentence end symbols include periods, question marks or exclamation marks.
3. The method for dynamically splitting text blocks based on RAG according to claim 1, characterized in that, The process of parsing the multi-level directory structure of a document includes the following steps: Identify hierarchical markup features in the document and extract explicit chapter, section or article numbering systems as the basis for trunk division; automatically generate independent text containers for each numbering node to form an initial tree-like skeleton structure; perform paragraph granularity analysis for each container unit in the skeleton structure: detect the visual isolation zone formed by blank lines as a paragraph separation mark, and identify the text alignment boundaries generated by indentation changes as secondary division clues; In the divided paragraph units, sentence-level parsing is implemented to construct the minimum semantic unit boundary through the sentence-end punctuation triples; the list item detection mode is automatically activated for paragraphs containing numbered sequences to identify the substructure formed by continuous numbering; when a missing area of the directory is detected, the compensation mechanism is automatically activated to retain the existing directory nodes as fixed anchor points and construct virtual nodes in the missing intervals using semantic continuity analysis; The initial segmentation scheme derived from the directory is aligned with the symbol analysis results, and areas that still have conflicts after alignment are marked with abnormal signs and transferred to the sliding window analysis process.
4. The method for dynamically segmenting text blocks based on RAG according to claim 3, characterized in that, When the deviation between the catalog segmentation boundary and the natural paragraph boundary exceeds the preset tolerance, the following processing solutions are selected according to the context density; High-density terminology area: keep directory boundaries first; Narrative text area: Use natural paragraph boundaries; Transition Zone: Inserts collapsible virtual segment markers.
5. The method for dynamically segmenting text blocks based on RAG according to claim 3, wherein The process of constructing the minimum semantic unit boundary includes the following steps: The paragraph unit obtained by using the paragraph separation mark is used as the input basic structure of the minimum semantic unit boundary; the text flow in each container unit is used as the initial processing object; Within a given paragraph unit, scan the text character sequence to locate the position coordinates of the end-of-sentence punctuation triples. The existence verification of the end-of-sentence punctuation triples is based on paragraph boundary constraints; the text intervals between the triples generate primary semantic blocks, and the length of the primary semantic blocks is restricted by the result of paragraph granularity analysis; For the primary semantic blocks containing number prefixes, trigger the list item detection mode. The triggering condition is to identify the similarity between the substructure formed by consecutive numbers and the established hierarchical marking features. The consecutive numbered items automatically form a substructure tree, and the root node of the substructure tree establishes a parent-child association with the paragraph container node; When the sequence of primary semantic blocks is interrupted, start the vacancy area compensation mechanism. Use the parsed normal structure as an anchor reference system. By analyzing the verb tense patterns and entity reference relationships of the semantic blocks on both sides of the interrupted area, deduce the potential structural continuity, and the generated virtual nodes inherit the same hierarchical attributes as the adjacent real nodes.
6. The method for dynamically segmenting text blocks based on RAG according to claim 1, wherein, The process of performing adaptive segmentation on the detected topic boundaries includes the following steps: Based on the text content within the sliding window, calculate the topic distribution vector of the current window. The topic distribution vector represents the probability weights of each topic within the window; if the text contains a table of contents structure, the window is initialized to align with the table of contents node boundaries; if there is no table of contents structure, the window size is generated according to the initial setting; Adopt a symmetric probability distribution difference metric to compare the topic distribution vectors of adjacent windows; If the similarity is lower than the preset threshold, it is determined as a topic mutation; if the similarity is higher than the threshold but shows a continuous downward trend, it is determined as a topic gradual change area; The window size is dynamically adjusted according to the detection result; Insert a segmentation mark at the mutation point, insert an expandable segmentation mark in the gradual change area, and retrieve the cross-boundary associated context; if the segmentation point conflicts with the table of contents node, give priority to maintaining the integrity of the table of contents, and insert a two-way index anchor at the conflict position; Structured output of the segmentation result: type mark, original table of contents position, associated topic distribution vector, and hierarchical integration.
7. The method for dynamically segmenting text blocks based on RAG according to claim 6, wherein The process of calculating the topic similarity includes the following steps: Obtain the topic distribution vectors of adjacent windows as input, where each topic distribution vector contains the probability weight values of each topic within the window; if it is the first calculation, initialize a reference set containing all known topics; if not the first time, inherit the valid topic set confirmed by the previous calculation; Normalize the topic distribution vectors of adjacent windows to ensure that the sum of each component is 1; calculate the symmetric difference value between the two distribution vectors, and the symmetric difference value satisfies the three basic properties of non-negativity, symmetry, and triangle inequality of the distance function; Map the distribution difference value to a similarity score through a monotonically decreasing function, set an adaptively adjustable conversion threshold, the default value range is based on historical statistical results, output the normalized similarity score, and the value range is fixed at [0,1]; based on the characteristics of the current document type, load the preset determination threshold; Output the final topic similarity score of adjacent windows, compare the final score with the determination threshold, and generate a determination result of mutation or gradual change.
8. The method for dynamically splitting text blocks based on RAG according to claim 7, wherein The process of outputting the normalized similarity score includes the following steps: Receive the symmetric difference value, which characterizes the difference between two topic distributions; verify that the input value conforms to three basic mathematical constraints; record the current difference value into the historical statistical sequence; establish the conversion relationship from the difference value to the similarity, dynamically determine the initial form parameters of the conversion curve according to the characteristics of historical data, and set the boundary conditions of the conversion curve: the minimum difference corresponds to the highest similarity, and the theoretical maximum difference corresponds to zero similarity; Substitute the current symmetric difference value into the mapping function to obtain the initial similarity, and calibrate the stability of the initial result according to the characteristics of recent data fluctuations; Determine the reference baseline based on the characteristics of the historical data distribution, adjust the calibrated topic similarity into the standard measurement interval, and truncate the result within the specified range through the upper and lower limits; Establish a dynamic baseline by integrating the historical similarity level and the current document characteristics, adjust the threshold considering the topic diversity index and the context consistency factor, and apply the boundary protection mechanism to prevent the threshold drift from exceeding the reasonable range.
9. The method for dynamically segmenting text blocks based on RAG according to claim 1, wherein Among them, Hierarchical block reorganization includes: First-level block: inherit the directory structure segmentation result; Second-level block: semantic segmentation boundary applied inside the directory node; Third-level block: perform dynamic granularity adjustment of the sliding window for complex paragraphs.
10. A text block dynamic segmentation system based on RAG, characterized in that, Include: The segmentation mode module is responsible for parsing the multi-level directory structure of the document, segmenting text blocks based on directory nodes; detecting paragraph separators and end-of-sentence symbols as secondary segmentation boundaries; For documents with a complete directory, generate a tree-like text block structure that strictly corresponds to the directory level. For documents without a directory or with an incomplete directory, switch to a mixed segmentation mode of rules and semantics, retain the existing directory nodes as the main segmentation points, and enable semantic segmentation in the missing directory area for supplementary segmentation; The marker confirmation module is responsible for performing latent Dirichlet allocation topic modeling on the continuous text stream, and calculating the topic distribution vector of each text segment in real time; calculate the topic similarity of adjacent windows through the sliding window, and determine the topic boundary when the topic similarity exceeds the threshold, and perform adaptive segmentation on the detected topic boundary; insert hard segmentation markers at the topic mutation points, and use soft segmentation for the gradual topic areas; The block reorganization module is responsible for setting the initial window size of the sliding window according to the document type, and monitoring the semantic density within the window in real time; hierarchical block reorganization; when the results of rule-based segmentation and semantic segmentation conflict: give priority to maintaining the integrity of the directory structure, insert bidirectional index anchors in the conflict area, and maintain the context association during retrieval.
Citation Information
Patent Citations
A text segmentation method, device and medium for large language model
CN118536497B
Text segmentation method and device for retrieval enhancement generation
CN119311723A
Text segmentation method and device, storage medium and electronic equipment
CN119311830A
Document knowledge base-oriented multi-granularity structured retrieval enhancement generation method and device
CN118585615A
Large language model RAG optimization method based on tree neighbor context
CN119293195A
Cited By
Semantic enhancement adaptive partitioning method and system for natural resource large model questions and answers
CN120822526A
Semantic analysis method and system based on history teaching
CN120952008A
Medical information retrieval method and system based on multi-modal vectorization and storage medium
CN120994812A
Knowledge slice analysis processing method and system based on context window semantic clustering
CN121009195A
Financial institution system information conflict detection method and device, equipment and medium
CN121119449A