Virtual chapter construction and long text recall method based on similarity clustering

By constructing a virtual chapter structure based on similarity clustering, the problems of semantic depth and connectivity in long text recall are solved, ensuring the accuracy and efficiency of long text queries, and it is suitable for complex queries of structured or semi-structured documents.

CN120597879APending Publication Date: 2025-09-05WIND INFORMATION TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510930595.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-07
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

Existing long text recall methods cannot effectively capture the full semantic depth and connectivity of the text, resulting in incomplete answers or waste of computational resources, especially when dealing with topic questions that require integrating knowledge across different texts.

Method used

By constructing a virtual chapter structure based on similarity clustering, a recursive tree structure is used to organize text content, including semantic segmentation, low-dimensional vector conversion, local and global clustering, and tree-like hierarchical relationships, to ensure that a sufficiently long context is recalled.

Benefits of technology

It achieves accurate and comprehensive answers to long texts, reduces resource waste, improves computing efficiency and retrieval performance, and is suitable for complex queries of structured or semi-structured documents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120597879A_ABST
    Figure CN120597879A_ABST
Patent Text Reader

Abstract

The invention provides a similarity clustering-based virtual chapter construction and long text recall method, which is characterized in that text similarity is utilized to divide article content into a plurality of virtual'chapters', and the'chapters' are not actual chapters of an article but are semantically related text block sets formed by clustering, so that the content of a long text can be recalled; and then a recursive tree structure is constructed to ensure that the context which is long enough can be recalled, the long text requirement of complex query is met, and the problems of semantic depth and connectivity in the prior art are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of Internet technology, and in particular to a method for constructing virtual chapters and recalling long texts based on similarity clustering. Background Art

[0002] As large language models (LLMs) continue to scale, they have become powerful tools for many tasks. They serve as effective knowledge stores, encoding factual information in their parameters. The performance of these models can be further improved through fine-tuning on downstream tasks. However, even the largest models today cannot fully capture the domain knowledge required for specific tasks, and the information in the models can become outdated over time. Updating the knowledge in these models is challenging, especially when processing large amounts of text resources.

[0003] One existing solution is proposed in open-domain question answering systems, which uses a separate information retrieval system to index large amounts of text and process it by splitting the text into paragraphs. In this way, the indexed information can be provided as context to the large language model when answering questions.

[0004] However, existing retrieval enhancement techniques have certain limitations. Most methods only retrieve short, continuous text fragments, which limits their ability to capture complex discourse structures, especially when dealing with thematic problems that require integrating knowledge across different texts, such as understanding the content of an entire book.

[0005] Currently, there are mainly the following methods:

[0006] 1) Fusion-in-Decoder (FiD) combines the Dense Passage Retrieval (DPR) and BM25 retrieval models. The key idea is to process multiple paragraphs independently in the encoder phase and then fuse their outputs in the decoder phase, allowing the model to extract knowledge from multiple sources. This continuous segmentation approach can prevent the model from capturing the full semantic depth of the text. The extracted fragments may lack important context and may even be misleading.

[0007] 2) The Retrieval-augment Transformer (RETRO) is a model that utilizes cross-module attention and block-level retrieval. It splits long text into multiple blocks and uses attention across these blocks during generation, enabling the model to better capture contextual information when generating retrieval-based text. This approach has difficulties handling long-range dependencies, especially when information needs to span multiple blocks, which can lead to missing information or inadequate understanding.

[0008] 3) Retrieval-Augmented Generation (RAG) technology is a model that combines a retriever and a generator. The retriever is responsible for finding text fragments related to the input query from a large document collection, and the generator uses these paragraphs to generate answers. For example, the existing patent CN119167921A - A RAG text processing method, device and medium based on multi-way recall. This approach also relies on the standard retrieval method, that is, dividing the corpus into blocks and encoding it using a BERT-based retriever, which cannot capture the full semantic depth of the text.

[0009] In the chatdoc project, some queries require a large amount of original text to be answered completely and correctly. However, traditional RAG methods have problems recalling long texts and cannot effectively capture complete contextual information, resulting in incomplete answers or wasted computing resources:

[0010] 1. The method of hitting a chunk and then expanding the context has the problem of unclear expansion boundaries; too much expansion will waste resources, while too little expansion will result in content loss.

[0011] 2. The Tree RAG method attempts to cluster related text blocks and then summarize them, but both the clustering and summarization processes may introduce errors and cannot effectively deal with scenarios that require long answers. Summary of the Invention

[0012] The purpose of the technical solution of the present invention is to propose a virtual chapter construction and long text recall method based on similarity clustering, which solves the problems of semantic depth and connectivity by constructing a recursive tree structure.

[0013] The technical solution of the present invention provides a method for constructing virtual chapters and recalling long texts based on similarity clustering, comprising the following steps:

[0014] Determine whether the sentence in the document contains a preset semantic boundary. If so, continue to determine whether it exceeds the preset fixed segmentation length. If not, segment according to the preset semantic boundary to obtain semantically segmented text blocks;

[0015] If the preset fixed segmentation length is exceeded, the text is segmented according to the preset fixed segmentation length to obtain a first segmented text block and a corresponding first segmentation marker, and the next adjacent sentence is judged to contain a preset semantic boundary. If not, the text is segmented according to the preset fixed segmentation length to obtain a second segmented text block and a corresponding second segmentation marker, until the next adjacent sentence is judged to contain a preset semantic boundary, and the text is segmented according to the preset semantic boundary to obtain a third segmented text block and a corresponding third segmentation marker, so as to preserve semantic integrity;

[0016] The semantic segmentation text block, the first segmentation text block, the second segmentation text block, and the third segmentation text block are converted into corresponding embedding vectors, and the corresponding features are retained to compress the embedding vectors into low-dimensional space vectors to obtain the semantic segmentation low-dimensional vector, the first segmentation text low-dimensional vector, the second segmentation text low-dimensional vector, and the third segmentation text low-dimensional vector to solve the high-dimensional distance failure problem;

[0017] A local sliding window is introduced, and similarity, the first segmentation marker, the second segmentation marker, and the third segmentation marker are calculated based on the adjacent window. The semantic segmentation low-dimensional vector, the first segmentation text low-dimensional vector, the second segmentation text low-dimensional vector, and the third segmentation text low-dimensional vector are initially clustered to obtain virtual sections to ensure text continuity within the virtual sections.

[0018] Use global clustering to divide coarse-grained topics to perform the first-level clustering of virtual sections. Use the Bayesian Information Criterion (BIC) to select the optimal number of clusters. If the span of the virtual sections in the cluster corresponding to the optimal number of clusters exceeds the preset threshold of the total document length, the cluster is recursively split until the span of the virtual sections exceeds the preset threshold of the total document length. Use global clustering to divide the virtual sections into coarse-grained topics, calculate the mean cosine similarity within the cluster as the topic consistency score, and if the topic consistency score is less than the preset score, recursively split the cluster until the topic consistency score is greater than or equal to the preset score. Continue to perform the second-level local clustering within the global cluster to refine the semantic units to complete the secondary clustering to generate virtual sections.

[0019] The semantic segmentation text block, the first segmentation text block, the second segmentation text block, and the third segmentation text block are defined as leaf nodes, the virtual section summary is defined as an intermediate node, and the virtual section is used as the parent node. Child nodes are distinguished according to semantic relationships and the child node ID list is saved to the parent node. It is determined whether the content overlap rate of the parent node summary and the child node content is greater than 80%. If so, the child node content is deleted and the parent node summary is retained to compress and construct a tree hierarchy.

[0020] The embedding vectors corresponding to leaf nodes, intermediate nodes, parent nodes, and child nodes are stored in the FAISS index, and the tree hierarchy, parent-child relationship, and text start and end positions are added.

[0021] When recalling long text, if the query vector hits a non-leaf node, the tree structure is traversed downward to the leaf node. The content of all child nodes related to the leaf node is recalled based on the tree hierarchy and parent-child relationships. If the query vector hits a leaf node, the query is traced back to the upper-level node, the high-level summary is supplemented, the cosine similarity between the query vector and all nodes is calculated, and the node content is recalled according to the threshold. If the query vector is a short query, the high-level nodes are recalled first. If the query vector is a long query, the leaf nodes and intermediate nodes are mixed.

[0022] Preferably, the preset semantic boundaries include natural paragraph boundaries such as punctuation marks and paragraph separators.

[0023] Preferably, the conversion of the semantic segmented text block, the first segmented text block, the second segmented text block and the third segmented text block into corresponding embedding vectors can adopt a pre-trained SBERT model.

[0024] Preferably, compressing the embedding vector to a low-dimensional space vector may adopt UMAP dimensionality reduction.

[0025] Preferably, the semantic segmentation low-dimensional vector, the first segmentation text low-dimensional vector, the second segmentation text low-dimensional vector and the third segmentation text low-dimensional vector are clustered using a Gaussian mixture model.

[0026] Preferably, the formula for selecting the optimal number of clusters is as follows:

[0027]

[0028] Among them, N is the number of text blocks, k is the number of model parameters, is the maximum likelihood value.

[0029] Preferably, the text start and end positions include a text start position and a text end position, the text start position is defined as the smallest text block start offset in the associated node, and the text end position is defined as the largest text block end offset in the associated node.

[0030] The technical solution of the present invention proposes a method for constructing virtual chapters and recalling long texts based on similarity clustering. By utilizing text similarity, the article content is divided into multiple virtual "chapters". These "chapters" are not the actual chapters of the article, but a collection of semantically related text blocks formed by clustering. Then, a recursive tree structure is constructed to ensure that a sufficiently long context can be recalled to meet the long text requirements of complex queries, thereby solving the problems of semantic depth and connectivity in the existing technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 A flowchart of a method for constructing virtual chapters and recalling long texts based on similarity clustering is provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0032] Below in conjunction with specific embodiment, further set forth the present invention.Should be understood that these embodiments are only used to illustrate the present invention and are not used in limiting the scope of the present invention.In addition, should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms fall equally within the scope limited by the appended claims of the application.

[0033] The embodiment of the present invention provides a method for constructing virtual chapters and recalling long texts based on similarity clustering, comprising the following steps:

[0034] Determine whether a sentence in the document contains a preset semantic boundary. If so, determine whether it exceeds a preset fixed segmentation length. If not, segment the text block according to the preset semantic boundary to obtain a semantically segmented text block. The preset semantic boundary includes natural paragraph boundaries such as punctuation and paragraph separators.

[0035] If it exceeds the preset fixed segmentation length, the text block is segmented according to the preset fixed segmentation length to obtain the first segmented text block and the corresponding first segmentation marker, and the next adjacent sentence is judged to contain the preset semantic boundary. If not, the text block is segmented according to the preset fixed segmentation length to obtain the second segmented text block and the corresponding second segmentation marker, until the next adjacent sentence is judged to contain the preset semantic boundary, and the text block is segmented according to the preset semantic boundary to obtain the third segmented text block and the corresponding third segmentation marker to retain semantic integrity.

[0036] The semantic segmentation text block, the first segmentation text block, the second segmentation text block, and the third segmentation text block are converted into corresponding embedding vectors, and the corresponding features are retained to compress the embedding vectors into low-dimensional space vectors to obtain semantic segmentation low-dimensional vectors, first segmentation text low-dimensional vectors, second segmentation text low-dimensional vectors, and third segmentation text low-dimensional vectors to solve the high-dimensional distance failure problem. The semantic segmentation text block, the first segmentation text block, the second segmentation text block, and the third segmentation text block are converted into corresponding embedding vectors using a pre-trained SBERT (Sentence-BERT) model, and the embedding vectors are compressed into low-dimensional space vectors using UMAP dimensionality reduction.

[0037] A local sliding window is introduced. Based on the window's adjacent similarity calculations, the first segmentation marker, the second segmentation marker, and the third segmentation marker, the semantic segmentation low-dimensional vector, the first segmentation text low-dimensional vector, the second segmentation text low-dimensional vector, and the third segmentation text low-dimensional vector are initially clustered to obtain virtual sections to ensure text continuity within the virtual sections. A Gaussian Mixed Model (GMM) is used to cluster the semantic segmentation low-dimensional vector, the first segmentation text low-dimensional vector, the second segmentation text low-dimensional vector, and the third segmentation text low-dimensional vector.

[0038] Use global clustering to divide coarse-grained topics to perform the first-level clustering on virtual sections, and select the optimal number of clusters using the Bayesian Information Criterion (BIC). The formula is as follows:

[0039]

[0040] Among them, N is the number of text blocks, k is the number of model parameters, is the maximum likelihood value.

[0041] If the span of the virtual subsections in the cluster corresponding to the optimal number of clusters exceeds 20% of the total document length, the clusters are recursively split until the span of the virtual subsections does not exceed 20% of the total document length. Global clustering is used to divide the virtual subsections into coarse-grained topics, and the mean cosine similarity within the cluster is calculated as the topic consistency score. If the topic consistency score is less than 0.6, the clusters are recursively split until the topic consistency score is greater than or equal to 0.6. The second layer of local clustering is continued within the global cluster to refine the semantic units to complete the secondary clustering to generate virtual chapters.

[0042] The semantic segmentation text block, the first segmentation text block, the second segmentation text block and the third segmentation text block are defined as leaf nodes, the virtual section summary is defined as the intermediate node, the virtual chapter is used as the parent node, the child nodes are distinguished according to the semantic relationship and the child node ID list is saved to the parent node, and it is determined whether the content overlap rate of the parent node summary and the child node content is greater than 80%. If so, the child node content is deleted and the parent node summary is retained to compress and construct a tree hierarchy.

[0043] The embedding vectors corresponding to leaf nodes, intermediate nodes, parent nodes, and child nodes are stored in the FAISS index, along with the tree hierarchy, parent-child relationship, and text start and end positions. The text start and end positions include the text start position and text end position. The text start position is defined as the smallest text block start offset in the associated node, and the text end position is defined as the largest text block end offset in the associated node.

[0044] When recalling long text, if the query vector hits a non-leaf node, the tree structure is traversed downward to the leaf node. The content of all child nodes related to the leaf node is recalled based on the tree hierarchy and parent-child relationships. If the query vector hits a leaf node, the query is traced back to the upper-level node, the high-level summary is supplemented, and the cosine similarity between the query vector and all nodes is calculated. The node content is recalled according to a threshold (such as Top-20). If the query vector belongs to a short query (<10 words), the high-level nodes are prioritized (weight = 0.7). If the query vector belongs to a long query (>20 words), the leaf nodes (weight = 0.4) and the intermediate nodes (weight = 0.6) are mixed.

[0045] The beneficial effects of the embodiments of the present invention are as follows:

[0046] 1) By constructing a virtual chapter structure, the incomplete recall problem of traditional RAG methods when processing long texts is effectively solved, ensuring that long text questions can be answered accurately and comprehensively.

[0047] 2) It reduces resource waste during context expansion and improves computational efficiency, especially when facing long text retrieval and answering, and can avoid recalling irrelevant content.

[0048] 3) The virtual chapter structure constructed by similarity clustering enables the system to efficiently organize and recall document content even when there is no clear article structure, improving retrieval performance and answer quality.

[0049] 4) It is suitable for vertical field documents, especially structured or semi-structured documents, and can effectively support long text retrieval tasks and improve the processing capabilities of complex queries.

Claims

1. A method for constructing virtual chapters and recalling long texts based on similarity clustering, characterized in that: The following steps are involved: Determine whether the sentence in the document contains a preset semantic boundary. If so, continue to determine whether it exceeds the preset fixed segmentation length. If not, segment according to the preset semantic boundary to obtain semantically segmented text blocks; If the preset fixed segmentation length is exceeded, the text is segmented according to the preset fixed segmentation length to obtain a first segmented text block and a corresponding first segmentation marker, and the next adjacent sentence is judged to contain a preset semantic boundary. If not, the text is segmented according to the preset fixed segmentation length to obtain a second segmented text block and a corresponding second segmentation marker, until the next adjacent sentence is judged to contain a preset semantic boundary, and the text is segmented according to the preset semantic boundary to obtain a third segmented text block and a corresponding third segmentation marker, so as to preserve semantic integrity; The semantic segmentation text block, the first segmentation text block, the second segmentation text block, and the third segmentation text block are converted into corresponding embedding vectors, and the corresponding features are retained to compress the embedding vectors into low-dimensional space vectors to obtain the semantic segmentation low-dimensional vector, the first segmentation text low-dimensional vector, the second segmentation text low-dimensional vector, and the third segmentation text low-dimensional vector to solve the high-dimensional distance failure problem; A local sliding window is introduced, and similarity, the first segmentation marker, the second segmentation marker, and the third segmentation marker are calculated based on the adjacent window. The semantic segmentation low-dimensional vector, the first segmentation text low-dimensional vector, the second segmentation text low-dimensional vector, and the third segmentation text low-dimensional vector are initially clustered to obtain virtual sections to ensure text continuity within the virtual sections. Use global clustering to divide coarse-grained topics to perform the first-level clustering of virtual sections. Use the Bayesian Information Criterion (BIC) to select the optimal number of clusters. If the span of the virtual sections in the cluster corresponding to the optimal number of clusters exceeds the preset threshold of the total document length, the cluster is recursively split until the span of the virtual sections exceeds the preset threshold of the total document length. Use global clustering to divide the virtual sections into coarse-grained topics, calculate the mean cosine similarity within the cluster as the topic consistency score, and if the topic consistency score is less than the preset score, recursively split the cluster until the topic consistency score is greater than or equal to the preset score. Continue to perform the second-level local clustering within the global cluster to refine the semantic units to complete the secondary clustering to generate virtual sections. The semantic segmentation text block, the first segmentation text block, the second segmentation text block, and the third segmentation text block are defined as leaf nodes, the virtual section summary is defined as an intermediate node, and the virtual section is used as the parent node. Child nodes are distinguished according to semantic relationships and the child node ID list is saved to the parent node. It is determined whether the content overlap rate of the parent node summary and the child node content is greater than 80%. If so, the child node content is deleted and the parent node summary is retained to compress and construct a tree hierarchy. The embedding vectors corresponding to leaf nodes, intermediate nodes, parent nodes, and child nodes are stored in the FAISS index, and the tree hierarchy, parent-child relationship, and text start and end positions are added. When recalling long text, if the query vector hits a non-leaf node, the tree structure is traversed downward to the leaf node. The content of all child nodes related to the leaf node is recalled based on the tree hierarchy and parent-child relationships. If the query vector hits a leaf node, the query is traced back to the upper-level node, the high-level summary is supplemented, the cosine similarity between the query vector and all nodes is calculated, and the node content is recalled according to the threshold. If the query vector is a short query, the high-level nodes are recalled first. If the query vector is a long query, the leaf nodes and intermediate nodes are mixed.

2. The method for constructing virtual chapters and recalling long texts based on similarity clustering according to claim 1, characterized in that: The preset semantic boundaries include natural paragraph boundaries such as punctuation marks and paragraph separators.

3. The method for constructing virtual chapters and recalling long texts based on similarity clustering according to claim 1, characterized in that: The conversion of the semantic segmentation text block, the first segmentation text block, the second segmentation text block and the third segmentation text block into corresponding embedding vectors can adopt a pre-trained SBERT model.

4. The method for constructing virtual chapters and recalling long texts based on similarity clustering according to claim 1, wherein: The compression of the embedding vector to a low-dimensional space vector may be performed by using UMAP dimensionality reduction.

5. The method for constructing virtual chapters and recalling long texts based on similarity clustering according to claim 1, characterized in that: The semantic segmentation low-dimensional vector, the first segmentation text low-dimensional vector, the second segmentation text low-dimensional vector and the third segmentation text low-dimensional vector are clustered using a Gaussian mixture model.

6. The method for constructing virtual chapters and recalling long texts based on similarity clustering according to claim 1, characterized in that: The formula for selecting the optimal number of clusters is as follows: Among them, N is the number of text blocks, k is the number of model parameters, is the maximum likelihood value.

7. The method for constructing virtual chapters and recalling long texts based on similarity clustering according to claim 1, characterized in that: The text start and end positions include a text start position and a text end position. The text start position is defined as the smallest text block start offset in the associated node, and the text end position is defined as the largest text block end offset in the associated node.

Citation Information

Patent Citations

  • RAG text processing method and device based on multi-path recall and medium

    CN119167921A