Long text processing method and system based on generative compression and two-stage retrieval
Through generative compression and two-stage search methods, the problems of insufficient metadata information density and retrieval noise are solved, the accuracy and efficiency of long text retrieval is improved, and the multi-hop reasoning and professional field needs are supported in complex scenarios.
Patent Information
- Application Number
- CN202510400936.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-18
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the existing natural language processing and information retrieval technologies, the metadata information density is insufficient, the retrieval noise and efficiency bottlenecks are serious, and the context is missing in the generation stage, resulting in low retrieval accuracy and high calculation cost.
Generative compression and two-stage search method are adopted to generate compressed metadata through a large language model, filter the Top-K candidate sets in the first round, enhance the original text layer in the second round, combine the thinking chain prompt generation, and dynamically block the semantic units to improve semantic retention and efficiency.
Improves the accuracy of long text retrieval, reduces noise interference, supports multi-hop reasoning and professional field needs, reduces calculation costs, and improves the quality of retrieval recall and generates summary.
Smart Images

Figure CN120336513A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing and information retrieval, and particularly to a long text processing method combining generative model compression and dynamic retrieval enhancement. Background Art
[0002] In scenarios such as large-scale text knowledge base construction, intelligent question answering systems, and professional field literature analysis, natural language processing and information retrieval technologies are required.
[0003] However, in the existing natural language processing and information retrieval technologies, the following problems exist:
[0004] 1. Insufficient metadata information density: Traditional retrieval systems rely on manually annotated metadata (such as titles, keywords), which have short fields and low information density, making it difficult to capture the deep semantics of long texts.
[0005] 2. Retrieval noise and efficiency bottleneck: When directly retrieving long texts, redundant information interference leads to a decrease in precision (e.g., the recall rate of the BM25 algorithm is only 52%); existing methods (such as DPR) need to encode the full text, resulting in high computational costs.
[0006] 3. Lack of context in the generation stage: Summaries generated based on short metadata are prone to missing key details (such as the clause relevance in legal contracts). Summary of the Invention
[0007] The present invention provides a long text processing method and system based on generative compression and two-stage retrieval to solve the problems raised in the above background art.
[0008] To achieve the above object, the present invention adopts the following technical solutions:
[0009] A long text processing method based on generative compression and two-stage retrieval includes the following specific steps:
[0010] S1, generative compression stage;
[0011] S101, input the original long text corpus into the generative compression module;
[0012] S102, conduct compression quality assessment. If it fails, return to the generative compression module for compression. If it passes, input it into the metadata index library;
[0013] S2, two-stage retrieval and generation;
[0014] S201, perform the first-round retrieval to screen the Top-K candidate set;
[0015] S202, perform a second-round enhancement on the original text layer and generate using chain-of-thought prompting.
[0016] As a further improvement of this technical solution: The original long text corpus includes scientific research papers and legal contracts.
[0017] As a further improvement of this technical solution: The generative compression module uses a large language model to generate compressed metadata, and the large language model includes GPT-4 and Claude-3.
[0018] As a further improvement of this technical solution: In S102, compression quality assessment is performed. Specifically: The semantic fidelity is evaluated through ROUGE-L and BERTScore, and the metadata with the semantic consistency between the compressed text and the original text ≥ 85% is retained.
[0019] As a further improvement of this technical solution: In S201, the top-K candidate set is retrieved and screened in the first round. Specifically: In the first-round retrieval, the cosine similarity between the query and the metadata is calculated using a dense retrieval model, and the top-K candidate set is screened.
[0020] As a further improvement of this technical solution: The dense retrieval model is DPR.
[0021] As a further improvement of this technical solution: In S202, the original text layer is enhanced in the second round, and the generation is performed using the Chain-of-Thought prompt. Specifically: The associated original text is dynamically segmented by a dynamic segmentation engine, and the summary is generated using the Chain-of-Thought prompt.
[0022] As a further improvement of this technical solution: The dynamic segmentation by the dynamic segmentation engine is specifically as follows: The original text is segmented by semantic units, and the mapping relationship is established by combining NER to identify key entities.
[0023] As a further improvement of this technical solution: The semantic units include chapters and paragraphs.
[0024] The present invention also proposes a long text processing system based on generative compression and two-stage retrieval, which is applied to the long text processing method based on generative compression and two-stage retrieval described in any one of the above, and includes a generative compression module, a compression quality assessment, a first-round retrieval module, and a second-round enhancement module;
[0025] The compression module is used to generate compressed metadata using a large language model;
[0026] The compression quality assessment is used to perform compression quality assessment. If it fails, it returns to the generative compression module for compression. If it passes, it inputs the metadata index library;
[0027] The first-round retrieval module is used to retrieve and screen the top-K candidate set in the first round;
[0028] The secondary enhancement module is used to perform secondary enhancement on the original text layer and generates prompts using the chain of thought.
[0029] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0030] 1. Improve the retrieval accuracy of long texts: Generate compressed semantic dense metadata through generative compression to reduce noise interference.
[0031] 2. Balance efficiency and semantic retention: Quickly screen candidate sets with the primary metadata in the first round and supplement details by associating with the original text in the second round.
[0032] 3. Dynamically adapt to complex scenarios: Support multi-hop reasoning (such as cross-paragraph association analysis) and professional field requirements (such as medical literature retrieval, legal clause verification).
[0033] The above description is only an overview of the technical solution of the present invention. In order to understand the technical means of the present invention more clearly and implement it according to the content of the specification, the following describes in detail with reference to the preferred embodiments of the present invention and the accompanying drawings. The specific implementation manners of the present invention are given in detail by the following embodiments and their accompanying drawings. Description of the Drawings
[0034] The drawings described herein are used to provide a further understanding of the present invention and form a part of this application. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0035] Figure 1 is the method flow chart of the long text processing method based on generative compression and two-stage retrieval proposed by the present invention;
[0036] Figure 2 is the structural schematic diagram of the meta-generative compression module proposed by the present invention;
[0037] Figure 3 is the schematic diagram of the two-stage retrieval logic proposed by the present invention. Detailed Description of the Preferred Embodiments
[0038] The principles and features of the present invention are described below with reference to the accompanying drawings. The examples given are only used to explain the present invention and are not used to limit the scope of the present invention. The present invention is described more specifically by way of example in the following paragraphs with reference to the accompanying drawings. It should be noted that the drawings are all in a very simplified form and use non-precise scales, only for the purpose of facilitating and clearly assisting in explaining the purpose of the embodiments of the present invention.
[0039] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this invention belongs. The terms used in the description of the present invention herein are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.
[0040] In an embodiment of the present invention, a long text processing method based on generative compression and two-stage retrieval includes the following specific steps:
[0041] S1, the generative compression stage;
[0042] S101, input the original long text corpus into the generative compression module, where the original long text corpus includes scientific research papers and legal contracts;
[0043] S102, perform compression quality assessment. Evaluate semantic fidelity through ROUGE-L and BERTScore, and retain the metadata with the semantic consistency of the compressed text and the original text ≥ 85%. If not passed, return to the generative compression module for compression. If passed, input it into the metadata index library, where the generative compression module uses a large language model to generate compressed metadata, and the large language model includes GPT-4 and Claude-3;
[0044] S2, two-stage retrieval and generation;
[0045] S201, the first-round retrieval filters the Top-K candidate set. Specifically: the first-round retrieval uses a dense retrieval model (such as DPR) to calculate the cosine similarity between the query and the metadata, and filters the Top-K candidate set;
[0046] S202, perform a second-round enhancement on the original text layer, and generate using the Chain-of-Thought prompt. Specifically: dynamically divide the associated original text through a dynamic chunking engine, and generate a summary using the Chain-of-Thought prompt.
[0047] The dynamic chunking engine performs dynamic chunking specifically as follows: divide the original text according to semantic units (including chapters and paragraphs), and establish a mapping relationship by combining NER to identify key entities.
[0048] The present invention also proposes a long text processing system based on generative compression and two-stage retrieval, which is applied to the long text processing method based on generative compression and two-stage retrieval described in any one of the above, and includes a generative compression module, a compression quality assessment, a first-round retrieval module, and a second-round enhancement module;
[0049] The compression module is used to generate compressed metadata using a large language model;
[0050] Compression quality assessment is used to conduct compression quality assessment. If it fails, it returns to the generative compression module for compression. If it passes, it inputs the metadata index library;
[0051] The first-round retrieval module is used to retrieve and screen the Top-K candidate set in the first round;
[0052] The second-round enhancement module is used to perform a second-round enhancement on the original text layer, and generate using chain-of-thought prompting.
[0053] In the test of the PubMed literature dataset of the present invention, the retrieval recall rate (Recall@10) is increased by 25%, and the ROUGE-2 / F1 value of the generated abstract is increased by 18%.
[0054] The computing cost is reduced by 40% compared with the full-text encoding scheme (only need to process compressed metadata and part of the original text).
[0055] The present invention can be widely applied to contract review and medical literature analysis: in contract review, it can accurately locate the "liability for breach of contract" clause and associate the exception conditions in the original text, and the error rate is reduced by 32%;
[0056] In medical literature analysis, it can cross-paragraph associate drug side effects and dosage data to generate an interpretable report.
[0057] The above are only the preferred embodiments of the present invention, and do not impose any form of limitation on the present invention; any ordinary technician in the industry can implement the present invention smoothly according to the instructions in the drawings and the above description; however, any equivalent changes made by those skilled in the art within the scope of the technical solution of the present invention by using the technical content disclosed above, such as minor changes, modifications and evolutions, are all equivalent embodiments of the present invention; at the same time, any equivalent changes, modifications and evolutions made to the above embodiments based on the essence of the present invention still fall within the protection scope of the technical solution of the present invention.
Claims
1. A long text processing method based on generative compression and two-stage retrieval, characterized in that It includes the following specific steps: S1, generative compression stage; S101, input the original long text corpus into the generative compression module; S102, conduct compression quality assessment. If it fails, return to the generative compression module for compression. If it passes, input it into the metadata index library; S2, two-stage retrieval and generation; S201, in the first-round retrieval, screen the Top-K candidate set; S202, perform a second-round enhancement on the original text layer, and generate using chain-of-thought prompting.
2. The long text processing method based on generative compression and two-stage retrieval according to claim 1, wherein The original long text corpus includes scientific research papers and legal contracts.
3. The long text processing method based on generative compression and two-stage retrieval according to claim 1, characterized in that The generative compression module uses a large language model to generate compression metadata. The large language model includes GPT-4 and Claude-3.
4. The long text processing method based on generative compression and two-stage retrieval according to claim 1, wherein In S102, when conducting compression quality assessment, specifically: evaluate semantic fidelity through ROUGE-L and BERTScore, and retain the metadata with the semantic consistency between the compressed text and the original text ≥ 85%.
5. The long text processing method based on generative compression and two-stage retrieval according to claim 1, wherein In S201, when screening the Top-K candidate set in the first-round retrieval, specifically: in the first-round retrieval, use a dense retrieval model to calculate the cosine similarity between the query and the metadata, and screen the Top-K candidate set.
6. The long text processing method based on generative compression and two-stage retrieval according to claim 1, characterized in that The dense retrieval model is DPR.
7. The long text processing method based on generative compression and two-stage retrieval according to claim 1, characterized in that In S202, when performing a second-round enhancement on the original text layer and generating using chain-of-thought prompting, specifically: dynamically segment the associated original text through a dynamic chunking engine, and generate a summary using chain-of-thought prompting.
8. The long text processing method based on generative compression and two-stage retrieval according to claim 1, characterized in that The dynamic chunking engine performs dynamic chunking specifically as follows: segment the original text according to semantic units, and establish a mapping relationship by combining NER to identify key entities.
9. The long text processing method based on generative compression and two-stage retrieval according to claim 1, characterized in that Semantic units include chapters and paragraphs.
10. A long text processing system based on generative compression and two-stage retrieval, which is applied to the long text processing method based on generative compression and two-stage retrieval according to claims 1-9, characterized in that, It includes a generative compression module, a compression quality assessment, a first-round retrieval module, and a second-round enhancement module; The compression module is used to generate compression metadata using a large language model; The compression quality assessment is used to conduct compression quality assessment. If it fails, return to the generative compression module for compression. If it passes, input it into the metadata index library; The first-round retrieval module is used to screen the Top-K candidate set in the first-round retrieval; The second-round enhancement module is used to perform a second-round enhancement on the original text layer and generate using chain-of-thought prompting.