面向文档知识库的多粒度结构化检索增强生成方法及装置

By parsing and maintaining the structural information between document fragments in the retrieval enhancement generation system, adaptive recall and reordering are achieved, solving the problems of difficulty in adaptive information granularity and inaccurate matching in existing systems, and realizing more accurate information services.

CN118585615BActive Publication Date: 2026-07-17INST OF SOFTWARE - CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
INST OF SOFTWARE - CHINESE ACAD OF SCI
Filing Date
2024-04-28
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing retrieval enhancement generation systems disrupt the multiple information granularities and structural relationships of documents during index building, leading to difficulties in information granularity adaptation and inaccurate matching of effective information, making it difficult to meet users' information needs at different granularities.

Method used

Before building the index, the system parses and maintains the structural information between document fragments. Through hierarchical combination relationships, referential relationships, and native reference relationships, it adaptively recalls information fragments of different granularities and uses structured information to reorder them, thereby improving the accuracy of information matching.

Benefits of technology

It enables the effective utilization of multi-granularity information and structural relationships within and outside the document, adaptively recalls information of appropriate granularity, improves the accuracy of information matching, avoids the risk of error propagation, and provides more accurate information services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118585615B_ABST
    Figure CN118585615B_ABST
Patent Text Reader

Abstract

本发明公开一种面向文档知识库的多粒度结构化检索增强生成方法及装置,属于信息检索和自然语言处理领域。所述方法包括:将原始文档数据中的每一原始文档Di切分为若干个叶节点粒度的文档片段并生成不同粒度层级的文档片段后,提取文档片段间的层次化组合关系;在同一粒度层级上抽取文档片段间的指代关系,并获取文档片段所涉及的原生引用关系;根据输入问题与文档片段的相似性,召回若干个文档片段;基于层次化组合关系、指代关系和原生引用关系,对召回的文档片段进行重排序;将输入问题和重排序的文档片段拼接成问答提示语,并结合生成式语言模型得到输入问题的答案。本发明可以提升检索过程中信息匹配的精确度。
Need to check novelty before this filing date? Find Prior Art

Citation Information

Patent Citations

  • Intelligent text information processing system

    CN115455935A

  • Questions and answers generation

    US20110125734A1