Hierarchical Text Database Construction to Reduce Retrieval Interference
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The existing information retrieval systems suffer from interference information, leading to less accurate retrieval results due to the use of large models that process unstructured and complex text data.
Innovation Solution
A method involving clustering and summarizing original sentences using a large model to create a target database with both original and summary sentences, enhancing the accuracy of information retrieval by distinguishing and organizing document information effectively.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If large models process unstructured and complex text data directly, then information retrieval can be performed, but interference information increases and retrieval accuracy decreases
Solution Approach 1:
The patent segments the text data into original sentences and summary sentences at different levels. Original sentences provide detailed information while summary sentences provide high-level overview, allowing the system to process information in structured segments rather than as unstructured complex text, thereby reducing interference and improving retrieval accuracy
Solution Approach 2:
The patent extracts key information from the original text by generating summary sentences that capture the essence of the document. This extraction process removes unnecessary details and interference information, retaining only the most relevant content for retrieval operations, thus improving the signal-to-noise ratio in the database
2Loss of information
If only original sentences are stored in the database, then detailed information is preserved, but retrieval efficiency and high-level understanding are reduced
Solution Approach 1:
The patent merges two types of information representations into a single database structure: original sentences for detailed information and summary sentences for high-level overview. This combination allows the system to simultaneously provide detailed retrieval results and efficient high-level searching, eliminating the need to choose between detail and efficiency
Solution Approach 2:
The patent adds a dimensional layer by introducing summary sentences as a higher-level representation dimension. Instead of storing only original sentences at one level of detail, the system creates a multi-dimensional structure where summary sentences provide overview information and original sentences provide detailed information, enabling retrieval operations to operate at different levels of abstraction
Data Source
AI summary
A method for establishing a database including obtaining a document; performing clustering and summarizing processing on original sentences of the document for at least one round by using a large model to obtain summary sentences for each round; determining the summary sentence of each round as a second information set corresponding to the document; and constructing a target database based on a first information set and the second information set corresponding to the document. The first information set includes at least one original sentence corresponding to the document, and the target database includes the first information set and the second information set corresponding to a plurality of documents respectively.


