Hierarchical Text Database Construction to Reduce Retrieval Interference

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The existing information retrieval systems suffer from interference information, leading to less accurate retrieval results due to the use of large models that process unstructured and complex text data.

Innovation Solution

A method involving clustering and summarizing original sentences using a large model to create a target database with both original and summary sentences, enhancing the accuracy of information retrieval by distinguishing and organizing document information effectively.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If large models process unstructured and complex text data directly, then information retrieval can be performed, but interference information increases and retrieval accuracy decreases

Engineering Contradiction:
Improveretrieval accuracyVSAvoidinterference information
Core Design Contradiction:
Measurement precisionVSObject-generated harmful factors

Solution Approach 1:

The patent segments the text data into original sentences and summary sentences at different levels. Original sentences provide detailed information while summary sentences provide high-level overview, allowing the system to process information in structured segments rather than as unstructured complex text, thereby reducing interference and improving retrieval accuracy

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent extracts key information from the original text by generating summary sentences that capture the essence of the document. This extraction process removes unnecessary details and interference information, retaining only the most relevant content for retrieval operations, thus improving the signal-to-noise ratio in the database

Inventive Principle:
Principle #2Taking out (Extraction)

2Loss of information

If only original sentences are stored in the database, then detailed information is preserved, but retrieval efficiency and high-level understanding are reduced

Engineering Contradiction:
Improvehigh-level detailsVSAvoidretrieval efficiency
Core Design Contradiction:
Loss of informationVSProductivity

Solution Approach 1:

The patent merges two types of information representations into a single database structure: original sentences for detailed information and summary sentences for high-level overview. This combination allows the system to simultaneously provide detailed retrieval results and efficient high-level searching, eliminating the need to choose between detail and efficiency

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent adds a dimensional layer by introducing summary sentences as a higher-level representation dimension. Instead of storing only original sentences at one level of detail, the system creates a multi-dimensional structure where summary sentences provide overview information and original sentences provide detailed information, enabling retrieval operations to operate at different levels of abstraction

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS20250328572A1Method for establishing database and information retrieval and related devices
Publication Date: 2025.10.23 LENOVO (BEIJING) LTD
  • US20250328572A1 patent drawing
  • US20250328572A1 patent drawing
  • US20250328572A1 patent drawing

AI summary

A method for establishing a database including obtaining a document; performing clustering and summarizing processing on original sentences of the document for at least one round by using a large model to obtain summary sentences for each round; determining the summary sentence of each round as a second information set corresponding to the document; and constructing a target database based on a first information set and the second information set corresponding to the document. The first information set includes at least one original sentence corresponding to the document, and the target database includes the first information set and the second information set corresponding to a plurality of documents respectively.