Enterprise Document Search Indexing via Contextual Passage Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current search technologies, particularly in enterprise data systems, face challenges in efficiently processing and retrieving unstructured data due to limitations in text indexing, variation indexing, word frequency, and co-occurrence indexing, which often require high computing power or significant human involvement, and struggle to address context-dependent word meanings.

Innovation Solution

A method and system for generating real-time search results by indexing documents, identifying stems of search terms, and determining passages of interest within a context window, using a network of computing nodes and a system management module to process and analyze documents, thereby improving search relevance without excessive computational or human resource demands.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If text indexing and variation indexing are used to search unstructured data, then search capability is improved, but computing power requirements and human involvement increase significantly

Engineering Contradiction:
Improvesearch capabilityVSAvoidcomputing power
Core Design Contradiction:
Ease of operationVSUse of energy by moving object

Solution Approach 1:

The patent segments the search process into multiple stages: initial broad search using simple text indexing, followed by targeted refinement using variation indexing only on promising results. This segmentation allows the system to use lightweight indexing methods initially, then apply more computationally intensive methods only when necessary, resolving the contradiction between search capability and computing power requirements.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies partial variation indexing rather than complete indexing of all documents. By using text indexing for initial retrieval and only applying variation indexing to a subset of relevant documents, the system achieves adequate search capability without the full computational burden of indexing every document with all its variations.

Inventive Principle:
Principle #16Partial or excessive action

2Measurement precision

If co-occurrence indexing is used to address context-dependent word meanings, then search relevance is improved, but device complexity and computational resources increase

Engineering Contradiction:
Improvesearch relevanceVSAvoidcomputational resources
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent performs preliminary text indexing and basic term matching before applying more complex co-occurrence analysis. This preliminary action filters out clearly irrelevant documents early, so that co-occurrence indexing is only applied to a reduced set of candidate documents, thereby improving search relevance without the full computational cost of analyzing all documents.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent applies different indexing strategies to different parts of the search process: simple text indexing for initial retrieval, and more complex variation indexing and co-occurrence analysis only for refining results. This local quality approach ensures high search relevance for final results while minimizing overall computational resource consumption.

Inventive Principle:
Principle #3Local quality

3Ease of operation

If metacoding is used to structure text passages, then text searchability is improved, but implementation time and human effort increase

Engineering Contradiction:
Improvetext searchabilityVSAvoidimplementation time
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The patent employs automated text indexing that processes unstructured text documents automatically without requiring manual metacoding. The system extracts terms and creates indexes autonomously, eliminating the time-consuming manual process of coding text passages while maintaining good searchability through automated term extraction and indexing algorithms.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS11321336B2Systems and methods for enterprise data search and analysis
Publication Date: 2022.05.03 SAVANTX
  • US11321336B2 patent drawing
  • US11321336B2 patent drawing
  • US11321336B2 patent drawing

AI summary

A system and method for enterprise searching of documents. The system comprises a computing system configured to receive one or more search terms, and responsively analyze a group of documents to return analysis results. A method for enterprise searching includes indexing the group of documents, determining relevant terms and measuring the context between terms. Relevant portions of documents, also called passages of interest, are determined as part of the analysis process. The analysis also uses a calculated importance value of terms as part of the analysis process.