Semantic-Based Data Analysis for Document Clustering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data analysis techniques, such as keyword-based searching, are inadequate for extracting useful information from large datasets as they fail to understand the meaning of words or derive inferences, limiting their ability to perform semantic searches.

Innovation Solution

A computer-implemented method and apparatus for semantic-based data analysis that extracts and weights semantic information from documents, assigns links between documents with similar information, and uses these weighted links and semantics to perform inferential analysis and clustering.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If keyword-based searching (bag of words) is used, then specific information can be found within a database, but the ability to understand the meaning of words and derive inferences is limited

Engineering Contradiction:
Improveinformation extraction accuracyVSAvoidsemantic understanding capability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent transforms the search parameter from simple keyword matching to semantic representation by converting text into weighted semantic vectors. This parameter change enables the system to understand word meanings and relationships rather than just matching literal keywords, thereby improving both information extraction accuracy and semantic understanding capability simultaneously

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces semantic representation as an intermediary layer between raw text and search results. By converting text into semantic vectors that capture meaning and relationships, this intermediary enables the system to bridge the gap between simple keyword matching and true semantic understanding, allowing for both precise information extraction and inferential analysis

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If traditional keyword searching is used, then the search process is simple and fast, but the ability to perform semantic searches and derive inferences is insufficient

Engineering Contradiction:
Improvesearch speedVSAvoidsemantic information loss
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The patent performs preliminary action by pre-processing text into semantic representations and pre-computing similarity relationships between documents. This advance preparation allows the system to perform fast semantic searches without computing everything in real-time, maintaining search speed while capturing semantic information that would otherwise be lost in traditional keyword searching

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If semantic information extraction is performed, then meaningful information and relationships can be identified, but the computational complexity increases

Engineering Contradiction:
Improveinformation qualityVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies segmentation by breaking down the complex semantic analysis process into distinct modules: text preprocessing, semantic feature extraction, weight calculation, and similarity computation. This modular segmentation reduces system complexity by making each component independently manageable while maintaining high information quality through specialized processing at each stage

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS9298818B1Method and apparatus for performing semantic-based data analysis
Publication Date: 2016.03.29 SRI INTERNATIONAL
  • US9298818B1 patent drawing
  • US9298818B1 patent drawing
  • US9298818B1 patent drawing

AI summary

A computer implemented method and apparatus for analyzing content of a plurality of documents. The method extracts semantic information from content of a plurality of documents; assigns weights to the semantic information; assigns links between documents containing similar semantic information; assigns a weight to each link; extracts information about the content of the plurality of documents by using the weighted links and weighted semantics to cluster the documents, perform inferential analysis, or both.