Dimensional Reduction and Quantum Clustering for Unstructured Data Search

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The increasing volume of user information on the Internet and in databases leads to large, high-dimensional data sets that consume excessive system resources and often include 'false positives,' making it difficult to retrieve relevant information from unstructured data sources.

Innovation Solution

A system comprising modules for searching unstructured data, including an input module, collector module, tokenizer module, data processing module, dimensional reduction module, quantum clustering module, and output module, which processes search queries, tokenizes data, determines eigenvectors, applies dimensional reduction, and performs quantum clustering to identify relevant information without loading the entire data set into memory.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If traditional search queries are used on large data sets, then comprehensive search coverage is achieved, but system resources are consumed excessively and false positives increase

Engineering Contradiction:
Improvesearch results quantityVSAvoidsystem resource consumption
Core Design Contradiction:
Quantity of substanceVSLoss of energy

Solution Approach 1:

The patent segments the search process into multiple stages: initial broad search to gather comprehensive results, dimensional reduction to process and filter data, and refined search to identify relevant information. This segmentation allows the system to handle large data sets by processing them in manageable portions through distinct operational phases, reducing overall system resource consumption while maintaining search comprehensiveness.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies dimensional reduction techniques to transform high-dimensional data into lower-dimensional representations. By reducing the dimensionality of search results through techniques like singular value decomposition and eigenvector computation, the system can process and search through data more efficiently, consuming fewer system resources while preserving the essential information needed for accurate search results.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Measurement precision

If additional search terms are added to reduce false positives, then search accuracy improves, but relevant documents may be excluded

Engineering Contradiction:
Improvesearch accuracyVSAvoidrelevant document exclusion
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The patent performs preliminary dimensional reduction and clustering actions before final search result selection. By pre-processing the data to identify patterns and group similar documents through quantum clustering algorithms, the system can accurately distinguish relevant documents from false positives without needing to manually add restrictive search terms, thereby maintaining both accuracy and comprehensive coverage.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces dimensional reduction and quantum clustering as intermediary processes between the search query and final results. These intermediary steps act as filters that automatically distinguish relevant information from false positives based on data patterns and semantic similarity, eliminating the need for manual search term adjustment and preventing relevant document exclusion.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Adaptability or versatility

If unstructured data is searched using traditional programs, then data flexibility is maintained, but search effectiveness is reduced

Engineering Contradiction:
Improvedata structure flexibilityVSAvoidsearch effectiveness
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The patent replaces traditional mechanical search programs with quantum-based algorithms that can process unstructured data more effectively. By using quantum clustering and dimensional reduction techniques, the system can automatically organize and search through unstructured data patterns, maintaining data flexibility while significantly improving search effectiveness and productivity.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS9165061B2Identifying information related to a particular entity from electronic sources, using dimensional reduction and quantum clustering
Publication Date: 2015.10.20 REPUTATION COM
  • US9165061B2 patent drawing
  • US9165061B2 patent drawing
  • US9165061B2 patent drawing

AI summary

Presented are systems and methods for identifying information about a particular entity including acquiring electronic documents having unstructured text, that are selected based on one or more search terms from a plurality of terms related to the particular entity. Tokenizing the acquired documents to form a data matrix and then calculating a plurality of eigenvectors, using the data matrix and the transpose of the data matrix. The variance is then acquired for determining the amount of intra-clustering between the documents and then the acquired documents are clustered using some of the eigenvectors and the variance.