Generalized Data Mining Using Term Tensors for Insight Discovery

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The rapid growth of enterprise and consumer databases due to reduced data storage costs poses a challenge in extracting meaningful insights from massive corpora, especially since analysts often lack awareness of what they are looking for, necessitating novel approaches for data mining that can uncover insights from both structured and unstructured data.

Innovation Solution

The Generalized Data Mining and Analytics Apparatus, Methods, and Systems (GDMA) utilize term tensors to identify novel trends, relationships, and insights by associating terms with contextually related data vectors, allowing for automatic discovery and presentation of new and interesting information, even in domains with unknown ontologies, through a corpus query processor that processes queries and provides results via a summary dashboard.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If traditional search engines are used to query databases, then users can obtain search results through query statements, but users cannot discover novel insights or trends that they are not already aware of looking for

Engineering Contradiction:
Improvediscovery of novel insightsVSAvoiduser input requirements
Core Design Contradiction:
Loss of informationVSEase of operation

Solution Approach 1:

The system performs self-service data mining by automatically analyzing the corpus to identify novel trends, relationships, and insights without requiring users to specify what they are looking for. The GDMA autonomously generates term tensors, identifies patterns, and presents discoveries to users, enabling the system to serve itself in the knowledge discovery process.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The term tensor acts as an intermediary data structure that bridges raw corpus data and user-facing insights. It systematically organizes terms with their contextual relationships and statistical properties, enabling the system to translate massive unstructured data into discoverable patterns without direct user intervention in the analysis process.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Quantity of substance

If data storage costs are reduced leading to exponential growth in database size, then more data can be stored and analyzed, but extracting meaningful insights from massive corpora becomes increasingly difficult

Engineering Contradiction:
Improvedata volumeVSAvoidanalysis complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The system segments the massive corpus into manageable term tensors, where each tensor focuses on specific terms and their contextual relationships. This segmentation allows the system to process and analyze large volumes of data by breaking them down into smaller, computationally tractable units that can be independently analyzed and then synthesized.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system changes the analytical parameters by moving from traditional search query parameters to statistical and contextual parameters embedded in term tensors. By analyzing term frequencies, co-occurrences, and contextual relationships rather than relying on user-defined search parameters, the system can effectively navigate and extract insights from massive corpora.

Inventive Principle:
Principle #35Parameter changes

3Loss of information

If analysts attempt to extract emerging trends from a corpus, then they can identify useful information, but they are not aware of exactly what they are looking for

Engineering Contradiction:
Improveemergent trend detectionVSAvoidunknown target identification
Core Design Contradiction:
Loss of informationVSDifficulty of detecting and measuring

Solution Approach 1:

The system performs preliminary analysis by pre-processing the corpus and constructing term tensors that capture contextual relationships and statistical patterns before any specific analysis query is made. This preliminary structuring of data enables the system to rapidly identify emerging trends and insights without requiring analysts to first define what they are seeking.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system implements feedback mechanisms where term tensors continuously update and refine their representations based on corpus analysis, allowing emerging trends to be detected and fed back into the analysis process. This iterative feedback enables the system to adaptively identify patterns that analysts may not have initially considered.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS9183203B1Generalized data mining and analytics apparatuses, methods and systems
Publication Date: 2015.11.10 QUANTIFIND
  • US9183203B1 patent drawing
  • US9183203B1 patent drawing
  • US9183203B1 patent drawing

AI summary

The GENERALIZED DATA MINING AND ANALYTICS APPARATUSES, METHODS AND SYSTEMS (“GDMA”), in various embodiments, may identify statistical relationships among query terms by analyzing a corpus of electronic documents. Inputs may be automatically generated automatically and/or user provided. In one embodiment, a method includes: accessing a term tensor associated with at least one term in a corpus of documents, wherein the term tensor comprises a plurality of data type vectors corresponding respectively to a plurality of term-correlated data types correlated with the at least one term in the corpus and each data type vector comprising a plurality of binned data type values with corresponding weighted occurrence values derived from the corpus; providing at least one of the plurality of term-correlated data types for selectable display; receiving at least one term-correlated data type selection; and providing data type values associated with the at least one term-correlated data type selection for display.