Generalized Data Mining Using Term Tensors for Insight Discovery
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The rapid growth of enterprise and consumer databases due to reduced data storage costs poses a challenge in extracting meaningful insights from massive corpora, especially since analysts often lack awareness of what they are looking for, necessitating novel approaches for data mining that can uncover insights from both structured and unstructured data.
Innovation Solution
The Generalized Data Mining and Analytics Apparatus, Methods, and Systems (GDMA) utilize term tensors to identify novel trends, relationships, and insights by associating terms with contextually related data vectors, allowing for automatic discovery and presentation of new and interesting information, even in domains with unknown ontologies, through a corpus query processor that processes queries and provides results via a summary dashboard.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If traditional search engines are used to query databases, then users can obtain search results through query statements, but users cannot discover novel insights or trends that they are not already aware of looking for
Solution Approach 1:
The system performs self-service data mining by automatically analyzing the corpus to identify novel trends, relationships, and insights without requiring users to specify what they are looking for. The GDMA autonomously generates term tensors, identifies patterns, and presents discoveries to users, enabling the system to serve itself in the knowledge discovery process.
Solution Approach 2:
The term tensor acts as an intermediary data structure that bridges raw corpus data and user-facing insights. It systematically organizes terms with their contextual relationships and statistical properties, enabling the system to translate massive unstructured data into discoverable patterns without direct user intervention in the analysis process.
2Quantity of substance
If data storage costs are reduced leading to exponential growth in database size, then more data can be stored and analyzed, but extracting meaningful insights from massive corpora becomes increasingly difficult
Solution Approach 1:
The system segments the massive corpus into manageable term tensors, where each tensor focuses on specific terms and their contextual relationships. This segmentation allows the system to process and analyze large volumes of data by breaking them down into smaller, computationally tractable units that can be independently analyzed and then synthesized.
Solution Approach 2:
The system changes the analytical parameters by moving from traditional search query parameters to statistical and contextual parameters embedded in term tensors. By analyzing term frequencies, co-occurrences, and contextual relationships rather than relying on user-defined search parameters, the system can effectively navigate and extract insights from massive corpora.
3Loss of information
If analysts attempt to extract emerging trends from a corpus, then they can identify useful information, but they are not aware of exactly what they are looking for
Solution Approach 1:
The system performs preliminary analysis by pre-processing the corpus and constructing term tensors that capture contextual relationships and statistical patterns before any specific analysis query is made. This preliminary structuring of data enables the system to rapidly identify emerging trends and insights without requiring analysts to first define what they are seeking.
Solution Approach 2:
The system implements feedback mechanisms where term tensors continuously update and refine their representations based on corpus analysis, allowing emerging trends to be detected and fed back into the analysis process. This iterative feedback enables the system to adaptively identify patterns that analysts may not have initially considered.
Data Source
AI summary
The GENERALIZED DATA MINING AND ANALYTICS APPARATUSES, METHODS AND SYSTEMS (“GDMA”), in various embodiments, may identify statistical relationships among query terms by analyzing a corpus of electronic documents. Inputs may be automatically generated automatically and/or user provided. In one embodiment, a method includes: accessing a term tensor associated with at least one term in a corpus of documents, wherein the term tensor comprises a plurality of data type vectors corresponding respectively to a plurality of term-correlated data types correlated with the at least one term in the corpus and each data type vector comprising a plurality of binned data type values with corresponding weighted occurrence values derived from the corpus; providing at least one of the plurality of term-correlated data types for selectable display; receiving at least one term-correlated data type selection; and providing data type values associated with the at least one term-correlated data type selection for display.


