Subdomain-Specific Graph Embeddings for Large Data Prediction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional predictive classification techniques are inefficient in large data prediction domains due to their inability to capture complex relationships and entity-level attributes, leading to skewed datasets and performance compromises.
Innovation Solution
The implementation of graph-based predictive modeling techniques that generate subdomain-specific graphs to capture diverse relationships across multiple information subdomains, processed using graph-based machine learning models to encode graph embeddings for improved predictive classifications.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If traditional machine learning classification models are used, then the system can process data in tabular formats, but the system cannot capture complex relationships and entity-level attributes in large data prediction domains
Solution Approach 1:
The patent transforms traditional tabular data into graph data structures, adding a dimensional shift from flat tables to multi-dimensional networks with nodes, edges, and hierarchical relationships. This enables the system to capture complex relationships and entity-level attributes that cannot be represented in traditional tabular formats.
Solution Approach 2:
The patent segments large-scale prediction domains into multiple subdomain-specific graphs, each capturing relationships within specific information domains. This segmentation allows the system to manage complexity by dividing the overall problem into manageable subdomains while preserving complex relationships within each segment.
2Adaptability or versatility
If relational databases with predefined tables are used, then data can be stored in structured formats, but the system cannot easily extend or augment the data model to capture all available dimensions
Solution Approach 1:
The patent implements dynamic graph schemas that can be extended and augmented without rigid predefined structures. The graph data model allows flexible addition of new nodes, edges, and relationships as needed, enabling the system to adapt to changing requirements without extensive redesign time.
Solution Approach 2:
The patent creates a universal graph-based data model that can represent multiple types of information and relationships across different domains. This multi-functional framework can accommodate various data types and relationships without requiring separate specialized schemas for each domain.
3Measurement precision
If supervised classification models trained on historic data are used, then the models can learn from labeled examples, but the performance is limited by data purity within each target class
Solution Approach 1:
The patent introduces graph embeddings as an intermediary representation that captures complex relationships and contextual information from the graph structure. These embeddings serve as enriched features that improve the quality and purity of training data by incorporating relational context that goes beyond simple labeled examples.
Solution Approach 2:
The patent transforms traditional tabular training data into graph-based representations, adding dimensional richness that captures entity relationships, attributes, and contextual information. This dimensional enhancement improves data purity by incorporating multiple sources of information that help distinguish true positive cases from false positives.
Data Source
AI summary
Various embodiments of the present disclosure provide data storage, processing, and prediction techniques for providing predictive insights within large data prediction domains. The techniques may include generating, using a plurality of source tables for a prediction domain, a plurality of subdomain-specific graphs for the prediction domain. The techniques may include generating a plurality of subdomain-specific embeddings for the plurality of subdomain-specific graphs and a composite graph embedding based on the plurality of graph embeddings and a designated predictive task. The techniques may include initiating the performance of the designated predictive task based on the composite graph embedding.


