Ontological Data Model for Predicting Parent Variables

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data analysis methods fail to effectively discover underlying relationships defined by unknown rules within large datasets, particularly from diverse and noisy data sources like the Internet, limiting their ability to accurately predict events.

Innovation Solution

A method that retrieves data strings, builds a statistical model of parent-child relationships using Bayesian probability algorithms, filters noise, and constructs an ontological representation to determine conditional probabilities, thereby identifying relevant child variables and predicting parent variable values based on aggregated conditional probabilities.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional data analysis methods are used on large datasets from diverse sources, then data processing capability is maintained, but the ability to discover underlying relationships and predict events accurately deteriorates

Engineering Contradiction:
Improveprediction accuracyVSAvoiddata analysis complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the data analysis process into distinct modules: data retrieval, noise filtering, statistical model building, and probability calculation. This segmentation allows complex data to be processed in manageable stages, improving prediction accuracy while maintaining system throughput

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary statistical model layer between raw data and prediction output. This model acts as a mediator that transforms diverse data into structured probabilistic relationships, enabling accurate predictions without directly processing the full complexity of the original data

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If noise filtering is applied to data from diverse sources, then data quality improves, but data processing time increases

Engineering Contradiction:
Improvedata qualityVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent applies noise filtering as a preliminary action during data retrieval and preprocessing stages. By filtering noise early in the process, the system ensures high data quality for subsequent analysis while minimizing the time impact through efficient filtering algorithms that operate in parallel with data collection

Inventive Principle:
Principle #10Preliminary action

3Loss of information

If statistical models are built to capture parent-child relationships, then relationship discovery improves, but computational resources required increase

Engineering Contradiction:
Improverelationship informationVSAvoidcomputational resources
Core Design Contradiction:
Loss of informationVSUse of energy by moving object

Solution Approach 1:

The patent applies local quality by focusing statistical modeling only on relevant parent-child relationships identified through preliminary data analysis. Rather than modeling all possible relationships uniformly, the system identifies and models only the locally significant relationships that contribute to prediction accuracy, reducing computational resource consumption

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS7516050B2Defining the semantics of data through observation
Publication Date: 2009.04.07 POULIN HOLDINGS LLC
  • US7516050B2 patent drawing
  • US7516050B2 patent drawing
  • US7516050B2 patent drawing

AI summary

Techniques for estimating the structure and meaning of data using probability are described. The techniques include retrieving data as data strings from a data source, producing a dataset from the retrieved data strings and building a statistical model of parent-child relationships from data strings in the dataset. Building the statistical model includes determining incidence values for the data strings in the dataset and concatenating the incident values with the data strings to provide child variables. The techniques include analyzing the child variables and the parent variables to produce statistical relationships between the child variables and a parent variable, determining probabilities values based on the determined parent child relationships and building an ontological representation of the data based on subsequent conditional probabilities values.