Automated Dataset Relevancy Scoring via Statistical Correlation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In the field of data analytics, identifying relevant data in massive datasets is complex due to the manual effort required to compare and join datasets, especially as the number of data sources and datasets increases, making it difficult to discover relevant information beyond the target dataset.

Innovation Solution

A method that categorizes datasets into direct and indirect related pools based on shared key fields, transforms them using statistical measures, and creates a relevancy data store to determine the strength of relationships between target and related datasets, leveraging metadata and correlation analysis to automate the discovery of relevant data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual comparison and joining of each possible related dataset is performed, then relevancy of data can be investigated, but the complexity and time required increases significantly as the quantity of data sources increases

Engineering Contradiction:
Improverelevancy investigation accuracyVSAvoiddata analysis time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary actions by pre-processing datasets to extract and store metadata including data types, key fields, and field descriptions before actual relevancy analysis. This preliminary organization enables rapid comparison and joining operations during data analysis without requiring manual investigation of each dataset, thus reducing analysis time while maintaining accuracy

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system introduces an intermediary mechanism in the form of a relevancy scoring system that automatically evaluates and ranks datasets based on their relevance to the core dataset. This intermediary scoring mechanism eliminates the need for manual relevancy investigation, providing objective measurements that maintain accuracy while dramatically reducing the time required to analyze multiple data sources

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If manual comparison and joining of each possible related dataset is performed, then relevancy of data can be investigated, but the complexity of the process increases with the number of data sources

Engineering Contradiction:
Improverelevancy investigation accuracyVSAvoiddata processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system segments the complex data analysis process into distinct automated stages: metadata extraction, field matching, correlation calculation, and relevancy scoring. Each stage handles specific tasks independently, reducing overall process complexity while maintaining investigation accuracy. The segmentation allows the system to manage multiple data sources systematically without manual intervention at each step

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system implements self-service through automated relevancy assessment that performs data comparison, joining, and scoring without manual intervention. The automated algorithms independently evaluate datasets, calculate correlations, and generate relevancy scores, eliminating the need for manual process management while maintaining accurate relevancy investigation across numerous data sources

Inventive Principle:
Principle #25Self-service

3Loss of information

If statistical transformation and correlation analysis are performed on massive datasets, then relevant data discovery is enabled, but computational resources and processing time are consumed

Engineering Contradiction:
Improverelevant information discoveryVSAvoidcomputational resource consumption
Core Design Contradiction:
Loss of informationVSUse of energy by moving object

Solution Approach 1:

The system extracts and utilizes metadata from datasets to guide the statistical transformation and correlation analysis process. By extracting key fields, data types, and field descriptions beforehand, the system focuses computational resources only on relevant data portions rather than processing entire massive datasets, thus enabling relevant information discovery while reducing computational resource consumption

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system performs partial action by calculating relevancy scores based on selected key fields and statistical measures rather than analyzing all possible data combinations. The relevancy scoring mechanism applies statistical transformation and correlation analysis selectively to identified candidate datasets, sufficient for discovering relevant information without the excessive computational cost of exhaustive analysis

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS9558245B1Automatic discovery of relevant data in massive datasets
Publication Date: 2017.01.31 MAPLEBEAR INC
  • US9558245B1 patent drawing
  • US9558245B1 patent drawing
  • US9558245B1 patent drawing

AI summary

An approach for discovery of relevant data in massive datasets. Compare datasets including compare key fields, compare data fields and a core dataset including target data field(s) and core field(s) are received. The compare datasets are categorized into direct and indirect related dataset pools based on the target data field(s) correlation strength with matching compare and core fields. The direct related dataset pool and the core dataset are transformed into reduction datasets based on statistical measure of values of target data fields, shared key fields and compare data fields. Target correlations of the reduction datasets are creating based on a reduction compare and target data fields. Statistical relationship strength of core dataset and the direct related dataset pool are created based on a statistical mean of target correlations and a relevancy data store is created.