Data Search Engine Using Statistical Profile Vectors

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional data search methods are inefficient, produce misleading or irrelevant results, and have limited scope when searching for datasets, as they rely on search terms rather than data profiles, failing to account for the schema and statistical metrics of datasets.

Innovation Solution

The system generates a sample data vector based on the data schema and statistical metrics of a sample dataset, which is used to search a data index of reference datasets, calculating similarity metrics to identify relevant datasets, allowing for a more targeted and effective search.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If conventional search-term based methods are used to search for datasets, then the search can be performed using simple keywords, but the results are inefficient, misleading, or irrelevant and have limited scope

Engineering Contradiction:
Improvesearch operation simplicityVSAvoidsearch efficiency
Core Design Contradiction:
Ease of operationVSProductivity

Solution Approach 1:

The patent transforms the search approach by changing the parameters from simple keyword matching to multi-dimensional data profiling. Instead of searching based on single-term keywords, the system generates comprehensive data profiles including statistical metrics (mean, variance, skewness, kurtosis), data schemas, and metadata for both query and reference datasets. This parameter transformation enables efficient and accurate dataset discovery by comparing datasets based on their structural and statistical characteristics rather than superficial keyword matches.

Inventive Principle:
Principle #35Parameter changes

2Loss of time

If search-term based methods are used, then the search can be performed quickly with simple queries, but the scope is limited to a small number of drives, databases, or online resources

Engineering Contradiction:
Improvesearch timeVSAvoidsearch scope
Core Design Contradiction:
Loss of timeVSAdaptability or versatility

Solution Approach 1:

The patent implements a universal data profiling framework that can operate across diverse data sources including local drives, databases, and online resources. The system generates standardized data profiles that capture essential characteristics regardless of the source type, enabling consistent comparison and search across heterogeneous environments. This multi-functional approach allows the search system to adapt to various data storage locations and formats while maintaining efficient performance through profile-based matching.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Quantity of substance

If conventional keyword search is used, then the search can return large numbers of results, but many results are irrelevant to the desired objective

Engineering Contradiction:
Improvenumber of search resultsVSAvoidresult relevance accuracy
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The patent replaces the mechanical keyword-matching system with a statistical and structural analysis approach. Instead of relying on simple text search algorithms that return numerous potentially irrelevant results, the system uses data profiling to compute statistical metrics and compare data schemas. This substitution transforms the search mechanism from superficial keyword matching to deep structural and statistical analysis, significantly improving result relevance while maintaining manageable result sets.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

4Device complexity

If search-term based approaches are used, then the search can be performed without understanding data profiles, but the search cannot account for data schema or statistical metrics

Engineering Contradiction:
Improvesearch system complexityVSAvoiddata profile information
Core Design Contradiction:
Device complexityVSLoss of information

Solution Approach 1:

The patent applies preliminary action by generating comprehensive data profiles for all reference datasets before the actual search occurs. During the preprocessing phase, the system computes statistical metrics, extracts data schemas, and creates metadata profiles for each dataset. This preliminary profiling enables the search phase to efficiently compare query datasets against pre-analyzed reference datasets using profile similarity, avoiding the need to perform complex analysis during the search itself and preserving rich data profile information.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11474978B2Systems and methods for a data search engine based on data profiles
Publication Date: 2022.10.18 CAPITAL ONE SERVICES LLC
  • US11474978B2 patent drawing
  • US11474978B2 patent drawing
  • US11474978B2 patent drawing

AI summary

Systems and methods for searching data are disclosed. For example, the system may include one or more memory units storing instructions and one or more processors configured to execute the instructions to perform operations. The operations may include receiving a sample dataset and identifying a data schema of the sample dataset. The operations may include generating a sample data vector that includes statistical metrics of the sample dataset and information based on the data schema of the sample dataset. The operations may include searching a data index comprising a plurality of stored data vectors corresponding to a plurality of reference datasets. The stored data vectors may include statistical metrics of the reference datasets and information based on corresponding data schema. The operations may include generating, based on the search and the sample data vector, one or more similarity metrics of the sample dataset to individual ones of the reference datasets.