Context Driven Data Profiling via Sampling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Processing large volumes of data is computationally resource-intensive and challenging, making it difficult to derive insights and ensure data quality, especially when dealing with disparate data sets and identifying high-quality datasets.

Innovation Solution

A data profiling process that includes validation, standardization, and processing through rules engines to generate insights and improve data quality, utilizing techniques such as machine learning and artificial intelligence to enhance data quality and identify duplicate or anonymous data attributes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If large volumes of data are retrieved and processed to derive insights, then data quality and insights are improved, but computational resource consumption increases

Engineering Contradiction:
Improvedata qualityVSAvoidcomputational resource consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent segments the data processing task by profiling individual attributes separately and using sampling techniques to process only representative portions of data. The system profiles attributes like name, address, and phone number independently, and uses sampled data to infer characteristics of the entire dataset, thereby reducing computational resource consumption while maintaining data quality assessment accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies partial action by processing only a sample of the data rather than the entire dataset. The system uses sampling to select representative records for profiling, which reduces the computational burden while still providing sufficient insights into data quality characteristics. This partial processing approach allows the system to derive meaningful insights without exhaustively processing all data records.

Inventive Principle:
Principle #16Partial or excessive action

2Measurement precision

If large volumes of data are processed to identify high-quality datasets, then data quality assessment is improved, but processing time increases

Engineering Contradiction:
Improvedata quality assessmentVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent segments the data processing task by profiling individual attributes separately and using sampling techniques to process only representative portions of data. The system profiles attributes like name, address, and phone number independently, and uses sampled data to infer characteristics of the entire dataset, thereby reducing computational resource consumption while maintaining data quality assessment accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies partial action by processing only a sample of the data rather than the entire dataset. The system uses sampling to select representative records for profiling, which reduces the computational burden while still providing sufficient insights into data quality characteristics. This partial processing approach allows the system to derive meaningful insights without exhaustively processing all data records.

Inventive Principle:
Principle #16Partial or excessive action

3Measurement precision

If data validation and standardization are performed to improve data quality, then data quality scores increase, but device complexity increases

Engineering Contradiction:
Improvedata quality scoresVSAvoidprocessing system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent implements a universal data profiling system that can handle multiple data types and attributes through a single integrated framework. The system uses a common set of validation rules and standardization techniques that apply across different data types (text, numeric, date), reducing the need for separate specialized processing systems for each data type while maintaining high data quality scores.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent applies parameter changes by dynamically adjusting validation rules and standardization parameters based on the specific characteristics of each data attribute. The system can modify validation criteria and standardization approaches according to the data type, distribution, and quality characteristics, allowing the same processing system to adapt to different data scenarios without increasing overall system complexity.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11966402B2Context driven data profiling
Publication Date: 2024.04.23 COLLIBRA BELGIUM BV
  • US11966402B2 patent drawing
  • US11966402B2 patent drawing
  • US11966402B2 patent drawing

AI summary

The present disclosure relates to methods and systems for processing data via a data profiling process. Data profiling can include modifying attributes included in source data and identifying aspects of the source data. The data profiling process can include processing an attribute according to a set of validation rules to validate information included in the attribute. The process can also include processing the attribute according to a set of standardization rules to modify the attribute into a standardized format. The process can also include processing the attribute according to a set of rules engines. The modified attributes can be outputted for further processing. The data profiling process can also include deriving a value score and usage rank of an attribute, which can be used in deriving insights into the source data.