Automated Data Enrichment via Entity and Statistical Matching

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Incomplete, unstructured, and inconsistent data sets pose challenges for organizations, leading to inefficient and inaccurate data insights and decision-making, as they hinder the application of artificial intelligence and machine learning models and result in inefficient manual analysis, potentially causing margin leakage, inconsistent prices, and poor customer satisfaction.

Innovation Solution

A system and method that ingests data sets from multiple sources, standardizes and cleanses them, and applies entity matching and statistical matching techniques to generate an enriched data set, which is then used to produce accurate data insights and recommendations through AI and machine learning models.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual analysis is used on incomplete, unstructured, and inconsistent data sets, then human expertise can be applied, but the process becomes very inefficient, costly, and subject to human error

Engineering Contradiction:
Improveaccuracy of data analysisVSAvoidefficiency of data analysis
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent introduces an automated data enrichment system as an intermediary between raw data and analysis processes. This system applies entity matching techniques and statistical methods to transform incomplete, unstructured data into enriched, structured data sets, enabling efficient automated analysis while maintaining reliability through systematic processing rather than manual intervention

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces manual mechanical analysis processes with automated computational systems. By substituting human analysts with algorithmic entity matching and statistical enrichment techniques, the system eliminates human error and inefficiency while processing data at scale, thereby improving both productivity and consistency

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If artificial intelligence and machine learning models are applied to incomplete, unstructured, and inconsistent data sets, then automated insights can be generated, but the models lack the quality and quantity of data needed to train accurately

Engineering Contradiction:
Improveautomation of data insightsVSAvoidaccuracy of machine learning models
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent applies preliminary data enrichment actions before feeding data to machine learning models. By pre-processing data through entity matching and statistical techniques to fill gaps, standardize formats, and enrich with external sources, the system ensures that ML models receive high-quality, complete data, thereby improving model accuracy while maintaining automation

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent transforms data parameters by changing incomplete values into enriched values through entity matching and statistical inference. This parameter transformation process converts unstructured data into structured formats with complete attributes, enabling ML models to operate effectively on enhanced data with improved precision

Inventive Principle:
Principle #35Parameter changes

3Reliability

If data sets are kept incomplete to maintain confidentiality of proprietary information, then security is preserved, but the data lacks the quality and quantity needed for accurate analysis

Engineering Contradiction:
Improvedata securityVSAvoidcompleteness of data
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent uses external data sources as intermediaries to enrich internal proprietary data without exposing the proprietary information itself. By matching entities between internal and external data sets and transferring only necessary attributes, the system enhances data completeness while maintaining security boundaries

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent segments the data enrichment process into distinct matching and transfer operations. By dividing the data sets into internal proprietary portions and external reference portions, and only transferring specific matched attributes, the system achieves data enrichment without compromising the confidentiality of core proprietary information

Inventive Principle:
Principle #1Segmentation

4Ease of manufacture

If traditional column-row database storage is used for unstructured data sets, then data can be stored simply, but the data is not easily searchable and difficult to analyze

Engineering Contradiction:
Improvesimplicity of data storageVSAvoidsearchability of data
Core Design Contradiction:
Ease of manufactureVSDifficulty of detecting and measuring

Solution Approach 1:

The patent transforms unstructured data parameters into structured parameters through entity matching and data enrichment. By converting free-text fields into standardized categories, adding structured attributes, and organizing data according to defined schemas, the system maintains storage simplicity while dramatically improving searchability and analytical accessibility

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20240232232A1Automated data set enrichment, analysis, and visualization
Publication Date: 2024.07.11 PWC PRODUCT SALES LLC
  • US20240232232A1 patent drawing
  • US20240232232A1 patent drawing
  • US20240232232A1 patent drawing

AI summary

Systems, methods, and graphical user interfaces (GUIs) for ingesting and enriching data regarding a plurality of entities are provided. A first data set comprising company data and a second data set comprising customer data are ingested. The first data set is processed to generate a processed data set. The first data set may be processed by applying an entity matching technique, wherein one or more data elements are generated based on whether an entity of the first data set and an entity of the second data set are commonly associated. The first data set may additionally or alternatively be processed by applying a statistical matching technique, wherein one or more predicted data elements are generated based on similarity between an entity of the first data set and one or more entities of the second data set.