Automated Data Enrichment via Entity and Statistical Matching
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Incomplete, unstructured, and inconsistent data sets pose challenges for organizations, leading to inefficient and inaccurate data insights and decision-making, as they hinder the application of artificial intelligence and machine learning models and result in inefficient manual analysis, potentially causing margin leakage, inconsistent prices, and poor customer satisfaction.
Innovation Solution
A system and method that ingests data sets from multiple sources, standardizes and cleanses them, and applies entity matching and statistical matching techniques to generate an enriched data set, which is then used to produce accurate data insights and recommendations through AI and machine learning models.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual analysis is used on incomplete, unstructured, and inconsistent data sets, then human expertise can be applied, but the process becomes very inefficient, costly, and subject to human error
Solution Approach 1:
The patent introduces an automated data enrichment system as an intermediary between raw data and analysis processes. This system applies entity matching techniques and statistical methods to transform incomplete, unstructured data into enriched, structured data sets, enabling efficient automated analysis while maintaining reliability through systematic processing rather than manual intervention
Solution Approach 2:
The patent replaces manual mechanical analysis processes with automated computational systems. By substituting human analysts with algorithmic entity matching and statistical enrichment techniques, the system eliminates human error and inefficiency while processing data at scale, thereby improving both productivity and consistency
2Productivity
If artificial intelligence and machine learning models are applied to incomplete, unstructured, and inconsistent data sets, then automated insights can be generated, but the models lack the quality and quantity of data needed to train accurately
Solution Approach 1:
The patent applies preliminary data enrichment actions before feeding data to machine learning models. By pre-processing data through entity matching and statistical techniques to fill gaps, standardize formats, and enrich with external sources, the system ensures that ML models receive high-quality, complete data, thereby improving model accuracy while maintaining automation
Solution Approach 2:
The patent transforms data parameters by changing incomplete values into enriched values through entity matching and statistical inference. This parameter transformation process converts unstructured data into structured formats with complete attributes, enabling ML models to operate effectively on enhanced data with improved precision
3Reliability
If data sets are kept incomplete to maintain confidentiality of proprietary information, then security is preserved, but the data lacks the quality and quantity needed for accurate analysis
Solution Approach 1:
The patent uses external data sources as intermediaries to enrich internal proprietary data without exposing the proprietary information itself. By matching entities between internal and external data sets and transferring only necessary attributes, the system enhances data completeness while maintaining security boundaries
Solution Approach 2:
The patent segments the data enrichment process into distinct matching and transfer operations. By dividing the data sets into internal proprietary portions and external reference portions, and only transferring specific matched attributes, the system achieves data enrichment without compromising the confidentiality of core proprietary information
4Ease of manufacture
If traditional column-row database storage is used for unstructured data sets, then data can be stored simply, but the data is not easily searchable and difficult to analyze
Solution Approach 1:
The patent transforms unstructured data parameters into structured parameters through entity matching and data enrichment. By converting free-text fields into standardized categories, adding structured attributes, and organizing data according to defined schemas, the system maintains storage simplicity while dramatically improving searchability and analytical accessibility
Data Source
AI summary
Systems, methods, and graphical user interfaces (GUIs) for ingesting and enriching data regarding a plurality of entities are provided. A first data set comprising company data and a second data set comprising customer data are ingested. The first data set is processed to generate a processed data set. The first data set may be processed by applying an entity matching technique, wherein one or more data elements are generated based on whether an entity of the first data set and an entity of the second data set are commonly associated. The first data set may additionally or alternatively be processed by applying a statistical matching technique, wherein one or more predicted data elements are generated based on similarity between an entity of the first data set and one or more entities of the second data set.


