Enterprise Data Normalization Using NLP Tag Mapping
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Integrating and standardizing enterprise data from disparate sources with diverse formats, protocols, and structures is challenging due to compatibility issues, leading to increased manual processing, excessive I/O operations, and computing latency.
Innovation Solution
Automatically mapping raw data to computer-readable tags in near real-time using natural language processing and machine learning, normalizing data formats, and translating protocols to achieve compatibility and reduce manual input.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If data from disparate sources with diverse formats and protocols is integrated manually, then data compatibility can be achieved, but manual processing time and labor increase significantly
Solution Approach 1:
The patent replaces manual mechanical data integration processes with an automated machine learning system. The system uses trained ML models to automatically map data from disparate sources to a target schema, eliminating the need for manual data transformation while maintaining high compatibility. The automated system processes data formats, protocols, and structures through algorithmic transformation rather than human intervention.
Solution Approach 2:
The patent introduces a machine learning model as an intermediary between source data and the target data warehouse. This intermediary automatically learns and applies transformation rules, acting as a smart mediator that handles format conversion, protocol adaptation, and schema mapping without requiring manual configuration for each data source.
2Reliability
If repetitive data queries and packet transmissions are performed to ensure data accuracy, then data compatibility is maintained, but wear on storage devices increases and computing latency increases
Solution Approach 1:
The patent performs data validation, transformation, and mapping operations in advance during the data ingestion phase, rather than repeatedly querying and transmitting packets later. The ML model pre-processes and normalizes data as it arrives, ensuring accuracy is established upfront, which eliminates the need for subsequent repetitive verification queries and reduces storage device wear.
3Device complexity
If traditional data integration methods are used without automation, then system complexity remains manageable, but computing latency increases due to manual processes
Solution Approach 1:
The patent replaces manual data integration mechanics with automated machine learning systems. The ML models automatically handle complex transformation logic, format conversion, and protocol adaptation, reducing computing latency by eliminating manual processing steps while the system manages complexity through automated rather than manual processes.
Data Source
AI summary
Various embodiments relate to normalizing raw data by mapping the raw data to a computer-readable tag. A computer-readable tag may be an identifier that at least partially represents a category (e.g., a department) and/or the raw data itself. In response to receiving the raw data, some embodiments perform the mapping by, for example, performing natural language processing (NLP) on each particular department's raw data to associate natural language words in the raw data to its corresponding computer-readable tag and then populating, at a data structure that includes the computer-readable tag, an entry with data (representing the raw data) in a standardized format. In this way, regardless of whether different sets of raw data come from disparate sources that have diverse formats, protocols, or structures relative to each other, the normalized data and standardized form makes the data compatible.


