Machine Learning Reference Data Identification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods fail to efficiently and systematically identify reference data values in unmanaged data sources, leading to data inconsistencies and integration failures, and require significant processing resources.
Innovation Solution
A computer-implemented method using a machine learning model to analyze a block of attribute values, determine presentation layout, reading direction, and identify inspection areas, which then uses tokenization to determine if these areas contain reference data values, thereby outputting the identified reference data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional methods are used to locate and analyze reference data sets, then data integration can be achieved, but significant processing resources are required and data inconsistencies occur
Solution Approach 1:
The system automatically identifies and extracts reference data from source data sets using machine learning models, eliminating the need for manual location and analysis. The model self-adjusts to different data formats and layouts, performing the entire reference data identification process autonomously without human intervention or extensive manual configuration.
Solution Approach 2:
The patent replaces conventional mechanical/manual methods of locating and analyzing reference data with an intelligent machine learning system. The model automatically analyzes data patterns, identifies reference data sets, and extracts values without requiring manual processing or conventional systematic approaches, significantly reducing processing resources.
2Productivity
If manual methods are used to identify reference data values, then data can be extracted, but the process is inefficient and time-consuming
Solution Approach 1:
The patent introduces a machine learning model as an intermediary between the source data set and the reference data extraction process. This model acts as a smart mediator that automatically analyzes data patterns, identifies reference data sets, and extracts values, replacing inefficient manual methods and significantly improving identification efficiency while reducing processing time.
Solution Approach 2:
The system dynamically adjusts processing parameters based on the analyzed data characteristics. The machine learning model adapts to different data formats, layouts, and structures by changing its analysis parameters automatically, enabling efficient reference data identification across various data types without manual reconfiguration or time-consuming manual processes.
3Reliability
If reference data is not systematically identified, then processing resources are saved, but data inconsistencies and integration failures occur
Solution Approach 1:
The patent performs preliminary identification and validation of reference data sets before the actual data integration process. The machine learning model proactively analyzes source data, identifies potential reference data sets, and extracts values in advance, ensuring data consistency and preventing integration failures before they occur, rather than reacting to problems during integration.
Solution Approach 2:
The system implements a feedback mechanism where the machine learning model continuously learns from identified reference data patterns and improves its identification accuracy. The model receives feedback from successful extractions and data consistency checks, adjusting its parameters and patterns to enhance reliability while maintaining manageable system complexity through iterative improvement.
Data Source
AI summary
A method, system, and computer program product for identifying reference data values in a source data set. The method may include inputting a block of attribute values to a predefined machine learning model. The method may also include receiving an indication of a presentation layout of the block of the attribute values and an associated reference data extraction method. The method may also include determining a reading direction of the block of values. The method may also include identifying one or more inspection areas in the reading direction of the block of values. The method may also include determining sets of the one or more inspection areas that share a common presentation feature. The method may also include identifying tokens in an inspection area. The method may also include determining if the inspection area includes reference data values. The method may also include outputting the reference data values.


