Columnar Data Labeling Using ML Split-and-Merge Structuring
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing database systems struggle with efficiently organizing and labeling data from various input sources, particularly in high-volume pharmacies, where data is often disorganized and lacks consistent product type information, leading to inefficiencies in data processing and management.
Innovation Solution
A computer-implemented method and system that utilizes machine learning models to recognize and label data by splitting input data into structured data structures, determining missing product types, and applying defined labels to columns, ensuring consistent data organization and presentation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If machine learning models are used to automatically recognize and label data, then data processing efficiency and accuracy are improved, but system complexity increases
Solution Approach 1:
The system uses machine learning models to automatically perform data recognition and labeling without requiring manual intervention. The models self-train on input data and autonomously generate labeled outputs, enabling the system to serve itself rather than requiring human operators to manually process each data point.
Solution Approach 2:
The patent replaces manual data processing mechanisms with automated machine learning algorithms. Instead of human analysts manually reviewing and labeling data, the system uses computational models that process data automatically, substituting mechanical human labor with automated digital processing.
2Manufacturing precision
If data is split into multiple data structures and merged into unified structures, then data organization accuracy is improved, but processing time increases
Solution Approach 1:
The system divides the input data into multiple separate data structures based on different criteria or characteristics. This segmentation allows for more precise organization and labeling of specific data elements, improving overall accuracy by handling complex data relationships in manageable segments.
Solution Approach 2:
After splitting data into multiple structures for processing, the system merges these structures back together into unified data outputs. This merging step consolidates the processed information while maintaining the accuracy gains from the segmentation process, creating organized and labeled data structures.
3Measurement precision
If machine learning models are trained on tabular data with header rows, then column labeling accuracy is improved, but data preparation complexity increases
Solution Approach 1:
The system performs preliminary processing of the input data to identify and extract header rows before the main data processing begins. By preparing the data structure in advance and identifying column headers upfront, the system simplifies subsequent processing steps while maintaining high labeling accuracy.
Solution Approach 2:
The machine learning models act as intermediaries between the raw input data and the final labeled output. These models process the tabular data, using header row information as input to generate accurate column labels, serving as a computational bridge that transforms raw data into structured, labeled information.
Data Source
AI summary
A computer-implemented method includes receiving input data that is organized into a set of rows and a set of columns. A column of the set of columns includes data associated with a set of product types. A first row of the set of rows includes data associated with an individual of a set of individuals and a first product type of the set of product types. A second row of the set of rows includes data associated with the individual and a second product type of the set of product types. The method includes splitting the input data into a first data structure associated with the first product type and a second data structure associated with the second product type. The method includes generating output data that is organized into rows and columns by merging the first data structure and the second data structure into a unified data structure.


