Hierarchical Taxonomy Classifier for Unlabeled Structured Datasets
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The surge in data size within large-scale data repositories, such as data lakes and warehouses, makes it difficult to locate useful or relevant data due to the lack of labeling or classification of unlabeled structured datasets.
Innovation Solution
A computer-program product that identifies a target hierarchical taxonomy, extracts taxonomy tokens, computes taxonomy vectors, and uses machine learning models to cluster and classify unlabeled structured datasets into predefined taxonomy categories, thereby converting them into taxonomy-labeled datasets.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If data is stored in large-scale data repositories without labeling or classification, then data storage capacity is maximized, but data retrieval efficiency deteriorates
Solution Approach 1:
The system performs preliminary classification and labeling of structured datasets using machine learning models and hierarchical taxonomies before data retrieval operations. By pre-organizing data into categorized clusters with meaningful labels, the system enables efficient data location and access without requiring manual intervention during retrieval operations.
2Productivity
If manual classification and labeling of datasets is performed, then data retrieval efficiency is improved, but processing time and operational complexity increase
Solution Approach 1:
The system implements self-service automated classification and labeling capabilities using machine learning models that independently process structured datasets without requiring manual human intervention. The hierarchical taxonomy classifier automatically assigns labels and categories to datasets, eliminating time-consuming manual classification operations while maintaining high retrieval efficiency.
Solution Approach 2:
The system replaces manual mechanical classification processes with automated machine learning-based classification systems. The hierarchical taxonomy classifier and clustering algorithms substitute human operators with computational models that process and categorize data rapidly, significantly reducing processing time while improving consistency and scalability.
3Loss of time
If automated machine learning classification is implemented, then processing time is reduced, but system complexity increases
Solution Approach 1:
The system segments the classification task into distinct hierarchical levels and functional components. The hierarchical taxonomy is divided into multiple levels of categorization, with each level handled by specialized clustering algorithms and machine learning models. This segmentation allows the complex classification problem to be broken down into manageable, modular components that can be independently optimized and maintained.
Solution Approach 2:
The system introduces hierarchical taxonomies and intermediate clustering layers as mediators between raw structured datasets and final classification results. These intermediary structures organize data progressively through multiple levels of abstraction, simplifying the overall classification process by breaking it into staged operations rather than requiring a single complex classification step.
Data Source
AI summary
A computer-implemented system includes identifying a target hierarchical taxonomy comprising a plurality of distinct hierarchical taxonomy categories; extracting a plurality of distinct taxonomy tokens from the plurality of distinct hierarchical taxonomy categories; computing a taxonomy vector corpus based on the plurality of distinct taxonomy tokens; computing a plurality of distinct taxonomy clusters based on an input of the taxonomy vector corpus; constructing a hierarchical taxonomy classifier based on the plurality of distinct taxonomy clusters; converting a volume of unlabeled structured datasets to a plurality of distinct corpora of taxonomy-labeled structured datasets based on the hierarchical taxonomy classifier; and outputting at least one corpus of taxonomy-labeled structured datasets of the plurality of distinct corpora of taxonomy-labeled structured datasets based on an input of a data classification query.


