Hierarchical Multi-Task Model for Automated Concept Labeling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Manual annotation of concept labels in tabular data is costly and laborious, especially for non-experts, and existing weak supervision methods are not well-suited for training Neural Network algorithms due to concept label scarcity in tabular data settings.
Innovation Solution
A method and system for generating a concept label model using a hierarchical multi-task machine learning approach, involving a first machine learning model for predicting concept labels and a second model for predicting class labels, with a concept label feed between them, utilizing user-defined annotations, labelling functions, and a two-stage training procedure to improve generalization and account for dependencies in tabular data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual annotation of concept labels is performed on large-scale tabular datasets, then concept-based explainability can be achieved, but the cost and labor required become prohibitively expensive
Solution Approach 1:
The patent applies preliminary action by pre-training concept label models on small manually-annotated datasets before deploying them to automatically generate concept labels for large-scale tabular datasets. This preliminary training phase captures domain knowledge and annotation patterns, enabling subsequent automatic labeling without requiring manual annotation of the entire large dataset, thus resolving the contradiction between label accuracy and annotation time
Solution Approach 2:
The patent introduces concept label models as an intermediary between manual annotations and large-scale data labeling. These models act as mediators that translate a small amount of manual annotation into large-scale automatic labeling, bridging the gap between the need for accurate concept labels and the prohibitive cost of manual annotation, thereby resolving the technical contradiction
2Productivity
If only a small sample of the dataset is manually labelled to reduce costs, then annotation expenses decrease, but Neural Network algorithms achieve bad performance due to insufficient training data
Solution Approach 1:
The patent applies dimensionality change by transitioning from a single-dimension approach (manual annotation of all data or none) to a multi-dimensional strategy involving multiple concept label models trained on different aspects of the data. This enables the system to leverage small annotated samples effectively by creating specialized models for different concept types, thereby maintaining model performance while improving annotation efficiency
Solution Approach 2:
The patent changes parameters by training multiple concept label models with different configurations, architectures, and training strategies on the small annotated dataset. By varying model parameters and ensembling multiple models, the system achieves better generalization and performance despite limited training data, resolving the contradiction between annotation efficiency and model performance
3Measurement precision
If concept labels are annotated for tabular data with temporal structure (e.g., fraud detection), then accurate concept-based explanations can be provided, but the annotation task becomes extremely challenging and requires domain experts
Solution Approach 1:
The patent applies preliminary action by pre-training concept label models on small datasets annotated by domain experts to capture complex temporal patterns and domain knowledge. Once trained, these models automatically annotate large-scale temporal tabular data without requiring continuous expert involvement, thereby maintaining high annotation accuracy while reducing annotation complexity for large datasets
Solution Approach 2:
The patent enables self-service by allowing the trained concept label models to automatically annotate temporal tabular data without requiring ongoing human expert intervention. The models independently handle the complexity of temporal patterns and domain-specific concepts, resolving the contradiction between annotation accuracy and annotation complexity
4Productivity
If existing weak supervision methods are used to generate concept labels automatically, then annotation costs decrease, but the methods are not well-suited for training Neural Network algorithms due to concept label scarcity in tabular data
Solution Approach 1:
The patent applies preliminary action by pre-training concept label models on small high-quality manually-annotated datasets before using them to generate weak labels for large-scale training. This preliminary phase ensures that the automatic labeling process is guided by learned patterns rather than simple rules, improving the quality of generated labels and making them suitable for Neural Network training while maintaining high labeling speed
Data Source
Figure 1
Figure 2
Figure 3
AI summary
The present document discloses a method for obtaining a concept label model for training a hierarchical multi-task machine learning model comprising, sequentially connected, a first machine learning model for receiving input records and for predicting concept labels, and a second machine learning model for predicting class labels from predicted concept labels corresponding to respective input records, said hierarchical multi-task machine learning model comprising a concept label feed arranged to feed between the first and second machine learning models, for receiving predicted concept labels from the concept label model, for training said hierarchical multi-task machine learning model. It is further disclosed a non-transitory storage media including program instructions for obtaining a concept label model for training a hierarchical multi-task machine learning model, and a system comprising an electronic data processor arrange to carry out the disclosed method.