Training Data Recommender for Dynamic Dataset Labeling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Recommender systems face challenges in handling dynamic data and providing accurate recommendations due to the inefficiency of historical labeled data, especially when dealing with new users or products without prior data, and the lack of timely and sufficient labeled data.

Innovation Solution

A system and method for generating labeled datasets using a training data recommender technique, which involves receiving an unlabeled dataset and a labeled dataset, extracting feature subsets, generating labeling functions, and using a snorkel generative model to construct a sparse matrix for labeling, while determining an adequate amount of labeled data required for training machine learning models based on a labeled data prediction threshold.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If historical labeled data is used for training machine learning models, then personalized recommendations can be provided to customers, but the system becomes ineffective when dealing with new users or products without prior data

Engineering Contradiction:
Improverecommendation accuracyVSAvoidhandling new users and products
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The system performs preliminary action by pre-training machine learning models on available historical labeled data before new users or products arrive. This creates a baseline model that can be quickly adapted through data augmentation techniques when new data becomes available, rather than waiting for sufficient historical data to accumulate.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces data augmentation techniques as an intermediary mechanism that bridges the gap between limited historical labeled data and the need for accurate recommendations on new users/products. By generating synthetic or augmented data, the system mediates the transition from cold-start conditions to personalized recommendations.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If large amounts of historical labeled data are collected for training, then model training accuracy improves, but the time required for data collection and model training increases significantly

Engineering Contradiction:
Improvemodel training accuracyVSAvoiddata collection and training time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system applies partial action by using data augmentation to generate additional training samples beyond the available historical labeled data. This allows the model to be trained on a more diverse dataset without requiring proportional increases in data collection time, effectively achieving better training accuracy with limited initial data.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The patent changes parameters by applying various data augmentation transformations (such as synthetic data generation, feature transformations, or sampling strategies) to the existing labeled data. These parameter changes create variations of the original data that expand the training dataset without requiring additional data collection time.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If more labeled data is used for training machine learning models, then recommendation personalization improves, but the system cannot adapt quickly to sudden disruptions in customer preferences

Engineering Contradiction:
Improvepersonalization qualityVSAvoidadaptation speed to disruptions
Core Design Contradiction:
ReliabilityVSSpeed

Solution Approach 1:

The system implements dynamics by creating a flexible training pipeline where data augmentation techniques can be dynamically applied when disruptions are detected. Rather than relying on static historical data, the system can quickly generate augmented data reflecting new customer preferences, enabling fast adaptation while maintaining personalization quality.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS20220092354A1Method and system for generating labeled dataset using a training data recommender technique
Publication Date: 2022.03.24 TATA CONSULTANCY SERVICES LTD
  • US20220092354A1 patent drawing
  • US20220092354A1 patent drawing
  • US20220092354A1 patent drawing

AI summary

This disclosure relates generally to a method and system for generating labelled dataset using a training data recommender technique. Recommender systems face major challenges in handling dynamic data on machine learning paradigms thereby rendering inaccurate unlabeled dataset. The method of the present disclosure is based on a training data recommender technique suitably constructed with a newly defined parameter such as the labelled data prediction threshold to determine the adequate amount of labelled training data required for training the one or more machine learning models. The method processes the received unlabeled dataset for labelling the unlabeled dataset based on a labelled data prediction threshold which is determined using a trained training data recommender technique. This labelling data threshold leads to a significant reduction in training time while performing the one or more machine learning models and thus recommender systems to quickly adapt disruptions thereby decreasing the reduction factor.