Template Regularization for ML Generalization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Machine learning models are not optimally trained on data representative of the distribution they will be applied to, leading to suboptimal performance when faced with new data not included in the training corpus.
Innovation Solution
Assigning regularization penalties to templates based on domain knowledge, such as historic data and user input, to emphasize features from certain templates during training, thereby improving model generalization and performance on unseen data distributions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If machine learning models are trained on existing training data, then the model can be generated and deployed, but the model does not perform optimally when applied to new data distributions not represented in the training corpus
Solution Approach 1:
The patent applies preliminary action by assigning regularization penalties to templates before the model is trained on new data. Historic data and domain knowledge are used to pre-establish penalty values that guide the model's preference for certain feature templates, preparing the model in advance to handle new data distributions more effectively
Solution Approach 2:
The patent changes parameters by dynamically adjusting regularization penalty values assigned to different feature templates. These penalty parameters control the model's preference for features from certain templates versus others, allowing the model to adapt its feature selection behavior based on domain knowledge and historic performance patterns
2Reliability
If regularization penalties are assigned to templates based on domain knowledge, then the model prefers using features from certain templates, but this requires additional complexity in the training process
Solution Approach 1:
The patent applies local quality by assigning different regularization penalty values to different feature templates based on their specific characteristics and domain knowledge. Each template receives a customized penalty that reflects its reliability and relevance, rather than applying a uniform regularization approach across all features
Data Source
AI summary
Systems and techniques are disclosed for training a machine learning model based on one or more regularization penalties associated with one or more features. A template having a lower regularization penalty may be given preference over a template having a higher regularization penalty. A regularization penalty may be determined based on domain knowledge. A restrictive regularization penalty may be assigned to a template based on determining that a template occurrence is below a stability threshold and may be modified if the template occurrence meets or exceeds the stability threshold.


