ML Model Semantics Preservation via Constraint-Based Asymmetrical Retraining
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Machine learning models deployed in a dependent manner face challenges in maintaining consistent outputs due to changes in data distributions over time, leading to the need for retraining both upstream and downstream models, which can impact interoperability and accuracy.
Innovation Solution
The implementation of a data specification with a set of constraints allows for the preservation of semantics in machine learning models, enabling asymmetrical retraining where the downstream model can be updated independently of the upstream model, thus maintaining consistent outputs.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If both upstream and downstream machine learning models are retrained simultaneously to maintain output consistency, then model accuracy is improved, but system complexity and retraining time increase
Solution Approach 1:
The patent segments the model update process into independent upstream and downstream model retraining operations. By dividing the simultaneous retraining into separate, manageable segments that can be executed independently, the system reduces complexity while maintaining accuracy through constraint-based coordination.
Solution Approach 2:
The patent applies preliminary action by establishing constraint specifications before model retraining begins. These constraints define the expected output distributions and semantics in advance, allowing models to be retrained independently while ensuring they remain compatible with each other through pre-defined compatibility criteria.
2Measurement precision
If both upstream and downstream machine learning models are retrained simultaneously to maintain output consistency, then model accuracy is improved, but retraining time increases
Solution Approach 1:
The retraining process is segmented into independent upstream and downstream model updates that can proceed separately rather than requiring synchronized simultaneous training. This segmentation enables parallel execution of retraining tasks, significantly reducing total retraining time while maintaining model compatibility through constraint verification.
Solution Approach 2:
Constraint specifications are established beforehand to define compatibility requirements. This preliminary action allows models to be retrained independently without requiring iterative coordination during the training process, thereby reducing the time lost to synchronization and coordination overhead.
3Ease of operation
If constraint specifications are established to enable independent model updates, then ease of operation is improved, but device complexity increases
Solution Approach 1:
The patent uses copying by creating explicit constraint specification copies that define model compatibility requirements. These constraint copies serve as reusable templates that can be applied when updating models, automating the verification process and reducing the operational burden of managing model dependencies despite the underlying complexity.
4Productivity
If asymmetrical retraining is enabled through data specifications, then productivity is improved, but measurement precision may be compromised
Solution Approach 1:
The patent implements feedback mechanisms where constraint specifications define expected output distributions and semantics. After independent model retraining, the system verifies that outputs conform to these constraints, providing feedback that ensures precision is maintained despite the productivity gains from asymmetrical retraining.
Solution Approach 2:
By establishing constraint specifications in advance that define acceptable output ranges and semantic requirements, the system performs preliminary action to set precision boundaries before independent retraining begins. This ensures that even with asymmetrical updates, the models remain within acceptable precision thresholds.
Data Source
AI summary
The subject technology receives assessment values determined by a first machine learning model deployed on a client electronic device, the assessment values being indicative of classifications of input data and the assessment values being associated with constraint data that comprises a probability distribution of the assessment values with respect to the classifications of the input data. The subject technology applies the assessment values determined by the first machine learning model to a second machine learning model to determine the classifications of the input data. The subject technology determines whether accuracies of the classifications determined by the second machine learning model conform with the probability distribution for corresponding assessment values determined by the first machine learning model. The subject technology retrains the first machine learning model when the accuracies of the classifications determined by the second machine learning model do not conform with the probability distribution.


