Heterogeneous AI Model Training via Cross-Model Supervision
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The introduction of AI models, such as transformer models, into new AI tasks is often hindered by the need for pre-training on large-scale datasets, leading to time-consuming training processes that fail to meet service requirements.
Innovation Solution
A method where a first AI model is trained using the output from a complementary second AI model as a supervision signal, allowing for iterative updates and eliminating the need for pre-training on large datasets, thereby accelerating convergence and improving training efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If AI models are pre-trained on large-scale datasets, then model performance and accuracy are improved, but training time becomes excessively long
Solution Approach 1:
The patent applies preliminary action by pre-processing the training data to generate enhanced supervision signals before the main training process. The data processing module pre-processes training data to obtain enhanced supervision signals that guide the model training more effectively, allowing the model to converge faster without requiring extensive pre-training on large-scale datasets.
Solution Approach 2:
The patent implements feedback mechanisms where the model's output is continuously evaluated and used to adjust training parameters. The training module receives feedback from the evaluation module and dynamically adjusts training strategies, enabling the model to learn more efficiently and reduce training time while maintaining accuracy.
2Productivity
If traditional training methods are used without pre-training, then training time is reduced, but model convergence and performance are compromised
Solution Approach 1:
The patent applies preliminary action by pre-processing the training data to generate enhanced supervision signals before the main training process. The data processing module pre-processes training data to obtain enhanced supervision signals that guide the model training more effectively, allowing the model to converge faster without requiring extensive pre-training on large-scale datasets.
Solution Approach 2:
The patent changes training parameters dynamically during the training process. The training module adjusts learning rates, batch sizes, and other hyperparameters based on model performance metrics, enabling the model to converge reliably while maintaining high training efficiency without traditional pre-training requirements.
Data Source
AI summary
An artificial intelligence (AI) model training method is provided, including: determining a to-be-trained first model and a to-be-trained second model, where the first model and the second model are two heterogeneous AI models; inputting training data into the first model and the second model, to obtain a first output obtained by performing inference on the training data by the first model and a second output obtained by performing inference on the training data by the second model; and iteratively updating a model parameter of the first model by using the second output as a supervision signal of the first model and with reference to the first output, until the first model meets a first preset condition.


