Model training and task execution method and device, equipment, chip and medium
By constructing high-contrast positive and negative sample pairs, the problems of low sample quality and efficiency in training multimodal large models are solved, improving training efficiency and model performance, and achieving faster and more stable model convergence.
Patent Information
- Application Number
- CN202510888710.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-11-14
AI Technical Summary
Existing reinforcement learning methods suffer from dual bottlenecks in training large multimodal models: sample quality and efficiency. This results in low training efficiency, slow convergence speed, and difficulty in meeting the performance requirements of complex inference scenarios.
By constructing multiple positive and negative sample pairs that are more valuable for guiding policy optimization, high-quality samples are selected using evaluation metrics, data redundancy is avoided, and the utilization rate of training resources and the model convergence speed are improved.
It significantly improves the training efficiency and stability of multimodal large models, enhances the model's generalization ability and performance, and achieves faster and more stable model convergence.