Model training and task execution method and device, equipment, chip and medium

By constructing high-contrast positive and negative sample pairs, the problems of low sample quality and efficiency in training multimodal large models are solved, improving training efficiency and model performance, and achieving faster and more stable model convergence.

CN120953752APending Publication Date: 2025-11-14XIAOMI TECH (WUHAN) CO LTD +2
0 Cites 0 Cited by

Patent Information

Application Number
CN202510888710.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing reinforcement learning methods suffer from dual bottlenecks in training large multimodal models: sample quality and efficiency. This results in low training efficiency, slow convergence speed, and difficulty in meeting the performance requirements of complex inference scenarios.

Method used

By constructing multiple positive and negative sample pairs that are more valuable for guiding policy optimization, high-quality samples are selected using evaluation metrics, data redundancy is avoided, and the utilization rate of training resources and the model convergence speed are improved.

Benefits of technology

It significantly improves the training efficiency and stability of multimodal large models, enhances the model's generalization ability and performance, and achieves faster and more stable model convergence.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The invention provides a model training and task execution method and device, equipment, a chip and a medium, and relates to the field of artificial intelligence, and the method comprises the steps: obtaining an output data set associated with multi-modal training data; wherein the output data set comprises a plurality of task execution results output when the multi-modal large model is adopted to execute the visual language processing task on the training data; determining a plurality of positive and negative sample pairs from the plurality of task execution results according to the evaluation indexes of the plurality of task execution results; wherein the evaluation index is used for indicating the output quality of the task execution result, and the evaluation index of the positive sample in the positive and negative sample pair is higher than the evaluation index of the negative sample; and training a multi-modal large model according to the plurality of positive and negative sample pairs. Therefore, by constructing a plurality of positive and negative sample pairs with higher guiding value for strategy optimization in the output data set, a clear direction can be provided for strategy optimization, faster and more stable model convergence can be realized, and the training effect and generalization ability of the model can be enhanced.
Need to check novelty before this filing date? Find Prior Art