Model training method, data processing method, device, equipment, medium and product
By employing a dual-training-path approach, combining reinforcement learning and error correction training with a teacher model, the problems of low efficiency and limited capability ceiling in large language model training are solved, achieving more efficient model training results.
CN122432674APending Publication Date: 2026-07-21MOORE THREADS TECH CO LTD
View PDF 0 Cites 0 Cited by
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- MOORE THREADS TECH CO LTD
- Filing Date
- 2026-04-28
- Publication Date
- 2026-07-21
Smart Images

Figure CN122432674A_ABST
Abstract
The present disclosure provides a model training method, a data processing method, an apparatus, a device, a medium and a product, and belongs to the technical field of computers. The model training method comprises: determining target reward values of a plurality of sample data of a training batch of a target model; determining a first path loss of the plurality of sample data for an exploration training path and an adjustment coefficient for adjusting an influence proportion of the exploration training path and a correction training path on the model training according to the target reward values of the plurality of sample data; determining a second path loss of the plurality of sample data for the correction training path; determining a fusion loss of the plurality of sample data according to the first path loss, the second path loss and the adjustment coefficient; and adjusting at least part of model parameters of the target model based on the fusion loss of the plurality of sample data. The present disclosure can improve the model training effect.
Need to check novelty before this filing date? Find Prior Art