用于训练模型的方法

By combining hybrid validation and intent inspection with a reward mechanism to optimize model parameters, the problem of model overfitting was solved, the model's ability to follow complex instructions and training efficiency were improved, and more efficient training results were achieved.

CN120911638BActive Publication Date: 2026-07-17SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT
Filing Date
2025-08-01
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing reinforcement learning methods for instruction tasks are prone to model overfitting and cannot effectively improve the model's ability to follow complex instructions. Furthermore, existing validators suffer from insufficient consistency and controllability in detecting whether the model follows instructions.

Method used

A hybrid validation method is adopted, which combines the semantic judgment of predefined rule scripts and large language model judges to constrain the model output in terms of format, boundary, content, style, tone, logical rationality and language fluency. The model parameters are optimized through intent inspection and reward mechanism, and an instruction honeypot component is introduced to detect whether the model is overfitting.

Benefits of technology

It improves the model's ability to follow complex instructions, reduces overfitting, enhances training robustness and efficiency, and significantly improves performance and versatility in instruction task evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120911638B_ABST
    Figure CN120911638B_ABST
Patent Text Reader

Abstract

本发明公开了用于训练模型的方法。一种用于训练模型的方法包括:接收训练数据集,训练数据集包括复杂指令数据和相关联的验证器;使待训练的模型基于复杂指令数据生成输出;基于验证器对输出执行混合验证,混合验证包括基于预定义规则脚本的验证和基于大语言模型裁判的语义判断;对输出执行意图检查,意图检查用于判断输出是否满足复杂指令数据中的指令的意图;以及基于意图检查的结果和混合验证的结果来更新待训练的模型的参数。根据本发明的方法克服了利用指令任务强化学习的技术导致被训练的模型对指令任务过拟合的问题,提升了指令任务强化学习过程的鲁棒性和训练效率。
Need to check novelty before this filing date? Find Prior Art