用于训练模型的方法
By combining hybrid validation and intent inspection with a reward mechanism to optimize model parameters, the problem of model overfitting was solved, the model's ability to follow complex instructions and training efficiency were improved, and more efficient training results were achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT
- Filing Date
- 2025-08-01
- Publication Date
- 2026-07-17
AI Technical Summary
Existing reinforcement learning methods for instruction tasks are prone to model overfitting and cannot effectively improve the model's ability to follow complex instructions. Furthermore, existing validators suffer from insufficient consistency and controllability in detecting whether the model follows instructions.
A hybrid validation method is adopted, which combines the semantic judgment of predefined rule scripts and large language model judges to constrain the model output in terms of format, boundary, content, style, tone, logical rationality and language fluency. The model parameters are optimized through intent inspection and reward mechanism, and an instruction honeypot component is introduced to detect whether the model is overfitting.
It improves the model's ability to follow complex instructions, reduces overfitting, enhances training robustness and efficiency, and significantly improves performance and versatility in instruction task evaluation.
Smart Images

Figure CN120911638B_ABST