A Fine-Grained Reward-Based Model Training Method for Secure Code Generation
By constructing a course-based dataset of real vulnerability code and implicit prompts, and employing hierarchical training and word-level reward mechanisms, the problem of security knowledge transfer and vulnerability identification in large language models during code generation was solved, achieving a synergistic improvement in security and functionality.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- FUJIAN FUYAO UNIVERSITY OF SCIENCE & TECHNOLOGY
- Filing Date
- 2026-02-25
- Publication Date
- 2026-06-02
AI Technical Summary
Existing large-scale language models struggle to transfer security knowledge to real-world, open programming environments when generating code, and lack fine-grained reward feedback, making it difficult to identify and eliminate security vulnerabilities. Furthermore, existing technologies lack a systematic training framework.
By constructing a course-based dataset of real vulnerability code and implicit prompts, a hierarchical training task and a word-level reward mechanism are adopted. Combined with a large language model based on Transformer and MoE architecture, vulnerability identification and remediation training are carried out, fine-grained advantage values are calculated, and model parameters are updated.
This improves the security of the model in real-world development scenarios, enhances its security generalization ability and robustness, and ensures that the security and functionality of the generated code are improved in tandem.
Smart Images

Figure CN122132836A_ABST