A Fine-Grained Reward-Based Model Training Method for Secure Code Generation

By constructing a course-based dataset of real vulnerability code and implicit prompts, and employing hierarchical training and word-level reward mechanisms, the problem of security knowledge transfer and vulnerability identification in large language models during code generation was solved, achieving a synergistic improvement in security and functionality.

CN122132836APending Publication Date: 2026-06-02FUJIAN FUYAO UNIVERSITY OF SCIENCE & TECHNOLOGY

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
FUJIAN FUYAO UNIVERSITY OF SCIENCE & TECHNOLOGY
Filing Date
2026-02-25
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing large-scale language models struggle to transfer security knowledge to real-world, open programming environments when generating code, and lack fine-grained reward feedback, making it difficult to identify and eliminate security vulnerabilities. Furthermore, existing technologies lack a systematic training framework.

Method used

By constructing a course-based dataset of real vulnerability code and implicit prompts, a hierarchical training task and a word-level reward mechanism are adopted. Combined with a large language model based on Transformer and MoE architecture, vulnerability identification and remediation training are carried out, fine-grained advantage values ​​are calculated, and model parameters are updated.

Benefits of technology

This improves the security of the model in real-world development scenarios, enhances its security generalization ability and robustness, and ensures that the security and functionality of the generated code are improved in tandem.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122132836A_ABST
    Figure CN122132836A_ABST
Patent Text Reader

Abstract

This invention relates to a fine-grained reward-based model training method for secure code generation, comprising: collecting real vulnerability code; generating repair code from the real vulnerability code; obtaining vulnerability repair pairs; generating implicit prompts using a large language model for the vulnerability repair pairs; obtaining samples containing vulnerability repair pairs and implicit prompts; evaluating and classifying the samples to construct a curriculum-based dataset; sequentially training the policy model on vulnerability code identification and vulnerability repair capabilities based on the curriculum-based dataset to generate code sequences; analyzing the code sequences using a teacher model and calculating positive and negative lexical-level rewards, combining the original rewards to obtain lexical-level immediate rewards; calculating the fine-grained advantage value of each lexical based on the lexical-level immediate rewards, updating the parameters of the policy model, and obtaining the trained policy model. This invention systematically improves the security of code generated by large language models.
Need to check novelty before this filing date? Find Prior Art