Large model code translation reinforcement learning training method based on test example generation
By constructing a comprehensive reward signal through a joint evaluation mechanism that generates mutated code instances and unit test cases, the problems of insufficient test case discrimination ability and insufficient reward signal in the existing technology are solved, thereby improving the accuracy and robustness of code translation.
Patent Information
- Application Number
- CN202610697288.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-20
- Publication Date
- 2026-07-24
- Estimated Expiration
- 2046-05-20
AI Technical Summary
Existing code translation technologies rely on limited and low-difficulty test examples during training, making it difficult to effectively distinguish between complex inputs and erroneous implementations under boundary conditions. Furthermore, the reinforcement learning reward signal fails to fully reflect the discriminative ability of the test examples, thus limiting the improvement of semantic consistency.
By generating multiple mutated code instances and combining them with a parameterized strategy model to generate unit test cases, a joint generation and evaluation mechanism for unit test cases and target code is introduced to construct a comprehensive reward signal. The model is then trained using a group-relative strategy optimization method.
It significantly improves the discriminative power of test cases and the effectiveness of reward signals, enhances the model's ability to identify complex inputs, boundary conditions and hidden semantic errors, and improves the accuracy and robustness of code translation.