Large model code translation reinforcement learning training method based on test example generation

By constructing a comprehensive reward signal through a joint evaluation mechanism that generates mutated code instances and unit test cases, the problems of insufficient test case discrimination ability and insufficient reward signal in the existing technology are solved, thereby improving the accuracy and robustness of code translation.

CN122242635BActive Publication Date: 2026-07-24HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)
2 Cites 0 Cited by

Patent Information

Application Number
CN202610697288.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-05-20
Publication Date
2026-07-24
Estimated Expiration
2046-05-20

AI Technical Summary

Technical Problem

Existing code translation technologies rely on limited and low-difficulty test examples during training, making it difficult to effectively distinguish between complex inputs and erroneous implementations under boundary conditions. Furthermore, the reinforcement learning reward signal fails to fully reflect the discriminative ability of the test examples, thus limiting the improvement of semantic consistency.

Method used

By generating multiple mutated code instances and combining them with a parameterized strategy model to generate unit test cases, a joint generation and evaluation mechanism for unit test cases and target code is introduced to construct a comprehensive reward signal. The model is then trained using a group-relative strategy optimization method.

Benefits of technology

It significantly improves the discriminative power of test cases and the effectiveness of reward signals, enhances the model's ability to identify complex inputs, boundary conditions and hidden semantic errors, and improves the accuracy and robustness of code translation.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The application provides a large model code translation reinforcement learning training method based on test sample generation, and relates to the technical field of artificial intelligence. The method comprises the following steps: performing mutation processing on source language code to be translated, and generating a plurality of mutated code instances using a large model; generating a unit test sample set and target language code based on the source language code using a parameterized strategy model; applying the unit test sample set to the source language code for execution verification and defining a unit test sample reward; comparing and executing the target language code and the source language code according to standard test samples and test samples that can pass the verified source language code, and defining a target code reward; combining the rewards to form a reinforcement learning reward signal, updating the model according to a group relative strategy optimization method, and completing code translation. The application realizes the continuous reinforcement of the model generation capability by constructing a training mechanism for the collaborative optimization of test generation and code translation.
Need to check novelty before this filing date? Find Prior Art