Code review comment generation method and system based on reinforcement learning

By using a reinforcement learning-based approach and leveraging a large language model (LLM) to generate code review comments, this method addresses the issues of time-consuming, labor-intensive, and low-quality code review comments found in existing tools. It achieves efficient, semantically relevant, and highly practical code review comment generation, thereby improving the accuracy and efficiency of code review.

CN120929351APending Publication Date: 2025-11-11SHANDONG INSPUR SCI RES INST CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511007519.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-22
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing code review tools are time-consuming and labor-intensive, highly subjective, and have poor scalability. They generate low-quality comments and lack learning and optimization mechanisms, making it difficult to deeply understand the semantic intent and context of code changes and thus unable to generate high-quality comments.

Method used

We employ a reinforcement learning-based approach, collecting and preprocessing data on code discrepancies, human comments, and real code corrections. We then use a Large Language Model (LLM) to generate code review comments and optimize the model through semantic similarity rewards and subsequent task rewards to produce high-quality code review comments.

Benefits of technology

It significantly improves the accuracy, efficiency, and scalability of code review. The generated comments are highly semantically relevant and practical, effectively guiding code correction, reducing the burden of manual review, shortening the review cycle, and possessing self-learning and optimization capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120929351A_ABST
    Figure CN120929351A_ABST
Patent Text Reader

Abstract

The invention discloses a code review comment generation method and system based on reinforcement learning, belongs to the technical field of crossing of software engineering and artificial intelligence, and aims to solve the technical problem of how to improve the accuracy, efficiency and expandability of code review and overcome the defects that an existing tool is low in comment quality and poor in practicability. According to the technical scheme, the method comprises the steps of collecting and preprocessing data of code differences, human comments and real correction codes, and constructing a data set; code review comments are generated, specifically, a pre-trained large language model LLM is finely adjusted through the data set, code difference data are collected, and the review comments are obtained; semantic similarity reward: calculating the semantic similarity between the generated comment and the real comment, and generating a reward signal Rsematic; the generated comment and code difference is input into a code optimization model, a corrected code patch is generated, the similarity between the generated patch and a real patch is evaluated, and a reward signal Rtask is generated; performing reinforcement learning and fine tuning on the large language model LLM; and deploying and generating comments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of software engineering and artificial intelligence, specifically to a method and system for generating code review comments based on reinforcement learning. Background Technology

[0002] Code review is a crucial step in modern software development processes, ensuring code quality, identifying potential defects, and disseminating knowledge. However, traditional manual code review suffers from significant bottlenecks. For example, it is time-consuming and labor-intensive, requiring substantial developer time, especially in large projects or frequent submission scenarios, becoming a bottleneck in the development process; it is subjective and inconsistent, with review quality heavily reliant on the reviewer's experience and subjective judgment, leading to variations in comment style, depth, and focus; it lacks scalability, as project size and complexity increase, making it difficult to ensure all code changes are fully and promptly reviewed. Meanwhile, existing automated code review tools primarily focus on static code analysis (detecting syntax errors, style violations, potential vulnerabilities, etc.) or generating simple, suggestive comments based on rules / templates. However, these methods have limited understanding capabilities, struggling to grasp the semantic intent and context of code changes, and unable to generate comments addressing deeper issues such as design flaws, logical errors, and maintainability; the quality of comments is low, often rigid, lacking contextual relevance, semantic ambiguity, or practicality, failing to reach the level of human experts; and it lacks learning and optimization mechanisms, lacking a mechanism to learn from high-quality human comments and continuously improve its own generation capabilities.

[0003] Therefore, improving the accuracy, efficiency, and scalability of code review, and overcoming the shortcomings of low-quality and poor usability of existing tools, are urgent technical problems that need to be solved. Summary of the Invention

[0004] The technical objective of this invention is to provide a code review comment generation method and system based on reinforcement learning, in order to address how to improve the accuracy, efficiency, and scalability of code review, and overcome the shortcomings of existing tools in terms of low comment quality and poor usability.

[0005] The technical objective of this invention is achieved as follows: a code review comment generation method based on reinforcement learning, the specific method of which is as follows:

[0006] Data collection: Collect and preprocess data on code differences, human comments, and actual code fixes to build a dataset;

[0007] Generate code review comments: Fine-tune a pre-trained large language model (LLM) using a dataset and collect code discrepancy data to obtain review comments;

[0008] Semantic similarity reward: Calculate the semantic similarity (e.g., cosine similarity) between the generated comment and the real comment, and generate a reward signal R_semantic;

[0009] Subsequent task rewards: Input the generated comments and code differences into the code optimization model to generate a corrected code patch, and generate a reward signal R_task by evaluating the similarity between the generated patch and the real patch (including task loss value and CrystalBLEU score);

[0010] Reinforcement learning fine-tuning of large language model LLM: Using the supervised fine-tuned large language model LLM as the initial policy network, comments are generated as actions based on code differences as states, and R_semantic and R_task are calculated based on a dual reward model. The total reward R_total of the reinforcement learning fine-tuning stage is obtained by summing them according to their weights. Then, a reinforcement learning algorithm (such as PPO) is used to maximize the expected cumulative reward and optimize the policy of the large language model LLM.

[0011] Deployment and comment generation: Integrate the fine-tuned large language model into the code review system to automatically generate code review comments.

[0012] As a preferred option, the data collection is as follows:

[0013] Extract historical code commit records from open-source projects or internal version control systems, focusing on commits that include detailed review comments that have been discussed and accepted.

[0014] Each sample contains input_diff, human_review, and fixed_diff. input_diff represents the difference text of code changes, using the Unified Diff format; human_review represents the high-quality comment text written by human reviewers for this submission that was ultimately adopted; fixed_diff is used for subsequent task rewards, and is the difference text of the final merged corrected code from this submission or a reference directly used to generate corrected code.

[0015] Cleaning: Remove irrelevant information, sensitive data, overly short comments, or invalid diffs;

[0016] Standardization: Unify Diff format (such as line number handling, spaces / tabs) and comment text (such as removing special characters, standardizing encoding);

[0017] Tokenization: Tokenize the comment text (Tokenizer needs to be compatible with LLM and SBERT);

[0018] Dataset partitioning: Divide the dataset into training set, validation set, and test set according to a ratio (e.g., 7:2:1).

[0019] As a preferred option, the code review comments are generated as follows:

[0020] Supervised fine-tuning phase: Fine-tuning the large language model LLM using code differences and human comment datasets. Specifically, using a dataset containing code differences and corresponding real comments written by human experts, supervised instruction fine-tuning is performed on the pre-trained large language model LLM to initially learn the mapping relationship between code difference features and comment content.

[0021] The reinforcement learning fine-tuning stage: Using a large language model (LLM) as the policy network, the generation policy is optimized based on reward signals. Specifically, based on supervised fine-tuning, a reinforcement learning algorithm based on proximal policy optimization (PPO) is used for further optimization. The core of reinforcement learning fine-tuning lies in using the reward signals provided by the designed reward model to guide the large language model (LLM) to learn to generate outputs that better meet the standards of high-quality comments. The large language model (LLM) acts as a policy network, and its comment generation behavior is regarded as the action taken by the agent in a specific state (input code difference). The reward model provides feedback from the environment.

[0022] As a preferred option, the semantic similarity reward is as follows:

[0023] Comment encoding: Using a pre-trained semantic encoding model, the generated comments and their corresponding real comments are encoded into fixed-dimensional dense vectors.

[0024] Similarity metric: Calculate the cosine similarity between the generated comment and the corresponding dense vector of the real comment, both of which are fixed-dimensional vectors.

[0025] Reward calculation: The calculated cosine similarity value (usually in the range of [-1,1]) is appropriately scaled and translated (e.g., mapped to the range of [0,1] or [-R,R]) and fed back as the reward signal R_semanti c to the reinforcement learning fine-tuning process; where the reward signal R_semantic directly guides the large language model LLM to generate text that is semantically closer to human comments.

[0026] Furthermore, the specific rewards for subsequent tasks are as follows:

[0027] Task input construction: The generated code review comments are combined with their corresponding original code difference data and used as input to the downstream code optimization model (a trained model that can generate corrected code based on code differences and review comments);

[0028] Optimization result generation and comparison: The code optimization model generates corrected code patches (Patch) based on the input and compares the generated patches with the human correction patches actually adopted in the code change history (Ground Truth);

[0029] Quality Assessment and Reward Calculation: Multi-dimensional key metrics are used to evaluate the similarity between the generated patch and the real patch, and the reward signal R_task is calculated. Key metrics include the task loss value and the code similarity score. The task loss value is used to calculate the loss function value (e.g., cross-entropy loss) of the code optimization model when generating the patch, reflecting the confidence or difficulty of the code optimization model in generating a specific modification. The code similarity score is used to calculate the code-level similarity between the generated patch and the real patch. The code-level similarity uses the CrystalBLEU score, which, by penalizing the commonness of n-grams in the reference corpus, better highlights meaningful code snippet matching specific to the current modification.

[0030] The reward signal R_task, reflecting the usefulness of the review, is calculated by combining key indicators (such as weighted average or formula designed according to business needs).

[0031] During the fine-tuning phase of reinforcement learning, the final reward R_total provided to the large language model LLM (policy network) is a weighted sum of the semantic similarity reward R_semantic and the reward R_task for subsequent tasks, as shown in the following formula:

[0032] R_total=α*R_semantic+β*R_task;

[0033] Here, α and β are adjustable hyperparameters (α>=0, β>=0, α+β>0), used to balance the relative importance of semantic similarity and task practicality in the optimization objective, adapting to different review needs and application scenarios.

[0034] A code review comment generation system based on reinforcement learning, the system comprising:

[0035] The data collection module is used to collect and preprocess data on code differences, human comments, and real code fixes, build a dataset, and divide the dataset into training, validation, and test sets in a ratio (e.g., 7:2:1).

[0036] A code review comment generation framework is used to fine-tune a pre-trained large language model (LLM) using a dataset and to collect code discrepancy data to generate review comments.

[0037] The semantic similarity reward module is used to calculate the semantic similarity (such as cosine similarity) between the generated comment and the real comment, and generate a reward signal R_semantic;

[0038] The subsequent task reward module is used to input the generated comments and code differences into the code optimization model, generate the corrected code patch, and generate a reward signal R_task by evaluating the similarity between the generated patch and the real patch (including the task loss value and CrystalBLEU score);

[0039] The reinforcement learning fine-tuning module for the large language model (LLM) is used to take the supervised fine-tuned LLM as the initial policy network, generate comments as actions based on code differences as states, and calculate R_semantic and R_task based on a dual-reward model. The total reward R_total of the reinforcement learning fine-tuning stage is obtained by summing them by weight. Then, a reinforcement learning algorithm (such as PPO) is used to maximize the expected cumulative reward and optimize the LLM policy.

[0040] The deployment and comment generation module is used to integrate the fine-tuned large language model into the code review system and automatically generate code review comments.

[0041] As a preferred approach, the code review comment generation framework employs instruction tuning to fine-tune the Large Language Model (LLM), constructing samples as the instruction: Generate a code review comment for the given code diff.\n Diff:<input_diff> Command: The target output is human_review; and the training parameters of the large language model LLM use the standard language model training objective (such as cross-entropy loss), train on the training set for several epochs, monitor the loss and generation quality (such as BLEU, ROUGE) on the validation set, and select the optimal model checkpoint to save; the training framework is selected as Hugging Face Transformers or DeepSpeed;

[0042] The semantic similarity reward model uses pre-trained sentence-transformers / all-mpnet-base-v2 (an SBERT model). For the generated comment gen_review and the real comment human_review, 768-dimensional sentence embedding vectors V_gen and V_human are obtained through pre-trained sentence-transformers / all-mpnet-base-v2, respectively. R_semantic = cosine_similarity(V_gen,V_human) is calculated. At the same time, according to the reward mapping R_semantic = (cosine_similarity(V_gen,V_human)+1) / 2, R_semantic is mapped to [0,1].

[0043] The reward model for subsequent tasks is as follows:

[0044] The code optimization model architecture uses a large language model (LLM) (such as T5, BART, or the aforementioned models); the input format of the code optimization model is: Fix the code based on the diff and review:\n Diff:<input_diff> Review:<human_review> Fixed Diff: The target output is fixed_diff; the code optimization model is trained on the (input_diff, human_rev iew, fixed_diff) triplet dataset, optimizing the cross-entropy loss.

[0045] Evaluation metrics are calculated, including task loss and CrystalBLEU score. The task loss (Loss) is specifically recorded as the loss value (e.g., cross-entropy) generated by the code optimization model during the generation of the predicted pred_fixed_diff when input_diff and generated gen_review. A smaller Loss indicates a more "confident" code optimization model or easier generation, positively correlated with review quality. R_loss = -γ*Loss (γ>0) or its normalization is used as part of the reward. The CrystalBLEU score (C-BLEU) is calculated as the CrystalBLEU score between pred_fixed_diff and fixed_diff. A higher CrystalBLEU score indicates a closer approximation between the generated corrected code and the actual corrected code, indirectly reflecting the stronger the guidance of the gen_review, i.e., R_bleu = δ*C-BLEU (δ>0).

[0046] The comprehensive reward R_task is calculated as follows: R_task = R_bleu + R_loss or R_task = w1*C-BLEU + w2*(1-Loss / MaxLoss); where w1 and w2 are weights, and MaxLoss is the maximum loss reference value set empirically.

[0047] More specifically, the process of fine-tuning the large language model LLM module using reinforcement learning is as follows:

[0048] ① Sample a batch of input_diff(s_t) from the dataset; where the state s_t is defined as the input_diff.

[0049] ② Generate a review gen_review(a_t) using the current strategy π_θ; where action a_t is the review gen_review generated by LLM;

[0050] ③ Calculate R_semantic using a semantic similarity reward model with fixed parameters;

[0051] ④ Calculate R_task using code optimization model with fixed parameters and subsequent task reward calculation logic;

[0052] ⑤Calculate R_total=α*R_semantic(gen_review,human_review)+β

[0053] *R_task(gen_review,input_diff);

[0054] ⑥ Use R_total as the immediate reward and apply the PPO algorithm to update the parameter θ of the policy network π_θ;

[0055] The reinforcement learning fine-tuning large language model (LLM) module performs multiple iterations on the training set, monitors the average value of R_total on the validation set and human evaluation metrics of generated comments (such as usefulness and clarity scores), and sets early stopping to prevent overfitting.

[0056] The deployment and comment generation module encapsulates the trained reinforcement learning-tuned Large Language LLM (CoRAL model) as an API service or integrates it into the review interface of CI / CD platforms and code hosting platforms. When a developer submits code (generating input_diff), the trained reinforcement learning-tuned Large Language LLM is invoked. The input format for the trained reinforcement learning-tuned Large Language LLM is: Generate a code review comment:\n Diff:<input_diff> Comment: The output of the trained reinforcement learning-tuned large language LLM is the generated code review comment, presented to the developer or reviewer.

[0057] An electronic device includes: a memory and at least one processor;

[0058] The memory contains computer programs;

[0059] The at least one processor executes the computer program stored in the memory, causing the at least one processor to perform the reinforcement learning-based code review comment generation method as described above.

[0060] A computer-readable storage medium storing a computer program that can be executed by a processor to implement the reinforcement learning-based code review comment generation method described above.

[0061] The code review comment generation method and system based on reinforcement learning of the present invention have the following advantages:

[0062] (I) This invention uses a large-scale language model as the core for generation. Through supervised fine-tuning, it initially grasps the mapping relationship between code differences and comments. Then, through a reinforcement learning phase, it introduces semantic similarity rewards (to evaluate the semantic closeness of generated comments to human expert comments) and subsequent task rewards (to quantify the actual effect of comments driving code correction). The two-stage training innovatively employs a weighted reward mechanism (R_total = α·R_semantic + β·R_task) to dynamically balance the naturalness of language expression with practical guidance. This enables a deep understanding of code change semantics, generates highly practical and continuously optimizable automated code review comments that approach human expert levels.

[0063] (II) This invention is the first to transform code correction results into reinforcement learning reward signals: 1) It breaks through the bottleneck of semantic understanding by comparing the CrystalBLEU scores of real correction patches, forcing the model to learn comments with operability; 2) It innovates a dual reward synergy mechanism, which avoids the lack of practicality caused by simply imitating human language, and solves the problem of stiff expression caused by only pursuing task effect. Based on the reverse optimization of code correction results, it fundamentally solves the two major defects of existing technology: "empty comments" and "weak correction guidance", aiming to improve the efficiency, consistency and practicality of software code review.

[0064] (III) This invention includes a reinforcement learning system architecture with a dual-reward model, a method for using the output of the code optimization model as the basis for reward calculation, a patch effectiveness evaluation technology based on CrystalBLEU, a weight adjustment mechanism for semantic rewards and task rewards, and a full-process implementation plan from data preprocessing to model deployment. It particularly emphasizes the protection of the closed-loop training system of "comment generation-code correction-result feedback", covering three dimensions: system design, algorithm implementation and engineering integration.

[0065] (iv) This invention can be directly integrated into code hosting platforms (such as GitLab plugins), continuous integration systems (such as Jenkins extensions), and IDE intelligent review plugins. The core components include cloud-based model inference services, a localized lightweight reward calculation engine, and a differentiated configuration management interface, forming an AI review technology barrier from a technical perspective and shortening the customer's code review cycle by more than 40%. From a commercial perspective, it can output intelligent review SaaS services, providing differentiated competitiveness for cloud development platforms. From a strategic perspective, the accumulated code change-review-correction closed-loop data will continuously strengthen the product's leading position in the DevOps field, solving the problems of low review quality and poor usability of existing tools, and significantly improving the accuracy, efficiency, and scalability of automated review.

[0066] (v) This invention can take code differences (Diff) as input and combine deep learning and reinforcement learning techniques to automatically generate high-quality, semantically relevant and practically instructive code review comments, which significantly improves the efficiency and quality of code review.

[0067] (vi) This invention significantly improves the quality of comments, specifically: ① High semantic relevance: Through R_semantic rewards, it ensures that the generated comments are semantically close to human expert comments, with natural and fluent expression, and easy to understand; ② Strong practicality: Through R_task rewards, it guides the generation of comments that can effectively guide actual code corrections, improving the operability and value of comments and solving the problem of hollow or difficult-to-implement comments in traditional methods; ③ High accuracy: Combining the powerful understanding ability of LLM and the guidance of dual reward signals, it can more accurately identify potential problems in code changes (including deep problems such as design, logic, and maintainability).

[0068] (vii) This invention significantly improves review efficiency: it automatically generates high-quality comments, significantly reduces the burden on human reviewers, shortens the review cycle, and improves the efficiency of the development process;

[0069] (viii) This invention has consistency and scalability: the model generation style is relatively consistent, reducing human differences, and can handle large-scale, high-frequency code review needs;

[0070] (ix) The present invention has self-learning and optimization capabilities: the reinforcement learning framework enables the system to continuously learn from human feedback (implicit in the reward model) and task results, and continuously optimize the comment generation strategy;

[0071] (x) This invention is flexible and adaptable: by adjusting the weights of α and β, the system’s focus can be flexibly configured (to be more inclined to imitate human language style or to pay more attention to the actual corrective effect brought by comments) to adapt to the review standards of different teams or projects. Attached Figure Description

[0072] The invention will be further described below with reference to the accompanying drawings.

[0073] Appendix Figure 1 This is a schematic diagram of the structure of a code review and comment generation system based on reinforcement learning;

[0074] Appendix Figure 2 A flowchart illustrating the process of enhancing the fine-tuning phase of learning;

[0075] Appendix Figure 3 This is a schematic diagram of the workflow of the semantic similarity reward model;

[0076] Appendix Figure 4 This is a schematic diagram of the workflow for the subsequent task reward model. Detailed Implementation

[0077] The reinforcement learning-based code review and comment generation method and system of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0078] Example 1:

[0079] This embodiment provides a code review comment generation method based on reinforcement learning, as detailed below:

[0080] S1. Data Collection: Collect and preprocess data on code differences, human comments, and actual code fixes to build a dataset;

[0081] S2. Generate code review comments: Fine-tune the pre-trained Large Language Model (LLM) using the dataset and collect code discrepancy data to obtain review comments;

[0082] S3, Semantic Similarity Reward: Calculate the semantic similarity (e.g., cosine similarity) between the generated comment and the real comment, and generate a reward signal R_semantic;

[0083] S4. Subsequent task rewards: Input the generated comments and code differences into the code optimization model to generate corrected code patches, and generate reward signals R_task by evaluating the similarity between the generated patches and the real patches (including task loss value and CrystalBLEU score);

[0084] S5. Reinforcement Learning Fine-tuning of Large Language Model (LLM): Using the supervised fine-tuned Large Language Model (LLM) as the initial policy network, comments are generated as actions based on code differences as states, and R_semantic and R_task are calculated based on a dual-reward model. The total reward R_total of the reinforcement learning fine-tuning stage is obtained by summing the weights. Then, a reinforcement learning algorithm (such as PPO) is used to maximize the expected cumulative reward and optimize the LLM policy.

[0085] S6. Deployment and Comment Generation: Integrate the fine-tuned large language model into the code review system to automatically generate code review comments.

[0086] The specific data collection in step S1 of this embodiment is as follows:

[0087] S101. Extract historical code commit records from open source projects or internal version control systems, focusing on commits that include detailed review comments that have been discussed and accepted.

[0088] S102. Each sample contains input_diff, human_review, and fixed_diff. input_diff represents the difference text of code changes, using the Unified Diff format; human_review represents the high-quality comment text written by human reviewers for this submission and ultimately adopted; fixed_diff is used for subsequent task rewards, the difference text of the final merged corrected code of this submission, or a reference directly used to generate corrected code.

[0089] S103. Cleaning: Remove irrelevant information, sensitive data, overly short comments, or invalid diffs;

[0090] S104. Standardization: Unify Diff format (such as line number handling, spaces / tabs) and comment text (such as removing special characters, unifying encoding);

[0091] S105, Tokenization: Tokenize the comment text (Tokenizer needs to be compatible with LLM and SBERT);

[0092] S106. Dataset partitioning: Divide the dataset into training set, validation set, and test set according to a ratio (e.g., 7:2:1).

[0093] The specific steps for generating code review comments in step S2 of this embodiment are as follows:

[0094] S201, Supervised Fine-tuning Stage: Fine-tuning the Large Language Model (LLM) using code differences and human comment datasets. Specifically, using a dataset containing code differences and corresponding real comments written by human experts, the pre-trained Large Language Model (LLM) is subjected to supervised instruction fine-tuning to initially learn the mapping relationship between code difference features and comment content.

[0095] S202, Reinforcement Learning Fine-Tuning Stage: Using a Large Language Model (LLM) as the policy network, the generation policy is optimized based on reward signals. Specifically, based on supervised fine-tuning, a reinforcement learning algorithm based on Proximal Policy Optimization (PPO) is used for further optimization. The core of reinforcement learning fine-tuning lies in using the reward signals provided by the designed reward model to guide the LLM to learn to generate outputs that better meet the standards of high-quality comments. The LLM acts as a policy network, and its comment generation behavior is regarded as the action taken by the agent in a specific state (input code difference). The reward model provides feedback from the environment.

[0096] The semantic similarity reward in step S3 of this embodiment is as follows:

[0097] S301. Comment Encoding: Using a pre-trained semantic encoding model, the generated comments and the corresponding real comments are encoded into fixed-dimensional dense vectors.

[0098] S302. Similarity metric: Calculate the cosine similarity between the generated comment and the corresponding dense vector of the real comment, both of which are fixed-dimensional vectors.

[0099] S303, Reward Calculation: The calculated cosine similarity value (usually in the range of [-1,1]) is appropriately scaled and translated (e.g., mapped to the range of [0,1] or [-R,R]) and fed back as a reward signal R_semantic to the reinforcement learning fine-tuning process; wherein, the reward signal R_semantic directly guides the large language model LLM to generate text that is semantically closer to human comments.

[0100] The specific rewards for subsequent tasks in step S4 of this embodiment are as follows:

[0101] S401, Task Input Construction: The generated code review comments are combined with their corresponding original code difference data and used as input to the downstream code optimization model (a trained model that can generate corrected code based on code differences and review comments);

[0102] S402. Optimization Result Generation and Comparison: The code optimization model generates corrected code patches based on the input and compares the generated patches with the human-made corrective patches (Ground Truth) actually adopted in the code change history.

[0103] S403. Quality Assessment and Reward Calculation: Multi-dimensional key metrics are used to evaluate the similarity between the generated patch and the real patch, and the reward signal R_task is calculated. Key metrics include the task loss value and the code similarity score. The task loss value is used to calculate the loss function value (e.g., cross-entropy loss) of the code optimization model when generating the patch, reflecting the confidence or difficulty of the code optimization model in generating a specific modification. The code similarity score is used to calculate the code-level similarity between the generated patch and the real patch. The code-level similarity uses the CrystalBLEU score, which, by penalizing the commonness of n-grams in the reference corpus, better highlights meaningful code snippet matching specific to the current modification.

[0104] S404. The reward signal R_task reflecting the practicality of the review is calculated by combining key indicators (such as weighted average or formula designed according to business needs);

[0105] S405. In the fine-tuning stage of reinforcement learning, the final reward R_total provided to the large language model LLM (policy network) is a weighted sum of the semantic similarity reward R_semantic and the reward R_task for subsequent tasks, as shown in the following formula:

[0106] R_total=α*R_semantic+β*R_task;

[0107] Here, α and β are adjustable hyperparameters (α>=0, β>=0, α+β>0), used to balance the relative importance of semantic similarity and task practicality in the optimization objective, adapting to different review needs and application scenarios.

[0108] Example 2:

[0109] This embodiment provides a code review comment generation method based on reinforcement learning, as detailed below:

[0110] (I) Data preparation, as detailed below:

[0111] (1) Data source: Extract historical code commit records from open source projects (such as GitHub, Gerrit) or enterprise internal version control systems (such as GitLab, Bitbucket), focusing on those commits that are accompanied by detailed review comments that have been discussed and accepted;

[0112] (2) Data Units: Each sample contains input_diff, the text of the code changes (Unified Diff format); human_review, the high-quality comment text written by human reviewers for this submission that was ultimately adopted; and fixed_diff (used for subsequent task rewards), the text of the changes in the final merged corrected code for this submission (or a reference that can be directly used to generate corrected code);

[0113] (3) Preprocessing, as follows:

[0114] ① Cleaning: Remove irrelevant information, sensitive data, overly short comments, or invalid diffs;

[0115] ②Standardization: Unify Diff format (such as line number handling, spaces / tabs) and comment text (such as removing special characters, standardizing encoding);

[0116] ③ Tokenization: Tokenize the comment text (Tokenizer needs to be compatible with LLM and SBERT);

[0117] ④ Dataset partitioning: Divide the dataset into training set, validation set, and test set according to a ratio (e.g., 7:2:1);

[0118] (II) Monitor and fine-tune the LLM, as follows:

[0119] (1) Model selection: Select large-scale pre-trained open-source language models (such as CodeLlama, StarCoder, GPT-NeoX, etc., which have strong code understanding capabilities);

[0120] (2) Fine-tuning method: The instruction tuning paradigm is adopted, and the sample is constructed as follows: Instruction: Generate a code review comment for the given code diff. Diff:<input_diff> Comment: The target output is human_review;

[0121] (3) Training parameters: Use standard language model training objectives (such as cross-entropy loss), train on the training set for several epochs, monitor the loss and generation quality on the validation set (such as BLEU, ROU GE), and select the optimal model checkpoint to save; training frameworks can include Hugging Face Transformers, DeepSpeed, etc.

[0122] (III) Semantic similarity reward model, as follows:

[0123] Encoder: Use pre-trained sentence-transformers / all-mpnet-base-v2 (an SBERT model);

[0124] Process: For the generated comment gen_review and the real comment human_review, obtain their 768-dimensional sentence embedding vectors V_gen and V_human respectively through the SBERT model;

[0125] Calculate: R_semantic = cosine_similarity(V_gen, V_human);

[0126] Reward mapping: R_semantic = (cosine_similarity(V_gen,V_human) + 1) / 2 maps it to [0,1];

[0127] (iv) Subsequent task reward model, as detailed below:

[0128] A) Train code to optimize the model:

[0129] Model architecture: Select another LLM (such as T5, BART or the aforementioned model).

[0130] Input format: Fix the code based on the diff and review:\n Diff:<input_diff> Review:<human_review> Fixed Diff: The target output is fixed d_diff.

[0131] Training: The model is trained on the (input_diff,human_review,fixed_diff) triple dataset, with optimized cross-entropy loss.

[0132] B) Calculation of evaluation indicators:

[0133] Task Loss: Records the loss (e.g., cross-entropy) incurred by the code optimization model during the generation of its predicted `pred_fixed_diff`, given the input `input_diff` and the generated `gen_review`. A smaller loss generally indicates a more "confident" model or easier generation, potentially positively correlated with review quality. R_loss can be designed as -γ*Loss (γ>0) or normalized and included as part of the reward.

[0134] CrystalBLEU score (C-BLEU): Calculates the CrystalBLEU score between pred_fixed_diff and fixed_diff. A higher CrystalBLEU score indicates that the generated correction code is closer to the actual correction, indirectly reflecting the stronger the guidance significance of gen_review. R_bleu = δ * C-BLEU (δ>0).

[0135] C) Overall reward R_task:

[0136] For example, R_task = R_bleu + R_loss or R_task = w1*C - BLEU + w2*(1 - Loss / MaxLoss), where w1 and w2 are weights, and MaxLoss is an empirically set maximum loss reference value. The specific formula should be determined through tuning on the validation set.

[0137] (iv) Fine-tuning LLM (CoRAL training) using reinforcement learning, as detailed below:

[0138] Environment: Define the state s_t as the input_diff. The action a_t is the comment gen_review generated by the LLM.

[0139] Reward: R_total=α*R_semantic(gen_review,human_review)+β*R_task(gen_review,input_diff)

[0140] Agent: The LLM, after supervised fine-tuning and saving, is used as the initial policy network π_θ.

[0141] Algorithm: The PPO algorithm is adopted (such as the implementation in the Hugging Face trl library).

[0142] process:

[0143] ① Sample a batch of input_diff(s_t) from the dataset.

[0144] ② Generate a review gen_review(a_t) using the current strategy π_θ.

[0145] ③ Calculate R_semantic using a semantic similarity reward model with fixed parameters.

[0146] ④ Calculate R_task using code optimization model with fixed parameters and subsequent task reward calculation logic.

[0147] ⑤ Calculate R_total.

[0148] ⑥ Using R_total as the immediate reward, the PPO algorithm is applied to update the parameter θ of the policy network π_θ.

[0149] Training: Perform multiple iterations on the training set, and monitor the average R_total on the validation set, as well as human evaluation metrics for generated reviews (such as usefulness and clarity scores). Implement early stopping to prevent overfitting.

[0150] (V) Deployment and reasoning, as detailed below:

[0151] ① Package the trained CoRAL model (LLM fine-tuned by reinforcement learning) into an API service or integrate it into the review interface of CI / CD platform or code hosting platform.

[0152] ②When the developer submits code (generating input_diff), the system calls the model.

[0153] ③ The model input format is: Generate a code review comment:\n Diff:<inpu t_diff> Comment:

[0154] ④ The model output is the generated code review comments, which are presented to the developers or reviewers.

[0155] Example 3:

[0156] As attached Figure 1 As shown, this embodiment provides a code review comment generation system based on reinforcement learning, which includes:

[0157] The data collection module is used to collect and preprocess data on code differences, human comments, and real code fixes, build a dataset, and divide the dataset into training, validation, and test sets in a ratio (e.g., 7:2:1).

[0158] A code review comment generation framework is used to fine-tune a pre-trained large language model (LLM) using a dataset and to collect code discrepancy data to generate review comments.

[0159] The semantic similarity reward module is used to calculate the semantic similarity (such as cosine similarity) between the generated comment and the real comment, and generate a reward signal R_semantic;

[0160] The subsequent task reward module is used to input the generated comments and code differences into the code optimization model, generate the corrected code patch, and generate a reward signal R_task by evaluating the similarity between the generated patch and the real patch (including the task loss value and CrystalBLEU score);

[0161] The reinforcement learning fine-tuning module for the large language model (LLM) is used to take the supervised fine-tuned LLM as the initial policy network, generate comments as actions based on code differences as states, and calculate R_semantic and R_task based on a dual-reward model. The total reward R_total of the reinforcement learning fine-tuning stage is obtained by summing them by weight. Then, a reinforcement learning algorithm (such as PPO) is used to maximize the expected cumulative reward and optimize the LLM policy.

[0162] The deployment and comment generation module is used to integrate the fine-tuned large language model into the code review system and automatically generate code review comments.

[0163] In this embodiment, the code review comment generation framework uses instruction tuning to fine-tune the Large Language Model (LLM). The sample is constructed as: Instruction: Generate a code review comment for the given code diff.\n Diff:<input_diff> Command: The target output is human_review; and the training parameters of the large language model LLM use the standard language model training objective (such as cross-entropy loss), train on the training set for several epochs, monitor the loss and generation quality (such as BLEU, ROUGE) on the validation set, and select the optimal model checkpoint to save; the training framework is selected as Hugging Face Transformers or DeepSpeed.

[0164] As attached Figure 3 As shown, the semantic similarity reward model in this embodiment uses the pre-trained sentence-transformers / all-mpnet-base-v2 (an SBERT model). For the generated comment gen_review and the real comment human_review, 768-dimensional sentence embedding vectors V_gen and V_human are obtained through the pre-trained sentence-transformers / all-mpnet-base-v2, and R_semantic = cosine_similarity(V_gen,V_human) is calculated. At the same time, according to the reward mapping R_semantic = (cosine_similarity(V_gen,V_human)+1) / 2, R_semantic is mapped to [0,1].

[0165] As attached Figure 4 As shown, the subsequent task reward model in this embodiment is as follows:

[0166] (1) The model architecture for the code optimization model uses a large language model (LLM) (such as T5, BART, or the aforementioned models); the input format for the code optimization model is: Fix the code based on the diff and review:\n Diff:<input_diff> Review:<human_review> Fixed Diff: The target output is fixed_diff; the code optimization model is trained on the (input_diff, hum an_review, fixed_diff) triple dataset, optimizing the cross-entropy loss.

[0167] (2) Calculate the evaluation metrics: including task loss and CrystalBLEU score; where, the task loss (Loss) is specifically: when input_diff and generated gen_review are input, record the loss value (such as cross-entropy) generated by the code optimization model in the process of generating the predicted pred_fixed_diff; the smaller the Loss, the more "confident" the code optimization model is or the easier it is to generate, which is positively correlated with the quality of the review, and design R_loss=-γ*Loss (γ>0) or normalize it as part of the reward; the CrystalBLEU score (C-BLEU) is specifically: calculate the CrystalBLEU score between pred_fixed_diff and fixed_diff; the higher the CrystalBLEU score, the closer the generated corrected code is to the real correction, which indirectly reflects the stronger guidance significance of gen_review, that is, R_bleu=δ*C-BLEU (δ>0);

[0168] (3) Comprehensive reward R_task: If R_task=R_bleu+R_loss or R_task=w1*C-BLEU+w2*(1-Loss / MaxLoss); where w1, w2 are weights, and MaxLoss is the maximum loss reference value set by experience.

[0169] As attached Figure 2 As shown, the specific working process of the reinforcement learning fine-tuning large language model (LLM) module in this embodiment is as follows:

[0170] ① Sample a batch of input_diff(s_t) from the dataset; where the state s_t is defined as the input_diff.

[0171] ② Generate a review gen_review(a_t) using the current strategy π_θ; where action a_t is the review gen_review generated by LLM;

[0172] ③ Calculate R_semantic using a semantic similarity reward model with fixed parameters;

[0173] ④ Calculate R_task using code optimization model with fixed parameters and subsequent task reward calculation logic;

[0174] ⑤Calculate R_total=α*R_semantic(gen_review,human_review)+β*R_task(gen_review,input_diff);

[0175] ⑥ Using R_total as the immediate reward, the PPO algorithm is applied to update the parameter θ of the policy network π_θ.

[0176] The reinforcement learning fine-tuning large language model (LLM) module summarized in this embodiment performs multiple iterations on the training set, monitors the average value of R_total and human evaluation metrics of generated comments (such as usefulness and clarity scores) on the validation set, and sets early stopping to prevent overfitting.

[0177] In this embodiment, the deployment and comment generation module encapsulates the trained reinforcement learning-tuned Large Language LLM (CoRAL model) as an API service or integrates it into the review interface of a CI / CD platform or code hosting platform. When a developer submits code (generating input_diff), the trained reinforcement learning-tuned Large Language LLM is invoked. The input format of the trained reinforcement learning-tuned Large Language LLM is: Generate a code review comment:\n Diff:<input_diff> Comment: The output of the trained reinforcement learning-tuned large language LLM is the generated code review comment, presented to the developer or reviewer.

[0178] Example 4:

[0179] This embodiment also provides an electronic device, including: a memory and a processor;

[0180] The memory stores the instructions executed by the computer.

[0181] The processor executes computer execution instructions stored in the memory, causing the processor to execute the reinforcement learning-based code review comment generation method in any embodiment of the present invention.

[0182] The processor can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor can be a microprocessor or any conventional processor.

[0183] Memory is used to store computer programs and / or modules. The processor implements various functions of the electronic device by running or executing the computer programs and / or modules stored in the memory, and by accessing data stored in the memory. Memory can mainly include a program storage area and a data storage area. The program storage area can store the operating system, at least one application program required for a function, etc.; the data storage area can store data created based on the use of the terminal, etc. In addition, memory can also include high-speed random access memory, and can also include non-volatile memory, such as hard disks, RAM, plug-in hard disks, smart memory cards (SMC), secure digital cards (SD cards), flash memory cards, at least one disk storage device, flash memory devices, or other volatile solid-state storage devices.

[0184] Example 5:

[0185] This embodiment also provides a computer-readable storage medium storing multiple instructions, which are loaded by a processor to cause the processor to execute the reinforcement learning-based code review and comment generation method of any embodiment of the present invention. Specifically, a system or apparatus equipped with a storage medium may be provided, on which software program code implementing the functions of any of the above embodiments is stored, and the computer (or CPU or MPU) of the system or apparatus may read and execute the program code stored in the storage medium.

[0186] In this case, the program code read from the storage medium can itself implement the function of any of the above embodiments, and therefore the program code and the storage medium storing the program code constitute part of the present invention.

[0187] Storage media embodiments for providing program code include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RYM, DVD-RW, DVD+RW), magnetic tapes, non-volatile memory cards, and ROMs. Alternatively, program code can be downloaded from a server computer via a communication network.

[0188] Furthermore, it should be clear that not only can the program code read by the computer be executed, but also the operating system or other components operating on the computer can be instructed based on the program code to perform some or all of the actual operations, thereby realizing the function of any of the embodiments described above.

[0189] Furthermore, it is understood that the program code read from the storage medium is written to the memory set in the expansion board inserted into the computer or to the memory set in the expansion unit connected to the computer. Then, based on the instructions of the program code, the CPU or other components installed on the expansion board or expansion unit execute some and all of the actual operations, thereby realizing the function of any of the embodiments described above.

[0190] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A code review comment generation method based on reinforcement learning, characterized in that, The method is as follows: Data collection: Collect and preprocess data on code differences, human comments, and actual code fixes to build a dataset; Generate code review comments: Fine-tune a pre-trained large language model (LLM) using a dataset and collect code discrepancy data to obtain review comments; Semantic similarity reward: Calculate the semantic similarity between the generated comment and the real comment, and generate a reward signal R_semantic; Subsequent task rewards: Input the generated comments and code differences into the code optimization model to generate a corrected code patch, and generate a reward signal R_task by evaluating the similarity between the generated patch and the real patch; Reinforcement learning fine-tuning of large language model LLM: Using the supervised fine-tuned large language model LLM as the initial policy network, comments are generated as actions based on code differences as states, and R_semantic and R_task are calculated based on a dual-reward model. The total reward R_total of the reinforcement learning fine-tuning stage is obtained by summing them according to their weights. Then, the reinforcement learning algorithm is used to maximize the expected cumulative reward and optimize the policy of the large language model LLM. Deployment and comment generation: Integrate the fine-tuned large language model into the code review system to automatically generate code review comments.

2. The code review comment generation method based on reinforcement learning according to claim 1, characterized in that, The specific data collected is as follows: Extract historical code commit records from open-source projects or internal version control systems, focusing on commits that include detailed review comments that have been discussed and accepted. Each sample contains input_diff, human_review, and fixed_diff. input_diff represents the difference text of code changes, using the Unified Diff format; human_review represents the high-quality comment text written by human reviewers for this submission that was ultimately adopted; fixed_diff is used for subsequent task rewards, and is the difference text of the final merged corrected code from this submission or a reference directly used to generate corrected code. Cleaning: Remove irrelevant information, sensitive data, overly short comments, or invalid diffs; Standardization: Unify the Diff format and comment text; Word segmentation: Segmenting the comment text into words; Dataset partitioning: Divide the dataset into training set, validation set, and test set according to a set ratio.

3. The code review comment generation method based on reinforcement learning according to claim 1, characterized in that, The code review comments are generated as follows: Supervised fine-tuning phase: Fine-tuning the large language model LLM using code differences and human comment datasets. Specifically, the pre-trained large language model LLM is fine-tuned using a dataset containing code differences and corresponding real comments written by human experts, and the mapping relationship between code difference features and comment content is initially learned. Reinforcement learning fine-tuning stage: Using a large language model LLM as the policy network, the generation policy is optimized based on reward signals; Specifically, based on supervised fine-tuning, further optimization is achieved using a proximal policy-based reinforcement learning algorithm. The core of reinforcement learning fine-tuning lies in using the reward signal provided by the designed reward model to guide the large language model (LLM) to learn and generate outputs that better meet the standards of high-quality comments. The Large Language Model (LLM) serves as a policy network. The behavior of the LLM in generating comments is regarded as the action taken by the agent in a specific state, while the reward model provides feedback from the environment.

4. The code review comment generation method based on reinforcement learning according to claim 1, characterized in that, The semantic similarity reward is as follows: Comment encoding: Using a pre-trained semantic encoding model, the generated comments and their corresponding real comments are encoded into dense vectors of fixed dimensions. Similarity metric: Calculate the cosine similarity between the generated comment and the corresponding dense vector of the real comment, both of which are fixed-dimensional vectors. Reward Calculation: The calculated cosine similarity value is appropriately scaled and translated, and fed back as the reward signal R_semantic to the reinforcement learning fine-tuning process; whereby the reward signal R_semantic directly guides the large language model LLM to generate text that is semantically closer to human comments.

5. The code review comment generation method based on reinforcement learning according to any one of claims 1-4, characterized in that, The specific rewards for subsequent tasks are as follows: Task input construction: The generated code review comments are combined with their corresponding original code difference data and used as input for the downstream code optimization model; Optimization Result Generation and Comparison: The code optimization model generates corrected code patches based on the input and compares the generated patches with the human correction patches that were actually adopted in the history of code changes; Quality Assessment and Reward Calculation: Multi-dimensional key metrics are used to evaluate the similarity between the generated patch and the real patch, and the reward signal R_task is calculated. Key metrics include the task loss value and the code similarity score. The task loss value is used to calculate the loss function value of the code optimization model when generating the patch, reflecting the confidence or difficulty of the code optimization model in generating a specific modification. The code similarity score is used to calculate the code-level similarity between the generated patch and the real patch. The code-level similarity uses the CrystalBLEU score, which, by penalizing the commonness of n-grams in the reference corpus, better highlights meaningful code snippet matching specific to the current modification. The reward signal R_task, reflecting the practicality of the review, is calculated based on key performance indicators. During the fine-tuning phase of reinforcement learning, the final reward R_total provided to the large language model LLM is a weighted sum of the semantic similarity reward R_semantic and the reward R_task for subsequent tasks, as shown in the following formula: R_total=α*R_semantic+β*R_task; Here, α and β are adjustable hyperparameters used to balance the relative importance of semantic similarity and task practicality in the optimization objective, adapting to different review needs and application scenarios.

6. A code review comment generation system based on reinforcement learning, characterized in that, The system includes: The data collection module is used to collect and preprocess data on code differences, human comments, and real code fixes, build a dataset, and divide the dataset into training, validation, and test sets according to a certain ratio. A code review comment generation framework is used to fine-tune a pre-trained large language model (LLM) using a dataset and to collect code discrepancy data to generate review comments. The semantic similarity reward module is used to calculate the semantic similarity between the generated comment and the real comment, and generate a reward signal R_semantic; The subsequent task reward module is used to input the generated comments and code differences into the code optimization model, generate a corrected code patch, and generate a reward signal R_task by evaluating the similarity between the generated patch and the real patch; The reinforcement learning fine-tuning module for the large language model LLM is used to generate comments as actions based on code differences as states, and calculates R_semantic and R_task based on a dual-reward model. The total reward R_total of the reinforcement learning fine-tuning stage is obtained by summing them by weight. Then, the reinforcement learning algorithm is used to maximize the expected cumulative reward and optimize the large language model LLM policy. The deployment and comment generation module is used to integrate the fine-tuned large language model into the code review system and automatically generate code review comments.

7. The code review and comment generation system based on reinforcement learning according to claim 6, characterized in that, The code review comment generation framework uses instruction tuning to fine-tune the Large Language Model (LLM). The sample is constructed as: Instruction: Generate a code review command for the given code diff.\nDiff:<input_diff> Comment: The target output is human_review; and the training parameters of the Large Language Model (LLM) use the standard language model training objective, train on the training set for several rounds, monitor the loss and generation quality on the validation set, and select the optimal model checkpoint for saving; the training framework is selected as Hugging Face Transformers or DeepSpeed; The semantic similarity reward model uses the pre-trained sentence-transformers / all-mpnet-base-v2. For the generated comment gen_review and the real comment human_review, 768-dimensional sentence embedding vectors V_gen and V_human are obtained through the pre-trained sentence-transformers / all-mpnet-base-v2, and R_semantic = cosine_similarity(V_gen,V_human) is calculated. At the same time, according to the reward mapping R_semantic = (cosine_similarity(V_gen,V_human)+1) / 2, R_semantic is mapped to [0,1]. The reward model for subsequent tasks is as follows: The code optimization model uses a large language model (LLM) as its architecture; the input format of the code optimization model is: Fix the code based on the diff and review:\n Diff:<input_dif f> Review:<human_review> Fixed Diff: The target output is fixed_diff; the code optimization model is trained on the (input_diff,human_review,fixed_diff) triple dataset, optimizing the cross-entropy loss. Evaluation metrics are calculated, including task loss and CrystalBLEU score. The task loss is specifically recorded when the code optimization model generates the predicted pred_fixed_diff from the input_diff and the generated gen_review. A smaller loss indicates a more "confident" code optimization model or easier generation, which is positively correlated with review quality. R_loss = -γ*Loss or its normalized form is used as part of the reward. The CrystalBLEU score is calculated between pred_fixed_diff and fixed_diff. A higher CrystalBLEU score indicates that the generated corrected code is closer to the actual correction, indirectly reflecting the stronger the guidance of the gen_review, i.e., R_bleu = δ*C - BLEU. The comprehensive reward R_task is calculated as follows: R_task = R_bleu + R_loss or R_task = w1*C-BLEU + w2*(1-Loss / MaxLoss); where w1 and w2 are weights, and MaxLoss is the maximum loss reference value set empirically.

8. The code review and comment generation system based on reinforcement learning according to claim 6 or 7, characterized in that, The specific working process of fine-tuning the large language model LLM module using reinforcement learning is as follows: ① Sample a batch of input_diff(s_t) from the dataset; where the state s_t is defined as the input_diff. ② Generate a review gen_review(a_t) using the current strategy π_θ; where action a_t is the review gen_review generated by LLM; ③ Calculate R_semantic using a semantic similarity reward model with fixed parameters; ④ Calculate R_task using code optimization model with fixed parameters and subsequent task reward calculation logic; ⑤Calculate R_total=α*R_semantic(gen_review,human_review)+β*R_task(gen_review,input_diff); ⑥ Use R_total as the immediate reward and apply the PPO algorithm to update the parameter θ of the policy network π_θ; The reinforcement learning fine-tuning large language model (LLM) module performs multiple iterations on the training set, monitors the average value of R_total and the human evaluation metrics of generated comments on the validation set, and sets early stopping to prevent overfitting. The deployment and comment generation module encapsulates the trained reinforcement learning-tuned Large Language LLM as an API service or integrates it into the review interface of CI / CD platforms or code hosting platforms. When developers submit code, they call the trained reinforcement learning-tuned Large Language LLM. The input format of the trained reinforcement learning-tuned Large Language LLM is: Generate acode review comment:\n Diff:<in put_diff> Comment: The output of the trained reinforcement learning-tuned large language LLM is the generated code review comment, presented to the developer or reviewer.

9. An electronic device, characterized in that, include: Memory and at least one processor; The memory contains computer programs; The at least one processor executes the computer program stored in the memory, causing the at least one processor to perform the reinforcement learning-based code review comment generation method as described in any one of claims 1 to 5.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that can be executed by a processor to implement the reinforcement learning-based code review comment generation method as described in any one of claims 1 to 5.

Citation Information

Cited By

  • Optical character recognition model reinforcement learning optimization method and device

    CN121904558A