Translation model training method, text translation method and generative model training method
By introducing a mixed reward mechanism for format inspection and translation evaluation indicators in the machine translation model, the strategy of the translation model is optimized, and the problem of low translation quality of existing machine translation models is solved, and a higher quality translation effect is achieved.
Patent Information
- Application Number
- CN202510464605.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-29
AI Technical Summary
The translation quality of existing machine translation models is not high, the existing improvement methods are complex and have limited results, making it difficult to achieve or exceed the performance of the current best model.
The machine translation model training framework based on reinforcement learning is adopted, combining a mixed reward mechanism of format inspection rewards and translation evaluation indicators, and optimizes the translation model through format rewards and metric rewards to form a mixed reward of rules-metrics, and guides the model to improve translation quality without structured CoT data or complex search algorithms.
It improves the translation quality of machine translation models, is suitable for open translation tasks, simplifies the training process, avoids dependence on complex technologies, and achieves performance that matches or even exceeds the existing best models.
Smart Images

Figure CN120387527A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technologies, and particularly to a method for training a translation model, a method for text translation, and a method for training a generation model. Background Art
[0002] A large language model (LLM) refers to a deep learning model trained using a large amount of text data, enabling the model to generate natural language text or understand the meaning of language text. It is commonly used to handle complex tasks such as text summarization, machine translation, sentiment analysis, dialogue generation, and content recommendation.
[0003] Traditionally, in the field of machine translation, various implementation solutions for using LLM models to handle machine translation tasks have emerged. However, currently, the machine translation models trained using LLM models still have the problem of low translation quality. Summary of the Invention
[0004] Based on this, to address the above technical problems, it is necessary to provide a method for training a machine translation model, a method for text translation, a method for training a generation model, an apparatus, a computer device, a computer-readable storage medium, and a computer program product that can improve the translation quality of a machine translation model.
[0005] In a first aspect, the present application provides a method for training a machine translation model, including:[[]]
[0006] Inputting training texts in a first language into an initial translation model to obtain a model output text, where the model output text includes a model translation text in a second language;
[0007] Obtaining a format reward corresponding to the model output text based on the format of the model output text, where the format reward is used to characterize whether the model output text conforms to a predetermined format;
[0008] Obtaining a metric reward corresponding to the model translation text based on the training text and the model translation text, where the metric reward is used to characterize the translation quality of the model translation text;
[0009] Obtaining a target reward according to the format reward and the metric reward;
[0010] Training the initial translation model based on the target reward and a preset reinforcement learning algorithm to obtain a target translation model.
[0011] In one embodiment, the predetermined format is determined based on a translation instruction, and the format reward includes a first value or a second value; the first value is used to characterize that the model output text conforms to the predetermined format, and the second value is used to characterize that the model output text does not conform to the predetermined format.
[0012] In one embodiment, the metric reward includes a first metric reward and / or a second metric reward. The metric reward corresponding to the model translation text obtained based on the training text and the model translation text includes:
[0013] Calculating a first metric reward for the model translation text based on the reference translation text corresponding to the training text and the model translation text; and / or
[0014] Calculating a second metric reward for the model translation text by using a translation evaluation model based on the training text and the model translation text.
[0015] In one embodiment, the method further includes:
[0016] Performing alignment processing on the first metric reward and / or the second metric reward;
[0017] Fusing the aligned first metric reward and / or the second metric reward to obtain the metric reward.
[0018] In one embodiment, the metric reward is obtained based on mapping processing and is within a predetermined interval, and the predetermined interval is determined based on the value of the format reward.
[0019] In one embodiment, the mapping processing includes:
[0020] Based on the comparison result between the metric reward and a predetermined discrete reward threshold, mapping the metric reward to a third value or a fourth value within the predetermined interval; or mapping the metric reward to the predetermined interval by performing a predetermined linear transformation on the metric reward.
[0021] In one embodiment, obtaining a target reward according to the format reward and the metric reward includes:
[0022] Performing piecewise fusion on the format reward and the metric reward to obtain the target reward; or
[0023] Performing weighted fusion on the format reward and the metric reward to obtain the target reward.
[0024] In one embodiment, training an initial translation model based on the target reward and a preset reinforcement learning algorithm to obtain a target translation model includes:
[0025] Constructing an objective function of the reinforcement learning algorithm based on the target reward;
[0026] Optimizing the parameters of the initial translation model by maximizing the objective function to obtain the target translation model.
[0027] In a second aspect, the present application further provides a text translation method, and the method includes:
[0028] Receive a translation request, where the translation request indicates translating the source text in the source language into the target language;
[0029] In response to the translation request, input the source text in the source language into the target translation model to obtain a translation text in the target language. The target translation model is trained by using the training method of the translation model in the first aspect.
[0030] In a third aspect, the present application also provides a training method for a generation model. The method includes:
[0031] Input training data into an initial generation model to obtain model output data; the model output data includes generated data corresponding to the training data; the initial generation model is used for open-ended generation tasks;
[0032] Obtain a format reward corresponding to the model output data based on the format of the model output data. The format reward is used to characterize whether the model output data conforms to a predetermined format;
[0033] Obtain a metric reward corresponding to the generated data based on the training data and the generated data. The metric reward is used to characterize the generation quality of the generated data;
[0034] Obtain a target reward according to the format reward and the metric reward;
[0035] Train the initial generation model based on the target reward and a preset reinforcement learning algorithm to obtain a target generation model.
[0036] In a fourth aspect, the present application also provides a training device for a machine translation model, including:
[0037] A processing module, configured to input training text in a first language into an initial translation model to obtain a model output text. The model output text includes a model translation text in a second language;
[0038] A format reward module, configured to obtain a format reward corresponding to the model output text based on the format of the model output text. The format reward is used to characterize whether the model output text conforms to a predetermined format;
[0039] A metric reward module, configured to obtain a metric reward corresponding to the model translation text based on the training text and the model translation text. The metric reward is used to characterize the translation quality of the model translation text;
[0040] A reward fusion module, configured to obtain a target reward according to the format reward and the metric reward;
[0041] A training module, configured to train the initial translation model based on the target reward and a preset reinforcement learning algorithm to obtain a target translation model.
[0042] In a fifth aspect, the present application also provides a text translation device, including:
[0043] A receiving module, configured to receive a translation request, where the translation request indicates to translate a source text in a source language into a target language;
[0044] A translation module, configured to, in response to the translation request, input the source text in the source language into a target translation model to obtain a translation text in the target language, where the target translation model is obtained by training using a training device adopting the translation model in the fourth aspect.
[0045] In a sixth aspect, the present application further provides a training device for a generation model, including:
[0046] A processing module, configured to input training data into an initial generation model to obtain model output data; the model output data includes generated data corresponding to the training data; the initial generation model is used for an open-ended generation task;
[0047] A format reward module, configured to obtain a format reward corresponding to the model output data based on the format of the model output data, where the format reward is used to represent whether the model output data conforms to a predetermined format;
[0048] A metric reward module, configured to obtain a metric reward corresponding to the generated data based on the training data and the generated data, where the metric reward is used to represent the generation quality of the generated data;
[0049] A reward fusion module, configured to obtain a target reward according to the format reward and the metric reward;
[0050] A training module, configured to train the initial generation model based on the target reward and a preset reinforcement learning algorithm to obtain a target generation model.
[0051] In a seventh aspect, the present application further provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the steps of the methods in the first aspect, the second aspect, and the third aspect are implemented.
[0052] In an eighth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the methods in the first aspect, the second aspect, and the third aspect are implemented.
[0053] In a ninth aspect, the present application further provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the methods in the first aspect, the second aspect, and the third aspect are implemented.
[0054] The training method of the above-mentioned translation model, the text translation method, the training method of the generation model, the device, the computer device, the storage medium and the computer program product input the training text in the first language into the initial translation model to obtain the model output text. Among them, the model translation text in the second language in the model output text is then obtained. Next, the format reward corresponding to the model output text is obtained based on the format of the model output text, and the format reward is used to represent whether the model output text conforms to the predetermined format; and the metric reward corresponding to the model translation text is obtained based on the training text and the model translation text, and the metric reward is used to represent the translation quality of the model translation text; and according to the format reward and the metric reward, the target reward is obtained; and then the initial translation model is trained based on the target reward and the preset reinforcement learning algorithm to obtain the target translation model. That is, in the training method proposed in the embodiments of the present application, a hybrid reward method combining format reward and metric reward is adopted. While forcing the model to output in a structured format through the format reward, the metric reward can also be used to measure the quality of the translation text output by the model, which can provide rich and effective guiding signals for reinforcement learning to be applicable to open-ended translation tasks, so as to facilitate the optimization of the model strategy through reinforcement learning and finally improve the translation quality of the translation model. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or the related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0056] Figure 1 It is an application environment diagram of the training method of the translation model in one embodiment;
[0057] Figure 2 It is a schematic flowchart of the training method of the translation model in one embodiment;
[0058] Figure 3 It is a schematic flowchart of the training method of the translation model in another embodiment;
[0059] Figure 4 It is a schematic flowchart of the training method of the translation model in another embodiment;
[0060] Figure 5 It is a schematic flowchart of the training method of the translation model in another embodiment;
[0061] Figure 6 It is a schematic flowchart of the training method of the translation model in another embodiment;
[0062] Figure 7 Schematic diagram of the training effect of model training using the training method of this application in an embodiment;
[0063] Figure 8 Schematic diagram of the performance comparison between the model obtained by using the training method of this application and other existing models in an embodiment;
[0064] Figure 9 Schematic flow chart of the text translation method in an embodiment;
[0065] Figure 10 Schematic flow chart of the training method of the generation model in an embodiment;
[0066] Figure 11 Structural block diagram of the training device of the translation model in an embodiment;
[0067] Figure 12 Structural block diagram of the text translation device in an embodiment;
[0068] Figure 13 Structural block diagram of the training device of the generation model in an embodiment;
[0069] Figure 14 Internal structure diagram of a computer device in an embodiment. Detailed implementation manners
[0070] In order to make the objectives, technical solutions and advantages of this application clearer, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application, but not to limit this application.
[0071] Large Language Model (LLM) has achieved great success in natural language processing tasks. However, for open-ended generation tasks, especially the training of large language models in the field of Machine Translation (MT), there are still challenges.
[0072] When training large language models, existing methods for improving MT inference ability usually rely on complex techniques, such as process reward models that require fine annotation, Monte Carlo tree search, or structured Chain-of-Thought (CoT) data that requires manual design / synthesis. Moreover, the performance of these techniques often lags behind the current best models and requires a large amount of manually annotated data.
[0073] Since existing MT inference enhancement methods are of high complexity and have limited training effects, it is an urgent technical problem to be solved how to provide a simpler (not relying on CoT structured data, MCTS) and more effective method to improve the translation quality of LLM so as to achieve or even exceed the performance of existing best models (including closed-source models).
[0074] Based on a series of problems faced by existing machine translation models in improving translation quality, the embodiments of this application propose a training framework for machine translation models based on reinforcement learning to implement the training of machine translation models, thereby improving the translation quality of machine translation models. Among them, in the process of training a machine translation model based on a reinforcement learning framework, an effective reward mechanism suitable for MT tasks is also proposed, that is, a rule-based format check reward and a translation quality reward based on translation evaluation metrics are combined to form a Rule-Metric-Mixed Reward, which can guide the LLM to improve translation quality through an emergent inference process without structured CoT data or complex search algorithms. Exemplarily, the translation evaluation metrics can be network model-based translation evaluation metrics such as COMETKiwi metrics, XCOMET metrics, etc., or mathematical model-based translation evaluation metrics such as BLEU metrics, ROUGE metrics, etc.
[0075] The training method for the translation model provided by the embodiments of this application can be applied to an application environment as Figure 1 shown. Among them, the computer device 102 can be a terminal or a server. Among them, the terminal can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, etc., and the server can be implemented by an independent server or a server cluster composed of multiple servers. The computer device 102 realizes the training of the translation model by executing the training method of the translation model to obtain a translation model with better translation quality.
[0076] Exemplarily, the text translation method provided by the embodiments of this application can also be applied to an application environment as Figure 1 shown, that is, the computer device 102 can also use the translation model obtained after training to translate the input source language text to obtain the translated target language text.
[0077] Exemplarily, the training method for the generation model provided by the embodiments of this application can also be applied to an application environment as Figure 1In the application environment shown, that is, the computer device 102 can also execute the training method of the generation model to implement the training of the generation model, so as to obtain a generation model with higher content generation quality. Among them, the generation model can be applied to any open-ended generation task, that is, the application field where the reference content / answer is not unique, that is, in the training process, the scenario where the reference content is not unique. For example, the open-ended generation task can include but is not limited to translation tasks, summary tasks, content creation tasks, intelligent reply tasks, dialogue tasks, etc.
[0078] In an exemplary embodiment, as Figure 2 shown, a training method for a translation model is provided. Taking the computer device in Figure 1 as an example, the following steps 201 to step 205 are included. Among them:
[0079] Step 21, input the training text in the first language into the initial translation model to obtain the model output text, and the model output text includes the model translation text in the second language.
[0080] Among them, the initial translation model can be a translation model obtained after pre-training. The pre-trained translation model can refer to a translation model pre-trained on a general dataset, which has the ability to learn general features and adapt to multiple tasks.
[0081] Exemplarily, the pre-trained translation model can include: versions with any parameter magnitudes of any translation model (such as 3B, 7B). In the embodiments of the present application, there are no more limitations on the pre-trained translation model. In this embodiment, the pre-trained translation model is used as the basic model when training the translation model, and the training text in the first language is used to fine-tune the pre-trained translation model, so as to implement the training of the translation model and obtain the trained translation model. The translation model is used to translate the text in the first language into the text in the second language, where the first language and the second language are two different languages. It should be noted that the translation model can also implement translation tasks between any other two languages except from the first language to the second language, such as translating the second language into the first language, or translating the third language into the fourth language, translating the third language into the first language, translating the third language into the second language, etc.
[0082] Exemplarily, the training text in the first language can be obtained through any public channel, for example, from parallel translation data. Herein, the parallel translation data may refer to sentence pairs that are corresponding or translated in content, usually including source language text and target language text, which express the same or similar information. Specifically, taking the first language as the source language and the second language as the target language as an example, the source language text can be used as the training text, and the target language text can be used as the reference translation text to jointly form the training data; the source language text can also be used alone as the training data without a corresponding reference translation text. As an optional implementation manner, after obtaining the training text, preprocessing can also be performed on the training text, and then the preprocessed training text can be used as the training text of the translation model. Herein, the preprocessing can include data filtering processing. For example, sentences with the number of characters less than a preset number threshold in the training text are removed, such as sentences with less than 20 characters are removed, etc.
[0083] When training the initial translation model with the training text, the training text can be input into the initial translation model for text translation to output the model translation text corresponding to the training text. Herein, the model translation text refers to the translation text obtained by translating the training text using the initial translation model. Exemplarily, it can be stipulated that the initial translation model outputs the text in a predetermined format, and the predetermined format can at least include the model translation text, that is, the model output text of the initial translation model can at least include the model translation text of the translated second language. As an optional implementation manner, the predetermined format can include the translation thinking (or reasoning) process and the model translation text, that is, the model output text can include the translation thinking (or reasoning) process and the model translation text of the second language.
[0084] Step 202: Obtain the format reward corresponding to the model output text based on the format of the model output text. The format reward is used to represent whether the model output text conforms to the predetermined format.
[0085] Step 203: Obtain the metric reward corresponding to the model translation text based on the training text and the model translation text. The metric reward is used to represent the translation quality of the model translation text.
[0086] Step 204: Obtain the target reward according to the format reward and the metric reward.
[0087] In this embodiment, after obtaining the model output text, reward calculation can be performed on the format of the model output text and the translation quality of the model translation text in the model output text to obtain the format reward for the format of the model output text and the metric reward for the translation quality of the model translation text. Furthermore, the target reward corresponding to the initial translation model can be obtained according to the format reward and the metric reward.
[0088] Exemplarily, the target reward can be obtained by calculating the reward based on the preset reward metrics, training text, model output text, and model translation text. Among them, the preset reward metrics can include format reward metrics and metric reward metrics. The format reward metric refers to the metric for judging the output format of the model output text output by the initial translation model, and the metric reward metric refers to the metric for evaluating the translation quality of the model translation text output by the initial translation model.
[0089] As an alternative implementation, the computer device can use the format reward metric to judge the output format of the model output text, that is, to judge whether the output format of the model output text conforms to the predetermined format, so as to obtain the format reward. Exemplarily, for a translation task, it is required to generate a response in a specific format according to the translation request. Among them, the translation request can include the source language text src_text, source language src_language, and target language tgt_language, etc. After obtaining the translation request, a translation instruction for the computer device can also be generated, and the translation instruction can at least target the predetermined format of the response. Among them, the predetermined format of the response can include two parts: <think>The internal thinking / reasoning process surrounded by tags, and <translate>The final translation result surrounded by tags. That is, for the format reward metric, it can be required that the model's output includes a clear thinking process and the final translation result. For example: The thinking process can be qualified by <think> ...< / think> tags and through <think>The label encloses the internal thinking / reasoning process, and the final translation result can be qualified by <translate> ...< / translate> the label, and through <translate>The label surrounds the final translation result.
[0090] When judging the output format of the model output text based on the format reward metric, a regular expression can be used to check whether the model output conforms to the structure of " <think> ...< / think> <translate> ...< / translate> ", that is, to check whether the output format of the model output text output by the initial translation model is this structural format. Exemplarily, if the output format of the model output text is the predetermined format (i.e., the above structural format), the format reward is set to the first value, such as 1, and this first value is used to indicate that the output format of the model output text is correct, that is, the model output text conforms to the predetermined format; if the output format of the model output text is not the predetermined format, the format reward is set to the second value, such as -1, and this second value is used to indicate that the output format of the model output text is incorrect, that is, the model output text does not conform to the predetermined format. The corresponding calculation method is as follows:
[0091]
[0092] where S format represents the format reward. If the format is correct, 1 point is obtained. If there is any deviation in the format (for example, missing tags, incorrect tag nesting, incorrect tag order, etc.), the format component S format is -1.
[0093] Through the format reward metric, the model can be forced to generate an output that conforms to the predefined structure, which is a prerequisite for subsequent meaningful quality assessment and application. In this embodiment, it is required that the model output includes a clear thinking process and the final translation result.
[0094] As an alternative implementation, the computer device can use the metric reward metric to evaluate the translation quality of the training text and the model translation text, and obtain the metric reward corresponding to the model translation text, so as to measure the translation quality of the model translation text through the metric reward. Exemplarily, for the metric reward metric, advanced, learning-based automated translation evaluation metrics such as the COMETKiwi metric, the XCOMET metric, etc. can be used. These metrics can evaluate the translation quality based on the training text and the training translation text without a reference translation text; in addition, in the case where there is a reference translation text for the training text, metrics such as BLEU can also be used to evaluate the translation quality based on the training text, the reference translation text, and the training translation text, so as to obtain the metric reward value. In the embodiments of the present application, the type of the metric reward metric is not specifically limited, and during the training process, one or more metric reward metrics can be selected according to the training text to evaluate the translation quality.
[0095] Exemplarily, the computer device may determine whether to perform the calculation of the metric reward based on the format reward value. If the format reward indicates that the model output text conforms to the predetermined format, the metric reward calculation is performed based on the metric reward index, the training text, and the model translation text to determine the metric reward; if the format reward value indicates that the model output text does not conform to the predetermined format, the preset metric reward may be determined as the metric reward. That is, on the basis of the correct model output format, it is evaluated with the help of the metric reward index <translate>The translation quality of the content within the tag (i.e., the text translated by the model).
[0096] Exemplarily, taking the metric reward metric including the automated MT evaluation metric COMETKiwi (such as COMETKiwi - 23) as an example, when using the advanced automated MT evaluation metric COMETKiwi - 23 to evaluate the translation quality, the calculation of its metric reward depends on the source text (i.e., the training text) and the generated translation text (i.e., the model translation text). This metric is learning - based and has been proven to have a high correlation with human judgment of translation quality. In addition, the method proposed in this embodiment can also be applied to the metric calculation of the reference translation (reference) that needs to be manually written. Therefore, the calculation of the metric reward can be expressed as:
[0097] S metric = Metric([source,translation,(reference)])
[0098] where source is the source text, i.e., the training text, translation is the generated translation text, i.e., the model translation text, reference is the reference translation text corresponding to the training text, and Metric() can represent the metric function corresponding to the metric reward metric. S metric represents the metric reward.
[0099] It should be noted that in the case of different metric reward metrics, the metric functions corresponding to the metric reward metrics can be different, that is, the calculation methods of the metric rewards are different.
[0100] Exemplarily, when the format reward is the first value, it indicates that the output format of the model output text conforms to the predetermined format, that is, the output format is correct. At this time, the metric reward metric can be used to evaluate the translation quality of the model translation text in the model output text to obtain the metric reward. On the contrary, when the format reward value is the second value, it indicates that the output format of the model output text does not conform to the predetermined format, that is, the output format is incorrect. At this time, a fixed value can be set for the metric reward, that is, the preset metric reward, such as - 2, - 3, etc.
[0101] In the case of obtaining the format reward and the metric reward, the final target reward can be determined based on the format reward and the metric reward for subsequent driving of the reinforcement learning (RL) algorithm. Exemplarily, the format reward and the metric reward can be segmented and fused to obtain the target reward; or, the format reward and the metric reward can also be weighted and fused to obtain the target reward.
[0102] Exemplarily, for segmented fusion, the target reward can be determined according to the following logic:
[0103]
[0104] where r is the target reward value. If the format of the model output is incorrect, i.e., S format =-1, then it is the metric reward value S metric Assign a fixed and large negative reward value (such as -2). This design strongly encourages the model to quickly learn and master the correct output format first, because incorrect formats will be significantly penalized. Only when the format reward value S format =1, will the metric reward value S metric be calculated, which ensures that quality scoring is only performed on outputs with valid structures. When the format is correct, the reward score r is continuous, enabling the model to learn from subtle differences in translation quality. The choice of this reward mechanism has a significant impact on the training dynamics and the final model performance.
[0105] Exemplarily, for weighted fusion, when determining the target reward based on the format reward and the metric reward, the target reward can also be determined according to the result of the weighted sum of the format reward and the metric reward. For example: multiply the format reward by a coefficient and then sum it with the metric reward, such as the target reward r = β * S format + S metric , where 0 < β < 1.
[0106] Step 205: Train the initial translation model based on the target reward and a preset reinforcement learning algorithm to obtain the target translation model.
[0107] In reinforcement learning (RL), the reward signal is the core that guides the optimization process and shapes the model's behavior. In this embodiment, combining the format (rule) reward and the metric reward provides an effective scalar reward signal for RL training, so as to optimize the policy of the initial translation model through reinforcement learning, that is, to optimize the policy model π θ of the LLM model. The policy model π θ of the LLM model includes multiple parameters of the LLM model, and policy optimization is also parameter optimization.
[0108] Exemplarily, the computer device may construct an objective function of the reinforcement learning algorithm based on the target reward, and optimize the parameters of the initial translation model by maximizing the objective function to obtain the target translation model; wherein, the reinforcement learning algorithm may include, but is not limited to, the Group Relative Policy Optimization (GRPO) algorithm, the Proximal Policy Optimization (PPO) algorithm, the Direct Preference Optimization (DPO) algorithm, the DQN (Deep Q-Network) algorithm, etc., where the DQN algorithm is a reinforcement learning algorithm that combines deep learning and Q-learning. In this embodiment, the type of the reinforcement learning algorithm is not limited.
[0109] Exemplarily, taking the GRPO algorithm as an example below, the policy optimization process of the initial translation model will be described in detail.
[0110] For each training text (denoted as the source text q), sample from the current policy model π θ to generate a set containing G outputs, i.e., {o1, o2,..., o G}. For each generated output o i , use the above "rule-metric hybrid reward" to calculate the reward r i , and obtain the intra-group rewards {r1, r2,..., r G}.
[0111] Next, based on the intra-group rewards, calculate the relative advantage A i of each output, that is, the Advantage Function. The Advantage Function A i can measure the relative quality of the output o i relative to other outputs in the same group. The calculation method is usually based on the normalization of the intra-group rewards:
[0112]
[0113] where mean and std respectively represent the mean and standard deviation of the intra-group rewards {r1, r2,..., r G}, that is, estimate the advantage of each output through the normalization of the intra-group rewards.
[0114] Use the Clipped Objective Function of PPO to update the policy model π θ , and maximize the expected return. The optimization objective of GRPO, that is, the objective function, can be expressed as:
[0115]
[0116] where E represents the expectation, is the importance sampling ratio, clip(x, min, max) is a function that truncates x within the range [min, max], and A i is the advantage function calculated above, ∈ is the truncation hyperparameter of PPO, β is the weight of the KL divergence penalty term (which can be set to 0 in experiments), and D KL (π θ ||π ref ) is the KL divergence of the current policy model π θ relative to the initial reference policy (usually referring to the model after supervised fine-tuning (SFT)), which is used to constrain the magnitude of policy updates. Among them, the KL divergence (Kullback-Leibler, KL) is used to measure the difference between two probability distributions.
[0117] Exemplarily, during the training process, the hyperparameters involved can include: learning rate, sampling temperature, maximum generation length, batch size, the number of rollouts per sample (G), PPO truncation range (clipping range, ∈), number of training epochs, etc. It should be noted that the training hyperparameters given in this embodiment are only used as an example and are not used to limit the specific types and specific values of the hyperparameters.
[0118] The current policy model π of the initial translation model is optimized by maximizing the above objective function θ , that is, when performing policy updates, the policy model π is gradually updated by optimizing this objective function θ , so as to achieve iterative training of the initial translation model and obtain the target translation model. The above uses the GRPO algorithm (optionally, without KL divergence constraint) for reinforcement learning fine-tuning. This process can guide the LLM to spontaneously generate intermediate reasoning steps while improving the translation quality (recorded in <think>After the training is completed, the target translation model can be obtained.
[0119] In the training method of the above-mentioned translation model, a training text in a first language is input into the initial translation model to obtain a model output text, wherein the model output text is a model translation text in the second language, and then, based on the format of the model output text, a format reward corresponding to the model output text is obtained, and the format reward is used to characterize whether the model output text conforms to a predetermined format; and a metric reward corresponding to the model translation text is obtained based on the training text and the model translation text, and the metric reward is used to characterize the translation quality of the model translation text; and according to the format reward and the metric reward, a target reward is obtained; and then the initial translation model is trained based on the target reward and a preset reinforcement learning algorithm to obtain a target translation model. That is, in the training method proposed in the embodiment of the present application, a hybrid reward method combining format reward and metric reward is adopted. While forcing the model to output a structured format through format reward, the quality of the translation text output by the model can also be measured with the help of metric reward, which can provide information-rich and effective guidance signals for reinforcement learning, so as to be applicable to open translation tasks, thereby facilitating the optimization of model strategies through reinforcement learning, and ultimately improving the translation quality of the translation model.
[0120] In an exemplary embodiment, the training text in the first language may include unsupervised training text, i.e., a reference translation text in the second language without manual annotations, or supervised training text, i.e., a reference translation text in the second language with manual annotations. In the case where the training text includes both unsupervised and supervised training texts, the MT quality assessment index used for the unsupervised training text may be different from the MT quality assessment index used for the supervised training text. Exemplarily, the metric reward index may include a first metric reward index and / or a second metric reward index, wherein the first metric reward index may be a metric reward index corresponding to the supervised training text, and the second metric reward index may be a metric reward index corresponding to the unsupervised training text. Accordingly, the metric reward may include the first metric reward and / or the second metric reward corresponding to the supervised training text. Then, the step 203 of "obtaining the metric reward corresponding to the model translation text based on the training text and the model translation text" may include:
[0121] Step A: Calculate a first metric reward of the model translation text based on the reference translation text and the model translation text corresponding to the training text.
[0122] If there is a reference translation text for the training text, then based on the first metric reward indicator, the reference translation text corresponding to the training text, and the model translation text, a metric reward calculation is performed to determine the first metric reward. Exemplarily, the first metric reward indicator may include evaluation indicators based on word overlap such as the BLEU (Bilingual Evaluation Understudy) indicator, the ROUGE (Recall-Oriented Understudy for Gisting Evaluation) indicator, the METEOR (Metric for Evaluation of Translation with Explicit Ordering) indicator, etc., may also include evaluation indicators based on word vectors such as Embedding Average Score, Greedy Matching Score, Vector Extrema Score, etc., may further include distance-based evaluation indicators such as TER (Translation Edit Rate), HTER (Human-targeted Translation Edit Rate), etc., and other evaluation indicators for supervised training texts. When performing the first metric calculation, the reference translation text corresponding to the training text and the model translation text can be input into the calculation formula or mathematical model corresponding to the first metric reward indicator to calculate the first metric reward of the model translation text.
[0123] It should be noted that when calculating the first metric reward, one first metric reward indicator can be used for the metric calculation, or multiple first metric reward indicators can be used for the metric calculation, and after fusing the multiple first metric rewards corresponding to the multiple first metric reward indicators, the first metric reward is obtained.
[0124] Step B, based on the training text and the model translation text, use the translation evaluation model to calculate the second metric reward of the model translation text.
[0125] If there is no reference translation text for the training text, then based on the second metric reward indicator, the training text, and the model translation text, a metric reward calculation is performed to determine the second metric reward. Exemplarily, the second metric reward indicator may include evaluation indicators for unsupervised training texts such as COMET Kiwi, XCOMET indicator, etc. When performing the second metric calculation, the training text and the model translation text can be input into the translation evaluation model corresponding to the second metric reward indicator, and the second metric reward of the model translation text is output. Among them, the translation evaluation model can be an evaluation model trained based on machine learning, deep learning, etc.
[0126] It should be noted that when calculating the second metric reward, a single second metric reward indicator can be used for metric calculation, or multiple second metric reward indicators can be used for metric calculation. After fusing the multiple second metric rewards corresponding to the multiple second metric reward indicators, the second metric reward is obtained.
[0127] As an optional case, when there is a misalignment of physical dimensions between the first metric reward obtained based on the first metric reward indicator and the second metric reward obtained based on the second metric reward indicator, for the training texts evaluated for translation quality using the first metric reward indicator and the training texts evaluated for translation quality using the second metric reward indicator, there will be calculation biases during the subsequent calculation of the advantage function in reinforcement learning (RL) and the optimization of the RL algorithm, resulting in problems in the optimization of the LLM policy based on the RL algorithm. Therefore, for this case, the metric reward values of the unsupervised training data and the supervised training data can be aligned in terms of physical dimensions.
[0128] Exemplarily, alignment processing can be performed on the first metric reward and / or the second metric reward. For example, the first metric reward of the supervised training data can be aligned to the physical dimension of the second metric reward of the unsupervised training data. For example, when calculating the first metric reward, first, based on the first metric reward indicator, the reference translation text corresponding to the training text, and the model translation text, the first candidate metric reward is calculated; then, based on the first mapping ratio, the first candidate metric reward is mapped to obtain the first metric reward; where the first mapping ratio is used to align the physical dimension of the first metric reward indicator to the physical dimension of the second metric reward indicator.
[0129] Exemplarily, the second metric reward of the unsupervised training data can also be aligned to the physical dimension of the first metric reward of the supervised training data. For example, when calculating the second metric reward, first, based on the second metric reward indicator, the training text, and the model translation text, the second candidate metric reward of the model translation text is calculated using the translation evaluation model; then, based on the second mapping ratio, the second candidate metric reward is mapped to obtain the second metric reward; where the second mapping ratio is used to align the physical dimension of the second metric reward indicator to the physical dimension of the first metric reward indicator.
[0130] Exemplarily, the first metric reward of the supervised training data and the second metric reward of the unsupervised training data can also be aligned to a preset physical dimension. For example, for the first metric reward, after obtaining the first candidate metric reward, based on the third mapping ratio, the first candidate metric reward is mapped to obtain the first metric reward corresponding to the preset physical dimension; and for the second metric reward, after obtaining the second candidate metric reward, based on the fourth mapping ratio, the second candidate metric reward is mapped to obtain the second metric reward corresponding to the preset physical dimension; where the third mapping ratio is used to align the physical dimension of the first metric reward index to the preset physical dimension, and the fourth mapping ratio is used to align the physical dimension of the second metric reward index to the preset physical dimension.
[0131] Exemplarily, the above mapping ratios can be obtained through theoretical calculations, experimental verification, or experience. The mapping ratios can be the same or different, and the present application does not make specific limitations in this regard.
[0132] As another alternative implementation, for the supervised training text, that is, when there is a reference translation text for the training text, the first metric reward index and the second metric reward index can also be used simultaneously to comprehensively evaluate the quality of the model translation text; Exemplarily, when performing translation quality evaluation, based on the reference translation text corresponding to the training text and the model translation text, the first metric reward of the model translation text can be calculated, and based on the training text and the model translation text, the second metric reward of the model translation text can be calculated using the translation evaluation model. Then, alignment processing is performed on the first metric reward and / or the second metric reward, and the aligned first metric reward and / or the second metric reward are fused to obtain the metric reward of the model translation text.
[0133] That is to say, both the first metric reward index and the second metric reward index are used to evaluate the translation quality of the model's translated text. Furthermore, the evaluation results of the two, namely the first metric reward and the second metric reward, are fused to obtain the final metric reward. However, since the physical dimensions corresponding to the first metric reward index and the second metric reward index may be different, when performing reward fusion, it is necessary to align the first metric reward and the second metric reward. It is possible to align the first metric reward to the second metric reward, or align the second metric reward to the first metric reward, or align both the first metric reward and the second metric reward to a preset physical dimension. For the first case, aligning the first metric reward to the second metric reward can obtain the aligned first metric reward. Then, the aligned first metric reward and the second metric reward are fused to obtain the metric reward. For the second case, aligning the second metric reward to the first metric reward can obtain the aligned second metric reward. Then, the first metric reward and the aligned second metric reward are fused to obtain the metric reward. For the third case, aligning the first metric reward to the preset physical dimension can obtain the aligned first metric reward, and aligning the second metric reward to the preset physical dimension can obtain the aligned second metric reward. Then, the aligned first metric reward and the aligned second metric reward are fused to obtain the metric reward.
[0134] In this embodiment, the metric reward index may include the first metric reward index for supervised training text and / or the second metric reward index for unsupervised training text. Correspondingly, the metric reward may include the first metric reward and / or the second metric reward. For supervised training text, the first metric reward can be obtained by performing metric calculation based on the first metric reward index, the reference translation text corresponding to the training text, and the model's translated text. For unsupervised training text, the second metric reward can be obtained by performing metric calculation based on the second metric reward index, the training text, and the model's translated text. Additionally, it is also possible to align the physical dimensions of the first metric reward index and the second metric reward index through a mapping ratio, and fuse the aligned first metric reward and / or the second metric reward to obtain the final metric reward, which can ensure the consistency of quality evaluation for supervised training text and unsupervised training text during the training process, thereby improving the training reliability of the translation model and the translation quality of the translation model.
[0135] In an exemplary embodiment, since the score range of the metric reward index may not be [0,1], in order to match its magnitude with the format reward index and be more suitable for the calculation of the advantage function in the subsequent RL algorithm, the original S obtained by using the metric function can be metric The fractional mapping is scaled to a standardized fractional range, usually within the predetermined interval of [0, 1]; that is to say, the above-mentioned metric reward can be obtained based on the mapping process and is within the predetermined interval, and this predetermined interval is determined based on the value of the format reward.
[0136] Exemplarily, the mapping process may include: mapping the metric reward to a third value or a fourth value within the predetermined interval based on the comparison result between the metric reward and the predetermined discrete reward threshold; or mapping the metric reward to the predetermined interval by performing a predetermined linear transformation on the metric reward.
[0137] The mapping processes of the first metric reward indicator and the second metric reward indicator will be described in detail below respectively.
[0138] For the first metric reward indicator, the first mapping process is adopted. The above step A "calculate the first metric reward of the model translation text based on the reference translation text and the model translation text corresponding to the training text" may include steps 301 to 303. Among them:
[0139] Step 301: Input the reference translation text and the model translation text corresponding to the training text into the discrete reward function to determine the first discrete reward.
[0140] Among them, the discrete reward function is a discrete reward function determined based on the first metric reward indicator.
[0141] Step 302: If the first discrete reward is greater than or equal to the first discrete reward threshold, determine the first metric reward of the model translation text as the third value.
[0142] Step 303: If the first discrete reward is less than the first discrete reward threshold, determine the first metric reward of the model translation text as the fourth value.
[0143] Exemplarily, taking the first metric reward indicator as the BLEU indicator as an example, the above-mentioned S metric expression can be used. Input the reference translation text and the model translation text corresponding to the training text required in this expression into the discrete reward function to determine the first discrete reward, that is, obtain the Metric calculation value. Then, when the first discrete reward is greater than or equal to the first discrete reward threshold (such as 40), it can be considered that the translation quality is good. At this time, the metric reward of the model translation text can be the third value, such as 1, that is, when BLEU≥40, S metric =1; when the first discrete reward is less than the first discrete reward threshold (such as 40), it can be considered that the translation quality is poor. At this time, the metric reward of the model translation text can be the fourth value, such as 0, that is, when BLEU<40, S metric =0.
[0144] In this embodiment, for supervised training texts, discrete rewards based on BLEU thresholds can be used for translation quality evaluation to obtain the first metric reward.
[0145] For the first metric reward metric, using the second mapping processing method, the above step A "calculating the first metric reward of the model translation text based on the reference translation text and the model translation text corresponding to the training text" may include steps 401 to 402. Wherein:
[0146] Step 401: Input the reference translation text and the model translation text corresponding to the training text into the continuous reward function to determine the first continuous reward.
[0147] Among them, the continuous reward function is a continuous reward function determined based on the first metric reward metric.
[0148] Step 402: Normalize the first continuous reward to a preset reward value interval by using a predetermined linear transformation, and determine the first metric reward of the model translation text as the normalized first continuous reward.
[0149] Exemplarily, still taking the first metric reward metric as the BLEU metric as an example, the above-mentioned S metric expression can be used. Input the reference translation text and the model translation text corresponding to the training text required in this expression into the continuous reward function to determine the first continuous reward, that is, obtain the Metric calculation value. Assuming that the metric value of BLEU is 40, at this time, this metric value 40 can be scaled to the interval [0, 1], such as 40 / 100 = 0.4, and 0.4 is used as the first metric reward of the model translation text, that is, S metric = 0.4.
[0150] In this embodiment, for supervised training texts, continuous rewards based on BLEU can be used for translation quality evaluation to obtain the first metric reward.
[0151] For the second metric reward metric, using the first mapping processing method, the above step B "calculating the second metric reward of the model translation text by using a translation evaluation model based on the training text and the model translation text" may include steps 501 to 503. Wherein:
[0152] Step 501: Input the training text and the training translation text into the discrete evaluation model to determine the second discrete reward.
[0153] Among them, the discrete evaluation model is a discrete evaluation model determined based on the second metric reward metric.
[0154] Step 502, if the second discrete reward is greater than or equal to the second discrete reward threshold, determine the second metric reward of the model's translated text as the fifth value.
[0155] Step 503, if the second discrete reward value is less than the second discrete reward threshold, determine the second metric reward of the model's translated text as the sixth value.
[0156] Exemplarily, taking the second metric reward metric as the COMETKiwi metric as an example, the above S metric expression can be used. Input the training text and the model's translated text required in this expression into the discrete evaluation model to determine the second discrete reward, that is, obtain the Metric calculation value. Then, when the second discrete reward is greater than or equal to the second discrete reward threshold (such as 0.7), it can be considered that the translation quality is good. At this time, the second metric reward of the model's translated text can be the fifth value, such as 1, that is, when COMETKiwi ≥ 0.7, S metric = 1; when the second discrete reward is less than the second discrete reward threshold (such as 0.7), it can be considered that the translation quality is poor. At this time, the second metric reward of the model's translated text can be the sixth value, such as 0, that is, when COMETKiwi < 0.7, S metric = 0.
[0157] In this embodiment, for unsupervised training text, discrete rewards based on the COMETKiwi threshold can be used to evaluate the translation quality, so as to obtain the second metric reward.
[0158] For the second metric reward metric, adopting the second mapping processing method, the above step B "Based on the training text and the model's translated text, use the translation evaluation model to calculate the second metric reward of the model's translated text" can include steps 601 to 602. Among them:
[0159] Step 601, input the training text and the model's translated text into the continuous evaluation model to determine the second continuous reward.
[0160] Among them, the continuous evaluation model is a continuous evaluation model determined based on the second metric reward metric.
[0161] Step 602, determine the second continuous reward as the second metric reward of the model's translated text.
[0162] Exemplarily, still taking the second metric reward metric as the COMETKiwi metric as an example, the above S metric For the expression, input the required training text and the model translation text in the expression into the continuous evaluation model to determine the second continuous reward, that is, obtain the Metric calculation value. Assume that the metric value of COMETKiwi is 0.7. Since 0.7 is within the interval [0, 1], therefore, 0.7 can be directly used as the second metric reward for the model translation text, that is, S metric = 0.7. Of course, after scaling 0.7 to the interval [0, 1], it is still 0.7.
[0163] In this embodiment, for the unsupervised training text, the continuous reward based on COMETKiwi can be used for translation quality evaluation, so as to obtain the second metric reward.
[0164] The above provides multiple forms of quality evaluation based on metric reward indicators. During the actual training process, for the supervised training text and the unsupervised training text, any corresponding form of quality evaluation can be arbitrarily selected, and the embodiments of the present application do not make specific limitations in this regard.
[0165] The training method of the translation model proposed in the embodiments of the present application applies the RL framework oriented by inference to the translation field, explores its potential to improve the translation quality of the LLM, and solves the drawback that the existing MT inference ability needs to rely on complex technologies. It can not only reduce the training complexity of the LLM translation model, but also train the translation model with a simpler model framework, improve the training efficiency of the translation model, and improve the translation quality of the translation model. In addition, in order to improve the applicability of reinforcement learning (RL) in the MT field, an effective reward mechanism suitable for MT tasks is also proposed to solve the problem that it is difficult to design a simple and effective reward signal for MT tasks because there is no clear translation as a reference (different from verifiable fields such as mathematics with clear answers), and the translation quality evaluation is more subjective and detailed. Through the effective reward mechanism proposed in the present application, that is, the rule-metric hybrid reward mechanism, a reward function that can both ensure the output format and effectively improve the translation quality can be obtained, so as to be able to guide the LLM to improve the translation quality through the emerging inference process without structured CoT data or complex search algorithms.
[0166] Refer to Figure 7 As shown, it shows a schematic diagram of the effect of the training method of the present application. Exemplarily, the training method of the present application can be named MT-R1-Zero. Using model parameters of 3B and 7B to train MT-R1-Zero, it can be seen that as the training progresses, the model recovery length and the translation performance index increase synchronously. Among them, Figure 8 In the left figure, the solid line represents the translation effect of Chinese (ZH)-English (EN), and the dashed line represents the translation effect of English (EN)-Chinese (ZH).
[0167] The training method of this application does not rely on complex CoT data or Monte Carlo tree search. It can significantly improve the translation quality of the LLM only through RL optimization and stimulate the model's "spontaneous thinking" ability, opening up a new way for the RL application of open-domain generation tasks. In addition, the reward function combined with format rule checking and advanced MT metrics (such as COMETKiwi) solves the problem of defining effective RL rewards in MT tasks. The format reward ensures the usability of the model output, and the continuous metric reward provides a fine-grained quality feedback signal for the model, effectively guiding the model to optimize the translation performance.
[0168] Reference Figure 8 As shown, it shows a schematic diagram of the effect comparison between the training method of this application and other existing methods. It can be seen that the translation model trained by the training method of this application (especially the 7B version) has achieved the current best (State-of-the-Art) performance in benchmark tests such as WMT24 EN-ZH. It significantly outperforms open-source models of the same scale or even larger scale (70B+), and is superior to closed-source models including GPT-4 and Claude3.5 Sonnet, demonstrating the great potential of this method in practical applications.
[0169] In addition, the training method of this application also has strong zero-shot generalization ability (OOD Generalization). OOD (Out-of-Distribution) refers to the situation where the test data has a different distribution from the training data. During training, a model trained only on a specific language pair (such as Chinese-English) also shows strong zero-shot translation ability on unseen OOD language pairs (such as English-Japanese, German-English, German-Chinese). This shows that the translation ability and potential reasoning patterns learned through this RL method have good transferability, improving the robustness and generality of the model and reducing the dependence on multilingual parallel data. That is to say, by using the training method of this application, good training effects can be achieved without manually labeled data, thus improving the quality of model translation.
[0170] At the same time, during the training process, it was also observed that the translation model trained by the training method of this application <think>Diverse and evolving thinking patterns are generated in the tags, sometimes even thinking in the target language. It can be understood that <think>The thinking / reasoning process in the label will gradually evolve from the source language to the target language to be translated. Even if the training process does not involve the training of other languages, with this training mechanism, during the actual use of the model, it can also perform translation thinking and reasoning based on the target language, thereby achieving high-quality translation of other languages, enhancing the generalization ability of the model, and thus providing valuable qualitative insights into understanding how the LLM solves complex translation tasks through RL learning, demonstrating the plasticity of the "thinking" process inside the model. As shown in Table 1, it shows the translation ability of the training method of the present application between languages not involved in the training text.
[0171] Table 1
[0172]
[0173]
[0174] In an exemplary embodiment, as Figure 9 shown, a text translation method is provided. Taking the computer device in Figure 1 as an example for illustration, it includes the following steps 901 to step 902.
[0175] Wherein:
[0176] Step 901, receive a translation request, where the translation request indicates translating the source text in the source language into the target language.
[0177] Exemplarily, the translation request may include the source text in the source language, the source language, and the target language. As an optional implementation, the user can select the source language and the target language and input the source text in the source language to generate a translation request; as another optional implementation, the user can also only input the source text in the source language, and the computer device determines the source language based on the source text in the source language and determines the target language according to the scenario type to generate a translation request. For example, in certain specific scenarios, when the user inputs Chinese, it is default translated into English, and when the user inputs English, it is default translated into Chinese, etc.
[0178] Step 902, in response to the translation request, input the source text in the source language into the target translation model to obtain the translation text in the target language.
[0179] Wherein, the target translation model is trained by using the training method of the translation model in any of the above embodiments.
[0180] Exemplarily, after receiving a translation request, the computer device can generate a translation instruction recognizable by the computer device according to the translation request, and based on the translation instruction, call a target translation model to translate the source text in the source language, thereby obtaining a translation text in the target language; wherein, the translation instruction may include the source text in the source language, the target language, and a predetermined format output by the model, and the predetermined format defines the output format of the target translation model, for example: using <think>Label the translation thinking / reasoning process using <translate>Tags surround the final translation result, etc., so that when the target translation model outputs the translation text in the target language, it can also output the reasoning process of the translation. Exemplarily, the predetermined format can also define the output order of different output contents, for example: the thinking / reasoning process first, and the final translation result second, etc.
[0181] In the text translation method proposed in this embodiment, the target translation model trained by using the training method of the translation model in any of the above embodiments can execute the translation task in the actual translation scenario and can improve the translation quality of the translation task.
[0182] In an exemplary embodiment, as Figure 10 shown, a training method of a generation model is provided. Taking the computer device in Figure 1 as an example, it includes the following steps 1001 to 1005. Wherein:
[0183] Step 1001, input the training data into the initial generation model to obtain the model output data; the model output data includes the generated data corresponding to the training data; the initial generation model is used for open-ended generation tasks.
[0184] Step 1002, obtain the format reward corresponding to the model output data based on the format of the model output data, and the format reward is used to characterize whether the model output data conforms to the predetermined format.
[0185] Step 1003, obtain the metric reward corresponding to the generated data based on the training data and the generated data, and the metric reward is used to characterize the generation quality of the generated data.
[0186] Step 1004, obtain the target reward according to the format reward and the metric reward.
[0187] Step 1005, train the initial generation model based on the target reward and the preset reinforcement learning algorithm to obtain the target generation model.
[0188] That is to say, the above training method of the translation model can also be applied to other open-ended generation tasks, including but not limited to abstract summarization tasks, content creation tasks, intelligent reply tasks, emotional chat tasks, etc. For these content generation fields without standard answers or standard outputs as training reference data, the training idea involved in the training method of the translation model proposed in this application can also be used to train the generation model of the open-ended generation task, so as to improve the content generation quality of the generation model and thus meet the accurate responses in different scenarios and different requirements.
[0189] The training method of the generation model proposed in this embodiment adopts the training idea of the above-mentioned training method of the translation model to train a generation model with higher quality and greater accuracy, so as to solve the problem that the content generation for existing open-ended generation tasks is inaccurate, improve the content generation quality, and enhance the user experience.
[0190] It should be understood that although each step in the flowcharts involved in the above-described embodiments is shown in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps does not have a strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same moment, but can be executed at different moments. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0191] Based on the same inventive concept, the embodiments of the present application also provide a training device for a translation model for implementing the training method of the translation model involved above, a text generation device for implementing the text generation method involved above, and a training device for a generation model for implementing the training method of the generation model involved above. The implementation solutions for solving problems provided by this device are similar to the implementation solutions described in the above methods. Therefore, the specific limitations in the embodiments of one or more training devices for translation models, one or more text generation devices, and one or more training devices for generation models provided below can refer to the limitations on the training method of the translation model, the text generation method, and the training method of the generation model in the above text, and will not be repeated here.
[0192] In an exemplary embodiment, as Figure 11 shown, a training device for a translation model is provided, including: a processing module 1101, a format reward module 1102, a metric reward module 1103, a reward fusion module 1104, and a training module 1105, where:
[0193] The processing module 1101 is configured to input the training text in the first language into the initial translation model to obtain a model output text, and the model output text includes a model translation text in the second language.
[0194] The format reward module 1102 is configured to obtain a format reward corresponding to the model output text based on the format of the model output text, and the format reward is used to characterize whether the model output text conforms to a predetermined format.
[0195] A metric reward module 1103 is configured to obtain a metric reward corresponding to the model translation text based on the training text and the model translation text, and the metric reward is used to characterize the translation quality of the model translation text.
[0196] A reward fusion module 1104 is configured to obtain a target reward according to the format reward and the metric reward.
[0197] A training module 1005 is configured to train an initial translation model based on the target reward and a preset reinforcement learning algorithm to obtain a target translation model.
[0198] In one embodiment, the predetermined format is determined based on a translation instruction, and the format reward includes a first value or a second value; the first value is used to characterize that the model output text conforms to the predetermined format, and the second value is used to characterize that the model output text does not conform to the predetermined format.
[0199] In one embodiment, the metric reward includes a first metric reward and / or a second metric reward, and the metric reward module 1103 includes: a first metric reward unit and / or a second metric reward unit, where:
[0200] The first metric reward unit is configured to calculate a first metric reward of the model translation text based on a reference translation text corresponding to the training text and the model translation text.
[0201] The second metric reward unit is configured to calculate a second metric reward of the model translation text based on the training text and the model translation text by using a translation evaluation model.
[0202] In one embodiment, the apparatus further includes:
[0203] An alignment module is configured to perform alignment processing on the first metric reward and / or the second metric reward.
[0204] The metric reward module 1103 is specifically configured to fuse the aligned first metric reward and / or the second metric reward to obtain the metric reward.
[0205] In one embodiment, the metric reward is obtained based on mapping processing and is within a predetermined interval, and the predetermined interval is determined based on the value of the format reward.
[0206] In one embodiment, the mapping processing includes: based on the comparison result between the metric reward and a predetermined discrete reward threshold, mapping the metric reward to a third value or a fourth value within the predetermined interval; or, mapping the metric reward to the predetermined interval by performing a predetermined linear transformation on the metric reward.
[0207] In one embodiment, the reward fusion module 1104 is configured to perform segmented fusion on the format reward and the metric reward to obtain a target reward; or perform weighted fusion on the format reward and the metric reward to obtain a target reward.
[0208] In one embodiment, the training module 1005 includes:
[0209] a construction unit configured to construct an objective function of the reinforcement learning algorithm based on the target reward;
[0210] an optimization unit configured to optimize the parameters of the initial translation model by maximizing the objective function to obtain a target translation model.
[0211] In an exemplary embodiment, as Figure 12 shown, a text translation device is provided, including: a receiving module 1201 and a translation module 1202, where:
[0212] The receiving module 1201 is configured to receive a translation request, where the translation request indicates translating a source text in a source language into a target language.
[0213] The translation module 1202 is configured to, in response to the translation request, input the source text in the source language into the target translation model to obtain a translated text in the target language, where the target translation model is trained by using the training device of the translation model in any of the above embodiments.
[0214] In an exemplary embodiment, as Figure 13 shown, a training device for a generation model is provided, including: a processing module 1301, a format reward module 1302, a metric reward module 1303, a reward fusion module 1304, and a training module 1305, where:
[0215] The processing module 1301 is configured to input training data into an initial generation model to obtain model output data; the model output data includes generated data corresponding to the training data; the initial generation model is used for an open-ended generation task.
[0216] The format reward module 1302 is configured to obtain a format reward corresponding to the model output data based on the format of the model output data, where the format reward is used to characterize whether the model output data conforms to a predetermined format.
[0217] The metric reward module 1303 is configured to obtain a metric reward corresponding to the generated data based on the training data and the generated data, where the metric reward is used to characterize the generation quality of the generated data.
[0218] The reward fusion module 1304 is configured to obtain a target reward according to the format reward and the metric reward;
[0219] A training module 1305 is configured to train an initial generation model based on a target reward and a preset reinforcement learning algorithm to obtain a target generation model.
[0220] Each module in the above-mentioned training device of the translation model, the text translation device, and the training device of the generation model can be implemented in whole or in part by software, hardware, and their combination. Each of the above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each of the above modules.
[0221] In an exemplary embodiment, a computer device is provided. The computer device can be a terminal or a server. Here, taking the computer device as a server as an example, its internal structure diagram can be as Figure 14 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data related to model training and data related to model application. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it realizes at least one of a method for training a translation model, a method for text translation, and a method for training a generation model.
[0222] Those skilled in the art can understand that Figure 11 the structure shown in
[0223] is only a block diagram of a part of the structure related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0224] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the training method of the translation model, the text translation method, and the training method of the generation model in any of the above embodiments are implemented.
[0225] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the steps of the training method of the translation model, the text translation method, and the training method of the generation model in any of the above embodiments are implemented.
[0226] It should be noted that the data involved in this application (including but not limited to data for analysis, stored data, displayed data, etc.) are all information and data that have been fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0227] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.
[0228] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0229] The above-described embodiments only represent several implementation manners of the present application. Their descriptions are relatively specific and detailed, but they should not be construed as limiting the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.< / translate> < / think> < / think> < / think> < / think> < / translate> < / translate> < / think> < / translate> < / think>
Claims
1. A training method for a translation model, characterized in that, The method includes: Inputting the training text in the first language into the initial translation model to obtain a model output text, where the model output text includes a model translation text in the second language; Obtaining a format reward corresponding to the model output text based on the format of the model output text, where the format reward is used to represent whether the model output text conforms to a predetermined format; Obtaining a metric reward corresponding to the model translation text based on the training text and the model translation text, where the metric reward is used to represent the translation quality of the model translation text; Obtaining a target reward according to the format reward and the metric reward; Training the initial translation model based on the target reward and a preset reinforcement learning algorithm to obtain a target translation model.
2. The method according to claim 1, wherein The predetermined format is determined based on a translation instruction, and the format reward includes a first value or a second value; the first value is used to represent that the model output text conforms to the predetermined format, and the second value is used to represent that the model output text does not conform to the predetermined format.
3. The method according to claim 1, wherein The metric reward includes a first metric reward and / or a second metric reward. Obtaining the metric reward corresponding to the model translation text based on the training text and the model translation text includes: Calculating a first metric reward of the model translation text based on a reference translation text corresponding to the training text and the model translation text; and / or, Calculating a second metric reward of the model translation text based on the training text and the model translation text by using a translation evaluation model.
4. The method according to claim 3, characterized in that The method further includes: Performing an alignment process on the first metric reward and / or the second metric reward; Fusing the aligned first metric reward and / or second metric reward to obtain the metric reward.
5. The method according to claim 1, characterized in that The metric reward is obtained based on mapping processing and is within a predetermined interval, and the predetermined interval is determined based on the value of the format reward.
6. The method according to claim 5, wherein The mapping processing includes: Based on the comparison result between the metric reward and a predetermined discrete reward threshold, mapping the metric reward to a third value or a fourth value within the predetermined interval; or, mapping the metric reward to the predetermined interval by performing a predetermined linear transformation on the metric reward.
7. The method according to claim 1, wherein Obtaining the target reward according to the format reward and the metric reward includes: Performing segmented fusion on the format reward and the metric reward to obtain the target reward; or, Performing weighted fusion on the format reward and the metric reward to obtain the target reward.
8. The method according to claim 1, characterized in that, Training the initial translation model based on the target reward and a preset reinforcement learning algorithm to obtain a target translation model includes: Constructing an objective function of the reinforcement learning algorithm based on the target reward; Optimizing the parameters of the initial translation model by maximizing the objective function to obtain a target translation model.
9. A text translation method, characterized in that, The method includes: Receiving a translation request, where the translation request instructs to translate a source text in a source language into a target language; In response to the translation request, input the source text in the source language into the target translation model to obtain the translation text in the target language, where the target translation model is trained by using the training method of the translation model described in any one of claims 1-8.
10. A training method for a generative model, characterized in that, The method includes: Input the training data into the initial generation model to obtain model output data; the model output data includes generated data corresponding to the training data; the initial generation model is used for open-ended generation tasks; Obtain a format reward corresponding to the model output data based on the format of the model output data, where the format reward is used to indicate whether the model output data conforms to a predetermined format; Obtain a metric reward corresponding to the generated data based on the training data and the generated data, where the metric reward is used to indicate the generation quality of the generated data; Obtain a target reward according to the format reward and the metric reward; Train the initial generation model based on the target reward and a preset reinforcement learning algorithm to obtain a target generation model.
11. A training device for a translation model, characterized in that, The device includes: A processing module, configured to input the training text in the first language into the initial translation model to obtain a model output text, where the model output text includes a model translation text in the second language; A format reward module, configured to obtain a format reward corresponding to the model output text based on the format of the model output text, where the format reward is used to indicate whether the model output text conforms to a predetermined format; A metric reward module, configured to obtain a metric reward corresponding to the model translation text based on the training text and the model translation text, where the metric reward is used to indicate the translation quality of the model translation text; A reward fusion module, configured to obtain a target reward according to the format reward and the metric reward; A training module, configured to train the initial translation model based on the target reward and a preset reinforcement learning algorithm to obtain a target translation model.
12. A text translation device, characterized in that, The device includes: A receiving module, configured to receive a translation request, where the translation request instructs to translate the source text in the source language into the target language; A translation module, configured to, in response to the translation request, input the source text in the source language into the target translation model to obtain the translation text in the target language, where the target translation model is trained by using the training device of the translation model described in claim 11.
13. A training device for a generative model, characterized in that, The device includes: A processing module, configured to input the training data into the initial generation model to obtain model output data; the model output data includes generated data corresponding to the training data; the initial generation model is used for open-ended generation tasks; A format reward module, configured to obtain a format reward corresponding to the model output data based on the format of the model output data, where the format reward is used to indicate whether the model output data conforms to a predetermined format; A metric reward module, configured to obtain a metric reward corresponding to the generated data based on the training data and the generated data, where the metric reward is used to indicate the generation quality of the generated data; A reward fusion module, configured to obtain a target reward according to the format reward and the metric reward; A training module, configured to train the initial generation model based on the target reward and a preset reinforcement learning algorithm to obtain a target generation model.
14. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 10 are implemented.
15. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.
16. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.
Citation Information
Cited By
Model optimization method and device based on reinforcement learning and electronic equipment
CN120893600A
Model optimization method and device based on reinforcement learning and electronic equipment
CN120893600B