Model training method and device, medium and product
By employing a reinforcement learning strategy with partial parameter training and a reward-adaptive gradient propagation mechanism, the poor performance and high training cost of large language models in precise reasoning tasks are addressed, achieving efficient and accurate model training and inference, and improving the model's performance in complex tasks.
Patent Information
- Application Number
- CN202511693195.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-18
- Publication Date
- 2026-02-10
AI Technical Summary
Existing autoregressive large language models based on the Transformer architecture perform poorly on tasks requiring precise reasoning, and the training cost of large model reinforcement learning is high, while the efficiency of parameter updates and transmission is low.
A reinforcement learning strategy with partial parameter training is adopted. The Transformer block is divided into fixed parameters and parameters to be updated by pre-setting hierarchical coefficients. The parameters of the output layer and the module to be updated are obtained by using the reward value. The updated parameters are then transmitted to the inference cluster. The training process is optimized by using the on-policy DPO algorithm to avoid gradient decay in traditional Critic estimation.
It significantly reduces training and communication costs, improves the accuracy and efficiency of the model in complex reasoning tasks, enhances the robustness of the model, enables it to quickly find the correct answer, and reduces training time and resource consumption.
Smart Images

Figure CN121503685A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of model training, in particular to a model training method and device, medium and product. BACKGROUND
[0002] With the rapid development of artificial intelligence technology, large language models have made significant progress in natural language processing. Autoregressive large language models based on the Transformer structure have been widely used in various scenarios, but still face challenges in tasks that require accurate reasoning, such as mathematics, logical reasoning, and code generation.
[0003] The concept of Chain-of-Thought (CoT) proposed by OpenAI has been proven by numerous experiments to greatly improve the performance of models on tasks that require reasoning. However, due to the multi-peak limitation of its own probability model and the excessive freedom of token-by-token prediction, LLMs (Large Language Models) still perform poorly on very difficult tasks such as graduate mathematics competitions and international programming competitions.
[0004] In addition, existing large model reinforcement learning methods have the problem of inaccurate Critic estimation, and the training cost is also very high. For example, DeepSeek R1 has a total parameter size of 670B, and the cost of full-scale training is very high (training a 16k-long data requires 7200 H800 GPUs, which costs approximately 80000 yuan RMB at market price); in addition, reinforcement learning requires the model to generate trajectories for specific problems, and for efficiency considerations, the model needs to be deployed on inference resources, which poses many problems for data transmission of super-large models.
[0005] In summary, the existing technical problems have the following main shortcomings: autoregressive large language models based on the Transformer structure perform poorly on tasks that require accurate reasoning; large model reinforcement learning has high training costs, and full-scale parameter updates and transmission are inefficient. Therefore, there is an urgent need for a model training method that can effectively improve the reasoning ability of the model, reduce the training cost, and optimize the parameter update strategy. SUMMARY
[0006] To solve the technical problems of existing autoregressive large language models based on the Transformer structure performing poorly on tasks that require reasoning such as mathematics, logic, and code, and the bias-variance trade-off and high training cost problems in reinforcement learning methods, a model training method, device, medium, and product are provided to improve reasoning accuracy, training efficiency, model robustness, and reduce training communication cost.
[0007] The application provides a model training method, comprising:
[0008] In the training cluster, input a to-be-processed task into a to-be-trained model, and determine a reward value corresponding to the to-be-processed task;
[0009] Based on a preset hierarchical coefficient, divide a plurality of Transformer blocks in an intermediate layer of the to-be-trained model into a first module with fixed parameters and a second module with to-be-updated parameters;
[0010] Based on the reward value, obtain a gradient, and update parameters of the second module and an output layer of the to-be-trained model by using the gradient;
[0011] Send the updated parameters of the second module and the output layer from the training cluster to an inference cluster, so as to update corresponding parameters of a model deployed in the inference cluster.
[0012] Optionally, the step of inputting a to-be-processed task into a to-be-trained model in the training cluster and determining a reward value corresponding to the to-be-processed task comprises:
[0013] In the training cluster, input a to-be-processed task into a to-be-trained model, so as to obtain an inference result output by the to-be-trained model;
[0014] Based on the inference result and a standard result corresponding to the to-be-processed task, determine the reward value corresponding to the to-be-processed task.
[0015] Optionally, the step of determining the reward value corresponding to the to-be-processed task based on the inference result and a standard result corresponding to the to-be-processed task comprises:
[0016] Match the inference result with the standard result corresponding to the to-be-processed task, and determine the reward value corresponding to the to-be-processed task according to a matching result.
[0017] Optionally, the step of dividing, based on a preset hierarchical coefficient, a plurality of Transformer blocks in an intermediate layer of the to-be-trained model into a first module with fixed parameters and a second module with to-be-updated parameters comprises:
[0018] According to a preset hierarchical coefficient a, divide N Transformer blocks connected in sequence in the intermediate layer of the to-be-trained model into a first module composed of the first K blocks and a second module composed of the last N-K blocks, the parameters of the first module are fixed, and the parameters of the second module are to be updated;
[0019] Wherein, K is an integer value of N x a, 0 < a < 1, and N is a positive integer.
[0020] Optionally, when the to-be-trained model comprises an input layer, the intermediate layer and the output layer in sequence; the obtaining the gradient based on the reward value, and updating parameters of the second module and the output layer of the to-be-trained model by using the gradient, comprises:
[0021] constructing a loss function according to the reward value;
[0022] calculating the gradient of the second module and the output layer based on the loss function;
[0023] updating parameters of the second module and the output layer by using the gradient.
[0024] Optionally, the calculating the loss function according to the reward value comprises:
[0025] calculating the loss function based on the reward value by using an On-policy DPO algorithm.
[0026] Optionally, the method further comprises:
[0027] monitoring an average reward value of the to-be-trained model on a validation set during the training process;
[0028] stopping the training when the average reward value does not increase in a continuous preset number of training periods.
[0029] The application further provides an electronic device, comprising: one or more processors; and a memory storing computer program instructions, which, when executed, cause the processor to perform the steps of the model training method.
[0030] The application further provides a computer readable medium storing computer program instructions, which can be executed by a processor to implement the model training method.
[0031] The application further provides a computer program product comprising computer program / instructions, which, when executed by a processor, implement the steps of the model training method.
[0032] The beneficial effects of the present application: a model training method inputs a to-be-processed task into a to-be-trained model in a training cluster to determine a reward value corresponding to the to-be-processed task; based on a preset hierarchical coefficient, a plurality of Transformer blocks in an intermediate layer of the to-be-trained model are divided into a first module with fixed parameters and a second module with to-be-updated parameters; based on the reward value, a gradient is obtained, and the parameters of the second module and the parameters of an output layer of the to-be-trained model are updated using the gradient; the updated parameters of the second module and the parameters of the output layer are sent from the training cluster to an inference cluster to update corresponding parameters of a model deployed in the inference cluster. The present application uses a reinforcement learning training strategy of partial parameter training, fixes the parameters in front of the model, and only trains and updates the parameters in the rear part of the model, thereby greatly reducing the training and communication costs. An adaptive gradient conduction mechanism of the reward is adopted, so that the reward signal can be conducted to the earlier word units in the inference process without loss, and the gradient attenuation of the earlier word units in the traditional estimation method is avoided. BRIEF DESCRIPTION OF DRAWINGS
[0033] One or more embodiments are illustrated by way of example in the drawings in which like reference numerals indicate like elements, and in which:
[0034] Figure 1 A method flowchart of one embodiment of the model training method described in the present application;
[0035] Figure 2 A method flowchart of another embodiment of the model training method described in the present application;
[0036] Figure 3 An exemplary structural diagram of an electronic device according to the present application. DETAILED DESCRIPTION
[0037] The advantages of the present application are further described below in conjunction with the drawings and specific embodiments.
[0038] The exemplary embodiments will be described in detail herein below with reference to the drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments are not meant to represent all implementations consistent with the present disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0039] The terminology used in the present disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of the present disclosure. As used in the present disclosure and the appended claims, the singular forms "a," "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the term "and / or," as used herein, refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0040] It is to be understood that, although the terms first, second, third, etc. can be used herein to describe various information, these terms are not intended to denote a temporal or chronological order. These terms are used only to distinguish one type of information from another. For example, a first information can be termed a second information, and similarly, a second information can also be termed a first information, without departing from the scope of the present disclosure. Depending on the context, the word "if' as used herein can be interpreted to mean "when" or "in response to determining" or "in response to a determination".
[0041] In the description of the present application, it should be understood that the numerical reference number before the step does not identify the front and rear order of executing the step, but is only used for the convenience of describing the present application and distinguishing each step, and therefore cannot be understood as a limitation of the present application.
[0042] The following terms are used herein:
[0043] The Transformer structure is stacked by multiple Transformer blocks.
[0044] OpenAI is an American artificial intelligence research laboratory founded in 2015. Its original mission is "to advance digital intelligence in the way most likely to benefit all of humanity, without being constrained by a need to generate financial returns", and the core goal of OpenAI is to ensure that general artificial intelligence (AGI) can develop safely and benefit all of humanity.
[0045] Chain-of-Thought (CoT) is a prompting technique used to improve the performance of large language models (LLMs) on complex reasoning tasks such as mathematical problems, common sense reasoning, and symbolic reasoning.
[0046] Critic estimation is a core concept in reinforcement learning (RL), particularly related to the PPO algorithm.
[0047] Scaling Law is an empirical rule that reveals a predictable power-law relationship between the performance of large-scale AI models (e.g., accuracy, fluency, and accuracy of answers in various tests) and three key input resources: model size, data volume, and computational resources. Scaling Law tells us that as we proportionally increase the model's parameters (making it larger), the amount of training data (making it learn more), and the computational resources used for training (making it calculate longer), the model's performance will steadily improve in a predictable manner without encountering performance bottlenecks or saturation.
[0048] RL stage, in full Reinforcement Learning stage based on human feedback, is a crucial step in the training process of large language models. It follows pre-training and supervised fine-tuning, and its core goal is to align the model's behavior with complex and ambiguous human values and preferences.
[0049] PPO (Proximal Policy Optimization) is a policy gradient algorithm in the field of reinforcement learning. Its core design goal is to improve policy performance while ensuring that each update does not deviate too far from the old policy, thus maintaining the stability of training.
[0050] O1 is the internal code name for a series of hybrid training techniques and infrastructure used by OpenAI in developing ChatGPT and GPT-4, and is not a publicly defined algorithm. It is a system engineering miracle rather than a single algorithm. It represents OpenAI's overall approach to combining supervised fine-tuning and reinforcement learning human feedback in a large-scale, efficient, and stable manner.
[0051] GAE (Generalized Advantage Estimation) is a technique used in policy gradient algorithms such as PPO to more efficiently and accurately estimate the advantage function. GAE provides a method to balance bias and variance in the calculation of advantage.
[0052] On-policy DPO algorithm is an online, self-iterative, and continuous optimization training method, which is different from the classic DPO application using a fixed data set. It represents an algorithm for training data generated by the current latest policy model itself interacting with the environment to optimize the strategy.
[0053] DPO algorithm is an algorithm that directly optimizes the strategy from preference data, completely avoiding the complex steps of traditional training reward models and using PPO.
[0054] A computation graph is a directed graph used to describe mathematical computations. In deep learning, it is used to represent the forward propagation (computing the prediction result) and back propagation (computing the gradient to update the parameters) processes of a neural network.
[0055] Referring to Figure 1 As shown in the drawings, the model training method provided in the present application comprises the following steps:
[0056] S1. In the training cluster, input the to-be-processed task into the to-be-trained model, and determine the reward value corresponding to the to-be-processed task;
[0057] It should be noted that the to-be-trained model in this embodiment is a large language model.
[0058] Further, step S1 can comprise the following steps:
[0059] S11. In the training cluster, input the to-be-processed task into the to-be-trained model to obtain an inference result output by the to-be-trained model;
[0060] S12. Based on the inference result and the standard result corresponding to the to-be-processed task, determine the reward value corresponding to the to-be-processed task.
[0061] Specifically, step S12 can comprise: matching the inference result with the standard result corresponding to the to-be-processed task, and determining the reward value corresponding to the to-be-processed task according to the matching result.
[0062] In this embodiment, when the inference result completely matches the standard result GT (Ground Truth), a reward value of 1 is assigned to the correct sample track of the to-be-processed task; when the inference result partially matches the standard result GT or the inference result completely does not match the standard result, a reward value of -1 is assigned to the error sample track of the to-be-processed task.
[0063] S2. Based on a preset hierarchical coefficient, divide a plurality of Transformer blocks in an intermediate layer of the to-be-trained model into: a first module with fixed parameters and a second module with to-be-updated parameters;
[0064] Further, step S2 can comprise: according to a preset hierarchical coefficient a, divide N Transformer blocks connected in sequence in the intermediate layer of the to-be-trained model into a first module composed of the first K blocks and a second module composed of the last N-K blocks, the parameters of the first module being fixed, and the parameters of the second module being to-be-updated;
[0065] Wherein, K is the integer value of N x a, 0 < a < 1, and N is a positive integer.
[0066] In this step, the parameters of the front of the model to be trained are fixed to provide protection for subsequent steps. Considering that the model to be trained includes at least an input layer, an intermediate layer, and an output layer. The intermediate layer is composed of N stacked Transformer blocks. Therefore, in this embodiment, the parameters of all layers before the intermediate layer (such as the input layer) and all parameters of the first K Transformer blocks in the intermediate layer are fixed.
[0067] By way of example but not limitation, the preset hierarchical coefficient a can be set to 0.3, at this time if the intermediate layer has a total of 12 Transformer blocks, then K = 12 x 0.3 = 3.6, rounded to K = 4, that is, the first 4 Transformer blocks form the first module, the parameters of which are fixed; the last 8 Transformer blocks form the second module, the parameters of which need to be updated.
[0068] In another example, the preset hierarchical coefficient a can be set to 0.5, at this time if the intermediate layer has a total of 12 Transformer blocks, then K = 12 x 0.5 = 6, that is, the first 6 Transformer blocks form the first module, the parameters of which are fixed; the last 6 Transformer blocks form the second module, the parameters of which need to be updated.
[0069] The O1 technical solution proposed by OpenAI in the prior art processes the problem of token prediction error in the reasoning process through continuous trial and error, thereby greatly improving the performance of LLM on difficult reasoning tasks. The starting point of the O1 technical solution is to release the computing power of the reinforcement learning stage, and to explore the scaling law in the post-training and reasoning stage. Generally speaking, the RL stage of LLM will use the Proximal Policy Optimization (PPO) algorithm, and the common problems of this algorithm are two points: the inaccuracy of reward (Reward) estimation and Critic estimation. The essence of the O1 technical route is to deal with the problem of inaccurate reward estimation in the original Reinforcement Learning (RL) stage: the reward in the original RL process is learned from a large amount of human preference data by a reward learning, and its output cannot represent whether the model's answer to the problem is correct, which is fatal to problems related to mathematics / reasoning that have correct solutions. The O1 solution uses the correct answer GT of a large number of problems as the reward of the RL learning stage, thereby encouraging the model to improve the model's performance in the mathematics / reasoning field through self-stimulating learning of the correct answer in the RL stage.
[0070] Existing Critic estimation methods can alleviate the problem of inaccurate reward estimation and address how to handle this issue. One approach is the Generalized Advantage Estimation (GAE) algorithm, which combines advantage estimates from multiple time steps to provide more stable and accurate estimates. Its core idea is to perform a weighted average of advantage estimates from different time steps, with the weights controlled by an exponentially decaying factor. Specifically, GAE uses a time-series penalty coefficient to reduce the impact of noise in reward estimates at future lexical units, balancing the bias and excessive variance of reward estimates across different samples.
[0071] However, the advantage of GAE estimation has a significant bias-variance tradeoff. If the currently estimated term is located early in the sequence, the bias is high because the advantage estimation is based on fewer steps and therefore heavily relies on the accuracy of the value function. This limits the model's search space when training on more difficult task data. This is because the LLM pre-training process models the distribution of internet corpora, reducing the selectivity of terms at each position and converging it to the few terms with the highest probabilities.
[0072] Therefore, to efficiently search the search space that the pre-trained model has a low probability of possessing, it is necessary to release the restriction of the Critic-based algorithm on the search space of the preceding words, that is, the gradient of the ground truth will not decay due to the position of the words, so that the model can quickly search for the correct path.
[0073] In the following steps of this embodiment, to address the problems of the Critic-based method, the GT reward is used to generate an adaptive gradient that is losslessly propagated to the preceding word, thereby significantly changing the model's output and enabling more efficient and accurate searching for the correct GT answer.
[0074] S3. Obtain the gradient based on the reward value, and use the gradient to update the parameters of the second module and the parameters of the output layer of the model to be trained;
[0075] Furthermore, when the model to be trained sequentially includes an input layer, the intermediate layer, and the output layer; step S3 may include: constructing a loss function based on the reward value; calculating the gradients of the second module and the output layer based on the loss function; and updating the parameters of the second module and the output layer using the gradients.
[0076] Specifically, the step of calculating the loss function based on the reward value may include: using the On-policy DPO algorithm to calculate the loss function based on the reward value.
[0077] In this embodiment, the on-policy DPO algorithm is used to directly optimize and update the parameters of the later part of the model to be trained by directly utilizing reward value pairs. Correct sample trajectory. The reward value was assigned to 1, and the error sample trajectory was... The reward value will be assigned as -1. Trajectory pairs with different reward values will be used to calculate the DPO loss, where the log probability (logprob) of the correct trajectory with a reward value of 1 is: for The minuend in the logprob of the erroneous sample trajectory, with a reward value of -1: for The minuend in:
[0078]
[0079] in, This represents the expectation, where x represents the input instructions to the network, and y represents the expected value. w y represents the correct sample in the network's response. l Let D be the erroneous sample, and D be the data distribution. It is the sigmoid function, that is , It is the result of the correct sample being forwarded once in the policy model. It is the result of the erroneous sample being forwarded once in the policy model. The result of forwarding the correct sample once in the reference model. This is the result of forwarding the erroneous sample once in the reference model.
[0080] The above method can cancel the gradients of the fixed parameters in the computation graph, and only retain the gradients of the parameters of the layers that need to be updated.
[0081] In this embodiment, the on-policy DPO algorithm directly optimizes the policy, allowing the reward signal to directly influence the model's behavior without relying on the traditional value function (Critic). In this application, the on-policy DPO algorithm employs a ground truth (GT) reward-based policy optimization, adaptively adjusting the model's output to enable the model to find the correct answer more quickly during inference. Unlike the traditional PPO algorithm, DPO does not rely on a Critic network for estimation; instead, it propagates the reward signal through a direct optimization policy, avoiding the bias and variance problems associated with Critic estimation. The DPO algorithm allows the reward signal to directly influence the selection of each term during inference, thus more accurately guiding the model towards the correct answer during training.
[0082] Further, when the model to be trained sequentially includes an input layer, the intermediate layer, the decoding layer, and the output layer; step S3 may include: constructing a loss function based on the reward value; calculating the gradients of the second module, the decoding layer, and the output layer based on the loss function; and updating the parameters of the second module, the decoding layer, and the output layer using the gradients.
[0083] Specifically, the step of calculating the loss function based on the reward value may include: using the On-policy DPO algorithm to calculate the loss function based on the reward value.
[0084] In actual training, optimization algorithms such as stochastic gradient descent and Adam optimizer can be used to update parameters. The learning rate can be set to a value between 0.0001 and 0.01, and adjusted according to the specific task and model size.
[0085] S4. The updated parameters of the second module and the parameters of the output layer are sent from the training cluster to the inference cluster to update the corresponding parameters of the model deployed in the inference cluster.
[0086] In this step, an adaptive gradient propagation mechanism is used to losslessly pass the reward signal to earlier terms, thus generating more accurate gradients and ensuring that the model under training can quickly converge to the correct answer during inference. The updated parameters are sent from the training cluster to the inference cluster. Since the model under training only updates the parameters of non-fixed layers, only the updated parameters need to be transmitted when transferring the model to the inference side, reducing transmission costs.
[0087] In this embodiment, this can be achieved through network transmission, parameter file sharing, or other methods. After the update is complete, the model in the inference cluster will use the latest parameters for inference, thereby improving the model's inference performance.
[0088] This embodiment utilizes an adaptive gradient propagation mechanism to effectively propagate the ground truth (GT) reward to earlier terms in the inference process, avoiding the gradient decay of earlier terms in traditional Critic estimation methods. Specifically, gradient propagation is optimized as follows: during inference, for each generated answer, a reward value is calculated based on its matching degree with the GT answer, and this reward value is used as the model's optimization objective. This reward signal can be losslessly propagated to previously generated terms through gradient backpropagation, ensuring that the selection of each term maximizes its approximation to the GT answer. By adaptively adjusting the GT reward, the gradient propagation is ensured not to decay due to different positions, enabling the model to make optimal inferences about the correct answer at every time step. During training, the on-policy DPO algorithm is used to optimize the post-training strategy of LLM. Compared with the traditional PPO method, the training process in this application does not rely on the Critic network for reward estimation, but instead adjusts the model parameters by directly optimizing the objective function.
[0089] The above model training method enables efficient training of large models. By fixing the parameters of the first module and updating only the parameters of the second module and the output layer, the number of parameters that need to be updated can be significantly reduced, lowering the computational complexity and storage requirements of training. Simultaneously, the reward-based training approach allows the model to better adapt to the needs of specific tasks, improving the model's inference performance. Furthermore, separating training and inference into different clusters can improve the overall system efficiency and scalability.
[0090] In this embodiment, the model training method involves inputting the task to be processed into the model to be trained in the training cluster and determining the reward value corresponding to the task. Based on a preset hierarchical coefficient, multiple Transformer blocks in the intermediate layer of the model to be trained are divided into a first module with fixed parameters and a second module with parameters to be updated. The gradient is obtained based on the reward value, and the parameters of the second module and the output layer of the model to be trained are updated using the gradient. The updated parameters of the second module and the output layer are sent from the training cluster to the inference cluster to update the corresponding parameters of the model deployed in the inference cluster. This application, through a reinforcement learning training strategy of partial parameter training, fixes the parameters of the earlier parts of the model and only trains and updates the parameters of the later parts of the model, significantly reducing training and communication costs. It solves the gradient decay problem caused by the current GAE-based Critic estimation method during inference; by adopting an adaptive gradient propagation mechanism based on rewards, the reward signal can be losslessly propagated to earlier words in the inference process, avoiding the gradient decay of earlier words in the traditional Critic estimation method. Compared with existing technologies, this application improves inference accuracy by adaptively optimizing the model's inference process to quickly find the correct ground truth (GT) answer, avoiding common erroneous inference paths in traditional methods; it improves training efficiency by removing the dependence on the Critic network, making the training process more efficient, especially in high-dimensional and complex inference tasks, significantly reducing training time; and it enhances the model's robustness by directly optimizing the policy rather than relying on the value function, reducing the model's volatility during the inference process and enhancing its robustness in complex inference tasks.
[0091] In this embodiment, by training only a subset of the model's parameters and transmitting only a portion of those parameters, performance similar to full training can be achieved, significantly reducing the training, communication, and time costs. This application significantly improves the model's performance in complex reasoning tasks by introducing a criterion-free strategy and a ground truth reward adaptive gradient propagation mechanism. This method has broad application prospects in mathematical reasoning, programming tasks, and other fields, enabling LLMs to quickly and accurately provide correct answers when faced with difficult reasoning tasks.
[0092] In a preferred embodiment, see [reference] Figure 2 The model training method shown may also include the following steps:
[0093] S5. Monitor the average reward value of the model to be trained on the validation set during training;
[0094] S6. When the average reward value does not increase within a preset number of consecutive training cycles, training is stopped.
[0095] S7. When the average reward value increases within a preset number of consecutive training cycles, return to step S1.
[0096] In this embodiment, a preset number of training epochs can be set to 5. That is, when the model's average reward value on the validation set does not improve for 5 consecutive training epochs, the model is considered to have converged, and the training process is stopped. This early stopping strategy can avoid model overfitting and save computational resources.
[0097] The steps of the various methods described above are only for clarity. In practice, they can be combined into one step or some steps can be split into multiple steps. As long as they include the same logical relationship, they are all within the scope of protection of this patent. Adding insignificant modifications or introducing insignificant designs to the algorithm or process, but without changing the core design of the algorithm and process, are also within the scope of protection of this patent.
[0098] Furthermore, some embodiments of this application also provide an electronic device. The electronic device can be various forms of digital computer, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, etc. The electronic device can also be various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices.
[0099] The electronic device includes: one or more processors; and a memory storing computer program instructions that, when executed, cause the processor to perform the steps of the methods provided in any one or more of the above embodiments. Figure 3 An exemplary structural diagram of the electronic device is disclosed. For example... Figure 3 As shown, the electronic device includes one or more processors 1101, a memory 1102, and interfaces for connecting the components, including high-speed interfaces and low-speed interfaces. The components are interconnected via different buses and can be mounted on a common motherboard or otherwise as required. The processors can process instructions executed within the electronic device, including instructions stored in or on memory to display graphical information of a GUI on an external input / output device (such as a display device coupled to the interface). In some other embodiments, multiple processors and / or multiple buses can be used with multiple memories and multiple memory modules, if desired. Similarly, multiple electronic devices can be connected, each providing some of the necessary operations (e.g., as a server array, a group of blade servers, or a multiprocessor system). The components, their connections and relationships, and their functions shown herein are merely examples and are not intended to limit the implementation of the present application described and / or claimed herein.
[0100] The electronic device may further include an input device 1103 and an output device 1104. The processor 1101, memory 1102, input device 1103, and output device 1104 may be connected via a bus or other means. Figure 3 Taking the example of a connection between China and Israel via a bus.
[0101] Input device 1103 can receive input numerical or character information, and generate key signal inputs related to user settings and function control of the electronic device, such as a touch screen, keypad, mouse, trackpad, touchpad, joystick, one or more mouse buttons, trackball, joystick, etc. Output device 1104 may include a display device, auxiliary lighting device (e.g., LED), and haptic feedback device (e.g., vibration motor). The display device may include, but is not limited to, a liquid crystal display (LCD), a light-emitting diode (LED) display, and a plasma display. In some embodiments, the display device may be a touch screen.
[0102] To provide interaction with the user, the electronic device can be a computer. The computer has: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0103] In this embodiment, a computer-readable medium stores a computer program / instructions that, when executed by a processor, implement the steps of the methods provided in any one or more of the above embodiments. This computer-readable medium may be included in the electronic device described in the above embodiments; or it may exist independently and not assembled into that device. The aforementioned computer-readable medium carries one or more computer-readable instructions.
[0104] The memory 1102 can serve as a non-transitory computer-readable storage medium, used to store non-transitory software programs, non-transitory computer-executable programs, and modules. The processor 1101 executes various functional applications and data processing of the server by running the non-transitory software programs, instructions, and modules stored in the memory 1102, thereby implementing the program instructions / modules corresponding to the methods provided in any one or more of the embodiments described above in this application.
[0105] The memory 1102 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the electronic device. Furthermore, the memory 1102 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, the memory 1102 may optionally include memory remotely located relative to the processor 1101, and these remote memories can be connected to the electronic device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0106] It should be noted that the computer-readable medium described in this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0107] Computer-readable media include permanent and non-permanent, removable and non-removable media, which can store information by any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, read-only optical disc (CD-ROM), digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.
[0108] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0109] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. For example, it can be implemented using an application-specific integrated circuit (ASIC), a general-purpose computer, or any other similar hardware device. In some embodiments, the software program of this application can be executed by a processor to implement the steps or functions described above. Similarly, the software program of this application (including related data structures) can be stored in a computer-readable recording medium, such as RAM memory, magnetic or optical drives, floppy disks, or similar devices. Furthermore, some steps or functions of this application can be implemented in hardware, for example, as circuitry that works with a processor to perform the various steps or functions.
[0110] The computer program product provided in this application includes one or more computer programs / instructions. When executed by a processor, these computer programs / instructions generate, in whole or in part, the processes or functions described in this application. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).
[0111] The flowcharts or block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of devices, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-specific system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0112] The scope of this application is defined by the appended claims rather than the foregoing description, and is therefore intended to encompass all variations falling within the meaning and scope of equivalents of the claims. No reference numerals in the claims should be construed as limiting the scope of the claims. Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices recited in a device claim may also be implemented by a single unit or device in software or hardware. Terms such as "first," "second," etc., are used only for distinguishing descriptions and do not indicate any particular order, nor should they be construed as indicating or implying relative importance.
[0113] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
Claims
1. A model training method, characterized in that, include: In the training cluster, the task to be processed is input into the model to be trained, and the reward value corresponding to the task to be processed is determined; Based on a preset layering coefficient, the multiple Transformer blocks in the intermediate layer of the model to be trained are divided into: a first module with fixed parameters and a second module with parameters to be updated. The gradient is obtained based on the reward value, and the gradient is used to update the parameters of the second module and the parameters of the output layer of the model to be trained. The updated parameters of the second module and the parameters of the output layer are sent from the training cluster to the inference cluster to update the corresponding parameters of the model deployed in the inference cluster.
2. The model training method according to claim 1, characterized in that, In the training cluster, the task to be processed is input into the model to be trained, and the reward value corresponding to the task to be processed is determined, including: In the training cluster, the task to be processed is input into the model to be trained in order to obtain the inference result output by the model to be trained; Based on the reasoning result and the standard result corresponding to the task to be processed, the reward value corresponding to the task to be processed is determined.
3. The model training method according to claim 2, characterized in that, The step of determining the reward value corresponding to the task to be processed based on the reasoning result and the standard result corresponding to the task to be processed includes: The reasoning result is matched with the standard result corresponding to the task to be processed, and the reward value corresponding to the task to be processed is determined based on the matching result.
4. The model training method according to claim 1, characterized in that, Based on a preset hierarchical coefficient, the multiple Transformer blocks in the intermediate layer of the model to be trained are divided into: a first module with fixed parameters and a second module with parameters to be updated, including: According to the preset layering coefficient α, the N Transformer blocks connected in sequence in the intermediate layer of the model to be trained are divided into a first module consisting of the first K blocks and a second module consisting of the last NK blocks. The parameters of the first module are fixed, and the parameters of the second module are to be updated. Where K = N × α takes integer values, 0 < α < 1, and N is a positive integer.
5. The model training method according to claim 1, characterized in that, When the model to be trained sequentially includes an input layer, the intermediate layer, and the output layer; the step of obtaining the gradient based on the reward value and updating the parameters of the second module and the parameters of the output layer of the model to be trained using the gradient includes: Construct a loss function based on the reward value; The gradients of the second module and the output layer are calculated based on the loss function; The gradient is used to update the parameters of the second module and the output layer.
6. The model training method according to claim 5, characterized in that, The calculation of the loss function based on the reward value includes: The DPO algorithm is used to calculate the loss function based on the reward value.
7. The model training method according to claim 1, characterized in that, Also includes: During training, monitor the average reward value of the model to be trained on the validation set; Training is stopped when the average reward value does not increase within a preset number of consecutive training cycles.
8. An electronic device, characterized in that, The electronic device includes: One or more processors; and A memory storing computer program instructions, which, when executed, cause the processor to perform the steps of the model training method as described in any one of claims 1 to 7.
9. A computer-readable medium having a computer program / instructions stored thereon, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the model training method according to any one of claims 1 to 7.
10. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the model training method according to any one of claims 1 to 7.