Model training optimization method and device and computing equipment

By unloading machine resources after the inference phase and reusing these resources in the forward propagation and model training phase, the problem of insufficient efficiency in machine resource use in the prior art is solved, and efficient model training in resource-constrained environments is achieved.

CN120068988APending Publication Date: 2025-05-30SHUXING TECH (BEIJING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510160927.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing reinforcement learning scheme based on human feedback requires a large amount of machine resources during model training, resulting in a high threshold for use and is difficult to effectively carry out in an environment where machine resources are limited.

Method used

Time-sharing reuse of machine resources is achieved by unloading machine resources after the inference phase and reusing these resources in the forward propagation and model training phases.

Benefits of technology

This reduces the number of machine resources required in the reinforcement learning process and lowers the threshold for use, so that model training can be efficiently carried out in an environment where machine resources are limited.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068988A_ABST
    Figure CN120068988A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a model training optimization method and device and computing equipment, and the method comprises the steps: obtaining a reinforcement learning model, the reinforcement learning process of the reinforcement learning model comprises a reasoning stage, a forward propagation stage and a model training stage, and the reasoning stage, the forward propagation stage and the model training stage are carried out in series; after the reasoning stage is finished, machine resources used in the reasoning stage are unloaded, in the forward propagation stage and the model training stage, the machine resources used in the reasoning stage are reused, and forward propagation and model training are carried out on the reinforcement learning model based on sample data obtained in the reasoning stage. The used machine resources are unloaded after the reasoning stage is finished, and the machine resources used in the reasoning stage are multiplexed in a time-sharing manner in the forward propagation stage and the model training stage, so that the number of the machine resources required in the reinforcement learning process is reduced, and the threshold of the reinforcement learning method is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this specification relate to the field of computer technologies, and particularly to a method, an apparatus, and a computing device for optimizing model training. Background Art

[0002] With the rapid development of computers, the Internet, and information technologies, various model learning methods have emerged accordingly. Reinforcement learning is an important branch of machine learning, which is a machine learning method in which a machine (also known as an intelligent agent or agent) interacts with the environment to learn how to make decisions to maximize the cumulative reward.

[0003] Currently, a common application of reinforcement learning in natural language processing tasks is reinforcement learning based on human feedback (RLHF), that is, by using the preferences of human feedback on the output of a language model to optimize the language model so that the output of the language model aligns with the preference feedback of humans. In the prior art, in the reinforcement learning based on human feedback, a large amount of machine resources are required during the entire learning process, resulting in a high usage threshold. Therefore, there is an urgent need for a model training optimization solution under reinforcement learning. Summary of the Invention

[0004] In view of this, the embodiments of this specification provide a method for optimizing model training. One or more embodiments of this specification also relate to a model training optimization apparatus, a computing device, a computer-readable storage medium, and a computer program product to solve the technical defects existing in the prior art.

[0005] According to the first aspect of the embodiments of this specification, a method for optimizing model training is provided, including: Obtain a reinforcement learning model, where the reinforcement learning process of the reinforcement learning model includes an inference stage, a forward propagation stage, and a model training stage, and the inference stage, the forward propagation stage, and the model training stage are performed serially; After the inference stage ends, unload the machine resources used in the inference stage. In the forward propagation stage and the model training stage, reuse the machine resources used in the inference stage, and perform forward propagation and model training on the reinforcement learning model based on the sample data obtained in the inference stage.

[0006] According to the second aspect of the embodiments of this specification, a model training optimization apparatus is provided, including: An obtaining module, configured to obtain a reinforcement learning model, where the reinforcement learning process of the reinforcement learning model includes an inference stage, a forward propagation stage, and a model training stage, and the inference stage, the forward propagation stage, and the model training stage are performed serially; The first training module is configured to unload the machine resources used in the inference phase after the end of the inference phase, reuse the machine resources used in the inference phase during the forward propagation phase and the model training phase, and perform forward propagation and model training on the reinforcement learning model based on the sample data obtained in the inference phase.

[0007] According to the third aspect of the embodiments of the present specification, a computing device is provided, including: A memory and a processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the above-mentioned model training optimization method are implemented.

[0008] According to the fourth aspect of the embodiments of the present specification, a computer-readable storage medium is provided, which stores computer-executable instructions. When the instructions are executed by the processor, the steps of the above-mentioned model training optimization method are implemented.

[0009] According to the fifth aspect of the embodiments of the present specification, a computer program product is provided, including a computer program / instructions. When the computer program / instructions are executed by the processor, the steps of the above-mentioned model training optimization method are implemented.

[0010] In the embodiments of the present specification, a model training optimization method is provided. A reinforcement learning model is obtained. The reinforcement learning process of the reinforcement learning model includes an inference phase, a forward propagation phase, and a model training phase. The inference phase, the forward propagation phase, and the model training phase are performed serially. After the end of the inference phase, the machine resources used in the inference phase are unloaded. During the forward propagation phase and the model training phase, the machine resources used in the inference phase are reused, and forward propagation and model training are performed on the reinforcement learning model based on the sample data obtained in the inference phase.

[0011] One embodiment of the present specification realizes that for the reinforcement learning of a reinforcement learning model, after the end of the inference phase, the machine resources used in the inference phase can be unloaded. During the forward propagation phase and the training phase, the machine resources used in the inference phase are reused, and forward propagation and model training are performed on the reinforcement learning model based on the sample data obtained in the inference phase. In this way, the used machine resources can be unloaded after the end of the inference phase, so as to time-share and reuse the machine resources used in the inference phase during the forward propagation phase and the model training phase. Machine resources can be time-shared and reused between different phases in the reinforcement learning process, reducing the number of machine resources required in the reinforcement learning process, lowering the threshold for using the reinforcement learning method, and being user-friendly to users with limited machine resources. Description of the Drawings

[0012] Figure 1It is a flowchart of a model training optimization method provided by an embodiment of this specification; Figure 2 It is a schematic diagram of the training process of a model training optimization method provided by an embodiment of this specification; Figure 3 It is a flowchart of the processing process of a model training optimization method provided by an embodiment of this specification; Figure 4 It is a schematic diagram of the structure of a model training optimization device provided by an embodiment of this specification; Figure 5 It is a block diagram of the structure of a computing device provided by an embodiment of this specification. Detailed implementation manners

[0013] Many specific details are set forth in the following description in order to provide a thorough understanding of this specification. However, this specification can be implemented in many other ways different from those described herein, and those skilled in the art can make similar generalizations without departing from the connotation of this specification. Therefore, this specification is not limited by the specific implementations disclosed below.

[0014] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a" and "the" used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more of the associated listed items.

[0015] It should be understood that although the terms first, second, etc. may be used in one or more embodiments of this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first can also be referred to as the second, and similarly, the second can also be referred to as the first. Depending on the context, the word "if" as used herein can be interpreted as "when" or "while" or "in response to determining".

[0016] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data that have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of the relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entrances are provided for the user to select authorization or rejection.

[0017] First, explain the noun terms involved in one or more embodiments of this specification.

[0018] VLLM (Vectorized Language Model Library): A library for accelerating the inference of large language models, mainly used to accelerate the text generation process, especially performing excellently in scenarios of long text generation and batch generation. It improves the sampling speed through vectorization and parallelization techniques, thereby significantly enhancing the generation efficiency and significantly improving the performance when processing long sequence generation.

[0019] PPO (Proximal Policy Optimization) training: Generally refers to an algorithm in the field of reinforcement learning. PPO is a method for training agents (such as robots or characters in games) to make decisions in complex environments. It belongs to a type of policy gradient method and is particularly suitable for solving tasks with sparse rewards or requiring long-term memory.

[0020] Reinforcement Learning (RL): A machine learning method in which an agent learns how to take actions to maximize a certain cumulative reward by interacting with the environment. The core of reinforcement learning lies in the process of the agent learning the best policy through trial-and-error.

[0021] RLHF (Reinforcement Learning from Human Feedback): A technology that combines machine learning and human judgment. This method uses human preferences to train the model to make it learn to perform tasks in a way that better conforms to human values or expectations. The core idea of RLHF is to collect human preferences for the output results of different policies to train a reward model, which can predict which behaviors or outputs are more popular. Then, this reward model is used as part of the reinforcement learning training to guide the agent to learn better policies.

[0022] The steps of RLHF can include: First, a series of example data need to be collected. This data can be the performance records of the agent on a certain task. Usually, this data is generated by the initial policy, which may be a random policy or a pre-trained policy. Then, the preferences of humans for these example data need to be collected. This usually involves having human evaluators watch or experience different behaviors of the agent and express which behavior they prefer more. Then, use the human preference data to train a reward model, which will learn to predict which behavior is more favored by humans in a given state. Finally, use this reward model to train a new policy, which will try to maximize the expected reward. This process is usually implemented using algorithms such as PPO.

[0023] In the RLHF-PPO stage, there are a total of four main models, namely: Actor Model: The execution model, which is the target language model to be trained; Critic Model: The critic model, whose role is to estimate the total return; Reward Model: The reward model, whose role is to calculate the immediate return; Reference Model: The reference model, whose role is to add some "constraints" to the language model in the RLHF stage to prevent the language model from being trained incorrectly (updating in an uncontrolled direction, and the effect may become worse and worse).

[0024] Among them, the Actor / Critic Model needs to be trained in the RLHF stage, while the Reward / Reference Model has its parameters frozen. The Critic / Reward / Reference Model jointly constitutes a "reward-loss" calculation system, and the loss is calculated by synthesizing their results for updating the Actor and Critic Model.

[0025] Offloading: In deep learning and reinforcement learning, especially when dealing with large models, since the size of the model may exceed the video memory capacity of a single GPU, some techniques need to be adopted to solve this problem. "Offloading" is a common strategy, which refers to moving some computational tasks or data from the main memory (such as GPU video memory) to other devices, such as CPU memory or hard disk, to save precious video memory resources.

[0026] It should be noted that the mainstream RLHF technical solutions, such as the NeMo Aligner framework, are relatively complex, so the development and maintenance costs are very high. For example, during the training of the Actor and Critic, they are parallel, and during the concurrent inference of the Actor, Critic, Reward, and Reference, compared with serial execution, the mental burden of understanding is greater. In addition, for the same training, more machines are required, resulting in a high usage threshold. For example, the Actor and Critic each need to run on separate machines.

[0027] There are mainly two contradictions in PPO training. One is that the inference stage takes a relatively long time (70%-80%), and the other is that the four models occupy a relatively large amount of video memory.

[0028] Therefore, the embodiments of this specification provide a model training optimization scheme, designing a simple RLHF training framework. Since the nemo aligner (a toolkit for natural language processing (NLP)) architecture is a bit complex, to obtain benefits, it is more dependent on the efficiency of the pipeline / overlap, and it is necessary to finely adjust the resource ratio of each component (Actor / Critic). The resource usage in the fine-tuning stage of the RLHF model is not too much either. Therefore, in the embodiments of this specification, a simple learning architecture is used, and each stage is completely serial. Specifically, VLLM can be used to accelerate sampling, and each stage of PPO training is performed serially. Through gpu offload, the machine resources (mainly GPU computing power and video memory) are multiplexed time-sharing in each stage.

[0029] In the embodiments of this specification, the inference stage is serial with the forward propagation stage and the training stage, which is easy to understand; and in the scenario of pursuing high efficiency, the machine used for training can be reused by the component VLLM used during inference to achieve the maximum throughput; in addition, the model training optimization method provided by the embodiments of this specification can be applied to the entire training process: namely Pretrain (pre-training), SFT (supervised fine-tuning), RLHF; furthermore, it reduces the number of machine resources required in the reinforcement learning process, reduces the usage threshold of the reinforcement learning method, and is more user-friendly to users with limited machine resources.

[0030] In this specification, a model training optimization method is provided. This specification also relates to a model training optimization device, a computing device, a computer-readable storage medium, and a computer program product, which will be described in detail one by one in the following embodiments.

[0031] See Figure 1 , Figure 1 shows a flowchart of a model training optimization method provided according to an embodiment of this specification, which specifically includes the following steps.

[0032] Step 102: Obtain a reinforcement learning model. The reinforcement learning process of the reinforcement learning model includes an inference stage, a forward propagation stage, and a model training stage, and the inference stage, the forward propagation stage, and the model training stage are performed serially.

[0033] It should be noted that the model training optimization method provided in the embodiments of this specification is applied to reinforcement learning from human feedback (RLHF). Generally, reinforcement learning from human feedback (RLHF) includes three processes: the first process is supervised fine-tuning, using instruction prompts and outputs as training data to train a base model to obtain a supervised fine-tuned model; the second process is to train a reward model, using human preference data as training data to train a scoring model as the reward model; the third process is the reinforcement learning process, using the reward model trained in the second process and applying a reinforcement learning algorithm to optimize the supervised fine-tuned model to obtain a final language model aligned with human preferences.

[0034] In actual implementation, one iteration training of the reinforcement learning process generally includes three stages: an inference stage, a forward propagation stage, and a model training stage. The reinforcement learning model generally can include an Actor Model, a Reference Model, a Critic Model, and a Reward Model.

[0035] The Actor Model is the language model to be trained, such as a question-and-answer model, a translation model, etc. Specifically, a supervised fine-tuned model obtained from the first process of the above RLHF can be used as the Actor Model, and then trained based on reinforcement learning to obtain the final target language model required. Therefore, a supervised fine-tuned model produced by the supervised fine-tuning process can be obtained as the Actor Model of reinforcement learning. And a reward model obtained by training in the second process can be obtained as the reward model of reinforcement learning.

[0036] In addition, the Critic Model is used to estimate the total reward. The Critic Model is a model that needs to be trained in the reinforcement learning algorithm. Therefore, a custom neural network model can be initialized, and its ability to pre-train and learn to evaluate text quality can be used as the Critic Model, and then reinforcement learning is performed.

[0037] The Reference Model is used to add some "constraints" to the Actor Model to be trained in the RLHF stage to prevent the Actor Model from updating in an uncontrolled direction. A pre-trained model that has been fully trained and performs well on related tasks can be selected as the Reference Model.

[0038] In the embodiments of this specification, a reinforcement learning model can be obtained, and reinforcement learning is performed on this reinforcement model to obtain the final target language model required.

[0039] In actual implementation, learning samples for performing reinforcement learning on the reinforcement learning model can also be obtained. These learning samples can be obtained from a database or other third-party platforms. For example, in a question-and-answer scenario, the learning sample can be a question obtained from a database; in a translation scenario, the learning sample can be the original text obtained from a translation database.

[0040] In the inference stage, the reinforcement learning model is used to generate sample answers for the learning samples. The learning samples and the sample answers are used as sample data for performing forward propagation and model training on the reinforcement learning model in subsequent forward propagation stages and model training stages.

[0041] It should be noted that the reinforcement learning model includes four models: an actor model, a reference model, a critic model, and a reward model. Among them, the actor model and the critic model need to be trained, while the reference model and the reward model have frozen parameters. The critic model, the reward model, and the reference model jointly form a "reward-loss" calculation system, and their results are integrated to calculate the loss for updating the model parameters.

[0042] In actual implementation, the goal of reinforcement learning is to enable the actor model to generate answers that conform to human preferences. In the inference stage, corresponding machine resources can be used to generate sample answers for the learning samples using the reinforcement learning model. The learning samples and the sample answers are used as sample data, which can be used in subsequent forward propagation stages and training stages. Among them, the machine resources can be computing or storage resources such as GPU computing power and GPU video memory.

[0043] Specifically, the learning samples can be input into the actor model in the reinforcement learning model, enabling the actor model to generate corresponding sample answers. Subsequently, the learning samples and the sample answers are used as a sample data for subsequent forward propagation stages and training stages, participating in the "reward-loss" calculation system, and integrating their results to calculate the loss value for updating the model parameters of the actor model and the critic model.

[0044] In an optional implementation manner of this embodiment, generating sample answers for learning samples using a reinforcement learning model includes: By setting a text generation acceleration method, sampling the learning samples using the reinforcement learning model to generate corresponding sample answers.

[0045] Among them, the set text generation acceleration method is any algorithm that can accelerate text generation, such as VLLM, SparseML (an open-source library that supports pruning and quantization of PyTorch models), etc.

[0046] In actual implementation, taking the set text generation acceleration method as VLLM as an example, in the inference stage, the PyTorch code in the inference stage can be adapted to use the open-source VLLM sampling. Specifically, first install VLLM, define the execution model (Actor Model) used in the inference stage in PyTorch, and convert it into a form that VLLM can load. Usually, the execution model (Actor Model) used in the inference stage can be saved as a weight file and the corresponding model configuration file is provided. Then use the Engine class of VLLM to load the execution model (Actor Model) used in the inference stage, and use the Engine object to perform text generation to obtain the sample answers of the sample data.

[0047] In the embodiments of this specification, an open-source set text generation acceleration method can be adopted to implement text generation in the inference stage, obtain corresponding sample answers, greatly improve the generation efficiency of sample answers, save the generation time of sample answers in the inference stage, and accelerate the learning process.

[0048] Step 104: After the inference stage ends, unload the machine resources used in the inference stage, and reuse the machine resources used in the inference stage in the forward propagation stage and the model training stage, and perform forward propagation and model training on the reinforcement learning model based on the sample data obtained in the inference stage.

[0049] In actual implementation, after the inference stage ends, the machine resources used in the inference stage can be directly unloaded, so that the machine resources used in the inference stage can be reused in the forward propagation stage and the model training stage. Or, the machine resources used in the inference stage can also be unloaded when the machine resources meet the limit conditions, so as to reuse the machine resources used in the inference stage in the forward propagation stage and the model training stage and reduce the demand for machine resources.

[0050] Among them, the limit condition refers to a condition set in advance to judge whether the machine resources are sufficient. If the machine resources meet the limit condition, it means that the machine resources are limited; if the machine resources do not meet the limit condition, it means that the machine resources are sufficient. For example, the limit condition is that the number of graphics cards is less than the set threshold.

[0051] It should be noted that when the machine resources meet the limit conditions, it means that the available machine resources are limited. When dealing with large models, since the size of the model may exceed the load capacity of the machine resources, therefore, in the embodiments of this specification, the inference stage, the forward propagation stage, and the model training stage are executed serially. After the inference stage ends, the machine resources used in the inference stage are unloaded. In the forward propagation stage and the model training stage, the machine resources used in the inference stage are reused. Based on the sample data obtained in the inference stage, forward propagation and model training are performed on the reinforcement learning model to complete the reinforcement learning process of the reinforcement learning model, realizing the time-sharing reuse of machine resources, maximizing the resource utilization efficiency, and reducing the user's usage threshold. The number of machine resources can be reduced by two times.

[0052] In an optional implementation manner of this embodiment, the machine resources are video memory; unloading the machine resources used in the inference stage includes: Unloading the model weights and / or key-value caches of the reinforcement learning model in the inference stage from the video memory and putting them into the memory.

[0053] Among them, video random access memory (VRAM) is a type of memory in a graphics processing unit (GPU) used to store graphic data. VRAM plays a crucial role in the work of the GPU, especially in high-performance computing tasks involving a large amount of graphic rendering, deep learning, scientific computing, etc.

[0054] Memory is a place in a computer system used to temporarily store data and programs. Memory is usually divided into two main types: RAM (random access memory) and ROM (read-only memory). In modern computer systems, RAM is the most commonly used type of memory. It allows fast read and write operations on data, and the data will be cleared after the system power-off. ROM is used to store fixed programs such as BIOS or firmware, and these data still remain after the system power-off. The memory in the embodiments of this specification may refer to CPU memory, usually called main memory or RAM, which is a volatile memory that directly interacts with the central processing unit (CPU). CPU memory is mainly used to store the current running state and data of programs for the CPU to access quickly.

[0055] Model weights are an important concept in machine learning and deep learning models. Weights are part of the model parameters and are used to measure the influence degree of input features on the model output. In a neural network, weights determine the way of transmission and combination of input data between different layers.

[0056] KV Cache (Key-Value Cache) is a technique used in deep learning to accelerate the inference of sequence models. Especially when dealing with long sequences, this technique is particularly suitable for the Transformer architecture and its variants, such as GPT-3, T5, etc. KV Cache reduces redundant calculations by caching the keys and values in the attention mechanism, thus significantly improving the inference speed. In the self-attention mechanism of the Transformer, queries, keys, and values are calculated at each time step. For the decoder part, especially in autoregressive generation tasks, the attention weights of the entire sequence need to be recalculated each time a new token is generated.

[0057] In actual implementation, during the inference stage, when inputting sample data into the Actor Model to generate sample answers using the Actor Model, video memory resources will be used. The model weights and key-value cache in the inference stage are stored in the video memory. After the inference stage ends, in order to enable the subsequent forward propagation stage and training stage to reuse the video memory resources of the inference stage, the model weights and / or key-value cache of the reinforcement learning model in the inference stage can be offloaded from the video memory and placed in the memory to release the video memory resources.

[0058] Specifically, first, the model weights of the Actor Model used in the inference stage can be loaded into the CPU memory; if a key-value cache is used in the inference stage (such as in the self-attention mechanism of the Transformer model), the key-value pairs can be offloaded from the GPU to the CPU memory. For example, when creating the key-value cache, it can be specified to be stored in the CPU memory, and after each new key-value pair is generated, they can be offloaded from the GPU to the CPU memory.

[0059] In the embodiments of this specification, the model weights and / or key-value cache of the reinforcement learning model in the inference stage can be offloaded from the video memory and placed in the memory to release the video memory resources of the inference stage, enabling the subsequent forward propagation stage and model training stage to reuse the video memory resources of this inference stage, realizing time-sharing reuse of the video memory, and reducing the number of required graphics cards.

[0060] In an optional implementation manner of this embodiment, forward propagation and model training of the reinforcement learning model based on the sample data obtained in the inference stage include: In the forward propagation stage, adopt the first resource optimization strategy to perform forward propagation on the reinforcement learning model based on the sample data to obtain the forward propagation result of the sample data; In the model training stage, adopt the second resource optimization strategy to update the model parameters of the reinforcement learning model according to the forward propagation result.

[0061] Among them, the first resource optimization strategy refers to a pre-configured resource optimization strategy adopted in the forward propagation stage. For example, the first resource optimization strategy can be to put the model weights of the currently used model into the video memory and unload the model weights of the currently unused model from the video memory and put them into the memory. The second resource optimization strategy refers to a pre-configured resource optimization strategy adopted in the model training stage. For example, the model parameters of the currently used model are put into the video memory, the model parameters of the currently unused model are unloaded from the video memory and put into the memory, and / or the gradient data and optimizer state parameters are unloaded from the video memory.

[0062] It should be noted that the RLHF-PPO stage is mainly divided into three stages: the inference stage, the forward propagation stage, and the model training stage. In the inference stage, the main focus is on using the execution model (Actor Model) to generate text or sequences as sample answers, which usually involves setting initial conditions (such as starting prompts) and making predictions using the execution model (Actor Model). In the forward propagation stage, the sample data obtained in the inference stage is propagated forward through the reinforcement learning model to obtain an evaluation result for the sample data, which can include calculating the loss function, evaluating the reward, etc. In the model training stage, the parameters of the reinforcement learning model can be updated according to the results of the forward propagation, which usually involves calculating the loss, gradient backpropagation, and updating of the optimizer, etc.

[0063] In actual implementation, in the forward propagation stage, the reinforcement learning model can be propagated forward based on the sample data to obtain the forward propagation result of the sample data. In the above forward propagation process, the first resource optimization strategy can be adopted to optimize the resource usage in the forward propagation stage; in the model training stage, the model parameters of the reinforcement learning model can be updated according to the forward propagation result. In the process of updating the model parameters, the second resource optimization strategy can be adopted to optimize the resource usage in the model training stage.

[0064] In the embodiments of this specification, in addition to serially performing the inference stage, the forward propagation stage, and the model training stage as described above and reusing the machine resources of the inference stage in the forward propagation stage and the model training stage, corresponding resource optimization strategies can be respectively adopted in the forward propagation stage and the model training stage to further optimize the usage of machine resources, further improve the resource utilization rate, and save machine resources.

[0065] In an optional implementation manner of this embodiment, the reinforcement learning model includes an execution model, a reference model, a critic model, and a reward model; the first resource optimization strategy includes putting the model weights of the currently used model into the video memory and unloading the model weights of the currently unused model from the video memory and putting them into the memory; Adopt the first resource optimization strategy to perform forward propagation on the reinforcement learning model based on sample data, including: During the process of performing forward propagation on the reinforcement learning model based on sample data, when using the execution model, put the model weights of the execution model into the video memory, unload the model weights of the reference model from the video memory and put them into the memory; when using the reference model, put the model weights of the reference model into the video memory, unload the model weights of the execution model from the video memory and put them into the memory; When using the critic model, put the model weights of the critic model into the video memory, unload the model weights of the reward model from the video memory and put them into the memory; when using the reward model, put the model weights of the reward model into the video memory, unload the model weights of the critic model from the video memory and put them into the memory.

[0066] It should be noted that taking the first resource optimization strategy including putting the model weights of the currently used model into the video memory and unloading the model weights of the currently unused model from the video memory and putting them into the memory as an example, the execution model (Actor Model) is the language model to be trained, and the reference model (Reference Model) is used to add some "constraints" to the execution model (Actor Model) to prevent the execution model (Actor Model) from updating in an uncontrolled direction. Therefore, the execution model (Actor Model) and the reference model (Reference Model) do not need to use the model weights simultaneously.

[0067] In actual implementation, during the process of performing forward propagation on the reinforcement learning model based on sample data, if the execution model (Actor Model) needs to be used currently, the model weights of the execution model (Actor Model) can be put into the video memory for quick reading directly from the video memory, and the model weights of the currently unused reference model (Reference Model) are unloaded from the video memory and put into the memory to release the video memory resources for storing the model weights of the execution model (Actor Model) that need to be used currently; if the reference model (Reference Model) needs to be used currently, the model weights of the reference model (Reference Model) are put into the video memory for quick reading directly from the video memory, and the model weights of the currently unused execution model (Actor Model) are unloaded from the video memory and put into the memory to release the video memory resources for storing the model weights of the reference model (Reference Model) that need to be used currently.

[0068] In addition, since the Critic Model is used to predict the expected total return, like the Actor Model, its parameters also need to be updated. The Reward Model is used to calculate the immediate return and is a trained Reward Model. During the RLHF process, its parameters are frozen. The Critic Model and the Reward Model do not need to use the model weights simultaneously. Therefore, when the Critic Model needs to be used currently, the model weights of the Critic Model can be put into the video memory for quick reading directly from the video memory, while the model weights of the Reward Model that are not needed currently are unloaded from the video memory and put into the memory to release the video memory resources for storing the model weights of the Critic Model that need to be used currently. When the Reward Model needs to be used currently, the model weights of the Reward Model can be put into the video memory for quick reading directly from the video memory, while the model weights of the Critic Model are unloaded from the video memory and put into the memory to release the video memory resources for storing the model weights of the Reward Model that need to be used currently.

[0069] In the embodiments of this specification, during the forward propagation stage, four models are involved, namely the Actor Model, the Reference Model, the Critic Model, and the Reward Model. The Actor Model and the Reference Model do not need to use the model weights simultaneously, and the Critic Model and the Reward Model do not need to use the model weights simultaneously either. Therefore, the model weights of the model that needs to be used currently can be put into the video memory, and the model weights of the model that are not needed currently can be unloaded from the video memory and put into the memory, so as to enable multiple models to share the video memory resources for storing model weights in a time-sharing manner, optimize the use of video memory resources, improve the utilization rate of video memory resources, and reduce the number of required graphics cards.

[0070] In an optional implementation manner of this embodiment, the second resource optimization strategy includes putting the model parameters of the currently used model into the video memory and unloading the model parameters of the currently unused model from the video memory and putting them into the memory; Adopting the second resource optimization strategy to update the model parameters of the reinforcement learning model according to the forward propagation results includes: In the process of updating the model parameters of the reinforcement learning model according to the forward propagation result, when updating the execution model, unload the model parameters of the critic model from the video memory and put them into the memory; when updating the critic model, unload the model parameters of the execution model from the video memory and put them into the memory; Among them, the execution model and the critic model are the models to be trained in the reinforcement learning model.

[0071] It should be noted that taking the second resource optimization strategy including putting the model parameters of the currently used model into the video memory and unloading the model parameters of the currently unused model from the video memory and putting them into the memory as an example, both the execution model (Actor Model) and the critic model (Critic Model) are the models that need to be trained in the model training stage and need to update the model parameters, but the execution model (Actor Model) and the critic model (Critic Model) will not update the model parameters at the same time.

[0072] Therefore, in actual implementation, in the process of updating the model parameters of the reinforcement learning model according to the forward propagation result, if the currently needed model to be updated is the execution model (Actor Model), the model parameters of the critic model (Critic Model) can be unloaded from the video memory and put into the memory to release the video memory resources for storing the model parameters of the execution model (Actor Model). If the currently needed model to be updated is the critic model (Critic Model), the model parameters of the execution model (Actor Model) can be unloaded from the video memory and put into the memory to release the video memory resources for storing the model parameters of the critic model (Critic Model).

[0073] In the embodiments of this specification, in the model training stage, it is involved to update the model parameters of the execution model (Actor Model) and the critic model (Critic Model). The model parameters of the execution model (Actor Model) and the critic model (Critic Model) will not be used at the same time. Therefore, the model parameters of the currently used model can be put into the video memory, and the model parameters of the currently unused model can be unloaded from the video memory and put into the memory to release the video memory resources for storing the model parameters of the currently needed model, which is convenient for quickly reading directly from the video memory and avoiding video memory overflow. Thus, multiple models can share the video memory resources to store the model parameters in a time-sharing manner, optimize the use of the video memory resources, improve the utilization rate of the video memory resources, and reduce the number of required graphics cards.

[0074] In an optional implementation manner of this embodiment, the second resource optimization strategy includes unloading gradient data and optimizer state parameters from the video memory; Adopt the second resource optimization strategy to update the model parameters of the reinforcement learning model according to the forward propagation results, including: During the model training phase, construct gradient data according to the forward propagation results, use the optimizer to update the model parameters of the reinforcement learning model based on the gradient data, complete the training steps of the current round, and unload the gradient data from the video memory to the memory; Before using the optimizer to update the model parameters of the reinforcement learning model, load the optimizer state parameters of the optimizer into the video memory; after using the optimizer to update the model parameters of the reinforcement learning model, unload the optimizer state parameters from the video memory and write them to the memory.

[0075] It should be noted that during the model training phase, when updating the model parameters of the reinforcement learning model according to the forward propagation results, optimization algorithms such as stochastic gradient descent and optimizers will be used. Therefore, the second resource optimization strategy in the model training phase can also include unloading gradient data and optimizer state parameters from the video memory.

[0076] Among them, gradient data refers to the partial derivatives of the loss function with respect to the model parameters calculated during the backpropagation process. These partial derivatives indicate the change trend of the loss function at the current parameter values, and thus can be used to guide the update direction of the parameters. An optimizer is a specific algorithm used to update the model parameters. It adjusts the model parameters according to the gradient data in order to find the parameter settings that minimize the loss function.

[0077] In deep learning and machine learning, the optimizer not only is responsible for updating the model parameters but also maintains some internal state parameters. These state parameters are crucial for the correct operation of the optimization algorithm. These state parameters usually contain historical information such as the history of gradients and momentum, etc. These information can help the optimizer better adjust the learning rate or determine the direction of parameter updates.

[0078] In actual implementation, during the model training phase, gradient data can be constructed according to the forward propagation results in the forward propagation phase. Based on this gradient data, use the optimizer to update the model parameters of the Actor Model and Critic Model in the reinforcement learning model, complete the iterative training steps of the current round, and then unload the gradient data from the video memory to the memory to release the video memory resources. In addition, before using the optimizer to update the model parameters of the reinforcement learning model, the optimizer state parameters of the optimizer can be loaded into the video memory to facilitate directly reading the state parameters from the video memory to implement the update of the model parameters; after using the optimizer to update the model parameters of the reinforcement learning model, unload the optimizer state parameters from the video memory and write them to the memory to release the video memory resources.

[0079] In the embodiments of this specification, during the model training phase, gradient data and an optimizer are involved. After the training step of the current round is completed, the gradient data used to update the model parameters can be unloaded from the video memory, and the optimizer state parameters can be loaded into the video memory when using the optimizer and unloaded from the video memory and written to the memory after use. In this way, after the training step of the current round is completed, the video memory resources can be fully released, facilitating subsequent iterative training steps, multiplexing the video memory resources in a time-sharing manner, avoiding video memory overflow, optimizing the use of video memory resources, improving the utilization rate of video memory resources, and reducing the number of required graphics cards.

[0080] It should be noted that after completing the training step of the current round, return to continue executing the above-mentioned inference phase and continue iterative training until the training stop condition is met. The obtained execution model (Actor Model) is the final target language model required. The trained target language model can be provided for downstream applications. For example, if the target language model is a question-and-answer model, it can be provided for the downstream automatic question-and-answer platform. The answers of the trained target language model can be closer to human preferences; if the target language model is a translation model, it can be provided for the downstream automatic translation platform. The translation results of the trained target language model can be closer to human preferences.

[0081] In an optional implementation manner of this embodiment, after the inference phase ends, when the machine resources meet the limit conditions, the machine resources used in the inference phase are unloaded, and in the forward propagation phase and the model training phase, the implementation manner of multiplexing the machine resources used in the inference phase. After obtaining the reinforcement learning model, it further includes: After the inference phase ends, when the machine resources do not meet the limit conditions, in the forward propagation phase and the model training phase, other machine resources are used to perform forward propagation and model training on the reinforcement learning model based on the sample data obtained in the inference phase, where the other machine resources are different from the machine resources used in the inference phase.

[0082] It should be noted that when training the execution model (Actor Model) and the critic model (CriticModel) during the model training phase, the machine resources in the inference phase may not be fully offloaded (such as the nccl process group). In addition, the use of machine resources such as memory and io will all interact with the VLLM in the inference phase. Therefore, exclusive training of the VLLM resources in the inference phase can also be selected.

[0083] In practical applications, during the inference stage, a sample answer of a learning sample is generated using a reinforcement learning model. After using the learning sample and the sample answer as sample data, if the machine resources meet the limit conditions, it indicates that the machine resources are limited, and the implementation method of the aforementioned machine resource reuse can be adopted. If the machine resources do not meet the limit conditions, it indicates that the machine resources are sufficient and there is no need for reuse. To avoid the impact of machine resources in different stages, during the forward propagation stage and the training stage, other machine resources different from those used in the inference stage can be used to perform forward propagation and model training on the reinforcement learning model based on the sample data.

[0084] In addition, during the forward propagation stage and the model training stage, when performing forward propagation and model training on the reinforcement learning model based on the sample data using other machine resources different from those used in the inference stage, the above-mentioned first resource optimization strategy and second resource optimization strategy can be used to optimize resources during the forward propagation stage and the model training stage, or the resource optimization during the forward propagation stage and the model training stage may not be performed. The embodiments of this specification do not limit this.

[0085] It should be noted that different machine resources are used in the inference stage, the forward propagation stage, and the model training stage. Each stage can exclusively use machine resources to achieve inference and training, avoiding the mutual influence of machine resources between different stages and ensuring the stability of the entire training stage.

[0086] Figure 2 It is a schematic diagram of the training process of a model training optimization method provided by an embodiment of this specification. As Figure 2 shown, it is the completed training process of one round of iterative training. The first stage is the inference stage, where a learning sample is obtained, the learning sample is input into the Actor Model, and the sample answer of the learning sample is obtained. The learning sample and the sample answer are used as sample data. The second stage is the forward propagation stage, where the sample data is input into the Actor Model, the Critic Model, the Reference Model, and the Reward Model for forward propagation, and the evaluation results obtained from the forward propagation are written into the Experiences Buffer. The third stage is the model training stage. Based on the evaluation results obtained from the forward propagation in the experience buffer, the model parameters of the Actor Model and the Critic Model are updated, and then it returns to the first stage to continue the next round of iterative training. Until the training stop condition is met, the Actor Model obtained in the third stage is the final target language model required.

[0087] It should be noted that the model training optimization method provided in the embodiments of this specification can be applied not only to the above-mentioned reinforcement learning based on human feedback (RLHF), but also to other reinforcement learning processes. If the reinforcement learning process includes at least two stages, after any stage is completed, the machine resources used in that stage can be unloaded and reused in other stages. Additionally, if any stage involves at least two models, the above-mentioned first resource optimization strategy and / or second resource optimization strategy can also be adopted to unload the machine resources used by the models, enabling the reuse of machine resources among multiple models.

[0088] In the embodiments of this specification, a model training optimization method is provided. For the reinforcement learning of a reinforcement learning model, after the inference stage ends, the machine resources used in the inference stage can be unloaded, and in the forward propagation stage and the training stage, the machine resources used in the inference stage can be reused. Based on the sample data obtained in the inference stage, forward propagation and model training of the reinforcement learning model are performed. In this way, the machine resources used can be unloaded after the inference stage ends, so that the machine resources used in the inference stage can be time-shared and reused in the forward propagation stage and the model training stage. Machine resources can be time-shared and reused among different stages in the reinforcement learning process, reducing the number of machine resources required in the reinforcement learning process and lowering the threshold for using the reinforcement learning method, which is user-friendly for users with limited machine resources.

[0089] The following combines the attached Figure 3 , taking the application of the model training optimization method provided in this specification in the question-and-answer scenario as an example, to further illustrate the model training optimization method. Among them, Figure 3 FIG. shows the processing procedure flowchart of a model training optimization method provided in an embodiment of this specification, which specifically includes the following steps.

[0090] Step 302: In the inference stage, obtain a sample question, input the sample question into the question-and-answer model, and use the VLLM method to accelerate to obtain a sample answer to the sample question.

[0091] Step 304: After the inference stage ends, unload the model weights and / or key-value caches of the question-and-answer model in the inference stage from the video memory and put them into the memory.

[0092] Step 306: In the forward propagation stage, input the sample question and the sample answer into the Q&A model, the critic model, the reference model, and the reward model respectively, reuse the machine resources used in the inference stage, perform forward propagation based on the sample question and the sample answer, the Q&A model, the critic model, the reference model, and the reward model, and write the evaluation results obtained from the forward propagation into the experience buffer; during the forward propagation process, if the question model is currently used, put the model weights of the Q&A model into the video memory, unload the model weights of the reference model from the video memory, and put them into the memory; if the reference model is currently used, put the model weights of the reference model into the video memory, unload the model weights of the Q&A model from the video memory, and put them into the memory; if the critic model is currently used, put the model weights of the critic model into the video memory, unload the model weights of the reward model from the video memory, and put them into the memory; if the reward model is currently used, put the model weights of the reward model into the video memory, unload the model weights of the critic model from the video memory, and put them into the memory.

[0093] Step 308: In the model training stage, reuse the machine resources used in the inference stage, construct gradient data according to the evaluation results obtained from the forward propagation in the experience buffer, update the model parameters of the Q&A model and the critic model based on the gradient data using the optimizer, complete the training steps of the current round, unload the gradient data from the video memory, and transfer it to the memory; before using the optimizer to update the model parameters, load the optimizer state parameters of the optimizer into the video memory, and after using the optimizer to update the model parameters, unload the optimizer state parameters from the video memory and write them into the memory; during the process of updating the model parameters, if the Q&A model is currently updated, unload the model parameters of the critic model from the video memory and put them into the memory; if the critic model is currently updated, unload the model parameters of the execution model from the video memory and put them into the memory.

[0094] After the training steps of the current round, return to execute the above Step 302, continue iterative training until the training stop condition is met, and execute the following Step 310.

[0095] Step 310: Obtain the target Q&A model that has completed training, and provide this target Q&A model to the downstream automatic Q&A platform.

[0096] In the embodiments of this specification, a method for optimizing model training is provided. By using reinforcement learning based on human feedback (RLHF) to train a question-and-answer model, corresponding machine resources can be used during the inference stage. The question-and-answer model is utilized to generate sample answers for sample questions. After the inference stage ends, the machine resources used in the inference stage can be unloaded. During the forward propagation stage and the training stage, the machine resources used in the inference stage are reused, and forward propagation and model training are performed based on the sample questions and sample answers. In this way, in the case of limited machine resources, the used machine resources can be unloaded after the inference stage ends, so as to time-share and reuse the machine resources used in the inference stage during the forward propagation stage and the training stage. Machine resources can be time-shared and reused between different stages during the reinforcement learning process, reducing the number of machine resources required during the reinforcement learning process and lowering the threshold for using the reinforcement learning method, which is user-friendly to users with limited machine resources.

[0097] Corresponding to the above method embodiments, this specification also provides embodiments of a model training optimization device. Figure 4 The structural schematic diagram of a model training optimization device provided by an embodiment of this specification is shown. As Figure 4 shown, the device includes: An acquisition module 402, configured to acquire a reinforcement learning model. The reinforcement learning process of the reinforcement learning model includes an inference stage, a forward propagation stage, and a model training stage, and the inference stage, the forward propagation stage, and the model training stage are performed serially. A first training module 404, configured to unload the machine resources used in the inference stage after the inference stage ends, and during the forward propagation stage and the model training stage, reuse the machine resources used in the inference stage, and perform forward propagation and model training on the reinforcement learning model based on the sample data obtained in the inference stage.

[0098] Optionally, the first training module 404 is further configured to: During the forward propagation stage, adopt a first resource optimization strategy to perform forward propagation on the reinforcement learning model based on the sample data, and obtain the forward propagation result of the sample data pair. During the model training stage, adopt a second resource optimization strategy to update the model parameters of the reinforcement learning model according to the forward propagation result.

[0099] Optionally, the reinforcement learning model includes an execution model, a reference model, a critic model, and a reward model; the first resource optimization strategy includes putting the model weights of the currently used model into the video memory and unloading the model weights of the currently unused model from the video memory and putting them into the memory. The first training module 404 is further configured to: During the forward propagation of the reinforcement learning model based on sample data, when using the execution model, the model weights of the execution model are put into the video memory, and the model weights of the reference model are unloaded from the video memory and put into the memory; when using the reference model, the model weights of the reference model are put into the video memory, and the model weights of the execution model are unloaded from the video memory and put into the memory; When using the critic model, the model weights of the critic model are put into the video memory, and the model weights of the reward model are unloaded from the video memory and put into the memory; when using the reward model, the model weights of the reward model are put into the video memory, and the model weights of the critic model are unloaded from the video memory and put into the memory.

[0100] Optionally, the second resource optimization strategy includes putting the model parameters of the currently used model into the video memory and unloading the model parameters of the currently unused model from the video memory and putting them into the memory; The first training module 404 is further configured to: During the process of updating the model parameters of the reinforcement learning model according to the forward propagation result, when updating the execution model, the model parameters of the critic model are unloaded from the video memory and put into the memory; when updating the critic model, the model parameters of the execution model are unloaded from the video memory and put into the memory; Wherein, the execution model and the critic model are the models to be trained in the reinforcement learning model.

[0101] Optionally, the second resource optimization strategy includes unloading the gradient data and the optimizer state parameters from the video memory; The first training module 404 is further configured to: In the model training stage, construct gradient data according to the forward propagation result, update the model parameters of the reinforcement learning model based on the gradient data using the optimizer, complete the training steps of the current round, unload the gradient data from the video memory and transfer it to the memory; Before updating the model parameters of the reinforcement learning model using the optimizer, load the optimizer state parameters of the optimizer into the video memory; after updating the model parameters of the reinforcement learning model using the optimizer, unload the optimizer state parameters from the video memory and write them into the memory.

[0102] Optionally, the device further includes a second training module, which is configured to: After the end of the inference stage, if the machine resources do not meet the limit conditions, in the forward propagation stage and the model training stage, use other machine resources to perform forward propagation and model training on the reinforcement learning model based on the sample data obtained in the inference stage, where the other machine resources are different from the machine resources used in the inference stage.

[0103] Optionally, the first training module 404 is further configured to: Unload the model weights and / or key-value cache of the reinforcement learning model from the video memory to the memory during the inference stage.

[0104] In the embodiments of this specification, a model training optimization device is provided, including an acquisition module, a generation module, and a first training module. When running each module in the model training optimization device, through the interaction and cooperation of each module, for the reinforcement learning of the reinforcement learning model, after the inference stage ends, it is stated that the machine resources are limited, and the machine resources used in the inference stage can be unloaded. In the forward propagation stage and the training stage, the machine resources used in the inference stage are reused, and the forward propagation and model training of the reinforcement learning model are performed based on the sample data obtained in the inference stage. In this way, in the case of limited machine resources, the used machine resources can be unloaded after the inference stage ends, so as to time-share and reuse the machine resources used in the inference stage in the forward propagation stage and the model training stage. The machine resources can be time-shared and reused between different stages during the reinforcement learning process, reducing the number of machine resources required during the reinforcement learning process, lowering the threshold for using the reinforcement learning method, and being more user-friendly to users with limited machine resources.

[0105] The above is a schematic solution of a model training optimization device according to this embodiment. It should be noted that the technical solution of this model training optimization device and the technical solution of the above model training optimization method belong to the same concept. For the details not described in detail in the technical solution of the model training optimization device, reference can be made to the description of the technical solution of the above model training optimization method.

[0106] Figure 5 The structural block diagram of a computing device provided according to an embodiment of this specification is shown. The components of the computing device 500 include but are not limited to a memory 510 and a processor 520. The processor 520 is connected to the memory 510 through a bus 530, and a database 550 is used to store data.

[0107] The computing device 500 also includes an access device 540, which enables the computing device 500 to communicate via one or more networks 560. Examples of such networks include the Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 540 may include one or more of any type of wired or wireless network interface (e.g., a network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, or a Near Field Communication (NFC) interface.

[0108] In one embodiment of the present specification, the above components of the computing device 500, as well as Figure 5 other components not shown, may also be connected to each other, for example, via a bus. It should be understood that Figure 5 the block diagram of the computing device shown is for illustrative purposes only and is not a limitation on the scope of the present specification. Those skilled in the art can add or replace other components as needed.

[0109] The computing device 500 can be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or a personal computer (PC). The computing device 500 can also be a mobile or stationary server.

[0110] Among them, the processor 520 is used to execute the following computer-executable instructions, which, when executed by the processor, implement the steps of the above model training optimization method.

[0111] The above is a schematic solution of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the above model training optimization method belong to the same concept. For the details not described in detail in the technical solution of the computing device, reference can be made to the description of the technical solution of the above model training optimization method.

[0112] An embodiment of this specification also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the steps of the above model training optimization method.

[0113] The above is a schematic solution of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium and the technical solution of the above model training optimization method belong to the same concept. For the details not described in detail in the technical solution of the storage medium, reference can be made to the description of the technical solution of the above model training optimization method.

[0114] An embodiment of this specification also provides a computer program, which, when executed on a computer, causes the computer to execute the steps of the above model training optimization method.

[0115] The above is a schematic solution of a computer program according to this embodiment. It should be noted that the technical solution of this computer program and the technical solution of the above model training optimization method belong to the same concept. For the details not described in detail in the technical solution of the computer program, reference can be made to the description of the technical solution of the above model training optimization method.

[0116] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain implementations, multitasking and parallel processing are also possible or may be advantageous.

[0117] Computer instructions include computer program code, which can be in the form of source code, object code, executable files, or some intermediate forms, etc. A computer-readable medium can include: any entity or device capable of carrying computer program code, recording media, USB flash drives, mobile hard disks, magnetic disks, optical disks, computer memories, read-only memories (ROMs), random access memories (RAMs), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0118] It should be noted that for the foregoing method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of this specification are not limited by the described action sequence, because according to the embodiments of this specification, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential for the embodiments of this specification.

[0119] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0120] The preferred embodiments of this specification disclosed above are only used to help explain this specification. The alternative embodiments do not describe all the details in detail, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and changes can be made according to the content of the embodiments of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the embodiments of this specification, so that those skilled in the art can understand and utilize this specification well. This specification is only limited by the claims and their full scope and equivalents.

Claims

1. A model training optimization method, characterized in that: include: Acquire a reinforcement learning model, wherein the reinforcement learning process of the reinforcement learning model includes an inference phase, a forward propagation phase, and a model training phase, and the inference phase, the forward propagation phase, and the model training phase are performed in series; After the inference stage is finished, the machine resources used in the inference stage are unloaded, and in the forward propagation stage and the model training stage, the machine resources used in the inference stage are reused, and the reinforcement learning model is forward propagated and model trained based on the sample data obtained in the inference stage.

2. The model training optimization method according to claim 1, characterized in that: The forward propagation and model training of the reinforcement learning model based on the sample data obtained in the reasoning stage includes: In the forward propagation stage, a first resource optimization strategy is adopted to forward propagate the reinforcement learning model based on the sample data to obtain a forward propagation result of the sample data; During the model training phase, a second resource optimization strategy is adopted to update the model parameters of the reinforcement learning model according to the forward propagation results.

3. The model training optimization method according to claim 2, characterized in that: The reinforcement learning model includes an execution model, a reference model, a critic model and a reward model; the first resource optimization strategy includes putting the model weight of the currently used model into the video memory, and unloading the model weight of the currently unused model from the video memory and putting it into the internal memory; The adopting the first resource optimization strategy to forward propagate the reinforcement learning model based on the sample data includes: In the process of forward propagating the reinforcement learning model based on the sample data, when the execution model is used, the model weights of the execution model are placed in the video memory, and the model weights of the reference model are unloaded from the video memory and placed in the internal memory; when the reference model is used, the model weights of the reference model are placed in the video memory, and the model weights of the execution model are unloaded from the video memory and placed in the internal memory; When the critic model is used, the model weights of the critic model are placed in the video memory, and the model weights of the reward model are unloaded from the video memory and placed in the internal memory; when the reward model is used, the model weights of the reward model are placed in the video memory, and the model weights of the critic model are unloaded from the video memory and placed in the internal memory.

4. The model training optimization method according to claim 2, characterized in that: The second resource optimization strategy includes placing model parameters of the currently used model into the video memory, and unloading model parameters of the currently unused model from the video memory and placing them into the internal memory; The adopting the second resource optimization strategy to update the model parameters of the reinforcement learning model according to the forward propagation result includes: In the process of updating the model parameters of the reinforcement learning model according to the forward propagation result, when the execution model is updated, the model parameters of the critic model are unloaded from the video memory and placed in the internal memory; when the critic model is updated, the model parameters of the execution model are unloaded from the video memory and placed in the internal memory; The execution model and the critic model are models whose parameters need to be updated in the reinforcement learning model.

5. The model training optimization method according to claim 2, characterized in that: The second resource optimization strategy includes unloading gradient data and optimizer state parameters from the video memory; The adopting the second resource optimization strategy to update the model parameters of the reinforcement learning model according to the forward propagation result includes: In the model training phase, gradient data is constructed according to the forward propagation result, model parameters of the reinforcement learning model are updated using an optimizer based on the gradient data, the training step of the current round is completed, and the gradient data is unloaded from the video memory and transferred to the internal memory; Before using the optimizer to update the model parameters of the reinforcement learning model, the optimizer state parameters of the optimizer are loaded into the video memory; after using the optimizer to update the model parameters of the reinforcement learning model, the optimizer state parameters are unloaded from the video memory and written into the internal memory.

6. The model training optimization method according to any one of claims 1 to 5, characterized in that: After obtaining the reinforcement learning model, the method further includes: After the reasoning stage is completed, if the machine resources do not meet the restriction conditions, in the forward propagation stage and the model training stage, other machine resources are used to perform forward propagation and model training on the reinforcement learning model based on the sample data obtained in the reasoning stage, wherein the other machine resources are different from the machine resources used in the reasoning stage.

7. The model training optimization method according to any one of claims 1 to 5, characterized in that: The machine resource is a video memory; and the unloading of the machine resource used in the inference phase includes: The model weights and / or key-value cache of the reinforcement learning model in the inference phase are unloaded from the video memory and placed into the internal memory.

8. A model training optimization device, characterized in that: include: An acquisition module is configured to acquire a reinforcement learning model, wherein the reinforcement learning process of the reinforcement learning model includes an inference phase, a forward propagation phase, and a model training phase, and the inference phase, the forward propagation phase, and the model training phase are performed in series; The first training module is configured to unload the machine resources used in the reasoning stage after the reasoning stage ends, reuse the machine resources used in the reasoning stage in the forward propagation stage and the model training stage, and perform forward propagation and model training on the reinforcement learning model based on the sample data obtained in the reasoning stage.

9. A computing device, characterized in that include: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the model training optimization method described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that: It stores computer executable instructions, which, when executed by a processor, implement the steps of the model training optimization method described in any one of claims 1 to 7.

11. A computer program product, characterized in that It includes a computer program / instruction, which, when executed by a processor, implements the steps of the model training optimization method described in any one of claims 1 to 7.