Model training system and method and electronic equipment

Through the asynchronous distributed RLHF framework and adaptive resource coordination mechanism, the problems of low sampling efficiency and low resource utilization in large-scale language model training are solved, efficient training and inference parallelism is achieved, and the efficiency of model training and alignment effect are improved.

CN120633753APending Publication Date: 2025-09-12HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410650697.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-03-12
Filing Date
2024-05-21
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing large language model (LLM) training methods suffer from low sampling efficiency, high coupling between training and inference processes, and low resource utilization, making them difficult to effectively expand and optimize.

Method used

Adopting the asynchronous distributed RLHF framework, the parallel operation of training nodes and inference nodes is achieved through asynchronous communication and coordination. Combined with the adaptive resource coordination mechanism and timestamp management, the computing resource allocation and sample management are optimized, thus improving the training efficiency and alignment effect.

Benefits of technology

It improves the training efficiency and alignment effect of large language models, improves resource utilization and system throughput, and supports stable convergence of large-scale distributed cluster training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120633753A_ABST
    Figure CN120633753A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a model training system and method and electronic equipment. The model training system comprises a reasoning module used for generating a training sample used for model training; the experience sample pool is used for storing the training samples; and the training module is used for carrying out model training according to the training samples stored in the experience sample pool, and the reasoning module and the training module operate asynchronously and parallelly. According to the model training system provided by the embodiment of the invention, an asynchronous framework in which the training module and the reasoning module are separated is adopted, so that the training process and the reasoning process are decoupled, good expandability is achieved, the problem of low sampling efficiency can be solved, and the training efficiency of model training is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a model training system, method and electronic device. Background Art

[0002] Large Language Models (LLMs) are widely used in the field of artificial intelligence. LLMs lay the blueprint for the development of general artificial intelligence.

[0003] LLM relies on supervised training to achieve alignment, but the supervised training paradigm requires a large amount of labeled data, which is costly and difficult to guarantee in terms of labeling quality and efficiency. Therefore, a training method for LLM is needed to improve its training efficiency and alignment effect. Summary of the Invention

[0004] In order to improve the training efficiency and alignment effect of LLM, the present application provides a model training system, method and electronic device, and also provides a computer-readable storage medium.

[0005] The embodiments of this application adopt the following technical solutions:

[0006] In a first aspect, the present application provides a model training system, the model training system comprising:

[0007] The inference module is used to generate training samples for model training;

[0008] Experience sample pool, which is used to store training samples;

[0009] The training module is used to train the model based on the training samples stored in the experience sample pool, wherein the inference module and the training module run asynchronously and in parallel.

[0010] According to the model training system of the first aspect, an asynchronous architecture is adopted in which the training module and the inference module are separated, so that the training and inference processes are decoupled, and existing optimization technologies can be flexibly applied to training and inference respectively. It has good scalability, can effectively solve the problem of low sampling efficiency, and improve the training efficiency of model training.

[0011] In one implementation of the first aspect, the model training system is used to implement a reinforcement learning fine-tuning (RLHF) framework based on human feedback.

[0012] In an implementation of the first aspect, the reasoning module includes multiple parallel reasoning nodes, and / or the training module includes multiple parallel training nodes.

[0013] In an implementation of the first aspect, the model system further includes:

[0014] A coordination controller is used to adjust the computing resource allocation of each component in the reasoning module according to the throughput of the reasoning module and the training module, and / or adjust the computing resource allocation of the reasoning module and the training module.

[0015] According to the above-mentioned implementation method of the first aspect, an adaptive resource coordination mechanism is designed for the model system, which provides the function of dynamically adjusting the ratio of computing resources within the inference node and the ratio of computing resources between the training and inference nodes. This can avoid waiting for training resources and waste of inference data, and effectively improve training throughput and resource utilization.

[0016] In one implementation of the first aspect, the reasoning module includes a reference model, a reward model, and a first policy model, and the training module includes a second policy model;

[0017] The coordination controller is used to optimize and solve the problem with the goal of maximizing the total training throughput and inference throughput based on the training throughput of the second strategy model, the inference throughput of the first strategy model, the reference model, and the reward model under unit resources, and obtain the computing resource allocation ratio of each model in the inference module, as well as the computing resource allocation ratio between the inference module and the training module.

[0018] In an implementation of the first aspect:

[0019] The inference module is also used to add timestamps to training samples;

[0020] The experience sample pool is also used to detect the timestamp of the training sample. When the timestamp of the training sample exceeds a preset threshold, the training sample is deleted.

[0021] According to the above-mentioned implementation method of the first aspect, timestamps are added to the training samples, and the experience sample pool automatically eliminates expired training samples based on the timestamps of the training samples, thereby limiting the version difference between the model for generating samples and the training sample model, ensuring the convergence ability of model training under large-scale distributed cluster training, and improving the alignment effect.

[0022] In an implementation of the first aspect:

[0023] The inference module is also used to add timestamps to training samples;

[0024] The training module is also used to continuously sample training samples from the experience sample pool, check the differences between the training samples and the current training strategy model version, and discard the training samples when the difference exceeds a preset threshold.

[0025] According to the above-mentioned implementation method of the first aspect, a timestamp is added to the training sample, and the training module automatically compares the version timestamp of the training sample with that of the current model when obtaining the training sample, thereby limiting the version difference between the model of the generated sample and the training sample model, ensuring the model training convergence ability under large-scale distributed cluster training, and improving the alignment effect.

[0026] In an implementation of the first aspect, the reasoning module is deployed on homogeneous or heterogeneous computing resources.

[0027] In one implementation of the first aspect, the reasoning module includes a reference model, a reward model, and a first strategy model, wherein:

[0028] The reference model and the reward model are deployed on the first type of computing resources, and the first strategy model is deployed on the second type of computing resources.

[0029] In one implementation of the first aspect, the reasoning module includes a reference model, a reward model, and a first strategy model, wherein:

[0030] The reference model is deployed on the first type of computing resources, the reward model is deployed on the second type of computing resources, and the first strategy model is deployed on the third type of computing resources.

[0031] According to the above implementation of the first aspect, the inference module adopts a heterogeneous framework, which can effectively utilize different types of computing resources, thereby avoiding waste of cloud resources.

[0032] In an implementation of the first aspect, the training module is deployed on homogeneous computing resources.

[0033] In a second aspect, the present application provides a model training method, which is applied to a model system. The model system includes an inference module, an experience sample pool, and a training module. The method includes:

[0034] Use the inference module to generate training samples for model training;

[0035] Save the training samples generated by the inference module to the experience sample pool;

[0036] The training module is used to train the model based on the training samples stored in the experience sample pool, wherein the inference module and the training module run asynchronously and in parallel.

[0037] In an implementation of the second aspect, the model system is used to implement the RLHF framework.

[0038] In an implementation of the second aspect, the method further includes:

[0039] According to the throughput of the inference module and the training module, the computing resource allocation of each component in the inference module is adjusted, and / or the computing resource allocation of the inference module and the training module is adjusted.

[0040] In an implementation of the second aspect, the reasoning module includes a reference model, a reward model, and a first policy model, and the training module includes a second policy model;

[0041] Adjusting the allocation of computing resources to components within the inference module and / or the allocation of computing resources to the inference module and the training module based on the throughput of the inference module and the training module, including:

[0042] Based on the training throughput of the second strategy model, the inference throughput of the first strategy model, the reference model, and the reward model under unit resources, an optimization solution is performed with the goal of maximizing the total training throughput and inference throughput. The computing resource allocation ratio of each model in the inference module and the computing resource allocation ratio between the inference module and the training module are obtained.

[0043] In an implementation of the second aspect, the method further includes:

[0044] Deploy an inference module, wherein the inference module includes multiple parallel inference nodes;

[0045] and / or,

[0046] Deploy a training module, where the training module includes multiple parallel training nodes.

[0047] In an implementation of the second aspect, the method further includes:

[0048] Deploy the reasoning module, where the reasoning module is deployed on homogeneous or heterogeneous computing resources.

[0049] In an implementation of the second aspect, the reasoning module includes a reference model, a reward model, and a first strategy model;

[0050] Deploy the inference module, including:

[0051] The reference model and the reward model are deployed on the first type of computing resources, and the first strategy model is deployed on the second type of computing resources.

[0052] In an implementation of the second aspect, the reasoning module includes a reference model, a reward model, and a first strategy model;

[0053] Deploy the inference module, including:

[0054] The reference model is deployed on the first type of computing resources, the reward model is deployed on the second type of computing resources, and the first strategy model is deployed on the third type of computing resources.

[0055] In an implementation of the second aspect, the method further includes:

[0056] Deploy the training module, where the training module is deployed on homogeneous computing resources.

[0057] In an implementation of the second aspect, the method further includes:

[0058] Before saving the training sample into the experience sample pool, add a timestamp to the training sample;

[0059] The timestamps of the training samples in the experience sample pool are detected, and when the timestamps of the training samples exceed a preset threshold, the training samples are deleted from the experience sample pool.

[0060] In an implementation of the second aspect, the method further includes:

[0061] Before saving the training sample into the experience sample pool, add a timestamp to the training sample;

[0062] Continuously sample training samples from the experience sample pool to perform model training fine-tuning. The difference between the training samples and the current training strategy model version is checked, and the training samples are discarded when the difference exceeds a preset threshold.

[0063] In a third aspect, the present application provides an electronic device, comprising a memory for storing computer program instructions and a processor for executing computer program instructions, wherein when the computer program instructions are executed by the processor, the electronic device is triggered to execute the method steps described in the second aspect.

[0064] In a fourth aspect, the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, which, when executed on a computer, enables the computer to execute the method described in the second aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] Figure 1 Shown is a schematic diagram of the structure of a model system according to an embodiment of the present application;

[0066] Figure 2 FIG2 is a flow chart of a model creation system according to an embodiment of the present application;

[0067] Figure 3 Shown is a schematic diagram of the structure of a model system according to an embodiment of the present application;

[0068] Figure 4 FIG2 is a flow chart of the RLHF method according to an embodiment of the present application;

[0069] Figure 5FIG. 1 is a schematic structural diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0070] To make the purpose, technical solutions, and advantages of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with the specific embodiments of this application and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0071] The terms used in the implementation section of this application are only used to explain the specific embodiments of this application and are not intended to limit this application.

[0072] The primary goal of Large Language Models (LLMs) is to train a human-centric (honest, useful, and harmless) intelligent assistant based on a large amount of corpus data. Therefore, it is crucial that LLMs align with human preferences.

[0073] Based on the above principles, a feasible solution for LLM training is to adopt the reinforcement learning fine-tuning method based on human feedback (RLHF).

[0074] Reinforcement learning is a paradigm and methodology in machine learning that describes and solves the problem of how intelligent agents learn strategies to maximize rewards or achieve specific goals during their interactions with their environment. RLHF, or reinforcement learning with human feedback, is a technique for aligning large language models with human preferences. This involves training a reward model based on human preference data to model human inclinations to respond to the model. Using reinforcement learning algorithms, the language model is optimized to maximize cumulative rewards, thereby aligning the language model output with human preferences.

[0075] RLHF leverages reward models to learn value orientations from human preference data, guiding the fine-tuning of language models. It is a key technical paradigm for achieving human preference alignment. RLHF leverages human feedback, abstracts subjective preferences, and automatically optimizes models to generate higher-quality text that better aligns with human preferences and mainstream values. This significantly improves the training efficiency and alignment of large language models.

[0076] RLHF offers significant advantages in aligning large language models, but it also introduces numerous challenges. First, sampling efficiency is low. The RLHF sampling process involves training and inference of multiple large models, including policy, value, reference, and reward models. These models interact with each other in serial logic, causing device latency. Second, training efficiency is low. The frequent training and inference switching of reinforcement learning results in training sample size being constrained by sampling efficiency. Furthermore, the encroachment of multiple large models on graphics memory resources limits training throughput.

[0077] To solve the above-mentioned problems of RLHF, a feasible solution is to optimize the RLHF framework from multiple aspects, such as video memory optimization and efficient switching between training and pushing.

[0078] For example, the trlX framework is a serial iterative solution for training and inference. It is a distributed RLHF framework built around the HuggingFace ecosystem. The trlX framework's training process is based on HuggingFace's Accelerate backend and supports distributed training of models up to 20B.

[0079] Another example is the ColossalChat framework, a serial iterative approach to training and inference. It uses Low-Rank Adaptation (LoRA) and Zero Redundancy Optimizer (ZeRO) + Gemini technologies to improve memory efficiency, reduce communication frequency, and offload optimizer state from graphics processing unit (GPU) memory to central processing unit (CPU) memory or hard disk space, enabling training larger models with fewer computing resources.

[0080] Another example is the DeepSpeedChat framework, a serial iterative approach to training and inference. It proposes a hybrid training-inference engine that seamlessly transitions between inference and training modes, leveraging DeepSpeed ​​optimizations for both training (such as ZeRO and LoRA) and inference (such as tensor parallelism and high-performance operators).

[0081] However, the above optimization scheme has the following problems:

[0082] 1) Poor scalability: The above framework does not fully consider mechanisms such as distributed training, making it difficult to effectively scale to multiple model sizes and computing resources;

[0083] 2) Serial training and inference iteration: The training and inference processes of the above framework are highly coupled, which limits the use of inference optimization techniques;

[0084] 3) Low resource utilization: There is a large difference in RLHF training and inference throughput. The above framework does not consider the optimal allocation and heterogeneous combination of training and inference resources.

[0085] In response to the above problems, an embodiment of the present application provides a model system for implementing a reinforcement learning fine-tuning (RLHF) method based on human feedback (RLHF framework).

[0086] Specifically, in one embodiment of the present application, the model system adopts asynchronous distributed RLHF.

[0087] Asynchronous distribution refers to the use of asynchronous communication and coordination to achieve efficient parallel processing across multiple nodes. Asynchronous distribution can improve system throughput and scalability while reducing performance degradation caused by communication bottlenecks or synchronization delays.

[0088] In one embodiment of the present application, the model system adopts an asynchronous distributed RLHF framework, that is, in the model system, the training nodes and inference nodes of the RLHF framework run independently and in parallel without synchronous communication and waiting.

[0089] Specifically, in one embodiment, the RLHF framework in the model system includes multiple parallel training nodes and multiple parallel inference nodes. The training nodes and the inference nodes run independently and in parallel without synchronous communication and waiting.

[0090] Figure 1 Shown is a schematic diagram of the model system structure according to an embodiment of the present application.

[0091] like Figure 1 As shown, in one embodiment, the RLHF framework of the model system includes an inference module 110 , a training module 120 (including multiple parallel training nodes), and an experience buffer 140 .

[0092] The inference module 110 is used to implement the self-sampling inference process of the RLHF. The inference module 110 includes parallel inference nodes, which are responsible for self-sampling and generating sample data used for model training.

[0093] The reasoning module 110 includes a reference model (Ref Model) 111 , a reward model (RM) 112 , and a policy model (Policy Model) 113 .

[0094] Reference model 111 refers to the original language model used to initialize the policy model in RLHF. It is used to provide a reference answer for a given user question, which is compared with the model response generated by policy model 113 to limit the range of change of policy model 113.

[0095] Reward model 112, RLHF is a reward-guided policy optimization process. The reward model 112 takes {user question, model output} as input and outputs a scalar score, which is used to evaluate the quality of the response generated by the policy model 113.

[0096] The policy model 113 is the large language model to be optimized in RLHF. It takes user questions as input and outputs model responses. The ultimate goal of RLHF is to optimize the policy model so that it outputs high-quality text that aligns with human preferences.

[0097] In the inference stage of the inference module 110, user questions are sampled from the data set and input into the strategy model 113. Forward reasoning is performed to obtain question-answer pairs, which are then transmitted to the reference model 111 and the reward model 112 to infer and calculate the KL penalty and reward values, respectively, to obtain training samples, which are then stored in the experience sample pool 140.

[0098] The training module 120 is used for the RLHF fine-tuning training process. The training module 120 includes multiple parallel training nodes for training and fine-tuning the model based on the reinforcement learning method.

[0099] During the training phase, the training module 120 loads the policy model, continuously samples from the experience sample pool 140 , and performs model training fine-tuning based on a reinforcement learning algorithm (such as the Proximal Policy Optimization (PPO) algorithm).

[0100] PPO is an actor-critic based reinforcement learning algorithm used to train policy models to complete complex tasks.

[0101] The experience sample pool 140 is stored in the computing resource memory, and is used to receive the sample data constructed by the reasoning module 110 and provide the training module 120 with the training data required for RLHF fine-tuning.

[0102] The asynchronous distributed RLHF framework provided in one embodiment of the present application adopts an asynchronous heterogeneous multi-node cluster architecture in which the training module 120 is separated from the inference module 110 (training and inference separation), has good scalability, effectively solves the problem of low sampling efficiency of large-model RLHF methods, and improves the training efficiency of model training.

[0103] According to one embodiment of the present application, the RLHF framework adopts an asynchronous distributed framework, which decouples the RLHF training and reasoning processes, and adopts an asynchronous parallel operation mode of training nodes and reasoning nodes. It has good scalability and can flexibly apply existing optimization technologies for training and reasoning respectively.

[0104] Optionally, in one embodiment, the RLHF framework of the model system adopts a heterogeneous distributed framework. A heterogeneous distributed framework refers to a framework composed of nodes of different types or platforms, and the heterogeneous distributed framework simultaneously utilizes the respective advantages of nodes of different types or platforms to perform efficient parallel computing.

[0105] Specifically, in one embodiment, the inference module 110 is compatible with homogeneous or heterogeneous computing resource configurations. For example, the inference node computing resources of the RLHF framework can simultaneously use GPUs and neural network processing units (NPUs) for parallel computing.

[0106] In one embodiment, in the inference module 110 , the policy model 113 is independently deployed on a computing resource, and the reference model 111 and the reward model 112 are deployed together on the same computing resource.

[0107] According to one embodiment of the present application, the RLHF framework adopts a heterogeneous distributed framework, which can effectively utilize different types of computing resources (such as GPU, NPU, CPU), thereby avoiding waste of cloud resources.

[0108] Furthermore, in one embodiment, the model system further includes a coordination controller (Controller) 130. The coordination controller (Controller) 130 is used to adjust the resource usage of each component of the RLHF framework in the model system.

[0109] Specifically, the coordination controller 130 adaptively and dynamically adjusts the computing resource allocation of each component in the reasoning module 110 according to the throughput of the training module 120 and the reasoning module 110 (for example, adjusts the computing resource allocation ratio of the reference model 111, the reward model 112, and the strategy model 113), and adjusts the computing resource allocation of the training module 120 and the reasoning module 110 to achieve a dynamic balance of data flow, ensure that the training module 120 and the reasoning module 110 run at full load, and maximize the utilization of computing resources.

[0110] The asynchronous distributed RLHF framework provided in one embodiment of the present application is designed with an adaptive resource coordination mechanism, which provides the function of dynamically adjusting the computing resource ratio within the inference node and the computing resource ratio between the training and inference nodes, effectively improving the RLHF training throughput and resource utilization.

[0111] Figure 2 Shown is a flow chart of a model creation system according to an embodiment of the present application.

[0112] In one embodiment, the electronic device performs the following Figure 2 The following process is shown to build and train Figure 1 The RLHF framework shown.

[0113] S210, constructing the inference module 110 based on the available computing resources, and setting the maximum available resource threshold N of the inference module 110 infer .

[0114] S220: Build a training module 120 based on available computing resources and set a maximum available resource threshold N for the training module 120. train .

[0115] S230 , constructing a coordination controller 130 , and using the coordination controller 130 to calculate the optimization of the resource allocation between the reasoning module 110 and the training module 120 .

[0116] Specifically, the coordination controller 130 trains the throughput according to the policy model under the unit resource Policy model inference throughput Reference / reward model inference throughput An optimization solution is performed with the goal of maximizing the total training throughput and inference throughput to obtain the optimal computing resource ratio between the policy model and the reference / reward model in the inference module 110, as well as the optimal computing resource ratio between the training module 120 and the inference module 110.

[0117] In the description of the embodiments of this application, training throughput refers to the maximum number of samples or word units that can be processed per unit time. Inference throughput refers to the maximum number of samples or word units that can be inferred per unit time.

[0118] Preferably, in one embodiment, the coordination controller 130 performs an optimization solution for computing resource allocation, where the optimization goal is to maximize the training and pushing throughput of the RLHF framework under given training and pushing resources.

[0119] Specifically, the optimization function is established as follows:

[0120]

[0121] In formula 1, Represents the number of training resources, Represents the number of strategy model inference resources, Represents the number of reference / reward model reasoning resources. train With T infer are the total throughput of the training module and the inference module respectively.

[0122] To avoid resource idleness caused by the training module waiting for the inference module, the total training throughput must not exceed the total inference throughput. The specific constraints are as follows:

[0123]

[0124] In formula 2, and They are the single-card inference throughput and number of cards of the strategy model, and The single-card inference throughput and number of cards for the reference / reward model, respectively. i The minimum number of compute cards required for model inference.

[0125] S240 , starting the training job according to the training and inference resource ratio optimized and solved by the coordination controller 130 , and monitoring the throughput of the inference module 110 and the training module 120 in real time.

[0126] Specifically, when the network status fluctuates greatly, the coordination controller 130 can dynamically optimize the resource allocation and quantity of the inference module 110 and the training module 120.

[0127] According to one embodiment of the present application, the RLHF framework adopts adaptive resource coordination to dynamically adjust the ratio of computing resources within the inference node, optimize the throughput of the strategy model and the reference / reward model; dynamically adjust the computing resource ratio between training and inference nodes to avoid waiting for training resources and waste of inference data.

[0128] Optionally, in one embodiment, a timestamp is added to the training samples generated during the inference phase of the inference module 110 .

[0129] The training samples with timestamps added are stored in the experience sample pool 140. Sample data format:

[0130] x S ={prompt,response,logprob,value,reward,timestamp}. (Formula 3)

[0131] In one embodiment, after the generation time of the sample data in the experience sample pool 140 exceeds a threshold Staleness (eg, 10 training cycles), the experience sample pool 140 will automatically delete the data.

[0132] According to one embodiment of the present application, the RLHF framework uses timestamps, and the experience sample pool automatically eliminates expired training samples based on the timestamps, thereby limiting the version difference between the model for generating samples and the training sample model, ensuring the RLHF training convergence capability under large-scale distributed cluster training, and improving the alignment effect.

[0133] In another embodiment, the training module 120 continuously samples training samples from the experience sample pool 140, checks the degree of difference between the timestamp of the training sample and the currently trained policy model version, and directly discards the data if it exceeds a given threshold Staleness (for example, 10 version iteration cycles).

[0134] According to one embodiment of the present application, the RLHF framework uses timestamps. When the training module obtains training samples, it automatically compares the version timestamps of the training samples with those of the current model, limits the version differences between the model of the generated samples and the training sample model, ensures the convergence ability of RLHF training under large-scale distributed cluster training, and improves the alignment effect.

[0135] In another embodiment, if the sample data in the experience sample pool 140 exceeds a staleness threshold (e.g., 10 training cycles), the experience sample pool 140 automatically deletes the data. Furthermore, the training module 120 continuously samples training samples from the experience sample pool 140 and verifies the degree of difference between the timestamp of the training sample and the currently trained policy model version. If the staleness exceeds a given threshold (e.g., 10 version iteration cycles), the data is directly discarded.

[0136] The RLHF framework proposed in the embodiment of the present application can be applied to RLHF training scenarios under single computing resources (GPU or NPU only) and hybrid computing resources (GPU+NPU) or (GPU+NPU+CPU) on the cloud. The present invention is described in detail below with reference to the LLM embodiment with a size of tens of billions of parameters. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be pointed out that for those of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present invention. These all fall within the scope of protection of the present invention.

[0137] Figure 3 Shown is a schematic diagram of the model system structure according to an embodiment of the present application.

[0138] like Figure 3 As shown, in one embodiment, the model system is composed of a reasoning module 310 , a training module 320 , an experience sample pool 340 and a coordination controller 330 .

[0139] The experience sample pool 340 is stored in the computing resource memory and is used to store reinforcement learning experience samples.

[0140] During the inference process, the inference module 310 generates training data from sampling and stores it in the experience sample pool 340 .

[0141] The experience sample pool 340 of the reasoning module 310 receives the sample data constructed by the reasoning module 310 and provides the training module 320 with the training data required for RLHF fine-tuning.

[0142] like Figure 3 As shown, the reasoning module 310 includes a policy model 313 , a reference model 311 , and a reward model 312 .

[0143] The inference process of the inference module 310 is as follows: the policy model 313 generates responses based on the prompts from the original sample pool 314 through autoregression. The prompts and responses are transmitted to the reference model 311 and the reward model 312 to calculate the KL penalty and reward values, respectively. The generated samples {prompt, response, logprob, value, reward, timestamp} are stored in the experience sample pool 340.

[0144] The training module 320 includes a policy model 321 and a setting model (SET Model) 322 .

[0145] The training module 320 uses a supervised fine-tuning model (SFTModel) 322 to initialize the policy model 321 .

[0146] The training module 320 continuously samples training samples from the experience sample pool 340 and checks the degree of difference between the sample timestamp and the currently trained policy model version. If the difference exceeds a given threshold Staleness=2, the data is directly discarded.

[0147] Furthermore, there will be certain differences between the policy models (policy model 313 and policy model 321) of the training module 320 and the reasoning module 310 in the distributed parallel RLHF framework. The timestamp in the sample stored in the experience sample pool 340 is the timestamp of the policy model 313 in the current reasoning module 310, which is used to record the source of the sampling sample.

[0148] In a single computing resource scenario, the training module 320 and the inference module 310 both use the same hardware resources, such as a GPU or an NPU.

[0149] In a hybrid computing resource scenario, the training module 320 uses a single type of computing resource, such as a GPU or NPU, while the inference module 310 can use multiple types of computing resources, such as a GPU+NPU+CPU.

[0150] Specifically, when constructing the RLHF framework, one or more computing resources are selected for the reasoning module 310. Based on the communication problem of parallel training, one computing resource is selected for the training module 320.

[0151] For example, considering the memory usage and computing requirements of exascale LLM inference, Huawei Cloud offers cloud computing resources including NVIDIA A800 GPUs and Ascend D910 NPUs. The minimum inference resource for both in exascale models is a single card. The Ascend D910 NPU is Huawei's proprietary AI computing chip. The NVIDIA A800 GPU is a graphics processing unit developed by NVIDIA. One or more computing resources can be selected for the inference module. Due to communication issues in parallel training, only one computing resource can be selected for the training module.

[0152] Optionally, in one embodiment, the policy model 313 exclusively occupies computing resources, and the reward model 312 shares computing resources with the reference model 311. In another embodiment, the reward model 312 and the reference model 311 exclusively occupies computing resources.

[0153] The coordination controller 330 is deployed on the CPU. Based on the throughput of the training module 320 and the inference module 310, the coordination controller 330 adaptively and dynamically adjusts the ratio of computing resources within the inference module 310 and between the training module 320 and the inference module 310 (adjusting the ratio of inference and training nodes). This achieves dynamic balancing of data flow, ensures that the training module 320 and the inference module 310 operate at full capacity, and maximizes computing resource utilization.

[0154] Specifically, in one embodiment, within the inference module 310 , the total throughput of the policy model 313 is required to be less than or equal to the total throughput of the reference model 311 / reward model 312 , referring to Formula 2, that is:

[0155]

[0156] In formula 4, and They are the single-card inference throughput and number of cards of the strategy model, and They are the single-card inference throughput and number of cards for the reference / reward model, and are the total inference throughput of the policy model and the reference / reward model, respectively.

[0157] For the training module 320 and the inference module 310, the total training throughput is required to be no greater than the total inference throughput, that is:

[0158]

[0159] In formula 5, T train With T infer Represent the total throughput of the training module and the inference module respectively, and They are the single-card throughput and number of cards for model training, respectively.

[0160] The optimization goal of the coordination controller 330 is to maximize the training throughput of the RLHF framework under given training resources and calculate the optimal number of strategy model inference resources. Reference / reward model reasoning resources and the number of training resources Refer to Formula 1.

[0161] Figure 4 FIG2 is a flow chart of a RLHF method according to an embodiment of the present application.

[0162] S410, select the required inference computing resource type and set the upper limit of the resource quantity (refer to S210).

[0163] S420, select the required training computing resource type and set the upper limit of the resource quantity (refer to S220).

[0164] S430: The RLHF is tested with the minimum training and pushing unit, and the coordination controller calculates the initial resource allocation according to the training and pushing throughput (refer to S230).

[0165] S440: Adjust the training nodes according to the calculated resource allocation plan, and continuously sample from the experience sample pool to fine-tune the model.

[0166] Specifically, in S440 , the initial resource allocation calculated in S430 is initially used.

[0167] S450, the coordination controller calculates the optimal resource ratio of the current training and inference nodes according to the node search throughput (refer to S230).

[0168] After S450 , the process returns to S440 , where the optimal resource allocation calculated in S450 is used in S440 .

[0169] The RLHF framework provided in the embodiments of the present application is not only applicable to large-model RLHF scenarios, but also to general reinforcement learning training scenarios under large-model scenarios, so as to realize flexible expansion of training and inference node resources in reinforcement learning and improve the efficiency of reinforcement learning training.

[0170] Those skilled in the art will appreciate that, in addition to implementing the system and its various devices provided by the present invention in purely computer-readable program code, it is entirely possible to implement the same functions of the system and its various devices provided by the present invention in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, embedded microcontrollers, etc. by logically programming the method steps. Therefore, the system and its various devices provided by the present invention can be considered a hardware component, and the devices included therein for implementing the various functions can also be considered as structures within the hardware component; the devices for implementing the various functions can also be considered as both software modules implementing the method and structures within the hardware component.

[0171] Disclosed herein are only preferred embodiments of the present invention. These embodiments are selected and described in detail in this specification to better explain the principles and practical applications of the present invention and are not intended to limit the present invention. Any modifications and variations made by those skilled in the art within the scope of this specification are intended to fall within the scope of protection of the present invention.

[0172] In the description of the embodiments of the present application, for the convenience of description, the description of the coloring device is divided into various modules according to their functions. The division of each module is only a division of logical functions. When implementing the embodiments of the present application, the functions of each module can be implemented in the same or one or more software and / or hardware.

[0173] Specifically, the device proposed in the embodiment of the present application can be fully or partially integrated into a physical entity (for example, a GPU or other type of processor) during actual implementation, or it can be physically separated. And these modules can all be implemented in the form of software calling through a processing element; they can also all be implemented in the form of hardware; some modules can also be implemented in the form of software calling through a processing element, and some modules can be implemented in the form of hardware. For example, the detection module can be a separately established processing element, or it can be integrated in a chip of an electronic device. The implementation of other modules is similar. In addition, these modules can be fully or partially integrated together, or they can be implemented independently. During the implementation process, each step of the above method or each of the above modules can be completed by an integrated logic circuit of hardware in the processor element or an instruction in the form of software.

[0174] For example, the above modules may be one or more integrated circuits configured to implement the above methods, such as one or more application-specific integrated circuits (ASICs), one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs). For another example, these modules may be integrated together and implemented in the form of a system-on-a-chip (SOC).

[0175] An embodiment of the present application further provides an electronic device, which is used to execute the method flow or part of the method flow described in the embodiment of the present application.

[0176] Figure 5 FIG. 1 is a schematic structural diagram of an electronic device according to an embodiment of the present application.

[0177] like Figure 5 As shown, the electronic device 2500 includes a memory 2502 for storing computer program instructions and a processor 2501 for executing program instructions, wherein, when the computer program instructions are executed by the processor 2501, the electronic device 2500 is triggered to execute the method steps as described in the embodiments of the present application.

[0178] Specifically, in one embodiment of the present application, the above-mentioned one or more computer programs are stored in the above-mentioned memory 2502, and the above-mentioned one or more computer programs include instructions. When the above-mentioned instructions are executed by the above-mentioned electronic device 2500, the above-mentioned electronic device 2500 executes the method steps described in the embodiment of the present application.

[0179] It is understood that the structural description of the electronic device 2500 in the embodiment of the present application does not constitute a specific limitation on the electronic device 2500. In other embodiments of the present application, the electronic device 2500 may include other components besides the processor 2501 and the memory 2502.

[0180] The processor 2501 may be a device on a chip (SOC), and the processor 2501 may include a central processing unit (CPU), and may further include other types of processors.

[0181] The processor involved in processor 2501 may include, for example, a CPU, a DSP, a microcontroller, or a digital signal processor, and may also include a GPU, an embedded neural network processor (NPU), and an image signal processor (ISP). The processor may also include necessary hardware accelerators or logic processing hardware circuits, such as ASICs, or one or more integrated circuits for controlling the execution of the program of the technical solution of this application. In addition, the processor may have the function of operating one or more software programs, and the software programs may be stored in a storage medium.

[0182] The processor 2501 may include one or more processing units. For example, the processor may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Different processing units may be independent components or integrated into one or more processors. In some embodiments, the electronic device 2500 may also include one or more processors 2501. The controller may generate an operation control signal based on the instruction opcode and the timing signal to complete the control of instruction fetching and execution.

[0183] In some embodiments, the processor 2501 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a SIM card interface, and / or a USB interface, etc. Among them, the USB interface is an interface that complies with the USB standard specification, and specifically can be a Mini USB interface, a Micro USB interface, a USB Type C interface, etc. The USB interface can be used to connect a charger to charge the electronic device, and can also be used to transmit data between the electronic device and peripheral devices.

[0184] Electronic device 2500 may also include an external memory interface for connecting an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device. The external memory card communicates with processor 2501 via the external memory interface to implement data storage. For example, files such as music and videos can be stored on the external memory card.

[0185] Memory 2502 may include a code storage area and a data storage area. The code storage area may store an operating system. The data storage area may store data created during the use of electronic device 2500. Furthermore, memory 2502 may include high-speed random access memory and non-volatile memory, such as one or more disk storage components, flash memory components, and universal flash storage (UFS).

[0186] The memory 2502 may be a read-only memory (ROM), other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), a magnetic disk storage medium or other magnetic storage device, or any computer-readable medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer.

[0187] The processor 2501 and the memory 2502 may be combined into one processing device, or more commonly, they may be independent components.

[0188] One embodiment of the present application further provides an electronic chip. The electronic chip is used to execute the method flow or part of the method flow described in the embodiment of the present application. For example, the electronic chip can be a GPU, a CPU, or an NPU.

[0189] Specifically, the electronic chip includes a processor for executing program instructions. When the computer program instructions are executed by the processor, the electronic chip is triggered to execute the method steps described in the embodiments of the present application. The processor of the electronic chip can refer to the processor of the above-mentioned electronic device.

[0190] Optionally, the devices, apparatuses, and modules described in the embodiments of the present application may be implemented by computer chips or entities, or by products having certain functions.

[0191] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, apparatus, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code.

[0192] In the several embodiments provided in this application, if any function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, or the part that contributes to the prior art, or part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of this application.

[0193] Specifically, an embodiment of the present application further provides a computer-readable storage medium, which stores a computer program. When the computer-readable storage medium is run on a computer, the computer executes the method provided in the embodiment of the present application.

[0194] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program product is run on a computer, it enables the computer to execute the method provided in the embodiment of the present application.

[0195] The description of the embodiments in this application is described with reference to the flowcharts and / or block diagrams of the methods, devices (apparatus), and computer program products according to the embodiments of the application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0196] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0197] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0198] It should also be noted that, in the embodiments of the present application, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent the existence of A alone, the existence of A and B at the same time, and the existence of B alone. Among them, A and B can be singular or plural. The character " / " generally indicates that the previous and subsequent associated objects are in an "or" relationship. "At least one of the following" and similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b and c can represent: a, b, c, a and b, a and c, b and c or a and b and c, where a, b, c can be single or multiple.

[0199] In the embodiments of the present application, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not preclude the presence of additional identical elements in the process, method, commodity, or apparatus comprising the element.

[0200] The present application may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communications network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.

[0201] The various embodiments in this application are described in a progressive manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences from other embodiments. In particular, the device embodiments are generally similar to the method embodiments, so the description is relatively simple. For relevant parts, refer to the partial description of the method embodiments.

[0202] Those skilled in the art will appreciate that the various units and algorithm steps described in the embodiments of the present application can be implemented using a combination of electronic hardware, computer software, and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0203] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described devices, apparatuses and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0204] The above description is merely a specific embodiment of the present application. Any person skilled in the art may easily conceive of variations or substitutions within the technical scope disclosed in this application, and such variations or substitutions shall be within the scope of protection of this application. The scope of protection of this application shall be subject to the scope of protection of the claims.

Claims

1. A model training system, characterized in that: The model training system includes: The inference module is used to generate training samples for model training; An experience sample pool, which is used to store the training samples; A training module is used to train the model according to the training samples stored in the experience sample pool, wherein the reasoning module and the training module run asynchronously and in parallel.

2. The model training system according to claim 1, characterized in that The model training system also includes: A coordination controller is used to adjust the computing resource allocation of each component in the reasoning module according to the throughput of the reasoning module and the training module, and / or adjust the computing resource allocation of the reasoning module and the training module.

3. The model training system according to claim 2, characterized in that The reasoning module includes a reference model, a reward model, and a first strategy model, and the training module includes a second strategy model; The coordination controller is used to optimize and solve the problem with the goal of maximizing the total training throughput and inference throughput based on the training throughput of the second strategy model, the inference throughput of the first strategy model, the reference model, and the reward model under unit resources, and obtain the computing resource allocation ratio of each model in the inference module, as well as the computing resource allocation ratio between the inference module and the training module.

4. The model training system according to claim 1, characterized in that: The inference module is further configured to add a timestamp to the training sample; The experience sample pool is further used to detect the timestamp of the training sample, and when the timestamp of the training sample exceeds a preset threshold, delete the training sample.

5. The model training system according to claim 1, characterized in that: The inference module is further configured to add a timestamp to the training sample; The training module is further configured to continuously sample the training samples from the experience sample pool, check the difference between the training samples and the current training strategy model version, and discard the training samples when the difference exceeds a preset threshold.

6. The model training system according to any one of claims 1 to 5, characterized in that: The reasoning module is deployed on homogeneous or heterogeneous computing resources.

7. The model training system according to claim 6, characterized in that The reasoning module includes a reference model, a reward model, and a first strategy model, wherein: The reference model and the reward model are deployed on a first type of computing resources, and the first policy model is deployed on a second type of computing resources.

8. The model training system according to claim 7, characterized in that: The reasoning module includes a reference model, a reward model, and a first strategy model, wherein: The reference model is deployed on a first type of computing resources, the reward model is deployed on a second type of computing resources, and the first policy model is deployed on a third type of computing resources.

9. The model training system according to any one of claims 1 to 5, characterized in that: The training modules are deployed on homogeneous computing resources.

10. A model training method, characterized in that: The method is applied to a model system, the model system including an inference module, an experience sample pool, and a training module, and the method includes: Using the inference module to generate training samples for model training; Saving the training samples generated by the inference module to the experience sample pool; The training module is used to train the model according to the training samples stored in the experience sample pool, wherein the reasoning module and the training module run asynchronously and in parallel.

11. The method according to claim 10, characterized in that The method also includes: According to the throughput of the reasoning module and the training module, the computing resource allocation of each component in the reasoning module is adjusted, and / or the computing resource allocation of the reasoning module and the training module is adjusted.

12. The method according to claim 11, characterized in that The reasoning module includes a reference model, a reward model, and a first strategy model, and the training module includes a second strategy model; The adjusting the computing resource allocation of each component in the reasoning module according to the throughput of the reasoning module and the training module, and / or adjusting the computing resource allocation of the reasoning module and the training module, includes: Based on the training throughput of the second strategy model, the inference throughput of the first strategy model, the inference throughput of the reference model and the reward model under unit resources, an optimization solution is performed with the goal of maximizing the total training throughput and inference throughput, and the computing resource allocation ratio of each model in the inference module, as well as the computing resource allocation ratio between the inference module and the training module are obtained.

13. The method according to claim 10, characterized in that The method further comprises: Before storing the training sample in the experience sample pool, adding a timestamp to the training sample; The timestamps of the training samples in the experience sample pool are detected, and when the timestamps of the training samples exceed a preset threshold, the training samples are deleted from the experience sample pool.

14. The method according to claim 10, characterized in that The method further comprises: Before storing the training sample in the experience sample pool, adding a timestamp to the training sample; The training samples are continuously sampled from the experience sample pool to perform model training fine-tuning, wherein the difference between the training samples and the current training strategy model version is checked, and the training samples are discarded when the difference exceeds a preset threshold.

15. The method according to any one of claims 10 to 14, characterized in that The method further comprises: The reasoning module is deployed, wherein the reasoning module is deployed on homogeneous or heterogeneous computing resources.

16. The method according to claim 15, characterized in that The reasoning module includes a reference model, a reward model and a first strategy model; The deploying the inference module includes: The reference model and the reward model are deployed on a first type of computing resources, and the first strategy model is deployed on a second type of computing resources.

17. The method according to claim 15, characterized in that The reasoning module includes a reference model, a reward model and a first strategy model; The deploying the inference module includes: The reference model is deployed on a first type of computing resources, the reward model is deployed on a second type of computing resources, and the first strategy model is deployed on a third type of computing resources.

18. The method according to any one of claims 10 to 14, characterized in that The method further comprises: The training module is deployed, wherein the training module is deployed on homogeneous computing resources.

19. An electronic device, characterized in that: The electronic device comprises a memory for storing computer program instructions and a processor for executing the computer program instructions, wherein when the computer program instructions are executed by the processor, the electronic device is triggered to perform the method steps according to any one of claims 10 to 18.

20. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed on a computer, enables the computer to execute the method steps according to any one of claims 10 to 18.