Large Model Online Inference and Training Integration Method and System
By using online SFT and RLHF to optimize the fine-tuning model in the integrated method of large model inference and training, the problem of user input data in the existing technology cannot be used in real time and manual participation is solved, and the real-time and efficiency of model iteration is improved.
Patent Information
- Application Number
- CN202311340706.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-16
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2043-10-16
AI Technical Summary
During the deployment of inference services by existing large models, the user input data cannot be used in real time, and each model iteration requires manual startup of the process, resulting in large workload and delay.
A method of integrated online inference and training for large models is proposed. By receiving user input data, it determines whether there is a correction situation for the model reply result, stores it in message middleware, and uses online SFT and RLHF for supervision and fine-tuning and intensive training to optimize the fine-tuning model.
Reusing process vectors and user result feedback in the forward inference stage is realized, reducing manual participation and improving the real-time and efficiency of model iteration.
Smart Images

Figure CN117668153B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of data processing, and in particular, to a method and system for integrating online inference and training of large models. Background Art
[0002] Currently, for the implementation of large models, i.e., large language models, after obtaining the base model, it can be divided into three stages, including supervised fine-tuning model training, reinforcement model training that conforms to human preferences, and inference deployment. Specifically, as shown in the appendix Figures 1-3 First, a base model is fine-tuned through supervised fine-tuning into a fine-tuned model with conversation capabilities. Then, different responses of a certain amount of instruction models are manually scored and annotated to train a reward model, and reinforcement training that conforms to human preferences is performed based on the reward model, imposing text distribution constraints on the target fine-tuned model. It should also be noted that during the process of deploying the inference service of existing large models, multiple rounds of conversations are carried out according to the input instruction text of the user. Among them, the above three stages are basically independent of each other, resulting in the need to manually start the process and detect the intermediate results of the steps for each model iteration. There is a problem of a large amount of manual work in iteration, and there is a certain delay, and the user input data cannot be fully utilized in real time.
[0003] Therefore, providing a method that can integrate the above three stages, reduce manual participation during model iteration, and utilize user input data in real time is an urgent problem to be solved currently. Summary of the Invention
[0004] To solve the above problems, the present application proposes a method and system for integrating online inference and training of large models to solve the above problems.
[0005] On the one hand, the present application proposes a method for integrating online inference and training of large models, including the following steps:
[0006] Receive user input data and input it into the fine-tuned model to obtain a model reply result;
[0007] Judge whether there is a correction situation in the model reply result;
[0008] When there is such a correction situation, store the corresponding user input data and the model reply result into the message middleware;
[0009] Obtain the training data of the message middleware, and use online SFT and RLHF to perform supervised fine-tuning and reinforcement training on the fine-tuned model respectively to obtain the fine-tuned model that meets the preset requirements.
[0010] As an optional implementation of the present application, optionally, the step of receiving user input data and inputting it into the fine-tuned model to obtain a model reply result includes:
[0011] Receive user input instructions and text;
[0012] Input the user input instructions and text into the fine-tuning model for multiple rounds of conversation to obtain corresponding model reply results;
[0013] Among them, the fine-tuning model can deploy conversation tasks.
[0014] As an alternative implementation of this application, optionally, when using online SFT to perform supervised fine-tuning on the fine-tuning model, the training data of the online SFT includes SFT data and pre-annotated supervised training data;
[0015] Among them, the SFT data is obtained based on the corrected semantic combination corresponding to the user input instruction text and the model reply result.
[0016] As an alternative implementation of this application, optionally, when using RLHF to perform reinforcement training on the fine-tuning model, it includes:
[0017] Obtain reward model training data;
[0018] Train the reward model based on the reward model training data and update the reward model;
[0019] Use the reward model to score the model reply result and reinforce the training of the fine-tuning model.
[0020] As an alternative implementation of this application, optionally, the reward model training data is formed based on the user input data, as well as the chosen reply and rejected reply in the model reply result.
[0021] As an alternative implementation of this application, optionally, after obtaining the reward model training data, it further includes:
[0022] Perform text distribution constraint on the fine-tuning model to improve the scoring result of the reward model.
[0023] As an alternative implementation of this application, optionally, it further includes:
[0024] Use a bypass model to update the parameters of the fine-tuning model.
[0025] On the one hand, this application provides a system for implementing the large model online inference and training integration method described in any one of the above, including:
[0026] A data acquisition module configured to receive user input data and input it into a fine-tuning model to obtain model reply results;
[0027] A data judgment module, configured to judge whether there is a correction situation in the model reply result;
[0028] A data storage module, when there is such a correction situation, stores the corresponding user input data and the model reply result into a message middleware;
[0029] A model training module, obtains the training data of the message middleware, and uses online SFT and RLHF to perform supervised fine-tuning and reinforcement training on the fine-tuning model respectively, to obtain the fine-tuning model that meets the preset requirements.
[0030] On the one hand, this application provides an electronic device, including:
[0031] A processor;
[0032] A memory for storing instructions executable by the processor;
[0033] Wherein, when the processor is configured to execute the executable instructions, it implements the integrated method for online inference and training of the large model described in any one of the above.
[0034] On the one hand, this application provides a non-volatile computer-readable storage medium, on which computer program instructions are stored, and characterized in that when the computer program instructions are executed by a processor, the integrated method for online inference and training of the large model described in any one of the above is implemented.
[0035] The technical effects of the present invention:
[0036] This application stores the user input data and the reply result of the fine-tuning model into the message middleware, thereby making full use of the process vector and user result feedback in the forward inference stage to achieve the purpose of reuse, thus effectively solving the problem that the manual start process is required for each model iteration. Specifically, it includes receiving user input data and inputting it into the fine-tuning model to obtain a model reply result; judging whether there is a correction situation in the model reply result; when there is a correction situation, storing the corresponding user input data and the model reply result into the message middleware; obtaining the training data of the message middleware, and using online SFT and RLHF to perform supervised fine-tuning and reinforcement training on the fine-tuning model respectively, to obtain the fine-tuning model that meets the preset requirements. That is, this application uses online SFT and RLHF to perform supervised fine-tuning and reinforcement training on the fine-tuning model according to the online supervised training data and online human preference-compliant reinforcement model training data saved in the middleware, further optimizing the fine-tuning model to make it more compliant with human preferences.
[0037] According to the following detailed description of the exemplary embodiments with reference to the accompanying drawings, other features and aspects of the present disclosure will become clear. Description of the Drawings
[0038] The accompanying drawings, which are included in and form a part of this specification, illustrate exemplary embodiments, features, and aspects of the present disclosure and are used to explain the principles of the present disclosure together with the specification.
[0039] Figure 1 Shown as an SFT flowchart;
[0040] Figure 2 Shown as a reward model and PPO flowchart in RLHF;
[0041] Figure 3 Shown as a large model inference flowchart;
[0042] Figure 4 Shown as a flowchart of the integrated method for online inference and training of the large model of the present invention;
[0043] Figure 5 Shown as a schematic diagram of the implementation process of the integrated method for online inference and training of the large model of the present invention;
[0044] Figure 6 Shown as a schematic diagram of the bypass model in the integrated method for online inference and training of the large model of the present invention. Detailed implementation manners
[0045] Various exemplary embodiments, features, and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. Like reference numerals in the drawings denote elements having the same or similar functions. Although various aspects of the embodiments are shown in the drawings, the drawings are not necessarily drawn to scale unless otherwise specified.
[0046] The specific term "exemplary" herein means "serving as an example, embodiment, or illustration". Any embodiment described herein as "exemplary" is not necessarily to be construed as superior to or better than other embodiments.
[0047] In addition, for a better illustration of the present disclosure, numerous specific details are given in the following detailed implementation manners. Those skilled in the art should understand that the present disclosure can be implemented without some specific details. In some instances, methods, means, elements, and circuits well known to those skilled in the art are not described in detail so as to highlight the gist of the present disclosure.
[0048] Embodiment 1
[0049] As Figure 4 and Figure 5 shown, on the one hand, the present application provides an integrated method for online inference and training of a large model, and the method includes the following steps:
[0050] S100. Receive user input data and input it into the fine-tuning model to obtain a model reply result;
[0051] S200. Determine whether there is a correction in the model's reply result;
[0052] S300. When there is such a correction situation, store the corresponding user input data and the model's reply result in the message middleware;
[0053] S400. Obtain the training data of the message middleware, and use online SFT and RLHF to perform supervised fine-tuning and reinforcement training on the fine-tuning model respectively, so as to obtain the fine-tuning model that meets the preset requirements.
[0054] In this embodiment, by storing the user input data and the reply result of the fine-tuning model in the message middleware, the process vector in the forward inference stage and the user result feedback are fully utilized to achieve the purpose of reuse, thus effectively solving the problem that the manual process needs to be started for each model iteration. At the same time, according to the online supervised training data and the online human-preference-compliant reinforcement model training data saved in the middleware, the fine-tuning model is correspondingly supervised and fine-tuned and reinforced trained by using online SFT and RLHF, further optimizing the fine-tuning model to make it more in line with human preferences.
[0055] Specifically, through step S100, receive the user input data and input it into the fine-tuning model to obtain the model's reply result. Here, it should be noted that before receiving the user input data, first train a fine-tuning model that can deploy the dialogue task. The fine-tuning model can be obtained from the base model through the SFT and RLHF processes, or directly obtain a chat fine-tuning model with dialogue capabilities from the open source community. Subsequently, receive the user input data to complete multiple rounds of conversations. Since the forward inference process for the user input during inference is the same as the forward inference process for the input instructions and texts during training, it can be reused.
[0056] As an optional implementation of this application, optionally, the receiving the user input data and inputting it into the fine-tuning model to obtain the model's reply result includes: S110. Accept the user input instructions and texts; S120. Input the user input instructions and texts into the fine-tuning model for multiple rounds of conversations to obtain the corresponding model's reply result; where the fine-tuning model can deploy the dialogue task.
[0057] After obtaining the model's reply result of the fine-tuning model, through step S200, determine whether there is a correction in the model's reply result. For example, train a semantic judgment model based on the BERT model, and judge whether there is a situation where the user corrects the model output result during the conversation according to 3 consecutive inputs of the user context. If it exists, the label is 1, and if it does not exist, the label is 0. The annotation unit of the training data is 3 sentences (user question 1, user question 2, and user question 3), and the sentences are separated by <sep>Separated by standard delimiters, as shown in Table 1:
[0058]
[0059] Table 1 Example of training data for the user semantic correction judgment model
[0060] As can be seen from Table 1, when the user inputs "Please introduce LLM in Chinese. No, what I mean is the large language model LLM. Then, what is the difference from the language model?" in sequence, there is a situation of correcting the model data results. Further, through step S300, when there is such a correction situation, the corresponding user input data and the model reply results are stored in the message middleware. Specifically, in the form of queue messages, the user input with correction and the corresponding model reply results are sent to the message middleware, such as Kafka, etc. At the same time, the same number of semantic contexts without correction are randomly sampled and sent to the middleware for backup, such as user inputs and corresponding model replies like "What are the interesting places and delicious foods in Beijing? What about Tianjin?" etc. It should be noted that the data with and without correction situations are distinguished by different middleware topics.
[0061] Further, after storing the user input data and the model reply results in the message component, through step S400, the training data of the message middleware is obtained, and the fine-tuning model is supervised and fine-tuned and strengthened respectively by using online SFT and RLHF to obtain the fine-tuning model that meets the preset requirements. Here, it should be noted that before online training of the fine-tuning model, the number of messages in the message component needs to be monitored in real time. When the number of messages with corrected semantics reaches the preset training batch size, for example, accumulates to 64, the data is extracted at one time to start the online training process.
[0062] As an optional implementation of this application, optionally, when using online SFT to supervise and fine-tune the fine-tuning model, the training data of the online SFT includes SFT data and pre-annotated supervised training data; among them, the SFT data is obtained by combining the user input instruction text and the corresponding corrected semantics in the model reply results.
[0063] The online training of SFT will be described in detail below.
[0064] The conversations with correction situations, that is, the original questions input by the user and the answers corrected by the model, are combined into SFT data and trained according to the existing SFT method.
[0065] For example:
[0066] User: Please introduce LLM in Chinese.
[0067] Model: LLM stands for Master of Laws, which is a one-year program mainly targeting students who have completed a law degree, such as JD or LLB. The LLM program allows students to specialize in specific legal fields, such as corporate law, intellectual property law, international law, etc. The LLM program is offered by many law schools around the world and is highly recognized and respected in the legal community. After completing the LLM program, students can deepen their knowledge and skills and gain a greater advantage in the legal field. In summary, LLM is a Master of Laws degree, mainly targeting students who have completed a law degree, and provides professional legal learning opportunities.
[0068] User: No, I'm talking about the large language model LLM.
[0069] Model: LLM (Large Language Model) is an artificial intelligence model that can generate human language output by training on a large amount of text data. It can be used for various natural language processing tasks, such as text generation, question answering, language translation, etc. Since it can generate high-quality text and its applications may change multiple industries, such as customer service, marketing, and content creation, it has received increasing attention.
[0070] The online SFT data thus composed is:
[0071] Instuction: Please introduce LLM in Chinese.
[0072] Response: LLM (Large Language Model) is an artificial intelligence model that can generate human language output by training on a large amount of text data. It can be used for various natural language processing tasks, such as text generation, question answering, language translation, etc. Since it can generate high-quality text and its applications may change multiple industries, such as customer service, marketing, and content creation, it has received increasing attention.
[0073] At the same time, historical supervised data is read from the database, and online SFT training starts with 80% manually labeled supervised training data and 20% online SFT data.
[0074] As an optional implementation of this application, optionally, when using RLHF to perform reinforcement training on the fine-tuned model, it includes: obtaining reward model training data; training the reward model based on the reward model training data and updating the reward model; using the reward model to score the model reply results and performing reinforcement training on the fine-tuned model.
[0075] As an optional implementation of this application, optionally, the reward model training data is formed based on the user input data, as well as the chosen reply and rejected reply in the model reply results.
[0076] As an alternative implementation of this application, optionally, after obtaining the reward model training data, it further includes: performing text distribution constraint on the fine-tuning model to improve the scoring result of the reward model.
[0077] It should be noted that during the online reinforcement training that conforms to human preferences, the reward model is updated synchronously, and PPO reinforcement learning is performed using the reward model scoring and text distribution constraint scoring to further improve the overall performance of the fine-tuning model.
[0078] The online RLHF data sampling and reward model update mechanism will be described in detail below.
[0079] First, component reward model training data is obtained from the conversations with correction situations.
[0080] For example:
[0081] {
[0082] "query":"Please introduce LLM in Chinese.",
[0083] "chosen":"LLM (Large Language Model) is an artificial intelligence model that can generate human language output by training on a large amount of text data. It can be used for various natural language processing tasks, such as text generation, question answering, language translation, etc. Due to its ability to generate high-quality text and its potential to change multiple industries, such as customer service, marketing, and content creation, it has received increasing attention.",
[0084] "rejected":"LLM is Master of Laws, that is, a master's degree in law. This is a one-year program mainly for students who have completed a law degree, such as JD or LLB. The LLM program allows students to focus on specific legal fields, such as corporate law, intellectual property law, international law, etc. The LLM program is offered by many law schools around the world and is highly regarded and respected in the legal community. After completing the LLM program, students can deepen their knowledge and skills and gain greater advantages in the legal field. In summary, LLM is a master's degree in law, mainly for students who have completed a law degree, providing professional legal learning opportunities.",
[0085] }
[0086] At the same time, some semantic data without correction is sampled, and randomly selected model responses with or without corrected semantics in the batch are used as rejected responses to form the reward model training data.
[0087] For example:
[0088] {
[0089] "query": "Are there any recommended interesting places in Beijing?
[0090] "chosen": "Okay, I will provide you with some recommendations for interesting places in Beijing. The Great Wall: The Great Wall is an ancient defensive project in China and one of the largest defensive projects in the world. You can travel around the Great Wall and experience the historical culture and natural scenery. The Forbidden City: The Forbidden City is the ancient imperial palace in China and one of the most important buildings in Chinese history. You can travel in the Forbidden City and learn about ancient Chinese culture and history. I hope these recommendations can provide you with some interesting activities and enjoyable times!
[0091] "rejected": "LLM (Large Language Model) is an artificial intelligence model that can generate human language output by training a large amount of text data. It can be used for various natural language processing tasks, such as text generation, question answering, language translation, etc. Due to its ability to generate high-quality text and its potential to change multiple industries, such as customer service, marketing, and content creation, it has received increasing attention.
[0092] }
[0093] Secondly, generally, the training of the reward model requires humans to sort or score multiple different answers to the same instruction. In this application, in the case where there are corrections for online inference, the multi-level scoring mechanism can be weakened into a binary scoring mechanism (correction after reply, no correction after reply), that is, the chosen scoring gets the highest score, the rejected gets a low score, or the chosen is ranked ahead of the rejected.
[0094] It should be noted that the size of the reward model is relatively small (in the order of hundreds of megabytes). Therefore, for its full-scale training, the training data is the online training data mixed with 4 times the number of randomly sampled labeled data. The training objective is to maximize the difference between the chosen reply and the rejected reply, as shown in the following loss function:
[0095]
[0096] In the formula, N is the batch size, σ is the activation function, r represents the reward model, and r(x, y) represents the scoring score of the reward model that returns y for the input x.
[0097] As an optional implementation scheme of this application, optionally, it further includes: updating the parameters of the fine-tuning model by using a bypass model.
[0098] In this embodiment, when adjusting the fine-tuning model parameters, we refer to the Gradient Boosting Decision Tree (GBDT), and adopt the method of superimposing multiple bypass models to correct the deviation between the original model and the expected vector. An adjustment of the bypass model generated by online training is added on the basis of the original fine-tuning model, so as to improve the overall performance. That is, instead of fully updating the original model, a 3-layer bypass model is added beside the original model. As Figure 6 shown, the input dimension and output dimension of the bypass layer 1 of the bypass model are 1 / 16 of the dimension of the original model, and a scaling layer / information fusion layer with a dimension of 1 / 64 of the dimension of the original model is added in the middle to enhance the non-linear ability of the model. After the input X vector is input into the original model and the bypass model respectively, when outputting, the parameters of the original model and the updated bypass model are superimposed and then normalized to obtain the final output vector. The output vector passes through, for example, a softmax classifier to obtain the result vector. During one online training process, assuming the data dimension is 5,
[0099] the output of the original model is [0.1, 0.2, 0.3, 0.2, 0.2],
[0100] the output of the bypass model 1 is [0.2, 0.1, 0.3, 0.1, 0.0],
[0101] the superimposed vector is [0.3, 0.3, 0.6, 0.3, 0.2],
[0102] the normalized output vector is [-0.29, -0.29, 1.92, -0.29, -1.03],
[0103] using the softmax classifier to obtain the result vector [0.08, 0.08, 0.72, 0.08, 0.04].
[0104] The online training loss calculated according to the result vector and the expected result vector is backpropagated to the current bypass model to ensure that the online updated parameters and the original model parameters can be distinguished. The model version can be easily rolled back by selecting whether to superimpose a certain bypass model.
[0105] Continuing from the previous example:
[0106] the expected result vector is [1, 0, 0, 0, 0],
[0107] the calculated difference is [0.921, -0.079, -0.724, -0.079, -0.038]
[0108] Based on this difference, only the parameters of the bypass model are adjusted, so that the output of the bypass model changes to
[0109] [0.3, 0.1, 0.1, 0.1, 0.0]
[0110] With the output of the original model remaining unchanged, repeat the above forward calculation steps. The final result vector is calculated to be [0.39, 0.10, 0.39, 0.10, 0.03], and the difference is reduced to [0.61, -0.10, -0.39,
[0111] -0.10, -0.03], which is closer to the target result [1, 0, 0, 0, 0] than before the adjustment.
[0112] During the online training process again, continue to add the second bypass model 2. The output of the bypass model 2 is [0.2, 0, -0.2, 0, 0]. The output result of the total bypass model becomes [0.5, 0.1, -0.1, 0.1, 0]. The final result vector is calculated to be [0.28, -0.09, -0.05, -0.09, -0.05], which is very close to the target result [1, 0, 0, 0, 0].
[0113] Since this method does not change the output of the original model, the adjustment only occurs in the small-scale bypass models. Since the number of samples during online training is much smaller than that during training, only a small number of parameters are needed for fitting, while reducing the number of training parameters and accelerating the online training speed of the model. That is, the main purpose of this method is to reduce the number of training parameters, reducing the originally required complete original model with 7 billion or even more than 176 billion parameters to only need to adjust the bypass models with parameters in the order of billions or tens of millions, improving the training efficiency.
[0114] In summary, after pre-training a fine-tuning model that can deploy dialogue tasks in this application, receive user input instructions and texts, complete multi-turn conversations, and when the fine-tuning model determines that there is a correction situation, send the conversation with the correction situation to the middleware for temporary storage. That is, use the middleware to save information such as process vectors and user result feedback in the forward inference stage for repeated use. When the data saved in the middleware reaches the preset quantity, mix a certain amount of supervised training data to form online supervised training data and online reinforcement model training data that conforms to human preferences respectively. Among them, during the online reinforcement training that conforms to human preferences, update the reward model synchronously, and perform PPO reinforcement learning using the reward model scoring and text distribution constraint scoring. Moreover, for the fine-tuning of the model parameters in this application, the method of adding an online training bypass model is adopted. Since the number of samples during online training is much smaller than that during training, only a small number of parameters are needed for fitting to achieve the purpose of updating the model parameters and improving the model performance.
[0115] Those skilled in the art can understand that to implement all or part of the processes in the above-described embodiment methods, it can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above-described embodiments of each control method. Among them, the storage medium can be a magnetic disk, an optical disc, a read-only memory (ROM), a random access memory (RAM), a flash memory, a hard disk drive (abbreviation: HDD), or a solid-state drive (SSD), etc.; the storage medium can also include a combination of the above-mentioned types of memories.
[0116] Embodiment 2
[0117] Furthermore, based on the implementation principle of Embodiment 1, the second aspect of the present application provides a system for implementing the integrated method of large model online inference and training described in any one of the above. Since the working principle of the device in the embodiments of the present disclosure is the same as or similar to the principle of the integrated method of large model online inference and training in the embodiments of the present disclosure, the repeated parts will not be elaborated here. The device in the disclosed embodiments of the present application includes:
[0118] A data acquisition module, configured to receive user input data and input it into a fine-tuning model to obtain a model reply result;
[0119] A data judgment module, configured to judge whether there is a correction situation in the model reply result;
[0120] A data storage module, when there is the correction situation, stores the corresponding user input data and the model reply result into a message middleware;
[0121] A model training module, obtains the training data from the message middleware, and respectively performs supervised fine-tuning and reinforcement training on the fine-tuning model by using online SFT and RLHF to obtain the fine-tuning model that meets the preset requirements.
[0122] Embodiment 3
[0123] Even further, in the third aspect of the present application, an electronic device is provided, including:
[0124] A processor;
[0125] A memory for storing processor-executable instructions;
[0126] Wherein, the processor is configured to implement the integrated method of large model online inference and training described in any one of the above when executing the executable instructions.
[0127] The control system of the embodiments of the present disclosure includes a processor and a memory for storing processor-executable instructions. Among them, the processor is configured to implement the integrated method for online inference and training of the large model described in any of the foregoing when executing the executable instructions.
[0128] Here, it should be noted that the number of processors can be one or more. At the same time, in the control system of the embodiments of the present disclosure, an input device and an output device may also be included. Among them, the processor, the memory, the input device, and the output device can be connected through a bus or in other ways, which are not specifically limited here.
[0129] As a computer-readable storage medium, the memory can be used to store software programs, computer-executable programs, and various modules, such as: the programs or modules corresponding to the integrated method for online inference and training of the large model of the embodiments of the present disclosure. The processor executes various functional applications and data processing of the control system by running the software programs or modules stored in the memory.
[0130] The input device can be used to receive input numbers or signals. Among them, the signal can be a key signal related to the user settings and function control of the device / terminal / server. The output device can include display devices such as a display screen.
[0131] Embodiment 4
[0132] In the fourth aspect of the present application, a non-volatile computer-readable storage medium is provided, on which computer program instructions are stored, and it is characterized in that the computer program instructions implement the integrated method for online inference and training of the large model described in any of the above when being executed by a processor.
[0133] The embodiments of the present disclosure have been described above. The above description is exemplary and not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art in the technical field without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, practical applications, or technical improvements to the technologies in the market, or to enable other ordinary technical personnel in the technical field to understand the disclosed embodiments.< / sep>
Claims
1. An integrated method for large model online inference and training, characterized in that, It includes the following steps: Receive user input data and input it into the fine-tuning model to obtain the model reply result; Judge whether there is a correction situation in the model reply result; When there is such a correction situation, store the corresponding user input data and the model reply result in the message middleware; Obtain the training data of the message middleware, and use online SFT and RLHF to perform supervised fine-tuning and reinforcement training on the fine-tuning model respectively to obtain the fine-tuning model that meets the preset requirements; Before online training of the fine-tuning model, monitor the number of messages in the message middleware in real time. When the number of messages with corrected semantics reaches the preset training batch size, extract the data at one time to start the online training process.
2. The integrated method for online inference and training of large models according to claim 1, wherein The step of receiving user input data and inputting it into the fine-tuning model to obtain the model reply result includes: Accept user input instructions and text; Input the user input instructions and text into the fine-tuning model for multi-round conversations to obtain the corresponding model reply result; Among them, the fine-tuning model can deploy dialogue tasks.
3. The integrated method for online inference and training of large models according to claim 1, wherein When performing supervised fine-tuning on the fine-tuning model using online SFT, the training data of the online SFT includes SFT data and pre-annotated supervised training data; Among them, the SFT data is obtained by combining the user input instruction text and the corresponding corrected semantics in the model reply result.
4. The integrated method for online inference and training of large models according to claim 1, wherein When performing reinforcement training on the fine-tuning model using RLHF, it includes: Obtain reward model training data; Train the reward model based on the reward model training data to update the reward model; Use the reward model to score the model reply result and perform reinforcement training on the fine-tuning model.
5. The integrated method for online inference and training of large models according to claim 4, characterized in that, The reward model training data is formed based on the user input data, as well as the chosen reply and rejected reply in the model reply result.
6. The integrated method for online inference and training of large models according to claim 4, characterized in that After obtaining the reward model training data, it further includes: Perform text distribution constraints on the fine-tuning model to improve the scoring result of the reward model.
7. The integrated method for online inference and training of large models according to claim 1, characterized in that It also includes: Update the parameters of the fine-tuning model using a bypass model.
8. A system for implementing the integrated method of large model online inference and training described in any one of the above claims 1-7, characterized in that, It includes: A data acquisition module configured to receive user input data and input it into the fine-tuning model to obtain the model reply result; A data judgment module configured to judge whether there is a correction situation in the model reply result; A data storage module, when there is such a correction situation, store the corresponding user input data and the model reply result in the message middleware; A model training module, obtain the training data of the message middleware, and use online SFT and RLHF to perform supervised fine-tuning and reinforcement training on the fine-tuning model respectively to obtain the fine-tuning model that meets the preset requirements.
9. An electronic device, characterized in that, It includes: A processor; A memory for storing processor-executable instructions; Among them, the processor is configured to implement the large model online inference and training integration method described in any one of claims 1 to 7 when executing the executable instructions.
10. A non-volatile computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, the large model online inference and training integration method described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Machine reading understanding model training method and device, electronic equipment and storage medium
CN111160568A
Conversation processing method and system, storage medium and terminal
CN116303949A
Reward model training method and device, storage medium and electronic equipment
CN116861259A