Industry large language model training method and device based on reinforcement learning
Patent Information
- Application Number
- CN202510351772.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2045-03-24
AI Technical Summary
然而,相关技术中的行业大语言模型的推理能力有限,难以应对真实行业应用场景中处理复杂任务的高精度要求,限制了模型的广泛应用
[0019]本申请提供的基于强化学习的行业大语言模型训练方法及装置,首先,利用目标行业的行业非结构化文本数据通过增量预训练的方式对基座大语言模型进行训练,得到第一模型;所述基座大语言模型为基于通用领域知识文本数据训练得到的;之后,利用所述目标行业的高质量推理数据通过指令精调的方式对所述第一模型进行一次训练,再使用强化学习方法进行二次训练,得到第二模型,并使用拒绝采样的方法,利用所述第二模型生成第一数据集;所述第一数据集包含行业知识图谱数据的多条推理路径;最后,使用所述第一数据集通过指令精调的方式对所述第二模型进行训练一次训练,再使用强化学习方法进行二次训练,得到第三模型,并基于任务向量运算利用所述第一数据集和所述高质量推理数据将所述基座大语言模型和所述第三模型进行融合,得到所述目标行业的推理型行业大语言模型。如此,通过优化模型的训练策略以及行业推理数据的构造方法,极大地提高了行业大语言模型针对特定领域的推理能力和泛化能力,降低了模型的迭代成本。
Smart Images

Figure CN120278270B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of large language model training technology, and in particular to an industry-specific large language model training method and apparatus based on reinforcement learning. Background Technology
[0002] Large Language Models (LLMs) are deep learning models trained on large amounts of text data, enabling them to generate natural language text or understand the meaning of language text. These models can provide in-depth knowledge and language production on a wide range of topics by being trained on massive datasets, and can simulate human language cognition and generation processes to some extent.
[0003] Industry-specific large language models are driving intelligent upgrades across various fields, with a future focus on specialization, multimodality, and cost control. However, the reasoning capabilities of these models are limited, making it difficult to meet the high-precision requirements of complex tasks in real-world industry applications, thus restricting their widespread application. This is mainly reflected in the following aspects: industry data often exists in the form of unstructured text and structured knowledge graphs, making it difficult to integrate this data to generate high-quality datasets that can be used to improve reasoning capabilities; and industry-specific large language models lack generalization ability, easily leading to incomplete reasoning paths and erroneous inferences, reducing the credibility of the reasoning results.
[0004] Therefore, there is an urgent need for a training method for industry-specific large language models to improve their reasoning and generalization capabilities for specific domains. Summary of the Invention
[0005] The purpose of this application is to provide a training method and apparatus for industry-specific large language models based on reinforcement learning. By optimizing the model training strategy and the construction method of industry inference data, the inference ability and generalization ability of the industry-specific large language model for specific domains are greatly improved, and the iteration cost of the model is reduced.
[0006] This application provides a method for training industry-specific large language models based on reinforcement learning, including: A first model is obtained by training a base large language model using unstructured text data from the target industry through incremental pre-training. This base large language model is trained based on general domain knowledge text data. A second model is obtained by training the first model once using high-quality inference data from the target industry through instruction fine-tuning, followed by a second training using reinforcement learning. A first dataset is generated using the second model using a rejection sampling method. This first dataset contains multiple inference paths from industry knowledge graph data. A third model is obtained by training the second model once using the first dataset through instruction fine-tuning, followed by a second training using reinforcement learning. Finally, the base large language model and the third model are fused using task vector operations based on the first dataset and the high-quality inference data to obtain a reasoning-based industry large language model for the target industry.
[0007] Optionally, the method of using rejection sampling to generate the first dataset using the second model includes: converting each industry knowledge graph data into a text expression in natural language using a rule-based method, and controlling the second model to generate multiple inference paths to be identified based on the converted industry knowledge graph data using a prompting method; extracting the final inference answer from the multiple inference paths to be identified, matching the obtained final inference answer with the corresponding correct answer in the industry knowledge graph data, and removing the inference paths that do not match the correct answer from the multiple inference paths to be identified, thereby obtaining the first dataset.
[0008] Optionally, training the second model once using the first dataset through instruction fine-tuning includes: using a hybrid method of dense and sparse retrieval to recall similar text fragments in the industry knowledge graph data that meet a preset threshold of similarity with the first dataset; optimizing the expression of the instruction data of the second model based on the similar text fragments, and completing the training of the second model once.
[0009] Optionally, the step of fusing the base large language model and the third model using the first dataset and the high-quality inference data to obtain the inference-based industry large language model for the target industry based on task vector operation includes: fine-tuning the base large language model using the high-quality inference data to obtain first model parameters, and fine-tuning the third model using the first dataset to obtain second model parameters; calculating the difference between the second model parameters and the first model parameters to obtain a task vector, and adding the task vector to the second model parameters after scaling by adjustment coefficients to obtain third model parameters; and using the third model parameters as the model parameters of the model after fusing the base large language model and the third model to obtain the inference-based industry large language model.
[0010] Optionally, after fusing the base large language model and the third model using the first dataset and the high-quality inference data based on task vector operations to obtain the inference-based industry large language model for the target industry, the method further includes: obtaining a second dataset generated during the use of the inference-based industry large language model; the second dataset is: a dataset obtained by manually annotating the answers to erroneous inference samples generated during the use of the inference-based industry large language model; and using the second dataset to perform reinforcement training on the inference-based industry large language model to obtain a refined inference-based industry large language model.
[0011] This application also provides a reinforcement learning-based industry-specific large language model training device, comprising: The model training module is used to train the base large language model using unstructured text data of the target industry through incremental pre-training to obtain a first model; the base large language model is trained based on general domain knowledge text data; the model training module is also used to train the first model once using high-quality inference data of the target industry through instruction fine-tuning, and then use reinforcement learning methods for a second training to obtain a second model; the data acquisition module is used to generate a first dataset using the second model using a rejection sampling method; the first dataset contains multiple inference paths of industry knowledge graph data; the model training module is also used to train the second model once using the first dataset through instruction fine-tuning, and then use reinforcement learning methods for a second training to obtain a third model; the model fusion module is used to fuse the base large language model and the third model using the first dataset and the high-quality inference data based on task vector operations to obtain the inference-type industry large language model of the target industry.
[0012] Optionally, the data acquisition module is specifically used to convert each industry knowledge graph data into a text expression in natural language using a rule-based method, and based on the converted industry knowledge graph data, to control the second model to generate multiple inference paths to be identified using a prompting method; the data acquisition module is further used to extract the final inference answer from the multiple inference paths to be identified, match the obtained final inference answer with the corresponding correct answer in the industry knowledge graph data, and remove the inference paths that do not match the correct answer from the multiple inference paths to be identified, thereby obtaining the first dataset.
[0013] Optionally, the model training module is specifically used to recall similar text fragments in the industry knowledge graph data that meet a preset threshold in similarity to the first dataset using a hybrid method of dense retrieval and sparse retrieval; the model training module is also specifically used to optimize the expression of the second model instruction data based on the similar text fragments, and complete one training of the second model.
[0014] Optionally, the model fusion module is specifically used to fine-tune the base large language model using the high-quality inference data to obtain first model parameters, and to fine-tune the third model using the first dataset to obtain second model parameters; the model fusion module is further used to calculate the difference between the second model parameters and the first model parameters to obtain a task vector, and to add the task vector to the second model parameters after scaling by adjustment coefficients to obtain third model parameters; the model fusion module is further used to use the third model parameters as model parameters of the model after fusing the base large language model and the third model to obtain the inference-based industry large language model.
[0015] Optionally, the device further includes: a model update module; the data acquisition module is further configured to acquire a second dataset generated during the use of the reasoning-based industry language model; the second dataset is: a dataset obtained by manually annotating answers to erroneous reasoning samples generated during the use of the reasoning-based industry language model; the model update module is configured to use the second dataset to perform reinforcement training on the reasoning-based industry language model to obtain a perfected reasoning-based industry language model.
[0016] This application also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the reinforcement learning-based industry large language model training method described above.
[0017] This application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the reinforcement learning-based industry large language model training method described above.
[0018] This application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the reinforcement learning-based industry large language model training method described above.
[0019] The industry-specific large language model training method and apparatus based on reinforcement learning provided in this application firstly trains a base large language model using unstructured text data of the target industry through incremental pre-training to obtain a first model; the base large language model is trained based on general domain knowledge text data. Then, the first model is trained once using high-quality inference data of the target industry through instruction fine-tuning, and then trained a second time using reinforcement learning to obtain a second model. A rejection sampling method is then used to generate a first dataset using the second model; the first dataset contains multiple inference paths from industry knowledge graph data. Finally, the second model is trained once using the first dataset through instruction fine-tuning, and then trained a second time using reinforcement learning to obtain a third model. Based on task vector operations, the base large language model and the third model are fused using the first dataset and the high-quality inference data to obtain the inference-based industry-specific large language model for the target industry. Thus, by optimizing the model training strategy and the construction method of industry inference data, the inference and generalization capabilities of the industry-specific large language model are greatly improved, while reducing the model iteration cost. Attached Figure Description
[0020] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1 This is a flowchart illustrating the industry-specific large language model training method based on reinforcement learning provided in this application. Figure 2 This is a schematic diagram of the structure of the industry-specific large language model training device based on reinforcement learning provided in this application; Figure 3 This is a schematic diagram of the structure of the electronic device provided in this application. Detailed Implementation
[0022] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0023] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0024] Currently, improving the reasoning capabilities of large language models in the industry faces four main technical challenges: ① The reinforcement learning mechanism is complex, and there is a lack of practical technical routes to improve industry reasoning capabilities; ② Industry data often exists in the form of unstructured text and structured knowledge graphs, making it difficult to integrate this data to generate high-quality datasets that can be used to improve reasoning capabilities; ③ Logical reasoning not only requires the model to provide correct answers, but also requires the model to generate clear, reasonable, and reproducible reasoning paths. However, industry models lack generalization ability, which easily leads to incomplete reasoning paths and incorrect reasoning, reducing the credibility of reasoning results; ④ In real-world application scenarios, users will discover reasoning errors during their interaction with large language models.
[0025] To address the aforementioned technical problems in related technologies, this application provides a reinforcement learning-based method for training industry-specific large language models. This method solves the four major technical challenges of training reasoning-based industry-specific large language models by optimizing the model's training strategy, proposing an industry-specific inference data construction method, a parallel enhancement method for the model's industry capabilities and generalization capabilities, and a low-cost model iteration method.
[0026] The following description, in conjunction with the accompanying drawings, details the industry-specific language model training method based on reinforcement learning provided in this application through specific embodiments and application scenarios.
[0027] like Figure 1As shown in the embodiment of this application, a method for training an industry-specific large language model based on reinforcement learning is provided. This method may include the following steps 101 to 103: Step 101: Train the base large language model using the unstructured text data of the target industry through incremental pre-training to obtain the first model.
[0028] The base language model is trained based on general domain knowledge text data.
[0029] For example, the target industry mentioned above can be any industry including manufacturing, insurance, advertising, and construction. Before training an industry-specific large language model for a specific industry (i.e., the target industry mentioned above), it is necessary to first collect relevant data, including: industry unstructured text data, industry knowledge graph data, and high-quality inference data. Then, the collected data is used to further train the base large language model (e.g., DeepSeek, GPT, Tongyi Qianwen, Wenxin Yiyan, etc.) for the specific industry, training it into an industry-specific large language model capable of solving specific industry problems.
[0030] It should be noted that, in this embodiment, three types of key data are first collected: industry unstructured text data, industry knowledge graph data, and high-quality inference data. This data is acquired through web crawling, industry databases, and collaboration with experts, covering everything from industry news to professional papers and expert-verified inference cases. The collected data undergoes preprocessing, such as content cleaning and structuring, to ensure accurate and efficient input resources for model training. This data foundation supports the model in deeply understanding industry-specific knowledge and accurately executing subsequent incremental pre-training, instruction fine-tuning, and reinforcement learning steps.
[0031] For example, training a base-based large language model using unstructured text data from a target industry through incremental pre-training to obtain a first model may include the following steps: First, collecting a large amount of unstructured text data specific to the industry. This data may include industry reports, technical articles, forum discussions, and other relevant documents, ensuring broad coverage of industry knowledge and terminology. Subsequently, this unstructured text data is used to incrementally pre-train a base-based large language model. Specifically, a fine-tuning approach is employed, where industry-specific text data is input into an existing large-scale general language model, and model parameters are adjusted to better adapt to the industry context and expertise. This process may involve adjusting training parameters such as the learning rate, optimization algorithm, and batch size to optimize the model's performance on industry-specific text. After this stage of incremental pre-training, the resulting first model will possess the initial ability to understand and process industry-specific information, providing a foundation for subsequent deep training and applications.
[0032] Step 102: Use the high-quality inference data of the target industry to train the first model once through instruction fine-tuning, and then use reinforcement learning to train it a second time to obtain the second model. Then use the rejection sampling method to generate the first dataset using the second model.
[0033] The first dataset contains multiple reasoning paths from industry knowledge graph data.
[0034] For example, in the step of training the first model to obtain the second model, the first model, which has already been obtained through incremental pre-training, is used in this embodiment of the application to perform instruction fine-tuning using high-quality inference data. This inference data is extracted from real-world scenarios in a professional field and has been precisely labeled by experts to ensure the quality and usability of the data.
[0035] For example, during instruction fine-tuning, the model learns how to find or deduce the correct answer in a given context based on the instructions and questions in the inference data, thereby improving its understanding and handling of complex problems. After instruction fine-tuning is completed, the model is further trained using reinforcement learning methods to optimize the model's decision-making process.
[0036] For example, in the reinforcement learning phase, the model is treated as an agent that performs tasks in a simulated environment, where the environment provides positive or negative rewards for each decision made by the model. In this way, the model learns how to maximize positive rewards across a series of decisions, thereby optimizing its long-term decision-making strategy. The key to this phase is designing an effective reward mechanism that ensures the model can improve inference accuracy while also optimizing the logic and efficiency of the solution process. For instance, correct inference steps are rewarded positively, while incorrect inferences or inefficient solution paths are penalized. Through continuous trial and error and the accumulation of rewards, the model gradually learns to generate more accurate and efficient inference paths. Ultimately, through this series of training steps, we obtain a second model that not only demonstrates higher accuracy in understanding and performing industry-specific tasks but also exhibits higher efficiency and intelligence in automated inference.
[0037] Specifically, step 102 above, which uses the rejection sampling method to generate the first dataset using the second model, may further include the following steps 102a1 and 102a2: Step 102a1: Use a rule-based method to convert each industry knowledge graph data into a text expression in natural language, and use a prompting method to control multiple inference paths to be identified in the second model to generate industry knowledge graph data based on the converted industry knowledge graph data.
[0038] For example, to facilitate subsequent data processing, the data in the industry knowledge graph first needs to be converted into natural language. This can be achieved through predefined rules, such as converting relationships and entities in the knowledge graph into easily understandable narrative text. For instance, if there is an entity relationship in the knowledge graph such as "Company A owns technology B," it can be converted into "Company A is the owner of technology B." Then, using a second model, prompt-based reasoning path generation is performed multiple times. This process involves constructing specific prompts to guide the model in reasoning along predetermined knowledge paths. For example, a question about a specific technology or market trend can be provided, requiring the model to explore possible causal relationships or historical context, thereby generating multiple reasoning paths.
[0039] Step 102a2: Extract the final reasoning answer from the multiple reasoning paths to be identified, and match the obtained final reasoning answer with the corresponding correct answer in the industry knowledge graph data. Remove the reasoning paths that do not match the correct answer from the multiple reasoning paths to be identified to obtain the first dataset.
[0040] For example, after obtaining multiple inference paths to be identified, a regular expression-based method is used to extract the final inference answer from the inference path. If it cannot match the corresponding correct answer in the knowledge graph data, it is discarded. After generating the inference path, the next task is to extract the final inference answer from the path and match it with the correct answer in the knowledge graph. This step can be implemented using regular expressions to accurately extract the answer from the inference path. If the generated answer does not match the correct answer in the knowledge graph, the corresponding inference path is discarded. This process helps ensure that the final constructed first dataset contains only accurate and valid inference paths.
[0041] Step 103: Train the second model once using the first dataset through instruction fine-tuning, and then train it a second time using reinforcement learning to obtain the third model. Based on task vector operations, fuse the base large language model and the third model using the first dataset and the high-quality inference data to obtain the inference-type industry large language model for the target industry.
[0042] For example, a first dataset obtained through a rejection sampling method, which contains correct reasoning paths from multiple industry knowledge graph data, is used to fine-tune and train a second model via instructions.
[0043] Specifically, step 103 above, which involves training the second model once using the first dataset through fine-tuning via instructions, may further include the following steps 103a1 and 103a2: Step 103a1: Use a combination of dense and sparse retrieval methods to recall similar text fragments in the industry knowledge graph data that meet a preset threshold in similarity to the first dataset.
[0044] Step 103a2: Optimize the expression of the second model instruction data based on the similar text fragments to complete the first training of the second model.
[0045] For example, during the instruction fine-tuning training process, the model is optimized based on the inference paths and corresponding correct results in the first dataset, thereby improving the model's reasoning ability on specific industry knowledge. After completing the instruction fine-tuning, reinforcement learning methods are then used to further train the fine-tuned second model.
[0046] For example, during the reinforcement learning training phase, the model learns through interaction with the environment and continuously adjusts its parameters through a reward mechanism to enhance its performance and adaptability in complex reasoning tasks. After this series of training steps, a more refined and efficient third model is obtained, which can more accurately handle and solve industry-specific reasoning problems.
[0047] Specifically, in step 103 above, the step of fusing the base large language model and the third model using the first dataset and the high-quality inference data based on task vector operations to obtain the inference-type industry large language model for the target industry may further include the following steps 103a1 to 103a3: Step 103a1: Use the high-quality inference data to fine-tune the base large language model to obtain the first model parameters, and use the first dataset to fine-tune the third model to obtain the second model parameters.
[0048] Step 103a2: Calculate the difference between the second model parameters and the first model parameters to obtain the task vector, and add the task vector to the second model parameters after scaling by adjustment coefficients to obtain the third model parameters.
[0049] Step 103a3: Use the third model parameters as the model parameters of the model after fusing the base big language model and the third model to obtain the reasoning industry big language model.
[0050] For example, in the process of model fusion, the base large language model is first selected. And the third model obtained after training through the previous steps This serves as the input model for this step. It provides a broad foundation for language comprehension, and This includes industry-specific knowledge and reasoning ability, and the combination of these two is to generate a more complex and efficient model.
[0051] For example, Fine-tuning of instructions is performed on high-quality inference data, which covers a variety of complex inference scenarios and aims to improve... The expert's domain-specific reasoning ability. After fine-tuning, the adjusted parameter set was obtained. These parameters will serve as the basis for comparison in subsequent steps. The third model... Fine-tuning is performed on the first dataset. This first dataset, generated in step S4, contains filtered, effective inference paths, which helps the model learn deeply on specific data. After fine-tuning, the adjusted parameter set for the third model is obtained. .
[0052] For example, after obtaining the parameter set adjusted from the base large language model (i.e., the parameters of the first model mentioned above), and the parameter set adjusted for the third model. After obtaining the second model parameters (as mentioned above), model fusion can proceed. First, the task vector needs to be calculated. Through calculation and The difference is obtained from _base, i.e. :. This vector This represents the parameter adjustments required for everything from general language processing to specific industry applications. It captured incremental changes in industry-specific knowledge and reasoning abilities, providing crucial targeted optimization guidance for model fusion. Subsequently, the fused model was computed. parameters (i.e., the third model parameter mentioned above), calculated as follows: .
[0053] For example, the fused model (i.e., the parameters of the aforementioned reasoning-based industry-wide language model) It is by... k times the task vector The values are calculated by addition. Here, k is a hyperparameter used to adjust the influence of the task vector on the final model parameters. In this way, Not only inherited Its professional reasoning ability, along with adjustments made to differentiate it from the base model, enhances its adaptability and responsiveness to specific industry scenarios.
[0054] Optionally, in this embodiment of the application, after obtaining the reasoning-based industry-wide language model, iterative processing can be performed based on user feedback.
[0055] For example, after step 103 above, the industry-specific large language model training method based on reinforcement learning provided in this application embodiment may further include the following steps 104 and 105: Step 104: Obtain the second dataset generated during the use of the inference-based industry large language model.
[0056] The second dataset is a dataset obtained by manually annotating the answers to erroneous inference samples generated during the use of the inference-based industry large language model.
[0057] Step 105: Use the second dataset to perform reinforcement training on the reasoning-based industry large language model to obtain a more complete reasoning-based industry large language model.
[0058] For example, erroneous reasoning samples generated by users during the use of the reasoning-based industry large language model are collected, and the answers are labeled by experts to construct a second dataset. The reasoning-based industry large language model is then periodically trained using reinforcement learning on the second dataset in order to update and iterate the model.
[0059] For example, the following is a brief introduction to the use of the model: User Question: In the current environment of global financial instability, how can blockchain technology be used to enhance the transparency and security of banking operations? Model's answer: "Yes, the user is asking how blockchain technology can enhance the transparency and security of the banking industry given the current global financial instability. First, I need to understand why they care about this issue. The user may be a banking professional or someone interested in fintech. They may have noticed the lack of transparency in traditional banking systems, such as opaque transactions and complex audits, or they may be concerned about security, such as data tampering and fraud risks. Blockchain, as a distributed ledger technology, does indeed have advantages in transparency and security, but how can it be specifically applied?" Next, I should consider the specific application scenarios that users might need. These include cross-border payments, identity verification, smart contracts, and data storage. How can transparency and security be improved in each scenario? For example, cross-border payments can be tracked in real time via blockchain, reducing intermediaries and lowering costs. Smart contracts execute automatically, reducing human error or fraud.
[0060] Next, consider the core advantages of blockchain technology: immutability, traceability, distributed storage, and cryptographic security. How can these characteristics address the pain points of the banking industry? For example, once transaction records are on the blockchain, they cannot be altered, making auditing easier and improving transparency. Distributed storage reduces single points of failure and enhances system security.
[0061] Users may also be concerned about the challenges of implementing blockchain, such as regulatory compliance, technology integration, and performance issues. However, their concerns tend to focus more on solutions, so the emphasis may need to be on viable applications rather than the challenges themselves. If the opportunity arises, it could be briefly mentioned, for example, the choice between private and consortium blockchains to comply with regulatory requirements.
[0062] It's also necessary to consider different types of blockchains—public, private, and consortium blockchains—and which is more suitable for the banking industry. The banking industry may prefer private or consortium blockchains because it needs to control access permissions while maintaining a certain level of transparency. For example, consortium blockchains allow multiple banks to share ledgers, improving the efficiency of cross-border cooperation, while protecting privacy by keeping data confidential.
[0063] The industry-specific large language model training method based on reinforcement learning provided in this application firstly trains a base large language model using unstructured text data of the target industry through incremental pre-training to obtain a first model. The base large language model is trained based on general domain knowledge text data. Then, the first model is trained once using high-quality inference data of the target industry through instruction fine-tuning, followed by a second training using reinforcement learning to obtain a second model. A rejection sampling method is then used to generate a first dataset using the second model. The first dataset contains multiple inference paths from industry knowledge graph data. Finally, the second model is trained once using the first dataset through instruction fine-tuning, followed by a second training using reinforcement learning to obtain a third model. Based on task vector operations, the base large language model and the third model are fused using the first dataset and the high-quality inference data to obtain a reasoning-based industry-specific large language model for the target industry. Thus, by optimizing the model training strategy and the construction method of industry inference data, the reasoning and generalization capabilities of the industry-specific large language model are greatly improved, while reducing the model's iteration cost.
[0064] It should be noted that the industry-specific large language model training method based on reinforcement learning provided in this application embodiment can be executed by an industry-specific large language model training device based on reinforcement learning, or by a control module within that device for executing the industry-specific large language model training method based on reinforcement learning. This application embodiment uses the execution of the industry-specific large language model training method based on reinforcement learning by an industry-specific large language model training device as an example to illustrate the industry-specific large language model training device based on reinforcement learning provided in this application embodiment.
[0065] It should be noted that, in the embodiments of this application, the methods illustrated in the accompanying drawings for training industry-specific large language models based on reinforcement learning are all exemplified by referring to one of the accompanying drawings in the embodiments of this application. In specific implementation, the methods illustrated in the accompanying drawings for training industry-specific large language models based on reinforcement learning can also be implemented in conjunction with any other accompanying drawings shown in the above embodiments, which will not be elaborated here.
[0066] The following describes the industry-specific large language model training device based on reinforcement learning provided in this application. The following description corresponds to the above description of the industry-specific large language model training method based on reinforcement learning.
[0067] Figure 2 A schematic diagram of the structure of the industry-specific large language model training device based on reinforcement learning provided in the embodiments of this application is shown below. Figure 2 As shown, it specifically includes: The model training module 201 is used to train the base large language model using unstructured text data of the target industry through incremental pre-training to obtain a first model; the base large language model is trained based on general domain knowledge text data; the model training module 201 is also used to train the first model once using high-quality inference data of the target industry through instruction fine-tuning, and then use reinforcement learning methods for a second training to obtain a second model; the data acquisition module 202 is used to generate a first dataset using the second model using a rejection sampling method; the first dataset contains multiple inference paths of industry knowledge graph data; the model training module 201 is also used to train the second model once using the first dataset through instruction fine-tuning, and then use reinforcement learning methods for a second training to obtain a third model; the model fusion module 203 is used to fuse the base large language model and the third model using the first dataset and the high-quality inference data based on task vector operations to obtain the inference-type industry large language model of the target industry.
[0068] Optionally, the data acquisition module 202 is specifically used to convert each industry knowledge graph data into a text expression in natural language using a rule-based method, and based on the converted industry knowledge graph data, to control the second model to generate multiple inference paths to be identified using a prompting method; the data acquisition module 202 is also specifically used to extract the final inference answer from the multiple inference paths to be identified, and to match the obtained final inference answer with the corresponding correct answer in the industry knowledge graph data, and to remove the inference paths that do not match the correct answer from the multiple inference paths to be identified, thereby obtaining the first dataset.
[0069] Optionally, the model training module 201 is specifically used to recall similar text fragments in the industry knowledge graph data that meet a preset threshold in similarity to the first dataset using a hybrid method of dense retrieval and sparse retrieval; the model training module 201 is also specifically used to optimize the expression of the second model instruction data based on the similar text fragments, and complete one training of the second model.
[0070] Optionally, the model fusion module 203 is specifically used to fine-tune the base large language model using the high-quality inference data to obtain first model parameters, and to fine-tune the third model using the first dataset to obtain second model parameters; the model fusion module 203 is further used to calculate the difference between the second model parameters and the first model parameters to obtain a task vector, and to add the task vector to the second model parameters after scaling by adjustment coefficients to obtain third model parameters; the model fusion module 203 is further used to use the third model parameters as model parameters of the model after fusing the base large language model and the third model to obtain the inference-based industry large language model.
[0071] Optionally, the device further includes: a model update module; the data acquisition module 202 is further configured to acquire a second dataset generated during the use of the reasoning-based industry language model; the second dataset is: a dataset obtained by manually annotating answers to erroneous reasoning samples generated during the use of the reasoning-based industry language model; the model update module is configured to use the second dataset to perform reinforcement training on the reasoning-based industry language model to obtain a perfected reasoning-based industry language model.
[0072] The reinforcement learning-based industry-specific large language model training device provided in this application first trains a base large language model using unstructured text data of the target industry through incremental pre-training to obtain a first model. The base large language model is trained based on general domain knowledge text data. Then, the first model is trained once using high-quality inference data of the target industry through instruction fine-tuning, followed by a second training using reinforcement learning to obtain a second model. A rejection sampling method is then used to generate a first dataset from the second model. The first dataset contains multiple inference paths from industry knowledge graph data. Finally, the second model is trained once using the first dataset through instruction fine-tuning, followed by a second training using reinforcement learning to obtain a third model. Based on task vector operations, the base large language model and the third model are fused using the first dataset and the high-quality inference data to obtain a reasoning-based industry-specific large language model for the target industry. Thus, by optimizing the model training strategy and the construction method of industry inference data, the reasoning and generalization capabilities of the industry-specific large language model are greatly improved, while reducing the model's iteration cost.
[0073] Figure 3 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 3As shown, the electronic device may include: a processor 310, a communications interface 320, a memory 330, and a communications bus 340, wherein the processor 310, the communications interface 320, and the memory 330 communicate with each other through the communications bus 340. The processor 310 can call logic instructions in the memory 330 to execute a reinforcement learning-based industry-specific large language model training method. This method includes: first, training a base large language model using unstructured text data from the target industry through incremental pre-training to obtain a first model; the base large language model is trained based on general domain knowledge text data; then, training the first model once using high-quality inference data from the target industry through instruction fine-tuning, followed by a second training using reinforcement learning to obtain a second model, and using a rejection sampling method to generate a first dataset; the first dataset contains multiple inference paths from industry knowledge graph data; finally, training the second model once using the first dataset through instruction fine-tuning, followed by a second training using reinforcement learning to obtain a third model, and fusing the base large language model and the third model using the first dataset and the high-quality inference data based on task vector operations to obtain a reasoning-based industry-specific large language model for the target industry. Thus, by optimizing the model's training strategy and the construction method of industry inference data, the reasoning and generalization capabilities of the industry-specific large language model are greatly improved, while reducing the model's iteration cost.
[0074] Furthermore, the logical instructions in the aforementioned memory 330 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0075] On the other hand, this application also provides a computer program product, which includes a computer program stored on a computer-readable storage medium. The computer program includes program instructions, and when the program instructions are executed by a computer, the computer can execute the reinforcement learning-based industry large language model training method provided by the above methods. The method includes: first, training a base large language model using unstructured text data of the target industry through incremental pre-training to obtain a first model; the base large language model is trained based on general domain knowledge text data; then, training the first model once using high-quality inference data of the target industry through instruction fine-tuning, and then performing a second training using reinforcement learning to obtain a second model, and using a rejection sampling method to generate a first dataset using the second model; the first dataset contains multiple inference paths of industry knowledge graph data; finally, training the second model once using the first dataset through instruction fine-tuning, and then performing a second training using reinforcement learning to obtain a third model, and fusing the base large language model and the third model using the first dataset and the high-quality inference data based on task vector operations to obtain the inference-based industry large language model of the target industry. In this way, by optimizing the model's training strategy and the construction method of industry inference data, the reasoning ability and generalization ability of the industry-specific large language model are greatly improved, and the iteration cost of the model is reduced.
[0076] In another aspect, this application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the aforementioned reinforcement learning-based industry-specific large language model training methods. The method includes: first, training a base large language model using unstructured text data of the target industry through incremental pre-training to obtain a first model; the base large language model is trained based on general domain knowledge text data; then, training the first model once using high-quality inference data of the target industry through instruction fine-tuning, and then performing a second training using reinforcement learning to obtain a second model, and using a rejection sampling method to generate a first dataset using the second model; the first dataset contains multiple inference paths from industry knowledge graph data; finally, training the second model once using the first dataset through instruction fine-tuning, and then performing a second training using reinforcement learning to obtain a third model, and fusing the base large language model and the third model based on task vector operations using the first dataset and the high-quality inference data to obtain a reasoning-based industry-specific large language model for the target industry. In this way, by optimizing the model's training strategy and the construction method of industry inference data, the reasoning ability and generalization ability of the industry-specific large language model are greatly improved, and the iteration cost of the model is reduced.
[0077] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0078] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0079] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A method for training industry-specific large language models based on reinforcement learning, characterized in that, include: The base large language model is trained using unstructured text data from the target industry through incremental pre-training to obtain the first model; The base language model is trained based on general domain knowledge text data; The first model is trained once using high-quality inference data from the target industry through instruction fine-tuning, and then trained a second time using reinforcement learning to obtain a second model. The second model is then used to generate a first dataset using a rejection sampling method. The first dataset contains multiple inference paths from industry knowledge graph data. The second model is trained once using the first dataset through instruction fine-tuning, and then trained a second time using reinforcement learning to obtain the third model. Based on task vector operations, the base large language model and the third model are fused using the first dataset and the high-quality inference data to obtain the inference-type industry large language model for the target industry. The method of using rejection sampling to generate the first dataset using the second model includes: A rule-based approach is used to convert each industry knowledge graph data into a textual expression in natural language. Based on the converted industry knowledge graph data, a prompting method is used to control multiple inference paths to be identified in the industry knowledge graph data generated by the second model. Extract the final reasoning answer from the multiple reasoning paths to be identified, and match the obtained final reasoning answer with the corresponding correct answer in the industry knowledge graph data. Remove the reasoning paths that do not match the correct answer from the multiple reasoning paths to be identified to obtain the first dataset. The process of fusing the base large language model and the third model to obtain the reasoning-based industry large language model for the target industry includes: The high-quality inference data is used to fine-tune the base large language model to obtain the first model parameters, and the first dataset is used to fine-tune the third model to obtain the second model parameters. The difference between the second model parameters and the first model parameters is calculated to obtain the task vector. The task vector is then scaled by an adjustment factor and added to the second model parameters to obtain the third model parameters. The third model parameters are used as model parameters for the model after fusing the base large language model with the third model to obtain the reasoning-based industry large language model.
2. The method according to claim 1, characterized in that, The step of training the second model once using the first dataset through instruction fine-tuning includes: A method combining dense and sparse retrieval is used to recall similar text fragments in the industry knowledge graph data that meet a preset threshold of similarity with the first dataset. The expression of the second model instruction data is optimized based on the similar text fragments, and the second model is trained once.
3. The method according to claim 1 or 2, characterized in that, After fusing the base large language model and the third model using the first dataset and the high-quality inference data based on task vector operations to obtain the inference-based industry large language model for the target industry, the method further includes: Obtain the second dataset generated during the use of the reasoning-based industry language model; the second dataset is: a dataset obtained by manually annotating the answers to the erroneous reasoning samples generated during the use of the reasoning-based industry language model. The reasoning-based industry-wide language model is reinforced and trained using the second dataset to obtain a more complete reasoning-based industry-wide language model.
4. A training device for a large industry language model based on reinforcement learning, characterized in that, The device includes: The model training module is used to train the base language model using unstructured text data of the target industry through incremental pre-training to obtain the first model; the base language model is trained based on general domain knowledge text data. The model training module is also used to train the first model once using high-quality inference data from the target industry through instruction fine-tuning, and then to train it a second time using reinforcement learning methods to obtain the second model. The data acquisition module is used to generate a first dataset using the second model with the method of rejection sampling; the first dataset contains multiple inference paths of industry knowledge graph data; The model training module is also used to train the second model once using the first dataset through instruction fine-tuning, and then to train it a second time using reinforcement learning methods to obtain the third model. The model fusion module is used to fuse the base large language model and the third model based on task vector operations using the first dataset and the high-quality inference data to obtain the inference-type industry large language model of the target industry; The data acquisition module is specifically used to convert each industry knowledge graph data into a text expression in natural language using a rule-based method, and based on the converted industry knowledge graph data, to control multiple inference paths to be identified in the second model by using a prompting method. The data acquisition module is further configured to extract the final reasoning answer from the multiple reasoning paths to be identified, match the obtained final reasoning answer with the corresponding correct answer in the industry knowledge graph data, and remove the reasoning paths that do not match the correct answer from the multiple reasoning paths to be identified to obtain the first dataset. The model fusion module is specifically used to fine-tune the base large language model using the high-quality inference data to obtain first model parameters, and to fine-tune the third model using the first dataset to obtain second model parameters; the model fusion module is also specifically used to calculate the difference between the second model parameters and the first model parameters to obtain a task vector, and to add the task vector to the second model parameters after scaling by adjustment coefficients to obtain third model parameters; the model fusion module is also specifically used to use the third model parameters as the model parameters of the model after fusing the base large language model and the third model to obtain the inference-based industry large language model.
5. The apparatus according to claim 4, characterized in that, The model training module is specifically used to recall similar text fragments in the industry knowledge graph data that meet a preset threshold in the first dataset using a hybrid method of dense retrieval and sparse retrieval. The model training module is further used to optimize the expression of the second model instruction data based on the similar text fragments, thereby completing one training of the second model.
6. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the reinforcement learning-based industry large language model training method as described in any one of claims 1 to 3.
7. A computer-readable storage medium, characterized in that, It stores a computer program that, when executed by a processor, implements the steps of the reinforcement learning-based industry large language model training method as described in any one of claims 1 to 3.
Citation Information
Patent Citations
Open domain natural language reasoning question-answering system and method driven by large language model
CN116932708A
Bridge field intelligent question-answering method, bridge field generation type large model training method and training device
CN118228822A