Industrial question answering model training method based on reinforcement learning and knowledge base matching

By building a knowledge base in the industrial field and using reinforcement learning training methods to optimize the strategy of industrial Q&A model, the difficulty of large language models in understanding professional knowledge in the industrial field is solved, and higher Q&A accuracy and adaptability are achieved.

WO2025148471A1PCT designated stage expired Publication Date: 2025-07-17NANJING UNIV OF SCI & TECH

Patent Information

Application Number
PCT/CN2024/126707
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-10
Filing Date
2024-10-23
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

Existing large language models are difficult to accurately understand expertise in the industrial field, resulting in problems of inauthentic or invalid outputs.

Method used

Using reinforcement learning and knowledge base matching methods, we use industrial knowledge bases to train industrial Q&A models using sorting loss functions and punishment terms, and combine Actor-Critic network and PPO algorithm for multiple iterative training to optimize model strategies.

Benefits of technology

The accuracy and accuracy of the industrial question-and-answer model for industrial expertise issues is improved, untrue or invalid output is reduced, and the adaptability of the model is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024126707_17072025_PF_FP_ABST
    Figure CN2024126707_17072025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed is an industrial question answering model training method based on reinforcement learning and knowledge base matching, comprising the following steps: S1, collecting professional knowledge questions and answers in an industrial field to construct an industrial knowledge base, training a reward model, carrying out, for industrial knowledge questions and answers, matching comparison on outputs of an industrial question answering model and content of the industrial knowledge base, and obtaining reward values on the basis of similarities; S2, sorting the reward values, and using a sorting loss function to train and update parameters of a reward model network; and S3, carrying out industrial question answering model training, incorporating a penalty term for the reward values, and using a reinforcement learning algorithm to train the industrial question answering model multiple times to obtain an optimal strategy. According to the industrial question answering model training method based on reinforcement learning and knowledge base matching of the present invention, the reinforcement learning algorithm is used, and iterative training is carried out multiple times, thereby helping the industrial question answering model to learn and understand industrial professional knowledge and improving the question answering accuracy of the industrial question answering model.
Need to check novelty before this filing date? Find Prior Art

Description

A training method for industrial question answering models based on reinforcement learning and knowledge base matching Technical Field

[0001] The present invention relates to the technical field of large-scale language training, and in particular to an industrial question-answering model training method based on reinforcement learning and knowledge base matching. Background Art

[0002] On the one hand, industrial knowledge is vast and complex, encompassing multiple disciplines and technical fields, such as mechanics, electronics, and chemical engineering. On the other hand, this knowledge often contains a large amount of specialized terminology and complex technical language, making it difficult to understand and communicate. For humans, understanding and mastering this specialized knowledge in industrial fields is a huge challenge.

[0003] To overcome these difficulties, reinforcement learning and industrial knowledge bases can be used to train large language models to create industrial question-answering models that serve the industrial sector, capture industry expertise, and answer specialized questions. Large language models are a type of model trained using deep learning techniques that can generate coherent, grammatically, and contextually sound text. However, large language models also have some issues and can sometimes produce output that is unrealistic, toxic, or unhelpful to users. This is because they do not accurately understand industrial expertise, but instead generate predictions based on patterns in the training data. During training, the model may amplify biases in its dataset.

[0004] Applying reinforcement learning to large-scale language models can make them more intelligent and adaptive. Traditional large-scale language models only make predictions based on historical text data, but reinforcement learning allows large-scale language models to adjust their predictions based on real-time feedback. In industrial knowledge question answering, industrial knowledge bases provide accurate industry domain expertise that matches large-scale language models. Reinforcement learning can help large-scale language models learn to generate more appropriate responses based on the conversation flow and feedback from the industrial knowledge base, thereby creating an industrial question answering model that serves the industrial sector.

[0005] Therefore, it is necessary to provide an industrial question answering model training method based on reinforcement learning and knowledge base matching to solve the above problems. Summary of the Invention

[0006] The purpose of this invention is to provide an industrial question-answering model training method based on reinforcement learning and knowledge base matching, which improves the accuracy and precision of the industrial question-answering model for industrial professional knowledge questions and provides important technical support for intelligent question-answering in the industrial field.

[0007] To achieve the above objectives, the present invention provides an industrial question-answering model training method based on reinforcement learning and knowledge base matching, comprising the following steps:

[0008] S1. Build an industrial knowledge base and train the reward model. For industrial knowledge question answering, match and compare the output of the industrial question answering model with the content of the industrial knowledge base, and derive the reward value based on the similarity.

[0009] S2. Arrange the reward values ​​in order and use the sorting loss function to train and update the parameters of the reward model network;

[0010] S3. Train the industrial question-answering model, add a penalty term to the reward value, and use the reinforcement learning algorithm to train the industrial question-answering model multiple times to obtain the optimal strategy.

[0011] Preferably, in step S1, professional knowledge questions and answers in the industrial field are collected to build an industrial knowledge base, the reward model is trained, the output of the industrial question-answering model is matched and compared with the content of the industrial knowledge base, and the reward value is obtained according to the similarity function. , The expression is as follows:

[0012] ;

[0013] in, For prior knowledge, Different responses for industrial question answering models.

[0014] Preferably, in step S2, the parameters of the reward model network are updated by training using a ranking loss function, and the ranking loss function expression is as follows:

[0015] ;

[0016] in, 、 is the reward value corresponding to different texts, and σ is the Sigmoid function.

[0017] Preferably, in step S3, a penalty term is added to the reward value, and reinforcement learning is used to perform multiple trainings until convergence to obtain the optimal strategy. The specific steps are as follows:

[0018] S31, calculate the reward value according to the reward model in step S1 , add a penalty term to the reward value, then the final reward value The expression is as follows:

[0019] ;

[0020] Among them, β is the penalty term coefficient, The strategy output of cascading a linear fully connected layer to the last layer of the current iterative industrial question answering model, The policy output of a linear fully connected layer cascaded to the last layer of the initial industrial question answering model;

[0021] S32. Use reinforcement learning to perform multiple trainings until the reward function converges, specifically:

[0022] Based on the Actor-Critic network, the reward model is optimized using the reinforcement learning PPO algorithm. The PPO algorithm limits the policy update amplitude by clipping the loss at the proximal ratio:

[0023] ;

[0024] in, is the update amplitude of the strategy:

[0025] ;

[0026] ∈ is a hyperparameter used to control the clipping amplitude, and clip is the clipping function.

[0027] Preferably, in step S32, the Actor-Critic network architecture includes an Actor neural network and a Critic neural network, wherein the Actor neural network is represented as θ and is used to learn the strategy; the Critic neural network is represented as ω and is used to estimate the value function V(S) of the current state S. The specific steps are as follows:

[0028] Initialize the weight parameters of the Actor network and the Critic network, select actions based on the Actor network, observe the rewards and next state returned by the environment, and output actions , calculate the advantage function :

[0029] ;

[0030] Where λ = 1, δt is the time difference error, which is calculated by the Critic neural network:

[0031]

[0032] use Update the parameters of the Actor and Critic neural networks:

[0033] ;

[0034] ;

[0035] in, Represent the learning rates of the Actor network and the Critic network respectively, and γ represents the discount factor;

[0036] Improve the parameters of the Actor network and Critic network through multiple iterations until the reward function converges and outputs the optimal strategy π, otherwise start the next training.

[0037] Therefore, the present invention adopts the above-mentioned industrial question-answering model training method based on reinforcement learning and knowledge base matching, and the beneficial effects are as follows:

[0038] (1) Unlike other language model training methods that require human feedback to label the quality of each model's answers, in the knowledge base-based industrial question-answering model, only the model output needs to be matched and compared with the content of the knowledge base;

[0039] (2) The present invention trains the deep reinforcement learning algorithm PPO algorithm until convergence and obtains an optimal strategy π;

[0040] (3) The data set used by the present invention to train the optimal strategy is professional knowledge in the industrial field, which is highly professional and practical;

[0041] (4) After the present invention obtains the optimal strategy π, it helps the industrial question-answering model learn and understand industrial professional knowledge, effectively improving the accuracy of the industrial question-answering model's questions and answers.

[0042] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] FIG1 is an overall flow chart of an embodiment of an industrial question-answering model training method based on reinforcement learning and knowledge base matching according to the present invention;

[0044] FIG2 is a flowchart of an embodiment of an industrial question-answering model training method based on reinforcement learning and knowledge base matching according to the present invention;

[0045] FIG3 is a diagram of iterative training of a question-answering model according to an embodiment of an industrial question-answering model training method based on reinforcement learning and knowledge base matching according to the present invention;

[0046] FIG4 is a diagram illustrating an implementation of a PPO algorithm based on an Actor-Critic network architecture in an embodiment of an industrial question-answering model training method based on reinforcement learning and knowledge base matching according to the present invention;

[0047] FIG5 is a diagram showing the training results of an embodiment of an industrial question-answering model training method based on reinforcement learning and knowledge base matching according to the present invention. DETAILED DESCRIPTION

[0048] The technical solution of the present invention is further described below with reference to the accompanying drawings and embodiments.

[0049] Unless otherwise defined, technical or scientific terms used in the present invention shall have the same meaning as commonly understood by one of ordinary skill in the art to which the present invention belongs.

[0050] As shown in Figure 1-2, the present invention provides an industrial question-answering model training method based on reinforcement learning and knowledge base matching, including the following steps:

[0051] S1. Collect a large amount of professional knowledge questions and answers in the industrial field to build an industrial knowledge base, train the reward model, match and compare the output of the industrial question-answering model with the content of the industrial knowledge base, and derive the reward value based on the similarity;

[0052] A large amount of professional knowledge questions and answers in industrial fields are collected through manual data collection, including but not limited to professional knowledge questions and answers in mechanical engineering, electrical engineering, materials science, industrial automation and control, production and manufacturing, etc.

[0053] The specific reward value ri is obtained based on the similarity function. The expression of ri is as follows:

[0054] ;

[0055] in, For prior knowledge, Different responses for industrial question answering models.

[0056] This method uses a large-scale language model invocation platform or a development framework for building large-scale language model applications to import an industrial knowledge base, and then uses the industrial knowledge base to query relevant information, expanding the model's knowledge. The industrial question-answering model can adopt various open-source large-scale language models from home and abroad, such as ChatGLM-6B, LLAMA, and Vicuna-13B.

[0057] S2. Arrange the reward values ​​in order and use the corresponding ranking loss function to train and update the parameters of the reward model network. The ranking loss function expression is as follows:

[0058] ;

[0059] in, 、 is the reward value corresponding to different texts, and σ is the Sigmoid function.

[0060] S3. Train the industrial question-answering model, add a penalty term to the reward value, and use the reinforcement learning algorithm to train the industrial question-answering model multiple times to obtain the optimal strategy.

[0061] The specific steps of reinforcement learning are as follows:

[0062] S31, calculate the reward value according to the reward model in step S1 , and in order to prevent the industrial question answering model from deviating from the initial base model, a penalty term is added to the reward value, so the final reward value The expression is as follows:

[0063] ;

[0064] Among them, β is the penalty term coefficient, The strategy output of cascading a linear fully connected layer to the last layer of the current iterative industrial question answering model, The policy output of a linear fully connected layer cascaded to the last layer of the initial industrial question answering model;

[0065] The penalty term is used to penalize the reinforcement learning policy for generating text that deviates significantly from the initial model in each training batch, ensuring that the model outputs reasonably coherent text. Removing this penalty term may cause the model to generate gibberish during optimization, fooling the reward model into providing high rewards.

[0066] S32. To maximize the reward, use reinforcement learning to perform multiple trainings until the reward function converges. Specifically:

[0067] As shown in Figure 3, based on the Actor-Critic network, the reinforcement learning PPO algorithm is used to optimize the reward model. PPO is a trust region optimization algorithm that limits the policy update amplitude by clipping the loss at the proximal ratio to achieve stable and efficient training results: where rt (θ) is the policy update amplitude, to achieve stable and efficient training results:

[0068]

[0069] in, is the update amplitude of the strategy:

[0070] ;

[0071] ∈ is a hyperparameter used to control the clipping amplitude, clip is the clipping function, when When the specified upper and lower limits are exceeded, the function will return the corresponding upper and lower limits;

[0072] As shown in Figure 4, in step S32, the PPO algorithm is implemented based on the Actor-Critic network architecture. The Actor-Critic network architecture includes an Actor neural network and a Critic neural network. The Actor neural network is represented by θ and is used to learn the strategy; the Critic neural network is represented by ω and is used to estimate the value function V(S) of the current state S. The two networks are continuously improved during the training process, and finally the Actor neural network is output as the optimal strategy π. The specific implementation steps are as follows:

[0073] Initialize the weight parameters of the Actor network (for learning strategies) and the Critic network (for estimating value functions), select actions based on the Actor network, observe the rewards and next states returned by the environment, and output actions , calculate the advantage function :

[0074] ;

[0075] Among them, usually is the time difference error, which can be calculated by the Critic neural network:

[0076] ;

[0077] use Update the parameters of the Actor and Critic neural networks:

[0078] ;

[0079] ;

[0080] in, Represent the learning rates of the Actor network and the Critic network respectively, and γ represents the discount factor;

[0081] During the training process, the penalty coefficient β=3 is selected, and the learning rate 、 , hyperparameter ∈ = 0.2, discount factor γ = 0.9;

[0082] As shown in Figure 5, Figure (a) is the reward function convergence curve, and Figure (b) is the loss function curve. Through multiple iterations, multiple interactions and updates, the parameters of the Actor network and the Critic network are gradually improved until the reward function converges and outputs the optimal strategy π, otherwise the next training begins.

[0083] Therefore, the present invention adopts the above-mentioned industrial question-answering model training method based on reinforcement learning and knowledge base matching, aiming to improve the accuracy and precision of the industrial question-answering model for industrial professional knowledge questions. By matching the industrial knowledge base to obtain a reward model, no human feedback is required to label the pros and cons of each answer of the model. Then, the reinforcement learning PPO algorithm is used. After multiple iterative training, the industrial question-answering model is helped to learn and understand industrial professional knowledge, effectively improving the accuracy of the industrial question-answering model's questions and answers.

[0084] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit the same. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that they can still modify or replace the technical solutions of the present invention with equivalents, and these modifications or equivalent replacements cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. An industrial Q&A model training method based on reinforcement learning and knowledge base matching, characterized in that, It includes the following steps: S1. Construct an industrial knowledge base, train the reward model, for industrial knowledge Q&A, match and compare the output of the industrial Q&A model with the content of the industrial knowledge base, and obtain the reward value according to the similarity; S2. Arrange the reward values in order, and use the ranking loss function to train and update the parameters of the reward model network; S3. Conduct industrial Q&A model training, add a penalty term to the reward value, and use the reinforcement learning algorithm to train the industrial Q&A model multiple times to obtain the optimal strategy; In step S1, professional knowledge questions and answers in the industrial field are collected to build an industrial knowledge base, the reward model is trained, the output of the industrial Q&A model is matched and compared with the content of the industrial knowledge base, and the reward value is obtained according to the similarity function , The expression is as follows: ; Among them, it is Prior knowledge, are different answers of the industrial Q&A model; In step S2, when using the ranking loss function to train and update the parameters of the reward model network, the expression of the ranking loss function is as follows: ; Among them, Reward values corresponding to different texts is the Sigmoid function; In step S3, the specific steps of the reinforcement learning are as follows: S31. Calculate the reward value according to the reward model in step S1 , a penalty term is added to the reward value, and then the final reward value The expression is as follows: ; Among them, is the penalty term coefficient, The policy output that cascades a linear fully connected layer for the last layer of the current iterative industrial question-answering model is the policy output of the last layer of the initial industrial Q&A model cascaded with a linear fully connected layer; S32. Use reinforcement learning to train multiple times until the reward function converges. Specifically, based on the Actor-Critic network, use the reinforcement learning PPO algorithm to optimize the policy of the reward model. The PPO algorithm limits the policy update amplitude through the proximal ratio clipping loss: ; Among them, is the update amplitude of the policy: ; is a hyperparameter used to control the clipping amplitude, and clip is the clipping function; In step S32, the Actor-Critic network architecture includes an Actor neural network and a Critic neural network, where the Actor neural network is represented as , for learning strategies; The Critic neural network is represented as , which is used to estimate the value function V(S) of the current state S. The specific steps are as follows: Initialize the weight parameters of the Actor network and the Critic network, select an action according to the Actor network, observe the reward and the next state returned by the environment, and output the action , calculate the advantage function : ; where λ = 1, is the temporal difference error, which is calculated by the Critic neural network: ; Utilize Update the parameters of the Actor and Critic neural networks: ; ; Among them, respectively represent the learning rates of the Actor network and the Critic network, represents the discount factor; Iteratively improve the parameters of the Actor network and the Critic network multiple times until the reward function converges to output the optimal policy , otherwise start the next training.

Citation Information

Patent Citations

  • Dialogue strategy optimization method combined with knowledge enhancement and deep reinforcement learning

    CN113704425A

  • Disturbance award-oriented deep reinforcement learning confrontation defense method

    CN114925850A

  • Method and device for training generative large language model based on knowledge base feedback

    CN117009490A

  • Crawler automatic driving method fusing human feedback information and deep reinforcement learning

    CN117032208A

  • Fine adjustment method, system and equipment based on large language model and medium

    CN117290480A

Cited By

  • Intelligent agent training method based on DDPG

    CN120781914A

  • Method for training model, computer readable storage medium and computer program product

    CN120911638A

  • Knowledge base question and answer vector generation method and system based on meta-learning

    CN121235063A

  • Navigation control method and equipment for underground inspection robot and medium

    CN121300356A

  • Invisible structure parameter optimization method and system based on reinforcement learning and layering strategy

    CN121302941A