Question and answer model training method and device, electronic equipment, storage medium and program product
By distinguishing the inference path of high and low accuracy in question-and-answer model training and calculating the corresponding loss value, the training process of the question-and-answer model is optimized, and the problem of low accuracy caused by a single loss function is solved, which improves the accuracy and reliability of the answer generation of the question-and-answer model.
Patent Information
- Application Number
- CN202510377496.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-07-11
AI Technical Summary
In the prior art, the training methods of question-answer models are usually trained by a single loss function, resulting in the question-answer models being less accurate in solving problems.
By determining the question sample carrying the answer label, multiple candidate inference paths are generated and classified into high-accuracy and low-accuracy paths, the loss values of different paths are calculated to optimize the training process of the question-answer model, including determining the loss values of the first inference path and the second inference path, and training the model based on these loss values.
It improves the performance of the Q&A model on different inference paths, enhances its accuracy and reliability when generating answers, and improves the overall performance of the Q&A model.
Smart Images

Figure CN120297414A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technologies, and in particular, to a method, an apparatus, an electronic device, a storage medium, and a program product for training a question answering model. Background Art
[0002] A question answering model is an artificial intelligence system that aims to understand and answer questions posed by users. Such a model is typically based on natural language processing (NLP) technology and can handle various types of questions, including factual questions, interpretive questions, and advisory questions, etc. The question answering model can extract information from various forms of data such as text, audio, and video, and generate answers that are understandable to humans. They are widely used in fields such as intelligent assistants, search engines, customer service, and educational tutoring to provide users with fast and accurate information acquisition and problem-solving services.
[0003] In the related art, for the training of a question answering model, it is usually to train the question answering model by determining a single loss, which results in poor performance of the trained question answering model in solving problems. Summary of the Invention
[0004] Embodiments of the present application provide a method, an apparatus, an electronic device, a computer-readable storage medium, and a computer program product for training a question answering model, which effectively improve the performance of the question answering model in solving problems.
[0005] The technical solution of the embodiments of the present application is implemented as follows:
[0006] Embodiments of the present application provide a method for training a question answering model, including:
[0007] Based on a first question sample carrying an answer label, determine multiple candidate inference paths, where the candidate inference paths are used to indicate the inference steps that need to be executed to solve the problem corresponding to the first question sample;
[0008] Classify the multiple candidate inference paths to obtain a first inference path belonging to a first category and a second inference path belonging to a second category, where the inference accuracy of the first inference path is greater than that of the second inference path;
[0009] Based on the first inference path and the answer label, determine a first loss value of the question answering model, and based on the first inference path and the second inference path, determine a second loss value of the question answering model;
[0010] Based on the first loss value and the second loss value, train the question answering model to obtain a target question answering model.
[0011] An embodiment of the present application provides a method for generating an answer, including:
[0012] Based on a problem to be solved, generate multiple candidate inference paths through a target question-and-answer model, and through the target question-and-answer model, determine a target inference path for processing the problem from the multiple candidate inference paths;
[0013] Based on the target inference path, answer the problem to be solved through the target question-and-answer model to obtain the answer to the problem;
[0014] Among them, the target question-and-answer model is trained by using the training method of the above-mentioned question-and-answer model.
[0015] An embodiment of the present application provides a training device for a question-and-answer model, including:
[0016] A determination module, configured to determine multiple candidate inference paths based on a first question sample carrying an answer label, where the candidate inference paths are used to indicate the inference steps that need to be executed to solve the problem corresponding to the first question sample;
[0017] A classification module, configured to classify the multiple candidate inference paths to obtain a first inference path belonging to a first category and a second inference path belonging to a second category, where the inference accuracy of the first inference path is greater than that of the second inference path;
[0018] A loss module, configured to determine a first loss value of the question-and-answer model based on the first inference path and the answer label, and determine a second loss value of the question-and-answer model based on the first inference path and the second inference path;
[0019] A training module, configured to train the question-and-answer model based on the first loss value and the second loss value to obtain a target question-and-answer model.
[0020] In the above solution, the training device for the above-mentioned question-and-answer model further includes: a pre-training module, configured to determine multiple third inference paths based on a second question sample through an initial question-and-answer model; determine a fourth inference path from the multiple third inference paths based on the answer label carried by the second question sample; generate a first answer corresponding to the fourth inference path through the initial question-and-answer model based on the fourth inference path; and train the initial question-and-answer model based on the first answer corresponding to the fourth inference path and the answer label carried by the second question sample to obtain the question-and-answer model. The above determination module is further configured to determine the multiple candidate inference paths based on the first question sample through the question-and-answer model.
[0021] In the above solution, the above pre-training module is further configured to, for each of the third inference paths, generate a first answer corresponding to the third inference path through the initial question-answering model based on the third inference path; determine a first similarity between each of the first answers and the answer label carried by the second question sample, and determine the third inference path corresponding to the first answer with the largest first similarity as the fourth inference path.
[0022] In the above solution, the above pre-training module is further configured to determine a third loss value of the initial question-answering model based on the difference between the first answer corresponding to the fourth inference path and the answer label carried by the second question sample; determine a fourth loss value of the initial question-answering model based on each inference step in the fourth inference path; and train the initial question-answering model based on the third loss value and the fourth loss value to obtain the question-answering model.
[0023] In the above solution, the above pre-training module is further configured to determine a first intermediate loss value of the question-answering model based on the first answer and the answer label carried by the second question sample, and train the initial question-answering model based on the first intermediate loss value to obtain a first question-answering model; generate an i-th answer of the second question sample through the (i - 1)-th question-answering model based on the second question sample; determine an i-th intermediate loss value of the question-answering model based on the i-th answer and the answer label carried by the second question sample, and train the (i - 1)-th question-answering model based on the i-th intermediate loss value to obtain an i-th question-answering model; traverse i until the difference between the i-th intermediate loss value and the (i - 1)-th intermediate loss value is less than a difference threshold, and determine the i-th question-answering model as the question-answering model.
[0024] In the above solution, the above pre-training module is further configured to generate multiple initial inference paths through the question-answering model based on the first question sample, where the path length of the initial inference path is less than the path length of the candidate inference path; for each of the initial inference paths, expand the path of the initial inference path through the question-answering model based on the first question sample to obtain a candidate inference path corresponding to the initial inference path.
[0025] In the above solution, the above classification module is further configured to perform perplexity evaluation on each of the candidate inference paths to obtain the perplexity of each of the candidate inference paths, where the size of the perplexity is negatively correlated with the inference accuracy of the candidate inference path; determine the candidate inference path with a perplexity lower than the perplexity threshold as the first inference path, and determine the candidate inference path with a perplexity greater than or equal to the perplexity threshold as the second inference path.
[0026] In the above solution, the above classification module is further configured to perform the following processing on each of the candidate inference paths: generate a candidate answer corresponding to the candidate inference path through the Q&A model, and determine a second similarity between the candidate answer and the answer label; when the second similarity is greater than the similarity threshold, determine the candidate inference path as the first inference path; when the second similarity is less than or equal to the similarity threshold, determine the candidate inference path as the second inference path.
[0027] In the above solution, the loss module is further configured to determine, based on the answer label, a reference inference path with the largest answer similarity to the answer label from the multiple first inference paths; generate a reference answer corresponding to the reference inference path through the Q&A model based on the reference inference path; determine a first target loss value of the Q&A model based on the difference between the reference answer and the answer label; determine a first reference loss value of the Q&A model based on each inference step in the reference inference path; perform a weighted sum of the first target loss value and the first reference loss value to obtain the first loss value.
[0028] In the above solution, the loss module is further configured to, for each of the first inference paths, generate a second answer corresponding to the first inference path through the Q&A model based on the first inference path; determine a first target loss value corresponding to the first inference path based on the difference between the second answer and the answer label; determine a first reference loss value corresponding to the first inference path based on each inference step in the first inference path; perform a weighted sum of the first target loss value and the first reference loss value to obtain the first loss value.
[0029] In the above solution, the loss module is further configured to obtain a first conditional probability of the problem model generating each of the first inference paths under the condition of inputting the first question sample, and obtain a second conditional probability of the problem model generating each of the second inference paths under the condition of inputting the first question sample; determine the sum of the first conditional probabilities as a third conditional probability, and determine the sum of the second conditional probabilities as a fourth conditional probability; divide the third conditional probability by the fourth conditional probability to obtain the second loss value.
[0030] In the above solution, the above loss module is further configured to obtain the first conditional probabilities of generating the respective first inference paths by the problem model under the condition of inputting the first problem sample, and obtain the second conditional probabilities of generating the respective second inference paths by the problem model under the condition of inputting the first problem sample; for each of the first conditional probabilities, divide the first conditional probability by each of the second conditional probabilities to obtain the first results corresponding to the respective second conditional probabilities, and sum the first results to obtain the second result corresponding to the first conditional probability; sum the second results corresponding to the first conditional probabilities to obtain the second loss value.
[0031] In the above solution, the above loss module is further configured to train the question-and-answer model based on the first loss value and the second loss value to obtain a reference question-and-answer model; through the reference question-and-answer model, generate a third answer for the first inference path based on the first inference path; replace the answer label carried in the first problem sample with the third answer to obtain a supplementary problem sample carrying the third answer; train the reference question-and-answer model based on the supplementary problem sample to obtain the target question-and-answer model.
[0032] An embodiment of the present application provides an answer generation device, including:
[0033] A generation module, through a target question-and-answer model, generates multiple candidate inference paths based on a problem to be solved, and determines a target inference path for processing the problem from the multiple candidate inference paths through the target question-and-answer model;
[0034] A solution module, configured to answer the problem to be solved based on the target inference path through the target question-and-answer model to obtain the answer to the problem; wherein, the target question-and-answer model is trained by using the training method of the above question-and-answer model.
[0035] An embodiment of the present application provides an electronic device, including:
[0036] A memory, configured to store computer-executable instructions or a computer program;
[0037] A processor, when executing the computer-executable instructions or the computer program stored in the memory, implements the answer generation method and the training method of the question-and-answer model provided by the embodiment of the present application.
[0038] An embodiment of the present application provides a computer-readable storage medium, storing computer-executable instructions, which are used to cause a processor to implement the answer generation method and the training method of the question-and-answer model provided by the embodiment of the present application when executed.
[0039] An embodiment of the present application provides a computer program product, which includes a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium, and the processor executes the computer-executable instructions, so that the electronic device executes the answer generation method and the question-and-answer model training method described above in the embodiments of the present application.
[0040] The embodiments of the present application have the following beneficial effects:
[0041] By determining multiple candidate inference paths based on the first question samples carrying answer tags, these candidate inference paths indicate the inference steps required to solve the problem. Classify the multiple candidate inference paths to obtain a first inference path belonging to the first category and a second inference path belonging to the second category. The inference accuracy of the first inference path is greater than that of the second inference path, enabling the question-and-answer model to distinguish the quality of different candidate inference paths, which helps the question-and-answer model to focus on higher-quality inference paths during training. Based on the first inference path and the answer tags, determine the first loss value of the question-and-answer model, ensuring that the gap between the output of the question-and-answer model and the correct answer is quantified and used as the basis for optimization. At the same time, based on the first inference path and the second inference path, determine the second loss value of the question-and-answer model, considering the performance of the question-and-answer model on different inference paths, which helps the question-and-answer model to balance the influence of different paths during training. Based on the first loss value and the second loss value, train the question-and-answer model to obtain the target question-and-answer model. By training the question-and-answer model with the first loss value and the second loss value in different dimensions, the question-and-answer model can learn how to generate answers more accurately and improve its performance on different inference paths, thus effectively improving the performance of the question-and-answer model in solving problems. Description of the Drawings
[0042] Figure 1 is a schematic diagram of the architecture of the training system of the question-and-answer model provided by the embodiments of the present application;
[0043] Figure 2 is a schematic diagram of the structure of an electronic device for training a question-and-answer model provided by the embodiments of the present application;
[0044] Figure 3 is a schematic diagram of the structure of an electronic device for generating answers provided by the embodiments of the present application;
[0045] Figure 4 is a schematic flowchart of the training method of the question-and-answer model provided by the embodiments of the present application Figure 1 ;
[0046] Figure 5 is a schematic flowchart of the training method of the question-and-answer model provided by the embodiments of the present applicationFigure 2 ;
[0047] Figure 6 is a flowchart showing the training method of the Q&A model provided by an embodiment of the present application Figure 3 ;
[0048] Figure 7 is a flowchart showing the training method of the Q&A model provided by an embodiment of the present application Figure 4 ;
[0049] Figure 8 is a flowchart showing the training method of the Q&A model provided by an embodiment of the present application Figure 5 ;
[0050] Figure 9 is a flowchart showing the training method of the Q&A model provided by an embodiment of the present application Figure 6 ;
[0051] Figure 10 is a flowchart showing the training method of the Q&A model provided by an embodiment of the present application Figure 7 ;
[0052] Figure 11 is a flowchart showing the training method of the Q&A model provided by an embodiment of the present application Figure 8 ;
[0053] Figure 12 is a flowchart showing the training method of the Q&A model provided by an embodiment of the present application Figure 9 ;
[0054] Figure 13 is a flowchart showing the generation process of the answer provided by an embodiment of the present application;
[0055] Figure 14 is a schematic diagram showing the principle of the training method of the Q&A model provided by an embodiment of the present application Figure 1 ;
[0056] Figure 15 is a schematic diagram showing the principle of the training method of the Q&A model provided by an embodiment of the present application Figure 2 . Detailed implementation manners
[0057] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be construed as limiting the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present application.
[0058] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it is understood that "some embodiments" may be the same subset or different subsets of all possible embodiments, and may be combined with each other without conflict.
[0059] In the following description, the terms "first / second / third" are only used to distinguish similar objects and do not represent a specific order for the objects. It is understood that "first / second / third" can be interchanged in a specific order or sequence when permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0060] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.
[0061] Before further elaborating on the embodiments of this application, the nouns and terms involved in the embodiments of this application are described. The nouns and terms involved in the embodiments of this application are subject to the following explanations.
[0062] 1) Large Language Model (LLM): A large language model refers to a deep learning model trained using a large amount of text data that can generate natural language text or understand the meaning of language text. Large language models can handle a variety of natural language tasks, such as text classification, question answering, dialogue, etc., and are an important approach to artificial intelligence. Large language models are designed to understand and generate human language. They are trained on a large amount of text data and can perform a wide range of tasks, including text summarization, translation, sentiment analysis, etc. The characteristics of large language models are their large scale, containing billions of parameters, which help them learn complex patterns in language data. These models are usually based on deep learning architectures, such as transformers, which help them achieve impressive performance in various natural language processing tasks.
[0063] 2) Natural Language Processing (NLP): It is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between humans and computers using natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, the research in this field will involve natural language, that is, the language people use in daily life. So it has a close connection with the research of linguistics, but there are also important differences. Natural language processing does not generally study natural language, but aims to develop computer systems, especially software systems, that can effectively achieve natural language communication. Thus, it is a part of computer science. Natural language processing is mainly applied in machine translation, public opinion monitoring, automatic summarization, opinion extraction, text classification, question answering, text semantic comparison, speech recognition, etc.
[0064] 3) Self-Improvement Training (SFT): It is an educational and training process aimed at helping individuals identify and improve their own deficiencies, and enhancing personal abilities and development potential. SFT emphasizes the active participation and self-reflection of individuals. Through a series of activities and exercises, it helps individuals recognize their strengths and weaknesses, and formulate specific plans to achieve personal goals.
[0065] 4) Direct Preference Optimization (DPO): It is a machine learning technique aimed at improving the performance of a model by directly optimizing the decision preferences of the model. The goal of DPO is to make the decision preferences of the model as consistent as possible with human preferences, thereby enhancing the practicality and user experience of the model in actual applications. DPO is usually trained based on preference data provided by humans. These preference data can be the direct evaluation of the model's decisions by humans, or the preference comparison between different decisions by humans. DPO optimizes the model by minimizing the difference between the model's decisions and human preferences. By learning human preferences, the model can better understand which decisions are preferred by humans and which are not. In this way, when making decisions, the model will tend to choose those options that conform to human preferences, thereby improving the quality and satisfaction of decisions.
[0066] 5) Question-Answering Model: A question-answering model is an artificial intelligence system that can understand and answer various types of questions. These questions can be statements in natural language form or broader problems to be solved, such as mathematical problems, logical reasoning problems, etc. Question-answering models typically rely on complex machine learning algorithms, especially deep learning techniques such as recurrent neural networks (RNNs), transformers, etc. For example, question-answering models need to answer various questions from a wide range of fields, such as science, history, culture, etc. For example, a question-answering model focuses on knowledge in a specific field, such as law, medicine, finance, etc. Question-answering models need to understand and answer users' questions in a conversation to provide personalized services. Question-answering models can be used as educational tools to help students answer questions and provide learning support. The development and application of question-answering models are of great significance for improving the efficiency of information retrieval, enhancing the human-computer interaction experience, and assisting professional decision-making. With the continuous progress of technology, the capabilities and application scope of question-answering models will continue to expand.
[0067] During the implementation process of the embodiments of this application, the applicant found the following problems in the related technologies:
[0068] In the related technologies, for the training of question-answering models, usually a single loss is determined to train the question-answering model, which results in a relatively low accuracy of the trained target question-answering model in solving problems.
[0069] The embodiments of this application provide a training method, device, electronic device, computer-readable storage medium, and computer program product for a question-answering model, which effectively improves the performance of the question-answering model in solving problems. The following describes an exemplary application of the training system for the question-answering model provided by the embodiments of this application.
[0070] See Figure 1 , Figure 1 is a schematic architecture diagram of the training system 100 for the question-answering model provided by the embodiments of this application. The terminal (exemplarily shows the terminal 400) is connected to the server 200 through the network 300. The network 300 can be a wide area network, a local area network, or a combination of the two.
[0071] The terminal 400 is used for the user to use the client 410 to display the answer on the graphical interface 410-1 (exemplarily shows the graphical interface 410-1). The terminal 400 and the server 200 are connected to each other through a wired or wireless network.
[0072] In some embodiments, the server 200 may be an independent physical server, or a server cluster or business system composed of multiple physical servers. It may also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. The terminal 400 may be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart TV, a smart watch, a vehicle-mounted terminal, etc., but is not limited thereto. The electronic device provided in the embodiments of the present application may be implemented as a terminal or as a server. The terminal and the server may be directly or indirectly connected through wired or wireless communication methods, which are not limited in the embodiments of the present application.
[0073] In some embodiments, based on the first question sample carrying the answer label, the server 200 determines multiple candidate inference paths, classifies the multiple candidate inference paths to obtain a first inference path and a second inference path, determines a first loss value based on the first inference path and the answer label, determines a second loss value based on the first inference path and the second inference path, and sends the first loss value and the second loss value to the terminal 400. The terminal 400 trains the question-and-answer model based on the first loss value and the second loss value to obtain a target question-and-answer model.
[0074] In some other embodiments, based on the first question sample carrying the answer label, the terminal 400 determines multiple candidate inference paths, classifies the multiple candidate inference paths to obtain a first inference path and a second inference path, determines a first loss value based on the first inference path and the answer label, determines a second loss value based on the first inference path and the second inference path, and sends the first loss value and the second loss value to the server 200. The server 200 trains the question-and-answer model based on the first loss value and the second loss value to obtain a target question-and-answer model.
[0075] In some embodiments, in the application scenario of medical diagnosis, the patient's symptom description, such as "I cough, have a fever, and body aches." The possible diagnostic paths generated by the question-and-answer model include the possibilities of various diseases. The paths with high accuracy may be based on common disease patterns and symptom associations, and the paths with low accuracy may include rare diseases or irrelevant symptoms. The question-and-answer model compares the inference path with the actual diagnosis (answer label) to calculate the first loss value; at the same time, it compares the first and second inference paths to calculate the second loss value. Optimize the accuracy of the model in medical diagnosis and improve the ability to identify common diseases.
[0076] In some embodiments, in the application scenario of financial investment, an investor asks for advice on stock investment, such as "Which technology stocks should I invest in?". The possible investment advice paths generated by the Q&A model include the analysis and prediction of different stocks. The paths with high accuracy may be based on historical data and market trend analysis, while the paths with low accuracy may contain inaccurate predictions or outdated information. The Q&A model compares the reasoning paths with the actual investment performance (answer labels) and calculates the loss value. This improves the accuracy and reliability of the Q&A model in financial investment advice.
[0077] See Figure 2 , Figure 2 FIG. is a schematic structural diagram of an electronic device for training a Q&A model provided by an embodiment of the present application. Among them, Figure 2 The electronic device 500 shown can be Figure 1 the server 200 or the terminal 400 in Figure 2 The electronic device 500 shown includes: at least one processor 430, a memory 450, and at least one network interface 420. Each component in the electronic device 500 is coupled together through a bus system 440. It can be understood that the bus system 440 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 440 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in Figure 2 all kinds of buses are labeled as the bus system 440.
[0078] The processor 430 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP, Digital Signal Processor), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0079] The memory 450 can be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard disk drives, optical disc drives, etc. The memory 450 optionally includes one or more storage devices that are physically located away from the processor 430.
[0080] The memory 450 includes volatile memory or non-volatile memory, and can also include both volatile and non-volatile memory. The non-volatile memory can be a read-only memory (ROM, Read Only Memory), and the volatile memory can be a random access memory (RAM, Random Access Memory). The memory 450 described in the embodiments of the present application is intended to include any suitable type of memory.
[0081] In some embodiments, the memory 450 is capable of storing data to support various operations. Examples of such data include programs, modules, and data structures, or subsets or supersets thereof, which are illustrated below.
[0082] The operating system 451 includes system programs for processing various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, driver layer, etc., for implementing various basic services and handling hardware-based tasks;
[0083] The network communication module 452 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 420. Exemplary network interfaces 420 include: Bluetooth, Wireless Fidelity (WiFi), and Universal Serial Bus (USB), etc.
[0084] In some embodiments, the training device of the question-and-answer model provided by the embodiments of the present application can be implemented in software. Figure 2 Shown is the training device 455 of the question-and-answer model stored in the memory 450, which can be software in the form of programs and plugins, etc., including the following software modules: determination module 4551, classification module 4552, loss module 4553, training module 4554. These modules are logical, and thus can be combined arbitrarily or further split according to the functions implemented. The functions of each module will be described below.
[0085] See Figure 3 , Figure 3 is a schematic structural diagram of an electronic device for generating answers provided by the embodiments of the present application, where Figure 3 The shown electronic device 600 can be Figure 1 the server 200 or the terminal 400 in Figure 3 The shown electronic device 600 includes: at least one processor 530, a memory 550, and at least one network interface 520. Each component in the electronic device 600 is coupled together through a bus system 540. It can be understood that the bus system 540 is used to realize the connection and communication between these components. The bus system 540 includes, in addition to the data bus, a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in Figure 3 all kinds of buses are labeled as the bus system 540.
[0086] The processor 530 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0087] The memory 550 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard disk drives, optical disk drives, etc. The memory 550 may optionally include one or more storage devices that are physically remote from the processor 530.
[0088] The memory 550 includes a volatile memory or a non-volatile memory, and may also include both volatile and non-volatile memories. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 550 described in the embodiments of the present application is intended to include any suitable type of memory.
[0089] In some embodiments, the memory 550 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplarily described below.
[0090] Operating system 551, including system programs for processing various basic system services and performing hardware-related tasks, such as framework layer, core library layer, driver layer, etc., for implementing various basic services and processing hardware-based tasks;
[0091] The network communication module 552 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 520. Exemplary network interfaces 520 include: Bluetooth, Wireless Fidelity (WiFi), and Universal Serial Bus (USB).
[0092] In some embodiments, the answer generation device provided in the embodiments of the present application can be implemented in software. Figure 3 The answer generation device 555 stored in the memory 550 is shown, which can be software in the form of a program and a plug-in, etc., including the following software modules: a generation module 5551, a solution module 5552, these modules are logical, and therefore can be arbitrarily combined or further split according to the functions implemented. The functions of each module will be explained below.
[0093] In some other embodiments, the training device for the question-and-answer model provided by the embodiments of the present application can be implemented in a hardware manner. As an example, the training device for the question-and-answer model provided by the embodiments of the present application can be a processor in the form of a hardware decoding processor, which is programmed to execute the training method for the question-and-answer model provided by the embodiments of the present application. For example, a processor in the form of a hardware decoding processor can adopt one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.
[0094] In some embodiments, the terminal or the server can implement the training method for the question-and-answer model provided by the embodiments of the present application by running a computer program or computer-executable instructions. For example, the computer program can be a native program in the operating system (such as a dedicated training program for the question-and-answer model) or a software module. For example, it can be a training module for the question-and-answer model embedded in any program (such as an instant messaging client, an album program, an electronic map client, a navigation client); for example, it can be a native application (APP), that is, a program that needs to be installed in the operating system to run. In short, the above computer program can be any form of application program, module, or plug-in.
[0095] The training method for the question-and-answer model provided by the embodiments of the present application will be described in combination with the exemplary applications and implementations of the server or terminal provided by the embodiments of the present application.
[0096] See Figure 4 , Figure 4 is a flowchart of the training method for the question-and-answer model provided by the embodiments of the present application Figure 1 will be described in combination with Figure 4 The steps 101 to 105 shown. The training method for the question-and-answer model provided by the embodiments of the present application can be implemented independently by the server or the terminal, or jointly implemented by the server and the terminal. The following will take the server implementing independently as an example for description.
[0097] In step 101, based on the first question sample carrying the answer label, multiple candidate inference paths are determined.
[0098] In some embodiments, the candidate inference path is used to indicate the inference steps that need to be executed to solve the problem corresponding to the first question sample.
[0099] In some embodiments, Candidate Inference Paths refer to, in a question - answering system, multiple possible solution paths generated through analysis and processing for a specific question sample. These paths represent different inference steps or logical processes, aiming to help the system find the answer to the question. Each candidate inference path is a series of steps or operations that need to be executed to solve the corresponding problem. The role of candidate inference paths in the question - answering system is to provide multiple possible solutions, enabling the system to flexibly handle different questions and improve the accuracy and reliability of the answers. By comparing and evaluating different candidate inference paths, the system can select the optimal path to solve the problem.
[0100] In some embodiments, Inference Steps refer to a series of ordered operations or thinking processes taken through logical analysis and knowledge application in the process of solving a problem. These steps aim to derive unknown conclusions or answers starting from known information or pre - conditions. In a question - answering system, inference steps are components of candidate inference paths, which guide how to transform a question into an answer.
[0101] As an example, suppose this application has a sample question: "What is photosynthesis?", and this application knows that the answer label is "Photosynthesis is the process by which plants use sunlight, carbon dioxide and water to synthesize organic matter and release oxygen." Based on this sample question, this application can determine multiple candidate reasoning paths, each of which represents a possible solution. The following are examples of several candidate reasoning paths: Path 1: Biological basis: Step 1: Identify the key information of the problem, namely "photosynthesis." Step 2: Review the basic knowledge of plant biology and understand the concept of photosynthesis. Step 3: Describe the chemical reactions of photosynthesis and its role in plant growth. Step 4: Summarize the definition and process of photosynthesis. Path 2: Ecological perspective, Step 1: Understand the ecological background of the problem, that is, the role of photosynthesis in the ecosystem. Step 2: Analyze how photosynthesis affects the carbon cycle and the production of oxygen. Step 3: Explore the contribution of photosynthesis to the earth's climate and biodiversity. Step 4: Summarize the definition of photosynthesis and its ecological significance. Path 3: Molecular biology perspective, Step 1: Understand photosynthesis from the perspective of molecular biology, involving the role of photosynthetic pigments and enzymes. Step 2: Explain the light and dark reaction phases of photosynthesis in detail. Step 3: Describe the generation and role of ATP and NADPH in photosynthesis. Step 4: Summarize the molecular mechanism of photosynthesis and its biological significance. Path 4: Environmental science perspective, Step 1: Consider the impact of environmental factors on photosynthesis, such as light intensity, temperature and moisture. Step 2: Analyze the efficiency and adaptability of photosynthesis under different environmental conditions. Step 3: Explore the role of photosynthesis in responding to climate change and environmental pollution. Step 4: Summarize the environmental adaptability of photosynthesis and its importance in sustainable development.
[0102] In some embodiments, the following processing may be performed before executing step 101 above: determine multiple third reasoning paths based on the second question sample through the initial question-answering model; determine a fourth reasoning path from the multiple third reasoning paths based on the answer label carried by the second question sample; generate a first answer corresponding to the fourth reasoning path through the initial question-answering model based on the fourth reasoning path; train the initial question-answering model based on the first answer corresponding to the fourth reasoning path and the answer label carried by the second question sample to obtain the question-answering model.
[0103] In some embodiments, the Question Answering Model can be a large language model, which aims to understand and answer questions raised by users. The Question Answering Model is usually based on Natural Language Processing (NLP) technology and can handle various types of questions, including factual questions, explanatory questions, and suggestive questions, etc. The Question Answering Model can extract information from various forms of data such as text, audio, and video, and generate answers that can be understood by humans. They are widely used in fields such as intelligent assistants, search engines, customer service, and educational tutoring, providing users with fast and accurate information acquisition and problem-solving services. The development of the Question Answering Model depends on a large amount of data, advanced algorithms, and powerful computing resources to continuously improve its ability to understand and answer questions.
[0104] In some embodiments, the server first needs to collect a large number of second question samples, which carry corresponding answer labels. This data may come from various sources, such as online Q&A platforms, educational databases, professional literature, etc. The preprocessing steps include cleaning the data, removing noise, standardizing the text format, etc., to ensure the quality and consistency of the data. The server uses a pre-trained initial Question Answering Model, which can generate multiple inference paths based on the question samples. These inference paths represent the inference steps that may need to be executed to solve the problem. The server uses the initial Question Answering Model to process the second question samples and generate multiple third inference paths. These paths are automatically generated by the model according to the question samples and may contain different solutions or ideas. The server determines the fourth inference path from multiple third inference paths based on the answer labels carried by the second question samples. This process may involve evaluating the accuracy, logical consistency, and relevance of each path to the answer labels. The server generates the first answer through the initial Question Answering Model based on the fourth inference path. This answer is obtained by the model according to the selected inference path. The server trains the initial Question Answering Model based on the first answer corresponding to the fourth inference path and the answer labels carried by the second question samples. The purpose of this step is to adjust the parameters of the model so that it can generate answers that are more consistent with the answer labels.
[0105] As an example, in the application scenario of the intelligent tutoring system of an online education platform, the server has collected a large number of student question samples, and these samples carry answer tags provided by teachers. For example, the question sample might be "What is the solution to a quadratic equation?", and the answer tag is "The solutions of a quadratic equation can be obtained through the quadratic formula." The server uses a pre-trained initial Q&A model that can generate multiple reasoning paths based on the student's question. For example, the model might generate the following reasoning paths: Path 1: Explain the basic concepts and forms of quadratic equations. Path 2: Introduce the derivation process of the quadratic formula. Path 3: Provide specific examples of quadratic equations and demonstrate the solution methods. The server uses the initial Q&A model to process the second question sample and generates multiple third reasoning paths. For example, for the question "What is the solution to a quadratic equation?", the model generates three reasoning paths. The server determines the fourth reasoning path from the multiple third reasoning paths based on the answer tag carried by the second question sample. For example, the server evaluates the three reasoning paths and finds that Path 2 best matches the answer tag, so it selects Path 2 as the fourth reasoning path. The server generates the first answer through the initial Q&A model based on the fourth reasoning path. For example, based on Path 2, the model generates the answer: "The solutions of a quadratic equation can be obtained through the quadratic formula, and the quadratic formula is x = [-b ± sqrt(b 2 - 4ac)] / (2a). " The server trains the initial Q&A model based on the first answer corresponding to the fourth reasoning path and the answer tag carried by the second question sample. For example, the server compares the answer generated by the model with the answer tag, calculates the loss value, and uses the backpropagation algorithm to adjust the parameters of the model to improve its ability to generate accurate answers. The server optimizes the trained model, which may include adjusting the model structure, improving the training algorithm, etc. Once the model reaches satisfactory performance, it can be deployed to the intelligent tutoring system of the online education platform for students to use.
[0106] In this way, by generating multiple third reasoning paths through the initial Q&A model, then determining the fourth reasoning path based on the answer tag and generating the corresponding first answer, and then using the first answer and the answer tag to train the initial Q&A model, this process can significantly improve the accuracy and generalization ability of the Q&A model. By generating multiple reasoning paths, the initial Q&A model increases the diversity and depth of the model's understanding of the question, which helps to capture multi-dimensional information of the question. Selecting the fourth reasoning path that best matches the answer tag ensures that the model learns the correct reasoning logic and answer generation strategy during the training process. The comparison between the generated first answer and the answer tag provides a direct feedback signal for the model. By calculating the loss value and backpropagating, the model can adjust its parameters and optimize its performance in reasoning and answer generation. This iterative training process not only enhances the model's understanding and solving ability for specific question samples but also improves the model's generalization ability and adaptability when facing new questions through continuous learning and optimization.
[0107] In some embodiments, to determine the fourth inference path from the multiple third inference paths based on the answer label carried by the second question sample, the following method may be adopted: for each of the third inference paths, use the initial Q&A model to generate a first answer corresponding to the third inference path based on the third inference path; determine the first similarity between each of the first answers and the answer label carried by the second question sample, and determine the third inference path corresponding to the first answer with the maximum first similarity as the fourth inference path.
[0108] In some embodiments, the server has collected a large number of second question samples, and these samples carry corresponding answer labels. For example, the question sample may be "What is photosynthesis?", and the answer label is "Photosynthesis is the process by which plants use sunlight, carbon dioxide, and water to synthesize organic matter and release oxygen." The server uses a pre-trained initial Q&A model, which can generate multiple inference paths based on the question sample. For example, the Q&A model may generate the following inference paths: Path 1: Define the basic concept of photosynthesis. Path 2: Explain the chemical reaction of photosynthesis. Path 3: Describe the role of photosynthesis in the ecosystem. The server uses the initial Q&A model to process the second question sample and generate multiple third inference paths. For example, for the question "What is photosynthesis?", the model generates three inference paths. The server uses the initial Q&A model to generate corresponding first answers based on each of the third inference paths. Then, the server calculates the first similarity between each first answer and the answer label carried by the second question sample. This can be achieved by using natural language processing techniques such as cosine similarity, Jaccard similarity, or a machine learning-based similarity model. The server determines the third inference path corresponding to the first answer with the maximum first similarity and determines it as the fourth inference path. For example, if the answer generated by Path 2 has the highest similarity to the answer label, then Path 2 is selected as the fourth inference path. The server trains the initial Q&A model based on the first answer corresponding to the fourth inference path and the answer label carried by the second question sample. The purpose is to adjust the parameters of the model so that it can generate answers that match the answer label more accurately. The server optimizes the trained model, which may include adjusting the model structure, improving the training algorithm, etc. Once the model reaches satisfactory performance, it can be deployed to the production environment for users to use.
[0109] As an example, the server has collected a large number of student question samples, and these samples carry answer labels provided by teachers. For example, the question sample may be "What is a quadratic equation?", and the answer label is "A quadratic equation is in the form of ax 2The equation \(ax^{2}+bx + c = 0\), where \(a\), \(b\), and \(c\) are constants and \(a\neq0\). The server uses a pre-trained initial Q&A model that can generate multiple reasoning paths based on the student's question. For example, the model may generate the following reasoning paths: Path 1: Define the basic concept of a quadratic equation. Path 2: Explain the solution method of a quadratic equation. Path 3: Provide application examples of quadratic equations. The server uses the initial Q&A model to process the second question sample and generates multiple third reasoning paths. For example, for the question "What is a quadratic equation?", the model generates three reasoning paths. The server, through the initial Q&A model, generates corresponding first answers based on each third reasoning path. Then, the server calculates the first similarity between each first answer and the answer label carried by the second question sample. The server determines the third reasoning path corresponding to the first answer with the largest first similarity and designates it as the fourth reasoning path. For example, if the answer generated by Path 1 has the highest similarity to the answer label, then Path 1 is selected as the fourth reasoning path. The server optimizes the trained model, which may include adjusting the model structure, improving the training algorithm, etc. Once the model reaches satisfactory performance, it can be deployed to the intelligent tutoring system of the online education platform for students to use.
[0110] In this way, the server can effectively determine the fourth reasoning path that best matches the answer label from multiple third reasoning paths. The server first uses the initial Q&A model to generate corresponding first answers based on each third reasoning path, and then calculates the similarity between these first answers and the answer label carried by the second question sample. By comparing these similarities, the server can identify the first answer that best matches the answer label and determine the third reasoning path corresponding to it as the fourth reasoning path. This not only improves the accuracy of the Q&A model but also enhances the model's in-depth understanding and reasoning ability for questions. The server can continuously optimize the Q&A model so that it can provide more accurate and expected answers when facing various questions.
[0111] In some embodiments, the above-mentioned training of the initial Q&A model based on the first answer corresponding to the fourth reasoning path and the answer label carried by the second question sample to obtain the Q&A model can be achieved in the following manner: Determine the third loss value of the initial Q&A model based on the difference between the first answer corresponding to the fourth reasoning path and the answer label carried by the second question sample; Determine the fourth loss value of the initial Q&A model based on each reasoning step in the fourth reasoning path; Train the initial Q&A model based on the third loss value and the fourth loss value to obtain the Q&A model.
[0112] In some embodiments, a third loss value is determined based on the difference between the first answer and the answer label. The first answer is the answer generated by the model through the fourth inference path, and the answer label is the correct answer provided in the question sample, which is usually measured by calculating the similarity or distance between the first answer and the answer label. For example, the mean squared error (MSE), cross-entropy loss, or more complex similarity metrics such as BLEU score (used to evaluate the similarity between machine-generated text and reference text) can be used. The third loss value reflects the gap between the answer generated by the model and the correct answer, and is a direct feedback on the model's performance in the answer generation task. Each step in the fourth inference path guides the model on how to transform from the question to the answer. The fourth loss value may be related to the accuracy of each step in the inference process of the model. For example, an attention mechanism can be used to detect the attention distribution of the model at each step to ensure that the model focuses on relevant information. The fourth loss value reflects the accuracy of the model in the inference process, ensuring that the model not only generates the correct answer, but also its inference process is reasonable and effective. The third loss value and the fourth loss value are combined to form a comprehensive loss function. This loss function can be the sum of the two, or different weights can be assigned according to their importance. An optimization algorithm (such as gradient descent, Adam, etc.) is used to minimize the loss function. Through the backpropagation algorithm, the gradient of the loss function with respect to the model parameters is calculated, and the parameters are adjusted to reduce the loss. The model not only learns how to generate the correct answer, but also learns how to perform reasonable reasoning, thereby improving its performance and reliability in practical applications.
[0113] As an example, the server collects a large number of second question samples, which carry corresponding answer tags. For example, the question sample might be "What is quantum mechanics?", and the answer tag is "Quantum mechanics is a branch of physics that studies the behavior of microscopic particles." The server uses a pre-trained initial question-answering model, which can generate multiple reasoning paths based on the question samples. For example, the model might generate the following reasoning paths: Path 1: Define the basic concepts of quantum mechanics. Path 2: Explain the historical development of quantum mechanics. Path 3: Describe the key experiments and discoveries in quantum mechanics. The server determines the fourth reasoning path from multiple third reasoning paths based on the answer tags carried by the second question samples. Then, through the initial question-answering model, based on the fourth reasoning path, a corresponding first answer is generated. For example, based on Path 1, the model generates the answer: "Quantum mechanics is a branch of physics that studies the behavior of microscopic particles." The server calculates the difference between the first answer and the answer tag to determine the third loss value of the initial question-answering model. In addition, the server also determines the fourth loss value of the initial question-answering model based on each reasoning step in the fourth reasoning path. These loss values reflect the errors of the model in the answer generation and reasoning processes. The server trains the initial question-answering model based on the third loss value and the fourth loss value. The server optimizes the trained model, which may include adjusting the model structure, improving the training algorithm, etc. Once the question-answering model reaches a satisfactory performance, it can be deployed to the production environment for users to use.
[0114] In some embodiments, the expression of the above fourth loss value can be:
[0115]
[0116] where L reasoning is used to indicate the fourth loss value, StepLoss(t) is the reasoning correctness loss corresponding to the reasoning step t in the reasoning path, CoherenceLoss(t) is used to indicate the loss measuring the coherence between reasoning steps, DepthLoss(t) is used to indicate the loss of the deeper reasoning process between reasoning steps, and α and β are used to indicate the corresponding weights.
[0117] In some embodiments, the expression of StepLoss(t), the single-step reasoning correctness loss, can be:
[0118] StepLoss(t) = -c - ∑Cy t,c log(p t,c )(2)
[0119] where y t,c is the one-hot encoding (category c) of the true label at the t-th step, p t,c is the probability distribution predicted by the question-answering model for the t-th step, and C: the number of possible reasoning result categories.
[0120] In some embodiments, CoherenceLoss(t), the coherence loss between steps, ensures the logical coherence of adjacent inference steps (such as the relevance between step t and step t-1), and the expression can be:
[0121] CoherenceLoss(t) = 1 - cos(h t , h t-1 ) (3)
[0122] where h t is used to indicate the hidden state vector of the t-th step, and h t-1 is used to indicate the hidden state vector of the (t - 1)-th step.
[0123] In some embodiments, the expression of DepthLoss(t), the loss in the depth inference process, can be:
[0124] DepthLoss(t) = max(0, Lexpected - Lactual) (4)
[0125] where Lexpected is the expected minimum number of inference steps (hyperparameter), and Lactual is the actual number of executed inference steps.
[0126] In this way, the server can effectively optimize the initial Q&A model and improve its accuracy and reasoning ability in the question answering task. The server first calculates the third loss value based on the difference between the first answer corresponding to the fourth inference path and the answer label carried by the second question sample. This loss value directly reflects the gap between the answer generated by the model and the correct answer. At the same time, the server also analyzes each inference step in the fourth inference path to determine the fourth loss value, which evaluates the accuracy of the model during the inference process. By combining the third loss value and the fourth loss value, the server can comprehensively measure the performance of the model and use an optimization algorithm to adjust the model parameters to minimize the combined loss function. This multi-dimensional loss function design not only ensures a high degree of consistency between the generated answer and the correct answer but also guarantees the rationality and effectiveness of the model's inference process. Therefore, through such a training process, the server can obtain a more intelligent and efficient Q&A model that can provide accurate and expected answers in various application scenarios.
[0127] In some embodiments, training the initial question-answering model based on the first answer corresponding to the fourth inference path and the answer label carried by the second question sample to obtain the question-answering model can be achieved in the following manner: determining the first intermediate loss value of the question-answering model based on the first answer and the answer label carried by the second question sample; training the initial question-answering model based on the first intermediate loss value to obtain the first question-answering model; generating the i-th answer of the second question sample through the (i - 1)-th question-answering model based on the second question sample; determining the i-th intermediate loss value of the question-answering model based on the i-th answer and the answer label carried by the second question sample, and training the (i - 1)-th question-answering model based on the i-th intermediate loss value to obtain the i-th question-answering model; traversing i until the difference between the i-th intermediate loss value and the (i - 1)-th intermediate loss value is less than the difference threshold, and determining the i-th question-answering model as the question-answering model.
[0128] In some embodiments, training the (i - 1)-th question-answering model based on the i-th intermediate loss value to obtain the i-th question-answering model can be achieved in the following manner: obtaining the learning rate corresponding to the i-th intermediate loss value, and training the (i - 1)-th question-answering model based on the product of the i-th intermediate loss value and the learning rate to obtain the i-th question-answering model.
[0129] In some embodiments, in machine learning, the learning rate is a hyperparameter used to control the step size of parameter updates in each iteration of the model. If the learning rate is too large, the model may skip the optimal solution, resulting in unstable training; if the learning rate is too small, the model may require a large number of iterations to converge, resulting in a long training time. Adaptive learning rate algorithms adjust the learning rate by detecting changes in the loss function. For example, when the loss function decreases slowly, the algorithm may increase the learning rate to accelerate the convergence speed; when the loss function approaches the minimum value, the algorithm may decrease the learning rate to avoid overfitting. Common adaptive learning rate algorithms include AdaGrad, RMSprop, and Adam, etc. The advantage of adaptive learning rate is that it can automatically adjust the learning rate, enabling the model to converge faster during training and not easily fall into local optimal solutions. Adaptive learning rate can also reduce the need for manual adjustment of the learning rate, making the training process more automated and efficient.
[0130] In some embodiments, the server starts the training process by first using an initial question-answering model. This model may be pre-trained or a randomly initialized model. The server uses the initial question-answering model to generate a first answer based on the fourth inference path. Then, the server compares the first answer with the answer label carried by the second question sample and calculates the first intermediate loss value. This loss value reflects the gap between the answer generated by the model and the correct answer. Next, the server uses this loss value to train the initial question-answering model, adjusting the model's parameters to reduce the loss. The trained model is called the first question-answering model. The server enters the iterative training phase. For each iteration step i (starting from i = 2), the server uses the (i - 1)-th question-answering model to generate the i-th answer based on the second question sample. Then, the server compares the i-th answer with the answer label and calculates the i-th intermediate loss value. Next, the server uses this loss value to train the (i - 1)-th question-answering model, adjusting the model's parameters to obtain the i-th question-answering model. After each iteration, it is checked whether the difference between the i-th intermediate loss value and the (i - 1)-th intermediate loss value is less than a preset difference threshold. This difference threshold is a very small positive number used to determine whether the model has converged, that is, whether the improvement in the model's performance is small enough to stop training. If the difference is less than the difference threshold, the server considers that the model has converged and the training process ends. At this time, the server determines the i-th question-answering model as the final question-answering model. If the difference is greater than or equal to the difference threshold, the server will continue the next iterative training. Once the model converges, the server may further optimize the model, such as adjusting the model structure, performing hyperparameter search, etc., to further improve the model's performance. The optimized model will be deployed to the production environment for users to use.
[0131] In some embodiments, the server gradually optimizes the accuracy of the Q&A model through an iterative training process. In each iteration, the model obtained from the previous training is used to generate new answers, which are compared with the answer labels to calculate the loss value. The Q&A model can continuously learn and improve, reduce prediction errors, and enhance the accuracy of the generated answers. The iterative training process helps the model learn a wider range of features and patterns, thereby improving its robustness in different questions and scenarios. Through multiple trainings and adjustments, the model can better handle various inputs, reduce the overfitting phenomenon, and improve its generalization ability. By gradually adjusting the model parameters, the server can optimize the performance of the model, including improving the inference speed, reducing resource consumption, etc. The iterative training process enables the model to improve its efficiency and practicality while maintaining high accuracy. The iterative training process enables the model to continuously optimize and improve through its own feedback. Each training is based on the previous performance and loss value, enabling the model to self-improve and enhance its performance in the question answering task. In practical applications, the data and requirements may change continuously. Through the iterative training process, the server can make the model adapt to these changes, maintain its accuracy and practicality, which helps the Q&A model maintain competitiveness and value in long-term operation.
[0132] As an example, the server starts training the Q&A model of the intelligent tutoring system. The initial model may be a pre-trained language model that can understand students' questions and generate corresponding answers. The server uses the initial model to generate the first answer based on the students' questions. Then, the server compares the generated answer with the answer label provided by the teacher and calculates the 1st intermediate loss value. This loss value reflects the gap between the answer generated by the model and the correct answer. Next, the server uses this loss value to train the initial model and adjusts the model parameters to reduce the loss. The trained model is called the 1st Q&A model. The server continues to use the 1st Q&A model to generate new answers based on the students' questions and calculates the 2nd intermediate loss value. Then, the server uses this loss value to train the 1st Q&A model to obtain the 2nd Q&A model. This process continues until the model converges. After each iteration, the server checks whether the difference in the loss values is less than the difference threshold. If it is less than, the training ends; otherwise, continue the iteration. Once the model converges, the server deploys the final Q&A model to the intelligent tutoring system for students to use. This model can accurately answer students' questions and provide personalized tutoring services.
[0133] In this way, the server can effectively optimize the initial question-and-answer model and improve its accuracy and reasoning ability in the question-answering task. Specifically, the server first calculates the first intermediate loss value based on the difference between the first answer corresponding to the fourth reasoning path and the answer carried by the second question sample. This loss value directly reflects the gap between the answer generated by the model and the correct answer. Then, the server uses this loss value to train the initial question-and-answer model, adjusts the parameters of the model to reduce the loss, and obtains the first question-and-answer model. Next, the server uses the (i - 1)th question-and-answer model to generate the ith answer based on the second question sample and calculates the ith intermediate loss value. In this way, the server can continuously optimize the model until the difference between the ith intermediate loss value and the (i - 1)th intermediate loss value is less than the difference threshold, and then determines the ith question-and-answer model as the final question-and-answer model. This iterative training process not only ensures a high degree of consistency between the generated answer and the correct answer but also guarantees the rationality and effectiveness of the reasoning process of the question-and-answer model, thereby improving the performance and reliability of the question-and-answer model in practical applications.
[0134] In some embodiments, step 101 above can be implemented in the following manner: Based on the first question sample, the multiple candidate reasoning paths are determined through the question-and-answer model.
[0135] In some embodiments, determining the multiple candidate reasoning paths based on the first question sample through the question-and-answer model can be implemented in the following manner: The following process is executed multiple times for the question sample: Based on the question sample, a candidate reasoning path is generated through the question-and-answer model.
[0136] In some embodiments, referring to Figure 5 , Figure 5 is a flowchart of the training method of the question-and-answer model provided by the embodiments of the present application Figure 2 , Figure 4 shown, step 101 can be implemented by Figure 5 steps 1011 to 1012 shown.
[0137] In step 1011, multiple initial reasoning paths are generated based on the first question sample through the question-and-answer model, and the path length of the initial reasoning path is less than the path length of the candidate reasoning path.
[0138] In some embodiments, a question-answering model, especially a deep learning-based model such as a Transformer model, has powerful language understanding and reasoning capabilities. These models can capture key information in the question and generate possible reasoning paths through their internal attention mechanisms and multi-layer neural networks. When the model receives a first question sample, it starts processing the question and extracts the semantic information therein. Based on this information, the model generates multiple initial reasoning paths. These paths are the model's initial guesses about the possible directions of answering the question. The path length of the initial reasoning paths is usually less than that of the candidate reasoning paths. This is because the initial paths are generated when the model initially processes the question, and they may not have fully unfolded all the reasoning steps. Controlling the path length helps the model to further refine and expand these paths in subsequent steps. The generated initial reasoning paths are then further processed and expanded to form candidate reasoning paths. This process may include adding more reasoning steps, considering more context information, and incorporating knowledge from external knowledge bases. Once multiple candidate reasoning paths are generated, the model evaluates the validity and rationality of each path. This may involve calculating the confidence score of each path or using other evaluation metrics to determine which paths are more likely to lead to the correct answer.
[0139] In step 1012, for each of the initial reasoning paths, through the question-answering model, based on the first question sample, the initial reasoning path is expanded to obtain the candidate reasoning path corresponding to the initial reasoning path.
[0140] In some embodiments, for each initial reasoning path, the model further expands the path. This involves adding more reasoning steps on the basis of the initial path to form a more complete reasoning process. The process of path expansion includes considering more context information, incorporating knowledge from external knowledge bases, or performing more in-depth logical reasoning. Through path expansion, the model obtains the candidate reasoning path corresponding to each initial reasoning path. These candidate paths are more detailed and complete than the initial paths and can better reflect the process of answering the question.
[0141] In some embodiments, the server starts processing the first question sample and uses the Q&A model to generate multiple initial reasoning paths. These initial paths are the model's preliminary guesses about the possible directions for answering the question, and the path lengths are relatively short. For each initial reasoning path, the server uses the Q&A model to perform path expansion. This involves adding more reasoning steps on the basis of the initial path to form a more complete reasoning process. The process of path expansion may include the following steps: Information retrieval, where the Q&A model queries an external knowledge base or database to obtain information related to the question. Logical reasoning, where the Q&A model may perform logical reasoning to derive new information or conclusions. Context expansion, where the Q&A model may consider more context information to enrich the reasoning path. The Q&A model may further refine the steps in the initial path to improve the detail and accuracy of the reasoning. Through path expansion, the server obtains candidate reasoning paths corresponding to each initial reasoning path. These candidate paths are more detailed and complete than the initial paths and can better reflect the process of answering the question.
[0142] As an example, the server receives a question from the user, such as: "Why hasn't my order been shipped yet?" The server's Q&A model processes this question and generates multiple initial reasoning paths, such as: Check the order status; Confirm the payment situation; Contact the logistics company. The initial paths may be just simple steps with relatively short path lengths, for example, each path has only 12 steps. For each initial path, the server will further perform path expansion, such as: Check the order status -> Confirm whether there is inventory -> Query the inventory management system; Confirm the payment situation -> Check whether the payment is successful -> Query the payment system; Contact the logistics company -> Query the logistics detection information -> Obtain the latest logistics status.
[0143] In this way, through the Q&A model, based on the first question sample, multiple initial reasoning paths are generated, and path expansion is performed for each initial reasoning path to obtain the corresponding candidate reasoning paths. First, the Q&A model generates multiple initial reasoning paths through the preliminary processing of the question sample. Although these paths have relatively short lengths, they provide a basis for subsequent reasoning. Subsequently, for each initial path, the model performs path expansion through further processing and analysis, which not only increases the length of the reasoning path but also enriches the content of the path, making the reasoning process more comprehensive and in-depth. This process of path expansion enables the model to better understand and answer questions, improving the accuracy and efficiency of the Q&A system. At the same time, by generating multiple candidate reasoning paths, the model can analyze the question from multiple perspectives, increasing the diversity and robustness of the reasoning, so that the Q&A system can provide more comprehensive and accurate answers when facing complex questions.
[0144] In step 102, classify the multiple candidate inference paths to obtain a first inference path belonging to the first category and a second inference path belonging to the second category.
[0145] In some embodiments, the inference accuracy of the first inference path is greater than that of the second inference path.
[0146] In some embodiments, in the inference path classification, the first category refers to those inference paths with higher inference accuracy and can answer questions more accurately. These paths are usually based on more comprehensive information, more reasonable logical reasoning, and more accurate knowledge application. In the inference path classification, the second category refers to those inference paths with lower inference accuracy and relatively poor accuracy in answering questions. These paths may be based on incomplete information, unrigorous logical reasoning, or improper knowledge application.
[0147] In some embodiments, the first inference path refers to the inference path belonging to the first category, and these paths show higher accuracy and reliability when answering questions. The first inference path is usually after the model is trained and optimized, and can more effectively capture the key information of the question and generate the correct answer. The second inference path refers to the inference path belonging to the second category, and the accuracy and reliability of these paths are relatively poor when answering questions. The second inference path may need further optimization and improvement to improve its performance in the question answering task. By classifying the multiple candidate inference paths, the inference process of the model can be better understood, and it can be identified which paths are more likely to lead to the correct answer, thereby improving the overall performance and accuracy of the question answering system.
[0148] In some embodiments, in the fields of natural language processing and artificial intelligence, inference accuracy refers to the degree of consistency between the inference path generated by the model when performing logical reasoning and question answering and the actual correct answer. It reflects the model's ability in aspects such as understanding questions, extracting relevant information, applying knowledge, and performing logical deduction. The higher the inference accuracy, the closer the inference path generated by the question answering model is to the correct answer, and the stronger the inference ability of the question answering model. Inference accuracy is usually evaluated by comparing the answer generated by the question answering model with the manually annotated correct answer, and can be quantified using indicators such as accuracy, recall rate, and F1 value. When training and optimizing the question answering model, improving inference accuracy is one of the important goals.
[0149] In some embodiments, refer to Figure 6 , Figure 6 is the flowchart of the training method of the question answering model provided by the embodiments of the present application Figure 3 , Figure 4 The step 102 shown can be implemented by Figure 6 the steps 1021A to 1022A shown.
[0150] In step 1021A, perplexity evaluation is performed on each of the candidate inference paths to obtain the perplexity of each candidate inference path, and the magnitude of the perplexity is negatively correlated with the magnitude of the inference accuracy of the candidate inference path.
[0151] In some embodiments, perplexity is a metric used in natural language processing to measure the performance of a language model. It reflects the prediction uncertainty of the model for a given text. The lower the perplexity, the more accurate the model's prediction of the text and the smaller the uncertainty; conversely, the higher the perplexity, the more uncertain the model's prediction and the lower the accuracy. In the evaluation of inference paths, the present application can use perplexity as a metric to measure the accuracy of inference paths.
[0152] In some embodiments, perplexity is a probability-based metric commonly used to evaluate the performance of a language model. For a given sentence, perplexity is the geometric mean of the reciprocals of the prediction probabilities of the model for each word. In the evaluation of inference paths, each inference step can be regarded as a "word" and the entire inference path as a "sentence." The prediction probability of the model for an inference step reflects the confidence of the model in this step. If the prediction probabilities of the model for each step are relatively high, then the perplexity of the entire inference path will be relatively low, indicating that the model is more certain about this path and the inference accuracy is higher. On the contrary, if the prediction probabilities of the model for some steps are relatively low, then the perplexity of the entire inference path will be relatively high, indicating that the model is more uncertain about this path and the inference accuracy is lower. For each candidate inference path, the present application can calculate its perplexity. The lower the perplexity of a path, the more certain the model's prediction of this path and the higher the inference accuracy; the higher the perplexity of a path, the more uncertain the model's prediction of this path and the lower the inference accuracy. Therefore, the magnitude of the perplexity is negatively correlated with the magnitude of the inference accuracy of the candidate inference path.
[0153] In some embodiments, in practical applications, perplexity can be used as a basis for screening candidate inference paths. For example, a threshold can be set to retain only the inference paths with perplexity lower than the threshold, and these paths are considered to have relatively high inference accuracy and can be used as the final answer. Alternatively, all candidate inference paths can be sorted according to perplexity, and the path with the lowest perplexity can be selected as the final answer.
[0154] In some embodiments, for each candidate inference path, the present application needs to calculate the probability of each inference step. This can be accomplished by a trained language model. For each step, the model outputs a probability value indicating the probability of this step occurring given the previous step. Based on the probabilities, the perplexity of each candidate inference path is determined through the calculation method of perplexity.
[0155] In step 1022A, the candidate inference path with a perplexity lower than the perplexity threshold is determined as the first inference path, and the candidate inference path with a perplexity greater than or equal to the perplexity threshold is determined as the second inference path.
[0156] In some embodiments, the server first needs to prepare a dataset containing a large number of annotated inference paths. This data will be used to train a language model so that it can learn the probability distribution of each step in the inference path. The server trains a language model using the prepared dataset. This model can be a neural network-based model, such as a recurrent neural network (RNN), a long short-term memory network (LSTM), or a Transformer model. The goal of training is to enable the model to learn the probability distribution of each step in the inference path. After receiving a question sample, the server uses the question-and-answer model to generate multiple candidate inference paths. These paths are the model's initial guesses about the possible directions of answering the question. For each candidate inference path, the server uses the trained language model to calculate the probability of each inference step. This involves the model's prediction of each step and the probability of occurrence given the previous step. The server calculates the perplexity of each candidate inference path based on the probability. The server sets a perplexity threshold. This threshold can be determined based on historical data, model performance, or business requirements. The server compares the perplexity of each candidate inference path with the set threshold: if the perplexity is lower than the threshold, the server determines this path as the first inference path. If the perplexity is greater than or equal to the threshold, the server determines this path as the second inference path. The server outputs the classified inference paths and usually selects the first inference path as the final answer because they have a lower perplexity and a higher inference accuracy.
[0157] As an example, the server receives a question raised by a user, such as: "Why hasn't my order been shipped yet?" The server's question-and-answer model generates multiple candidate reasoning paths based on the question, such as: - Check the order status -> Confirm if there is inventory -> Query the inventory management system - Confirm the payment situation -> Check if the payment is successful -> Query the payment system - Contact the logistics company -> Query the logistics detection information -> Obtain the latest logistics status. For each candidate reasoning path, the server uses the trained language model to calculate the probability of each step. For example, the model may consider the probability of the step "Check the order status" to be relatively high because this is a common first step in dealing with unshipped problems. The server calculates the perplexity of each candidate reasoning path using the perplexity formula. For example, the perplexity of the first path may be relatively low because the probability of each step is relatively high, indicating that the model is very certain about this path prediction. The server sets a perplexity threshold based on historical data and model performance. For example, the threshold may be set to 150. The server compares the perplexity of each candidate reasoning path with the threshold: If the perplexity is below 150, the server determines this path as the first reasoning path. If the perplexity is greater than or equal to 150, the server determines this path as the second reasoning path. The server selects the first reasoning path with the lowest perplexity as the final answer path and generates the final answer based on this path, such as: "Your order is being processed. This application will arrange the shipment as soon as possible and provide logistics detection information."
[0158] In this way, by evaluating the perplexity of each candidate reasoning path, the server can obtain a measure of the uncertainty of each path, which is negatively correlated with the reasoning accuracy, that is, the lower the perplexity, the higher the accuracy of the reasoning path. Based on this, the server sets a perplexity threshold and classifies the paths below this threshold as the first reasoning paths. These paths have high accuracy and reliability and are suitable for generating the final answer. While the paths above or equal to the threshold are classified as the second reasoning paths. These paths have lower accuracy and may require further verification or optimization. This not only improves the accuracy and efficiency of the question-and-answer system but also enhances the robustness of the system, enabling it to better handle complex and changing questions and provide more accurate and satisfactory answers to users.
[0159] In some embodiments, refer to Figure 7 , Figure 7 is a flowchart showing the training method of the question-and-answer model provided by the embodiments of the present application Figure 4 , Figure 4 The step 102 shown can be implemented by respectively executing Figure 7 the steps 1021B to 1023B shown
[0160] In step 1021B, through the Q&A model, a candidate answer corresponding to the candidate inference path is generated, and a second similarity between the candidate answer and the answer label is determined.
[0161] In some embodiments, the server uses the Q&A model to process the question sample to generate multiple candidate inference paths. These paths are the server's initial guesses about the possible directions for answering the question. For each candidate inference path, the server uses the Q&A model to generate a corresponding candidate answer. This can be done through the decoder part of the model, which generates a natural language answer based on the information in the inference path. The server needs to prepare the answer labels corresponding to the question sample. These labels are usually the manually annotated correct answers and are used to evaluate the quality of the candidate answers generated by the model. The server selects a suitable similarity calculation method to measure the similarity between the candidate answer and the answer label. Commonly used methods include cosine similarity, Jaccard similarity, BLEU score, etc. For each candidate answer, the server calculates the similarity score between it and the answer label. For example, if cosine similarity is used, the server represents the candidate answer and the answer label as vectors and then calculates the cosine value of these two vectors. The server takes the calculated similarity score as the second similarity between the candidate answer and the answer label. This similarity score reflects the closeness of the candidate answer to the correct answer. The server records the second similarity score of each candidate answer and sorts or filters the candidate answers based on these scores. The higher the score, the more similar the candidate answer is to the answer label and the better the quality.
[0162] In step 1022B, when the second similarity is greater than the similarity threshold, the candidate inference path is determined as the first inference path.
[0163] In some embodiments, if the second similarity score of the candidate answer is greater than the similarity threshold, it indicates that the similarity between the candidate answer and the answer label is high enough, and the quality of the answer generated by the model is high. Therefore, the server determines this candidate inference path as the first inference path. These paths have high accuracy and reliability and are suitable for generating the final answer.
[0164] In step 1023B, when the second similarity is less than or equal to the similarity threshold, the candidate inference path is determined as the second inference path.
[0165] In some embodiments, if the second similarity score of a candidate answer is less than or equal to the similarity threshold, it indicates that the similarity between the candidate answer and the answer label is not high enough, and the quality of the answer generated by the model is low. Therefore, the server determines this candidate reasoning path as the second reasoning path. The accuracy of these paths is relatively low and may require further verification or optimization. The similarity score reflects the proximity of the candidate answer to the correct answer. When the similarity score is higher than the threshold, it indicates that the answer generated by the model is very close to the correct answer, and it can be considered that the reasoning of the model for this question is accurate. Therefore, it is reasonable to determine this type of path as the first reasoning path.
[0166] As an example, the server receives a question raised by a student, such as: "What is Newton's First Law?" The server's Q&A model generates multiple candidate reasoning paths based on the question, such as: explaining the definition of Newton's First Law -> providing the physical meaning of the law -> giving examples of the application of the law - reviewing Newton's contributions -> introducing Newton's First Law -> analyzing the position of the law in modern science - comparing Newton's three laws -> elaborating on Newton's First Law in detail -> discussing the limitations of the law. For each candidate reasoning path, the server uses the Q&A model to generate the corresponding candidate answer. For example, the candidate answer for the first path may be: "Newton's First Law, also known as the law of inertia, states that an object will remain at rest or in uniform motion in a straight line unless acted upon by an external force." The server prepares the answer label corresponding to the question, such as: "Newton's First Law, that is, the law of inertia, states that an object will remain at rest or in uniform motion in a straight line if it is not acted upon by an external force." The server selects the cosine similarity as the similarity calculation method, represents the candidate answer and the answer label as vectors, and then calculates the cosine value between them. For example, the cosine similarity score between the candidate answer and the answer label of the first path is 0.85. The server sets the similarity threshold to 0.8. For the first path, the similarity score of 0.85 is greater than the threshold of 0.8, and the server determines this path as the first reasoning path. For other paths, if the similarity score is less than or equal to 0.8, the server determines these paths as the second reasoning paths. The server selects the candidate answer corresponding to the first reasoning path as the final answer and provides it to the student: "Newton's First Law, that is, the law of inertia, states that an object will remain at rest or in uniform motion in a straight line if it is not acted upon by an external force."
[0167] In this way, candidate answers corresponding to the candidate inference paths are generated by the Q&A model, and the second similarity between these candidate answers and the answer labels is determined, enabling the server to quantitatively evaluate the quality of the answers generated by the model. When the second similarity is greater than the similarity threshold, it indicates that the candidate answer is highly semantically consistent with the correct answer, and the inference path of the model is accurate and reliable. Therefore, such paths are determined as the first inference paths, which can ensure that the generated answers have high accuracy and credibility. On the contrary, when the second similarity is less than or equal to the similarity threshold, it shows that the similarity between the candidate answer and the correct answer is insufficient, and the inference path of the Q&A model may be biased or incomplete. Therefore, such paths are determined as the second inference paths, which require further verification or optimization. This classification method of inference paths based on the similarity threshold not only improves the accuracy and efficiency of the Q&A system but also enhances its robustness, enabling it to better handle complex and changing questions and provide more accurate and satisfactory answers to users.
[0168] In step 103, based on the first inference path and the answer label, the first loss value of the Q&A model is determined.
[0169] In some embodiments, the first loss value refers to a quantitative metric used to measure the difference between the first inference path generated by the model and the actual answer label during the training of the Q&A model. This loss value is usually calculated through a loss function, and the role of the loss function is to evaluate the gap between the model's prediction result and the true result. In the Q&A task, the first loss value reflects the performance of the model in generating the correct inference path. By minimizing the first loss value, the model can learn how to generate an inference path that is more consistent with the answer label, thereby improving the model's prediction ability and the accuracy of answering questions.
[0170] In some embodiments, refer to Figure 8 , Figure 8 is the flowchart of the training method of the Q&A model provided by the embodiments of the present application Figure 5 , when the number of the first inference paths is multiple, Figure 4 The step 103 shown in Figure 8 can be implemented by executing the steps 1031A to 1035A shown in
[0171] In step 1031A, based on the answer label, a reference inference path with the largest answer similarity to the answer label is determined from the multiple first inference paths.
[0172] In some embodiments, the answer similarity refers to the semantic similarity between the answer generated by the model and the answer label. This can be calculated by various methods, such as cosine similarity, Jaccard similarity, BLEU score, or ROUGE score, etc. For each answer generated by the first inference path, the server calculates the answer similarity between it and the answer label. The server compares the answer similarity scores of all the first inference paths and selects the path with the highest score as the reference inference path. The similarity between the answer of this path and the answer label is the largest, indicating that the model performs best when generating this path.
[0173] In some embodiments, the above-mentioned method of determining the reference inference path with the largest answer similarity with the answer label from the multiple first inference paths can be implemented as follows: for each first inference path, the answer corresponding to the first inference path is determined through the question-and-answer model based on the first inference path, and the similarity between the answer corresponding to the first inference path and the answer label is determined as the answer similarity of the first inference path; the first inference path with the largest answer similarity is determined as the reference inference path.
[0174] In some embodiments, the server first prepares multiple first inference paths, which are generated by the question-and-answer model and are considered to be inference paths with relatively high accuracy. For each first inference path, the server uses the question-and-answer model to generate the corresponding answer. This involves the decoding process of the model, and the question-and-answer model generates a natural language answer according to the information in the inference path. The server calculates the similarity between each generated answer and the answer label. This can be achieved by various methods, such as cosine similarity, Jaccard similarity, BLEU score, or ROUGE score, etc. The specific method to be selected depends on the specific requirements of the application and the data characteristics. The server compares the answer similarity scores of all the first inference paths and finds the path with the highest score. This score reflects the semantic similarity between the answer generated by the model and the answer label. The server determines the first inference path with the highest answer similarity score as the reference inference path.
[0175] In step 1032A, the reference answer corresponding to the reference inference path is generated through the question-and-answer model based on the reference inference path.
[0176] In some embodiments, the server first determines a reference inference path. This is typically done by comparing answer similarity scores among multiple inference paths and selecting the path with the highest similarity to the answer label. The server uses the decoder part of the Q&A model to generate a corresponding reference answer based on the reference inference path. The task of the decoder is to convert the information in the inference path into an answer in natural language form. Before generating the reference answer, the server needs to convert the reference inference path into an input format that the model can understand. This may involve encoding the steps or information in the path as vectors or sequences. The decoder of the Q&A model receives the encoded reference inference path as input and then generates the reference answer through a series of computational steps. This process may involve techniques such as attention mechanisms, recurrent neural networks (RNNs), or transformers. The reference answer generated by the decoder is a natural language string that should be consistent with the information in the reference inference path and as close as possible to the answer label. The server can further evaluate the quality of the reference answer by comparing it with the answer label and calculating the similarity score to ensure the accuracy and quality of the reference answer. The reference answer can be used in multiple aspects, such as model evaluation, user feedback, educational applications, etc. It provides users with a high-quality example answer to help them understand the correct solution to the problem.
[0177] In step 1033A, based on the difference between the reference answer and the answer label, determine a first target loss value of the Q&A model.
[0178] In some embodiments, the first target loss value of the question-answering model can be determined through cross-entropy loss. Cross-entropy loss is a commonly used loss function, especially in classification and sequence generation tasks. It measures the difference between the probability distribution predicted by the model and the probability distribution of the true labels. For the question-answering task, cross-entropy loss can be used to evaluate the similarity between the answers generated by the model and the actual answer labels. Before calculating the cross-entropy loss, the reference answer and the answer labels need to be represented as probability distributions. This usually involves converting the text answers into vector representations, such as using word embeddings or one-hot encoding. The reference answers generated by the question-answering model are converted into a probability distribution, representing the occurrence probability of each possible vocabulary in the answer. The answer labels are also converted into a probability distribution, usually a one-hot vector, representing the probability of each vocabulary in the correct answer. The calculated cross-entropy loss value reflects the difference between the answers generated by the model and the answer labels. The smaller the loss value, the closer the answers generated by the model are to the answer labels. During the training process, the goal of the model is to minimize the cross-entropy loss value by adjusting the model parameters. This is usually achieved through backpropagation algorithms and optimization algorithms (such as gradient descent). The first target loss value is one of the main optimization objectives for model training. By minimizing this loss value, the model can learn how to generate answers that are more consistent with the answer labels more accurately.
[0179] In step 1034A, based on each inference step in the reference inference path, determine the first reference loss value of the question-answering model.
[0180] In some embodiments, each inference step in the reference inference path needs to be represented in a form that the model can process. This may involve converting the inference steps into vectors or sequences and using word embedding techniques to capture the semantic information in the steps. The question-answering model predicts the next inference step or generates a complete inference path based on the input inference steps. This involves the encoder and decoder parts of the model. The encoder processes the input inference steps, and the decoder generates the predicted inference steps. Select an appropriate loss function to measure the difference between the inference steps predicted by the model and the actual inference steps. Commonly used loss functions include cross-entropy loss, mean squared error loss, etc. The specific selection depends on the nature of the task and the characteristics of the data. For each inference step, calculate the loss value between the step predicted by the model and the actual step. For example, if cross-entropy loss is used, the cross-entropy between the probability distribution of the predicted step and the probability distribution of the actual step can be calculated.
[0181] In step 1035A, perform a weighted sum of the first target loss value and the first reference loss value to obtain the first loss value.
[0182] In some embodiments, the server first obtains multiple first inference paths, which are generated by the question-answering model and are considered to be inference paths with relatively high accuracy. For each first inference path, the server calculates the similarity between its corresponding answer and the answer label. This can be achieved by methods such as cosine similarity, Jaccard similarity, BLEU score, or ROUGE score. The server compares the answer similarity scores of all the first inference paths and selects the path with the highest score as the reference inference path. The similarity between the answer of this path and the answer label is the largest, indicating that the model performs best when generating this path. The server uses the decoder part of the question-answering model to generate the corresponding reference answer according to the reference inference path. The server calculates the difference between the reference answer and the answer label and uses a loss function such as cross-entropy loss to measure this difference. The cross-entropy loss reflects the similarity between the answer generated by the model and the actual answer label. The smaller the loss value, the closer the answer generated by the model is to the answer label. This loss value is determined as the first target loss value of the question-answering model. The server evaluates the accuracy of the model when generating these steps based on each inference step in the reference inference path. This is achieved by calculating the loss value between the predicted inference step of the model and the actual inference step, using an appropriate loss function such as cross-entropy loss or mean squared error loss. The loss values of all the inference steps are accumulated to obtain the total loss value of the entire inference path, that is, the first reference loss value. The first loss value comprehensively reflects the performance of the model in both generating answers and inference steps and is one of the main objectives for model training and optimization.
[0183] As an example, assume that this application has an intelligent education system, with the server as the execution entity responsible for providing accurate answers and explanations to students. A student asks a question, and the server generates multiple first inference paths through a question-and-answer model. For example, the question might be "What is Newton's First Law?" The inference paths generated by the server may include different explanations and perspectives. The server calculates the similarity between the answers generated by each inference path and the answer label (the pre-determined correct answer), such as using cosine similarity. Suppose the answer generated by one of the paths has the highest similarity to the answer label, and the server determines it as the reference inference path. The server uses the question-and-answer model to generate the corresponding reference answer based on the reference inference path. For example, the reference inference path may include steps such as "Newton's First Law is the law regarding the unchanging state of motion of an object", and the server converts it into a complete natural language answer: "Newton's First Law, also known as the law of inertia, states that if an object is not acted upon by an external force, it will remain at rest or move in a straight line at a constant speed. The server calculates the difference between the reference answer and the answer label, using a loss function such as cross-entropy loss to measure this difference. For example, if the answer label is "Newton's First Law states that an object remains at rest or moves in a uniform straight line without the action of an external force", the server calculates the cross-entropy loss between the reference answer and the answer label to obtain the first target loss value. The server evaluates the accuracy of the model when generating these steps based on each inference step in the reference inference path. For example, the inference steps may include "Newton's First Law", "inertia", "external force action", etc. The server calculates the loss value between each step predicted by the model and the actual step, such as using cross-entropy loss, and then accumulates these loss values to obtain the first reference loss value. The server performs a weighted sum of the first target loss value and the first reference loss value to obtain the first loss value. For example, assume the first target loss value is 0.2, the first reference loss value is 0.3, and the weights are 0.6 and 0.4 respectively. Then the first loss value is 0.2 * 0.6 + 0.3 * 0.4 = 0.12 + 0.12 = 0.24.
[0184] In this way, the server can effectively evaluate and optimize the performance of the question-answering model. First, based on the answer label, the reference inference path with the highest answer similarity to the answer label is determined from multiple first inference paths. This step calculates the answer similarity score, such as cosine similarity, to identify the most accurate inference path generated by the model. Then, through the question-answering model, based on the reference inference path, the corresponding reference answer is generated. This step uses the decoding ability of the model to convert the inference path into a natural language answer. Next, based on the difference between the reference answer and the answer label, the first target loss value of the question-answering model is determined. Through loss functions such as cross-entropy loss, the accuracy of the answer generated by the model is quantified. At the same time, based on each inference step in the reference inference path, the first reference loss value of the question-answering model is determined. This step evaluates the accuracy of the model when generating inference steps. Finally, the first target loss value and the first reference loss value are weighted and summed to obtain the first loss value. This step comprehensively considers the performance of the model in generating answers and inference steps, providing a comprehensive loss signal for model optimization. Through this rigorous technical derivation process, the server can effectively improve the accuracy and efficiency of the question-answering model, providing higher-quality answers for users.
[0185] In some embodiments, referring to Figure 9 , Figure 9 is a flowchart showing the training method of the question-answering model provided by the embodiments of the present application Figure 6 , when the number of the first inference paths is multiple, Figure 4 the step 103 shown in Figure 9 can be implemented by executing the steps 1031B to 1034B shown in
[0186] In step 1031B, for each of the first inference paths, through the question-answering model, based on the first inference path, a second answer corresponding to the first inference path is generated.
[0187] In some embodiments, the server first obtains multiple first inference paths, which are generated by the question-answering model when processing questions. Each path represents a possible solution process of the model to the question. For each first inference path, the server needs to convert it into an input format that the model can understand. This may involve converting the inference steps into vectors or sequences and using word embedding techniques to capture the semantic information in the steps. The encoder part of the question-answering model receives the processed first inference path as input and encodes it into a fixed-length vector or sequence. This encoding process captures the key information and semantic relationships in the inference path. The encoded inference path is passed to the decoder part of the model. The task of the decoder is to generate a corresponding second answer based on the information output by the encoder. The second answer generated by the decoder is a natural language string, which should be consistent with the information in the first inference path and as close as possible to the correct answer to the question. The generated second answer may need to be post-processed to ensure its grammatical correctness and fluency. This may involve optimization of the language model, error correction, or polishing of the answer. The generated second answer can be stored for subsequent evaluation or optimization, or directly output to the user as the answer to the question.
[0188] In step 1032B, based on the difference between the second answer and the answer label, determine the first target loss value corresponding to the first inference path.
[0189] In some embodiments, based on the difference between the second answer and the answer label, to determine the first target loss value corresponding to the first inference path, the server first obtains the second answer and the corresponding question answer label. The second answer is generated by the question-answering model based on the first inference path, and the answer label is the pre-determined correct answer. The server calculates the similarity between the second answer and the answer label. This can be achieved by various methods, such as cosine similarity, Jaccard similarity, BLEU score, or ROUGE score, etc. The specific method to be selected depends on the specific requirements of the application and the characteristics of the data. Select a suitable loss function to measure the difference between the second answer and the answer label. Commonly used loss functions include cross-entropy loss, mean squared error loss, etc., and the specific selection depends on the nature of the task and the characteristics of the data. Using the selected loss function, calculate the loss value between the second answer and the answer label. For example, if cross-entropy loss is used, the cross-entropy between the probability distribution of the second answer and the probability distribution of the answer label can be calculated. The calculated loss value reflects the difference between the second answer and the answer label. The smaller the loss value, the closer the second answer is to the answer label. The server records the calculated first target loss value for subsequent evaluation and optimization of the question-answering model.
[0190] In step 1033B, based on each inference step in the first inference path, determine the first reference loss value corresponding to the first inference path.
[0191] In some embodiments, the server first obtains each inference step in the first inference path. These steps are generated by the model when processing the problem and represent the logical process of the model's answer to the problem. For each inference step, the server needs to convert it into a form that the model can process. This involves converting the inference step into a vector or sequence and using word embedding techniques to capture the semantic information in the step. Select an appropriate loss function to measure the difference between the inference steps predicted by the model and the actual inference steps. Commonly used loss functions include cross-entropy loss, mean squared error loss, etc., and the specific selection depends on the nature of the task and the characteristics of the data. Using the selected loss function, calculate the loss value between the inference steps predicted by the model and the actual inference steps. For example, if cross-entropy loss is used, the cross-entropy between the probability distribution of the predicted steps and the probability distribution of the actual steps can be calculated. Accumulate the loss values of all inference steps to obtain the total loss value of the entire inference path. This can be achieved by summing or taking the average, etc. Take the accumulated total loss value as the first reference loss value, and this value reflects the accuracy of the model when generating the first inference path. The smaller the loss value, the closer the inference steps generated by the model are to the actual steps.
[0192] In step 1034B, perform a weighted sum of the first target loss value and the first reference loss value to obtain the first loss value.
[0193] In some embodiments, the server first generates, for each first inference path, a second answer corresponding to the first inference path through the question-and-answer model based on the first inference path. After generating the second answer, the server determines the first target loss value corresponding to the first inference path based on the difference between the second answer and the answer label. Calculate the similarity or difference degree between the second answer and the answer label, and usually use the cross-entropy loss function to quantify this difference. The smaller the loss value, the closer the second answer is to the answer label. The server also determines the first reference loss value corresponding to the first inference path based on each inference step in the first inference path. To evaluate the accuracy of the model when generating inference steps, usually the cross-entropy loss function is also used to calculate the difference between the inference steps predicted by the model and the actual inference steps. The server performs a weighted sum of the first target loss value and the first reference loss value to obtain the first loss value. The selection of the weight depends on the specific task and model performance and is usually determined through experimental tuning. The first loss value comprehensively reflects the performance of the model in two aspects: generating answers and inference steps, and is one of the main objectives of model training and optimization.
[0194] As an example, there is an intelligent customer service system. The server serves as the execution entity, responsible for providing accurate answers and support to users. A user asks a question, such as "How do I reset my password?" The server uses a question-and-answer model to generate a corresponding second answer based on a predefined first reasoning path. For example, the reasoning path may include steps such as "Check account settings," "Find the password reset option," "Follow the instructions," etc. The server converts them into a complete natural language answer: "Please log in to your account, click on 'Settings', then select 'Reset', and complete the operation according to the system prompts." The server calculates the difference between the second answer and the answer label, and uses a loss function such as cross-entropy loss to measure this difference. For example, if the answer label is "Log in to the account, click on Settings, select Password Reset, and follow the system prompts," the server calculates the cross-entropy loss between the second answer and the answer label to obtain a first target loss value. The server evaluates the accuracy of the model when generating these steps based on each reasoning step in the first reasoning path. For example, the reasoning steps may include "Log in to the account," "Click on Settings," "Select Password Reset," etc. The server calculates the loss value between the predicted step of the model and the actual step for each step, such as using cross-entropy loss, and then accumulates these loss values to obtain a first reference loss value. The server performs a weighted sum of the first target loss value and the first reference loss value to obtain a first loss value. For example, assume the first target loss value is 0.2, the first reference loss value is 0.3, and the weights are 0.6 and 0.4 respectively. Then the first loss value is 0.2 * 0.6 + 0.3 * 0.4 = 0.12 + 0.12 = 0.24.
[0195] In this way, for each first reasoning path, a corresponding second answer is generated through the question-and-answer model, and the reasoning path is converted into a natural language answer using the decoding ability of the model. Then, based on the difference between the second answer and the answer label, the first target loss value is determined to quantify the accuracy of the answer generated by the model. At the same time, based on each reasoning step in the first reasoning path, the first reference loss value is determined to evaluate the accuracy of the model when generating the reasoning steps. Finally, a weighted sum of the first target loss value and the first reference loss value is performed to obtain a first loss value, comprehensively considering the performance of the model in generating answers and reasoning steps, providing a comprehensive loss signal for the optimization of the model. The server can effectively improve the accuracy and efficiency of the question-and-answer model and provide higher-quality answers to users.
[0196] In step 104, based on the first reasoning path and the second reasoning path, the second loss value of the question-and-answer model is determined.
[0197] In some embodiments, the second loss value is a loss metric used to measure the consistency and accuracy of the model when generating the first inference path and the second inference path during the training of the question-and-answer model. The second loss value is calculated by comparing the differences between the first inference path and the second inference path generated by the model, and usually loss functions such as cross-entropy loss and mean squared error loss are used to quantify this difference. The second loss value reflects the stability of the model when generating different inference paths, that is, whether the model can generate similar inference paths under different input conditions. By minimizing the second loss value, the generalization ability and robustness of the model can be improved, enabling it to generate more consistent and accurate answers when facing different questions and inputs.
[0198] In some embodiments, refer to Figure 10 , Figure 10 is a schematic flowchart of the training method of the question-and-answer model provided by the embodiments of the present application Figure 7 , Figure 10 The step 103 shown can be implemented by executing Figure 10 the steps 1041A to 1042A shown.
[0199] In step 1041A, obtain the first conditional probabilities of the question model generating each of the first inference paths respectively under the condition of inputting the first question sample, and obtain the second conditional probabilities of the question model generating each of the second inference paths respectively under the condition of inputting the first question sample.
[0200] In some embodiments, the first conditional probability refers to the probability that the model generates each first inference path under the condition of given the first question sample as the input of the question model. Specifically, it is the possibility that the model outputs each first inference path when processing the first question sample. This probability reflects the model's understanding of the question and the diversity of the answer paths.
[0201] In some embodiments, the second conditional probability refers to the probability that the model generates each second inference under the condition of given the first question sample as the input of the question model. The second inference path is another answer path generated by the model based on the first inference path. This probability reflects the flexibility and innovation ability of the model when generating different answer paths.
[0202] In some embodiments, the server inputs the first question sample into the question model, and the model starts the inference process. During the inference process, the model generates multiple possible inference paths, which represent different ways for the model to answer the question. The first inference path generated by the model is the initial answer of the model to the question, and the second inference path is another answer path generated by the model based on the first inference path. These inference paths can be steps in text form or other forms of representation. The server calculates the first conditional probability of the model generating each first inference path under the condition of inputting the first question sample. This can be achieved through the output probability distribution of the model, usually using the softmax function to convert the output of the model into a probability distribution. Similarly, the server calculates the second conditional probability of the model generating each second inference path under the condition of inputting the first question sample. To ensure the rationality and comparability of the probabilities, the server normalizes the calculated probabilities. This step ensures that the sum of the probabilities of all inference paths is 1, making the probability distribution meaningful. The server records the calculated first conditional probability and second conditional probability for subsequent analysis and use. These probability information can be used to evaluate the uncertainty and diversity of the model, as well as to guide the optimization and improvement of the model. The calculated conditional probabilities can be used for various purposes, such as selecting the most likely inference path as the final answer, or for generating diverse answer paths to improve the robustness of the model.
[0203] In step 1042A, the sum of the first conditional probabilities is determined as the third conditional probability, and the sum of the second conditional probabilities is determined as the fourth conditional probability; the third conditional probability is divided by the fourth conditional probability to obtain the second loss value.
[0204] In some embodiments, the server obtains the first conditional probabilities of generating each of the first inference paths under the condition that the problem model inputs the first problem sample, and obtains the second conditional probabilities of generating each of the second inference paths under the condition that the problem model inputs the first problem sample. The server inputs the first problem sample into the problem model, and the Q&A model starts the inference process. The model generates multiple possible first inference paths and second inference paths. These paths represent different ways for the model to answer the question. The server calculates the first conditional probabilities of the model generating each first inference path under the condition of inputting the first problem sample. This can be achieved through the output probability distribution of the model, usually using the softmax function to convert the output of the model into a probability distribution. Similarly, the server calculates the second conditional probabilities of the model generating each second inference path under the condition of inputting the first problem sample. To ensure the rationality and comparability of the probabilities, the server normalizes the calculated probabilities. This step ensures that the sum of the probabilities of all inference paths is 1, making the probability distribution meaningful. The server determines the sum of the first conditional probabilities as the third conditional probability. Similarly, the sum of the second conditional probabilities is determined as the fourth conditional probability. The server divides the third conditional probability by the fourth conditional probability to obtain the second loss value, which reflects the consistency and accuracy of the Q&A model in generating the first inference path and the second inference path.
[0205] As an example, assume there is an intelligent customer service system, with the server as the execution entity responsible for providing accurate answers and support to users. A user poses a question, such as "How do I reset my password?" The server inputs this question as the first question sample into the question model. The model generates multiple possible first inference paths, such as: Inference path 1: Check account settings -> Find the password reset option -> Follow the instructions. Inference path 2: Contact customer service -> Provide authentication -> Reset password. At the same time, the model generates multiple second inference paths, such as: Inference path 3: Forgot password -> Click on the forgot password link -> Enter email address -> Follow the instructions in the email. Inference path 4: Use the mobile app -> Enter settings -> Select password reset -> Follow the prompts. The server calculates the first conditional probability of each first inference path generated by the model under the condition of inputting the first question sample. For example, the probability of inference path 1 is 0.6, and the probability of inference path 2 is 0.4. Similarly, the server calculates the second conditional probability of each second inference path generated by the model under the condition of inputting the first question sample. For example, the probability of inference path 3 is 0.7, and the probability of inference path 4 is 0.3. The server sums up each of the first conditional probabilities and determines it as the third conditional probability. For example, the third conditional probability is 0.6 + 0.4 = 1.0. Similarly, the sum of each of the second conditional probabilities is determined as the fourth conditional probability. For example, the fourth conditional probability is 0.7 + 0.3 = 1.0. The server divides the third conditional probability by the fourth conditional probability to obtain the second loss value. For example, the second loss value is 1.0 / 1.0 = 1.0.
[0206] In this way, the server can effectively evaluate and optimize the performance of the Q&A model. Specifically, the server first obtains the first conditional probability of each first inference path generated by the question model under the condition of inputting the first question sample, and obtains the second conditional probability of each second inference path generated by the model. Then, the server sums up each of the first conditional probabilities and determines it as the third conditional probability, and sums up each of the second conditional probabilities and determines it as the fourth conditional probability. Finally, by dividing the third conditional probability by the fourth conditional probability, the second loss value is obtained. This loss value reflects the consistency and accuracy of the model when generating the first and second inference paths, providing an important feedback signal for the optimization of the model. By minimizing the second loss value, the generalization ability and robustness of the model can be improved, enabling it to generate more consistent and accurate answers when facing different questions and inputs.
[0207] In some embodiments, refer to Figure 11 , Figure 11 which is the flowchart of the training method of the Q&A model provided by the embodiments of the present application Figure 8 , Figure 11 The step 104 shown in Figure 11It is implemented by steps 1041B to 1042B shown.
[0208] In step 1041B, obtain the first conditional probability of each of the first inference paths generated by the problem model under the condition of inputting the first problem sample, and obtain the second conditional probability of each of the second inference paths generated by the problem model under the condition of inputting the first problem sample.
[0209] In some embodiments, the first conditional probability refers to the probability that the model generates each first inference path under the condition that the problem model is given the first problem sample as input. Specifically, it is the likelihood that the model outputs each first inference path when processing the first problem sample. This probability reflects the model's understanding of the problem and the diversity of the solution paths.
[0210] In some embodiments, the second conditional probability refers to the probability that the model generates each second inference under the condition that the problem model is given the first problem sample as input. The second inference path is another solution path generated by the model based on the first inference path. This probability reflects the flexibility and innovation ability of the model when generating different solution paths.
[0211] In step 1042B, for each of the first conditional probabilities, divide the first conditional probability by each of the second conditional probabilities to obtain a first result corresponding to each of the second conditional probabilities, and sum the first results to obtain a second result corresponding to the first conditional probability; sum the second results corresponding to each of the first conditional probabilities to obtain the second loss value.
[0212] In some embodiments, the server calculates the first conditional probability that the model generates each first inference path under the condition of inputting the first problem sample. This can be achieved through the output probability distribution of the model, usually using the softmax function to convert the output of the model into a probability distribution. Similarly, the server calculates the second conditional probability that the model generates each second inference path under the condition of inputting the first problem sample. For each of the first conditional probabilities, the server divides the first conditional probability by each of the second conditional probabilities to obtain a first result corresponding to each of the second conditional probabilities. This step aims to compare the relationship between the first conditional probability and the second conditional probability, reflecting the relative uncertainty of the model when generating the first inference path and the second inference path. The server sums the first results to obtain a second result corresponding to the first conditional probability. Combining the first results forms an overall metric. The server sums the second results corresponding to each of the first conditional probabilities to obtain the second loss value. This loss value comprehensively reflects the uncertainty difference of the question-and-answer model when generating the first inference path and the second inference path, providing an important feedback signal for the optimization of the question-and-answer model.
[0213] As an example, the user raises a question, such as "How do I reset my password?" The server inputs this question as the first question sample into the question model. The Q&A model generates multiple possible first inference paths, such as: Inference path 1: Check account settings -> Find the password reset option -> Follow the instructions; Inference path 2: Contact customer service -> Provide authentication -> Reset password. At the same time, the Q&A model generates multiple second inference paths, such as: Inference path 3: Forgot password -> Click the forgot password link -> Enter the email address -> Follow the instructions in the email; Inference path 4: Use the mobile app -> Enter settings -> Select password reset -> Follow the prompts. The server calculates the first conditional probability of each first inference path generated by the model under the condition of inputting the first question sample. For example, the probability of inference path 1 is 0.6, and the probability of inference path 2 is 0.4. Similarly, the server calculates the second conditional probability of each second inference path generated by the model under the condition of inputting the first question sample. For example, the probability of inference path 3 is 0.7, and the probability of inference path 4 is 0.3. For each of the first conditional probabilities, the server divides the first conditional probability by each of the second conditional probabilities to obtain the first results corresponding to each of the second conditional probabilities. For example, the first result of inference path 1 is 0.6 / 0.7 ≈ 0.857, and the first result of inference path 2 is 0.4 / 0.3 ≈ 1.333. The server sums up each of the first results to obtain the second result corresponding to the first conditional probability. For example, the second result is 0.857 + 1.333 ≈ 2.190. The server sums up the second results corresponding to each of the first conditional probabilities to obtain the second loss value. In this example, the second loss value is 2.190.
[0214] Thus, the server first obtains the first conditional probabilities of generating each first inference path under the condition that the question model inputs the first question sample, and obtains the second conditional probabilities of the model generating each second inference path. Then, for each of the first conditional probabilities, the server divides the first conditional probability by each of the second conditional probabilities to obtain a first result corresponding to each of the second conditional probabilities. Comparing the relationship between the first conditional probability and the second conditional probability reflects the relative uncertainty of the model when generating the first inference path and the second inference path. Next, the server sums up each of the first results to obtain a second result corresponding to the first conditional probability. Combining each of the first results forms an overall metric. Finally, by summing up the second results corresponding to each of the first conditional probabilities, the second loss value is obtained. This loss value comprehensively reflects the uncertainty difference of the model when generating the first inference path and the second inference path, providing an important feedback signal for the optimization of the model. By minimizing the second loss value, the generalization ability and robustness of the question-and-answer model can be improved, enabling it to generate more consistent and accurate answers when facing different questions and inputs.
[0215] In step 105, based on the first loss value and the second loss value, the question-and-answer model is trained to obtain the target question-and-answer model.
[0216] In some embodiments, the target question-and-answer model is used to answer the question to be solved to obtain the answer to the question. The target question-and-answer model is a question-and-answer model optimized through a specific training process, aiming to improve the accuracy and efficiency of the model when answering questions. The training process usually involves using a loss function to guide the learning of the model, where the first loss value and the second loss value are two important metrics for measuring the performance of the model. During the training process, the question-and-answer model will receive a large number of question samples and attempt to generate corresponding answers. The first loss value may be used to measure the accuracy of the model in generating answers, while the second loss value may be used to measure the diversity and rationality of the model in generating answers. By optimizing these two loss values, the question-and-answer model can learn how to understand and answer questions more accurately. Once the training is completed, the target question-and-answer model can be deployed to actual applications to answer questions raised by users. When a user inputs a question, the model will process the input, generate one or more possible answers, and select the most appropriate answer to return to the user according to the learning results of the model. The application scope of the target question-and-answer model is very wide, including fields such as intelligent customer service, search engines, and virtual assistants.
[0217] As an example, refer to Figure 14 , Figure 14 which is a schematic diagram of the principle of the training method of the question-and-answer model provided by the embodiments of the present application.Figure 1 , based on the first question sample carrying answer tags, determine multiple candidate inference paths, where the candidate inference paths are used to indicate the inference steps required to solve the problem corresponding to the first question sample; classify the multiple candidate inference paths to obtain a first inference path belonging to the first category and a second inference path belonging to the second category, and the inference accuracy of the first inference path is greater than that of the second inference path; based on the first inference path and the answer tags, determine a first loss value of the question-answering model, and based on the first inference path and the second inference path, determine a second loss value of the question-answering model; based on the first loss value and the second loss value, train the question-answering model to obtain a target question-answering model.
[0218] In some embodiments, referring to Figure 12 , Figure 12 is a flowchart showing the training method of the question-answering model provided by the embodiments of the present application Figure 9 , Figure 12 shown in step 105 can be implemented by executing Figure 12 the steps 1051 to 1053 shown.
[0219] In step 1051, based on the first loss value and the second loss value, train the question-answering model to obtain a reference question-answering model.
[0220] In some embodiments, it is necessary to define two loss functions, corresponding to the first loss value and the second loss value respectively. The first loss value may be used to measure the accuracy of the model in generating answers. For example, cross-entropy loss is used to evaluate the difference between the model output and the true answer. The second loss value may be used to measure the diversity and rationality of the model in generating answers. For example, KL divergence is used to measure the difference between the distribution generated by the model and the expected distribution. Use an optimization algorithm (such as gradient descent, Adam, etc.) to train the question-answering model. During the training process, the model will receive a large number of question samples and try to generate corresponding answers. After each answer is generated, the first loss value and the second loss value will be calculated, and then the parameters of the model will be adjusted according to the gradient information of these two loss values. When the performance of the question-answering model on the validation set reaches a satisfactory level, the model will be saved as a reference question-answering model. This question-answering model can be used in actual applications to answer questions raised by users.
[0221] In step 1052, through the reference question-answering model, based on the first inference path, generate a third answer for the first inference path.
[0222] In some embodiments, a first question sample is input into a reference question-answering model. This question sample may be a specific question or a piece of text related to the question. After receiving the question sample, the reference question-answering model generates multiple possible reasoning paths through its internal mechanism. These reasoning paths represent different understandings and answering methods of the model for the question. In this process, the model may utilize natural language processing techniques, such as semantic analysis and entity recognition, to understand the meaning of the question. After generating multiple reasoning paths, the reference question-answering model selects a most suitable reasoning path as the first reasoning path according to certain criteria (such as the confidence of the path, the relevance to the question, etc.). Based on the selected first reasoning path, the model generates a specific answer. This answer may be directly extracted from the text or generated through the model's generation mechanism. The generated answer should be closely related to the question and be able to accurately answer the question. The reference question-answering model outputs the generated third answer as the answer to the question. This answer can be in text form or other forms, such as voice, image, etc., depending on the application scenario and requirements.
[0223] In step 1053, the answer label carried in the first question sample is replaced with the third answer to obtain a supplementary question sample carrying the third answer; based on the supplementary question sample, the reference question-answering model is trained to obtain the target question-answering model.
[0224] In some embodiments, the server first collects a large number of question samples and provides answer labels for each sample. Then, the server uses these samples to train the question-answering model. During the training process, the server calculates a first loss value (e.g., cross-entropy loss) between the answer generated by the model and the answer label, as well as a second loss value (e.g., KL divergence) corresponding to the diversity and rationality of the answer generated by the model. By optimizing these two loss values, the server adjusts the parameters of the model, the accuracy and generalization ability of the model, and finally obtains a reference question-answering model. The server uses the reference question-answering model to process the first question sample and generates multiple inference paths. Then, the server selects the first inference path among them and generates a third answer based on this path. This process may involve natural language processing techniques such as semantic analysis, entity recognition, and context understanding to ensure that the generated answer is closely related to the question. The server replaces the original answer label carried in the first question sample with the third answer generated by the reference question-answering model, thereby obtaining a supplementary question sample carrying the third answer, aiming to increase the diversity of the training data and improve the adaptability of the model to different answer forms. The server uses the supplementary question sample to further train the reference question-answering model. During this process, the server calculates the loss value between the answer generated by the model and the new answer label again and optimizes the parameters of the model. Through this iterative training process, the server finally obtains a target question-answering model, which has higher accuracy and robustness when answering questions.
[0225] As an example, assume that this application has an online education platform, and the server is the execution entity responsible for providing personalized learning support for students. The server has collected a large number of questions raised by students and corresponding answer tags. Then, the server uses this data to train the question-and-answer model. During the training process, the server calculates the first loss value (e.g., cross-entropy loss) between the answer generated by the model and the answer tag, as well as the second loss value (e.g., KL divergence) corresponding to the sex and rationality of the answer generated by the model. By optimizing these two loss values, the server adjusts the parameters of the model, the accuracy and generalization ability of the model, and finally obtains a reference question-and-answer model. The server uses the reference question-and-answer model to process a question raised by a student, such as "What is photosynthesis?" The model generates multiple inference paths, including explaining the chemical process of photosynthesis, the biological significance of photosynthesis, etc. Then, the server selects the first inference path among them and generates a third answer based on this path, such as "Photosynthesis is the process by which plants use sunlight, carbon dioxide, and water to synthesize organic matter and release oxygen." The server replaces the answer tag carried in the original question sample with the third answer generated by the reference question-and-answer model, thereby obtaining a supplementary question sample carrying the third answer. For example, the answer tag that the original question sample may carry is "Photosynthesis is the process by which plants make food." After replacement, the answer tag becomes "Photosynthesis is the process by which plants use sunlight, carbon dioxide, and water to synthesize organic matter and release oxygen." The server uses the supplementary question sample to further train the reference question-and-answer model. During this process, the server calculates the loss value between the answer generated by the model and the new answer tag again and optimizes the parameters of the model. Through this iterative training process, the server finally obtains a target question-and-answer model, which has higher accuracy and robustness when answering students' questions.
[0226] In this way, by training the question-and-answer model based on the first loss value and the second loss value, this application can obtain a reference question-and-answer model, which has relatively high accuracy and generalization ability when answering questions. Then, through the third answer generated by the reference question-and-answer model, the answer tag in the first question sample can be replaced to obtain a supplementary question sample carrying the third answer. This not only increases the diversity of training data but also improves the model's adaptability to different answer forms. Finally, based on the supplementary question sample, the reference question-and-answer model is further trained to obtain a target question-and-answer model, which has higher accuracy and robustness when answering questions. This can not only improve the performance of the target question-and-answer model but also provide an important feedback signal for the optimization of the target question-and-answer model, and can train a more accurate and intelligent question-and-answer model to provide better services and support for users.
[0227] Thus, based on the first question samples carrying answer tags, multiple candidate reasoning paths are determined. These candidate reasoning paths indicate the reasoning steps required to solve the problem. The multiple candidate reasoning paths are classified to obtain the first reasoning path belonging to the first category and the second reasoning path belonging to the second category. The reasoning accuracy of the first reasoning path is greater than that of the second reasoning path, enabling the question-and-answer model to distinguish the quality of different candidate reasoning paths, which helps the question-and-answer model to focus on higher-quality reasoning paths during the training process. Based on the first reasoning path and the answer tags, the first loss value of the question-and-answer model is determined to ensure that the gap between the output of the question-and-answer model and the correct answer is quantified and used as the basis for optimization. At the same time, based on the first reasoning path and the second reasoning path, the second loss value of the question-and-answer model is determined, which considers the performance of the question-and-answer model on different reasoning paths and helps the question-and-answer model to balance the influence of different paths during the training process. Based on the first loss value and the second loss value, the question-and-answer model is trained to obtain the target question-and-answer model. By training the question-and-answer model with the first loss value and the second loss value in different dimensions, the question-and-answer model can learn how to generate answers more accurately and improve its performance on different reasoning paths, thus effectively improving the performance of the question-and-answer model in solving problems.
[0228] See Figure 13 , Figure 13 is a schematic flowchart of the generation of the answer provided by the embodiment of the present application, which will be described in conjunction with Figure 13 the steps 201 to 202 shown. The answer generation method provided by the embodiment of the present application can be implemented independently by the server or the terminal, or jointly implemented by the server and the terminal. Hereinafter, the case where the server implements it independently will be taken as an example for description.
[0229] In step 201, through the target question-and-answer model, multiple candidate reasoning paths are generated based on the problem to be solved, and through the target question-and-answer model, the target reasoning path for processing the problem is determined from the multiple candidate reasoning paths.
[0230] In some embodiments, the problem to be solved is input into the target question-answering model. This problem can be in the form of natural language or pre-processed structured data. After receiving the problem, the target question-answering model generates multiple candidate inference paths through its internal mechanism. These paths represent different understandings and solution methods of the model for the problem. The process of generating candidate inference paths may involve natural language processing techniques such as semantic analysis, entity recognition, relationship extraction, etc. to understand the meaning and context of the problem. After generating multiple candidate inference paths, the target question-answering model evaluates each path. The evaluation may be based on factors such as the confidence of the path, the relevance to the problem, and the coherence of the path. The target question-answering model may use a probability distribution to represent the confidence of each path, or use other metrics to measure the quality of the path. After evaluating the candidate inference paths, the target question-answering model selects the most appropriate path as the target inference path according to the evaluation results. This path should be the one most likely to correctly answer the question. Based on the selected target inference path, the model generates a specific answer. This answer may be directly extracted from the text or generated through the model's generation mechanism (such as sequence-to-sequence model, transformer model, etc.). The generated answer should be closely related to the question and be able to accurately answer the question. The target question-answering model outputs the generated answer as the answer to the question. This answer can be in text form or other forms such as voice, image, etc., depending on the application scenario and requirements.
[0231] In step 202, through the target question-answering model, based on the target inference path, the problem to be solved is answered to obtain the answer to the problem.
[0232] In some embodiments, the target question-answering model is trained using the training method of the above question-answering model.
[0233] In some embodiments, the server receives a problem to be solved, which can be a natural language problem input by the user through an interface or a problem to be answered automatically generated by the system. The server uses the target question-answering model to process the input problem and generates multiple candidate reasoning paths. These paths are automatically generated by the model based on the problem content and its trained knowledge base, and each path represents a possible way of answering. The server evaluates each generated candidate reasoning path. The evaluation process may involve calculating factors such as the confidence of each path, the relevance to the problem, and the logical coherence of the path. The model may use machine learning algorithms to automatically evaluate the quality of these paths. The server selects the most suitable candidate reasoning path as the target reasoning path according to the evaluation results. This path should be the one most likely to correctly answer the problem. The server uses the target question-answering model to answer the problem to be solved based on the selected target reasoning path. This process involves retrieving relevant information from the knowledge base, performing logical reasoning, or using natural language generation techniques to generate an answer. The server returns the generated answer to the user or the system as an answer to the problem. This answer can be in text form or other forms, such as voice, image, etc., depending on the application scenario and requirements. The server returns the generated answer to the user or the system as an answer to the problem. This answer can be in text form or other forms, such as voice, image, etc., depending on the application scenario and requirements.
[0234] As an example, in the application scenario of an intelligent customer service scenario, the server receives a user's consultation about product usage problems. Through the target question-answering model, multiple candidate reasoning paths are generated, such as consulting the product manual, searching for frequently asked questions answers, contacting technical support, etc. The server evaluates the feasibility of these paths and determines the best path, such as consulting the product manual. Then, based on the target reasoning path, the server retrieves relevant information from the product knowledge base and generates a detailed answer to help the user solve the problem.
[0235] As an example, in the application scenario of a medical diagnosis scenario, the server receives a doctor's inquiry about a patient's condition. Through the target question-answering model, multiple candidate reasoning paths are generated, such as consulting the content, performing imaging examinations, consulting expert opinions, etc. The server evaluates the urgency and effectiveness of these paths and determines the best path, such as performing imaging examinations. Then, based on the target reasoning path, the server analyzes the image data and generates a diagnostic report to support the doctor's decision-making.
[0236] As an example, in the application scenario of financial investment, the server receives investment advice requests from investors regarding stock investments. Through the target question-and-answer model, multiple candidate reasoning paths are generated, such as analyzing historical data, making market trend predictions, consulting investment advisors, etc. The server evaluates the risks and returns of these paths and determines the best path, such as analyzing historical data. Then, based on the target reasoning path, the server uses machine learning algorithms to analyze the historical performance of stocks and generates investment advice to help investors make informed decisions.
[0237] As an example, in a consulting scenario. The server receives consultations from customers regarding legal issues. Through the target question-and-answer model, multiple candidate reasoning paths are generated, such as referring to relevant laws and regulations, analyzing similar cases, consulting professional lawyers, etc. The server evaluates the accuracy and practicality of these paths and determines the best path, such as referring to relevant laws and regulations. Then, based on the target reasoning path, the server retrieves the legal regulations database and generates a detailed answer to provide legal support to the customers.
[0238] In this way, through the target question-and-answer model, multiple candidate reasoning paths are generated based on the problem to be solved, and through the target question-and-answer model, the target reasoning path for processing the problem is determined from the multiple candidate reasoning paths; through the target question-and-answer model, based on the target reasoning path, the problem to be solved is answered to obtain the answer to the problem. By generating multiple candidate reasoning paths, the model can comprehensively consider all aspects of the problem, improving the accuracy and comprehensiveness of the answer. By evaluating and selecting the best reasoning path, the model can ensure that the generated answer is not only correct but also conforms to the context and requirements of the problem. It can improve the robustness of the model, enabling it to still provide effective answers when faced with complex and changing problems.
[0239] Next, an exemplary application of the embodiments of the present application in an actual question-and-answer application scenario will be described.
[0240] This application relates to a reinforcement learning large language model training mechanism based on Self-Improving Training (SFT) and Direct Preference Optimization (DPO). Self-Improving Training (SFT): Through the Self-Improving Training (SFT) method, the system uses a small amount of high-quality demonstration data to guide the large language model to generate long-chain reasoning processes, simulating the human thinking mode, enabling the model to perform in-depth reasoning and efficient problem-solving in complex tasks. SFT helps the model gradually learn and optimize in reasoning tasks, thereby enhancing its thinking depth and precision. Direct Preference Optimization (DPO): Direct Preference Optimization (DPO) helps the model identify and generate higher-quality reasoning paths by optimizing positive and negative instance pairs. During the training process, the DPO method enables the model to effectively distinguish and exclude low-quality answers, thereby improving reasoning ability and accuracy. This process strengthens the model's reasoning ability and prompts it to generate more accurate solutions. Exploratory Reasoning and Self-Optimizing Training Mechanism: Combining the exploration and self-improving mechanisms with reinforcement learning, the system generates multiple possible reasoning paths through multiple "explorations" and self-improves by selecting the optimal path. The introduction of reinforcement learning enables the model to gradually optimize in each training iteration, enhancing its ability to solve complex tasks. The model continuously enhances its ability to handle problems in a wider range of fields by exploring unknown tasks, and accumulates more problem-solving strategies with training iterations, thereby continuously improving its performance. Through the above key technical points, this application provides an efficient and sustainable optimization training system for large language models, greatly enhancing the accuracy and reliability of the model when facing complex reasoning tasks.
[0241] By combining Self-Improvement Training (SFT) and Direct Preference Optimization (DPO), as well as an enhanced exploratory reasoning mechanism, this application proposes an improved training method to address the deficiencies of the prior art, thus solving the following problems: By adopting the methods of exploratory reasoning and self-generating high-quality reasoning paths, this application reduces the dependence on a large amount of labeled data. The model can gradually improve its own reasoning ability during training by generating multiple possible solutions, and thus no longer simply relies on manual annotation or fixed demonstration data. This self-generated data method greatly reduces the cost of data collection while improving the diversity and adaptability of data. By introducing the DPO optimization method, this application can effectively improve the model's adaptability to different fields and different types of tasks. During the training process, the DPO method enhances the model's judgment when facing different tasks by optimizing the difference between high-quality and low-quality answers, thereby improving its performance in unknown tasks. This enables the model to not only perform excellently in specific fields but also demonstrate powerful reasoning abilities in multiple fields. The exploratory reasoning mechanism proposed in this application encourages the model to actively generate multiple reasoning paths when facing complex problems and improves the reasoning effect by selecting the optimal path. This mechanism enables the model to make flexible adjustments during training to adapt to the requirements of different problems, overcomes the problem of a single reasoning path in the traditional SFT method, and enhances the model's ability to handle unknown and complex tasks. By combining the self-improvement mechanisms of SFT and DPO, this application enables the model to dynamically optimize the reasoning process in each training and self-improve based on new exploratory data. This mechanism enables the model to automatically generate new reasoning data and self-optimize when facing more complex tasks, avoiding the performance bottleneck caused by limited data in traditional technologies. Through continuous exploration and optimization, the model can continuously improve over a long period and solve increasingly complex problems. In summary, the technical solution of this application combines the self-improvement mechanisms of SFT, DPO, and reinforcement learning to solve the problems of data dependence, insufficient generalization ability, single reasoning process, and limited self-improvement in the prior art, providing a more efficient and flexible solution for the training of large language models.
[0242] In some embodiments, refer to Figure 15 , Figure 15 is a schematic diagram of the principle of the training method of the question-and-answer model provided by the embodiments of this application Figure 2, aiming to improve the capabilities of large language models (LLMs) in complex reasoning tasks, based on SFT and preference optimization DPO, construct an efficient inference reinforcement training framework. Among them, it is mainly used to evaluate the inference ability of the model to ensure its effectiveness in business scenarios. Feedback loop To ensure the stability of the optimized model in actual tasks, this application constructs a feedback loop: the optimized model continues to generate inference data, and the newly generated data undergoes product measurement and evaluation, and the model is adjusted according to the inference quality. Through this process, the performance of the model in complex reasoning tasks will continue to be iteratively improved. Automated inference optimization: Reduce manual intervention and improve inference ability through imitation learning + preference optimization. Interpretability enhancement: The model not only gives answers but also shows the complete reasoning path to improve credibility. Efficient closed-loop training: Continuously improve the inference ability through the feedback mechanism of product measurement to ensure that the model is applicable to various complex tasks. Ensure that the inference ability of the question-answering model can be continuously optimized, and finally provide more stable and accurate inference ability in business scenarios. Business data passes through the initial question-answering model to generate thought chain data (that is, the reasoning path described above), and the question-answering model is trained through the thought chain data. Business data passes through the question-answering model to generate thought chain data, obtaining the first reasoning path and the second reasoning path. The question-answering model is trained in the SFT manner through the first reasoning path, and the question-answering model is trained in the DPO manner through the first reasoning path and the second reasoning path.
[0243] The embodiments of this application adopt self-improving training, direct preference optimization, and exploratory reasoning self-optimization mechanisms. The system uses a small amount of high-quality demonstration data for initial training, and then optimizes the model's inference ability during the training process by generating multiple reasoning paths (i.e., reasoning rollback), thereby improving the inference accuracy and efficiency in complex tasks. The self-improving training (SFT) session is the basic part of improving the model's inference ability. The goal of this stage is to enable the model to gradually understand and generate long-chain reasoning processes through imitation learning, so as to be able to solve complex reasoning problems. In this stage, the model not only needs to learn the direct mapping from input to output, but also needs to learn how to organize thinking and generate reasoning processes (i.e., reasoning chains).
[0244] In some embodiments, the construction of a long-chain reasoning dataset. The long-chain reasoning dataset is an important resource for training the model. The construction of the dataset needs to include a large number of high-quality reasoning examples and ensure that each example contains a detailed reasoning process and the final solution. This dataset will help the model learn how to generate multi-step reasoning processes, thereby simulating the human thinking process.
[0245] In some embodiments, the data set sources are as follows: Manual annotation: Domain experts manually generate long-chain reasoning data. Language model generation: Existing large language models are used to generate the reasoning process. External reasoning system: For example, the generated long-chain reasoning data is extracted from systems such as Qwen and used as training data after appropriate processing and formatting. Format of the demonstration data: Long-chain reasoning data is usually divided into two parts: the reasoning process (thought) and the final solution (solution).
[0246] <|begin_of_thought|>
[0247] [Step 1: Reasoning]
[0248] [Step 2: Reasoning] ...
[0250] <|end_of_thought|>
[0251] <|begin_of_solution|>
[0252] [Detailed final solution]
[0253] <|end_of_solution|>
[0254] Each example not only contains specific reasoning steps but also shows how to derive the final solution from these steps. The reasoning part can be step-by-step, demonstrating problem analysis, hypothesis verification, exploration, and summarization, etc.
[0255] In some embodiments, the SFT fine-tuning model is used
[0256] Self-improving training (SFT) is performed using the constructed long-chain reasoning data set. In this process, the model fine-tunes the existing reasoning steps through supervised learning to enable it to generate outputs that are more in line with the real reasoning process.
[0257] Fine-tuning process:
[0258] In this step, the model is trained based on the input problem and the corresponding reasoning process. Through imitation learning, the model not only learns how to generate the correct answer but also learns how to generate the thinking process and makes the reasoning process and the final answer correlated with each other.
[0259] In some embodiments, for the loss function, during the SFT fine-tuning process, the loss function will consist of two parts: Cross-Entropy Loss: used to measure the difference between the answer generated by the model and the true answer. Reasoning Loss: used to optimize the accuracy of the reasoning process generated by the model. The Reasoning Loss helps improve the model's ability to handle complex thinking steps during the reasoning process.
[0260] As an example, the expression of the above loss function can be:
[0261]
[0262] where L SFT is used to indicate the loss function (i.e., the first loss value described above), λ is used to indicate the adjustment coefficient, controlling the impact of the reasoning loss on the final training, and Loss reasoning is used to indicate the loss of the reasoning process.
[0263] In some embodiments, the design of the reasoning loss: The reasoning loss not only focuses on the correctness of each reasoning step but also considers the coherence between steps. For example, during the reasoning process, the model needs to use the reasoning result of the previous stage to guide the thinking of the next stage, avoiding mental breaks or incorrect reasoning paths.
[0264] To strengthen the coherence of the reasoning process, the reasoning loss usually includes the following parts: Step Correctness Loss: ensuring the logical correctness of each reasoning step. Coherence Loss between Steps: ensuring that the connections between reasoning steps are clear, avoiding incorrect jumps or repetitions. Reasoning Depth Loss: encouraging the model to think deeply rather than relying on shallow reasoning.
[0265] As an example, the expression of the above reasoning loss can be:
[0266]
[0267] where L reasoning is used to indicate the reasoning loss, StepLoss(t) is the reasoning correctness loss corresponding to step t in the reasoning path, CoherenceLoss(t) is used to indicate the loss measuring the coherence between reasoning steps, and DepthLoss(t) is used to indicate the loss of a deeper reasoning process.
[0268] In some embodiments, to further enhance the inference ability of the model, the following optimization strategies can be adopted: Adaptive learning rate: Dynamically adjust the learning rate according to the performance of the loss function to avoid overfitting and accelerate convergence. Early stopping strategy: By detecting the loss value on the validation set, if the loss does not improve significantly in multiple training epochs, stop training early to avoid overtraining. During the training process, update the model parameters through gradient descent (such as the Adam optimizer) to minimize the total loss function. The update formula during training is:
[0269]
[0270] where θ t is the parameter of the question-answering model at the t-th iteration, μ is the learning rate, is the gradient of the loss function. θ t+1 is the parameter of the question-answering model at the (t + 1)-th iteration.
[0271] In some embodiments, in the exploration stage, the model generates multiple inference paths and optimizes the inference ability by selecting the best path. Inference path generation and exploration generate multiple candidate answers through strategies such as rollout and evaluate the quality of each inference path. The exploration strategies can include: Beam Search: Generate multiple candidate paths and select the optimal solution. Rollout Search: Roll back multiple times to generate different inference paths and finally select the best answer. By comparing the multiple generated inference paths, select the path that best matches the true answer. The evaluation criterion is usually the degree of match with the correct answer:
[0272] Score i = sim(y i , y true ) (8)
[0273] where Sim is the similarity calculation function, yi is the output of the candidate path, and y true is the true answer.
[0274] In some embodiments, the self-improvement phase is a core part of the model's capacity improvement. In this phase, the model utilizes the previously generated reasoning paths and solutions, and gradually enhances its reasoning ability by continuously optimizing the reasoning process and selecting higher-quality solutions. The self-improvement phase combines self-improving training (SFT) and direct preference optimization (DPO) methods. These two methods complement each other. By comparing high-quality and low-quality reasoning paths, the model continuously improves its reasoning ability. In the self-improvement phase, the model enhances its reasoning ability through iterative training and an ever-updating training dataset. The iterative training proceeds through the following steps: Generating new reasoning paths: Using the outputs of the model during the reasoning process (i.e., reasoning paths and final answers) to generate new reasoning data. These new data are candidate answers generated through the exploratory reasoning phase and are confirmed for their correctness by comparing with the true answers. Dataset update: Adding the newly generated high-quality reasoning paths and solutions to the original training dataset to form a new training set. This process is repeated to ensure that the dataset is gradually optimized as the model iteratively improves. Formula representation:
[0275] D t+1 = D t ∪{(x i , y i )|Output from exploration and comparison} (9)
[0276] where D t is the dataset after the t-th iteration, D t+1 is the dataset after the (t + 1)-th iteration, (x i , y i ) is a new reasoning example, x i is the input question, and y i is the corresponding reasoning path and answer.
[0277] In some embodiments, by strictly screening and filtering the newly generated reasoning paths, low-quality answers are eliminated. The quality of the data can be evaluated through metrics such as Perplexity. A low perplexity indicates that the reasoning path is relatively accurate and suitable for training. The expression for the above perplexity can be:
[0278]
[0279] where Perlexity(y i ) is used to indicate the reasoning path generated by the question-answering model, T is the length of the reasoning path, and P(y i,t |x i ) is used to indicate generating the t-th step result y under the condition of the given input x i i,t The probability.
[0280] In some embodiments, Direct Preference Optimization (DPO) is a key technique in the self-improvement phase. It helps the model learn how to select the best inference path by optimizing the contrast between high-quality and low-quality inference paths. This process is crucial for improving the accuracy and efficiency of inference. The optimization process of DPO: In DPO, the model is trained by contrasting positive instances (i.e., correct inference paths) and negative instances (i.e., low-quality inference paths). High-quality inference paths are regarded as positive instances, while low-quality inference paths are negative instances. By optimizing this contrast, the model can continuously improve its ability to select inference paths.
[0281] In some embodiments, for the selection of positive and negative instance pairs, the core idea of DPO is to train the model to better distinguish between the two by selecting positive and negative instance pairs (positive instances are correct inference paths, and negative instances are incorrect inference paths). Negative instances can be inference paths that do not reach the correct answer or paths with incoherent logic and shallow reasoning. Criteria for selecting positive and negative instances: Positive instances: Correctly generated inference paths or answers that match the true answer and have a rigorous reasoning process. Negative instances: Paths with errors in the reasoning process, such as jumps in the reasoning process, lack of coherence, or reasoning steps that fail to effectively lead to the correct answer. Optimization objective: Through the following formula, the question-answering model optimizes the selection of positive and negative instance pairs during training:
[0282]
[0283] where P(y positive |x i ) is used to indicate the conditional probability of the positive instance (correct inference path) given the input x i and P(y negative |x i ) is used to indicate the conditional probability of the negative instance (incorrect inference path) given the input x i .
[0284] In some embodiments, SFT and DPO can be used in combination to further enhance the model's inference ability in the following ways:
[0285] Alternating training (large architecture with SFT first and then DPO): In one training cycle, first use SFT for self-improvement training to optimize the inference process, and then use DPO for optimization to select high-quality inference paths and exclude low-quality paths. By alternately performing SFT and DPO optimizations, the model gradually strengthens its inference ability.
[0286] Combined Loss Function (Second Application of SFT during DPO Training Process): The loss functions of SFT and DPO can be combined to optimize both the inference process and the inference path selection simultaneously during the same training process. By introducing the combined loss, the model not only learns how to generate long-chain inferences but also learns how to select the best path from multiple inference paths.
[0287] The combined loss function can be expressed as:
[0288] L = L SFT + λL DPO (12)
[0289] where λ is a hyperparameter that controls the balance between SFT and DPO.
[0290] In some embodiments, during the training phase (imitation learning) (SFT), the input is a long-chain inference dataset. The process is to fine-tune using self-improving training (SFT), and the output is a model with inference capabilities. In the exploration phase, the input is a set of complex problems. The process is to generate multiple inference paths (rollback search), and the output is to select the best inference path through evaluation. In the self-improvement phase (DPO), the input is the improved inference path. The process is to optimize the model through SFT and DPO, and the output is an efficient inference model.
[0291] In this way, significant performance improvements have been achieved in multiple benchmark tests. The following are some key improvement situations: MATH-OAI: When trained with 3900 distilled examples, the model achieved an accuracy of 90.2%, showing a significant improvement compared to the base model (80.0%). Even with only 1100 examples, the model was improved to 86.0%, an increase of 7.5%. AIME: When using 3900 distilled examples, the accuracy of the model was 46.7%, a 251.1% improvement compared to the base model (13.3%). When using 1100 examples, the accuracy of the model was 33.3%, an increase of 153.8%. GPQA: For GPQA problems, when trained with 3900 examples, the accuracy of the model was 55.1%, a 27% improvement compared to the base model (43.4%). Even with 1100 examples, the model had an improvement of 10.6% and the accuracy reached 48%. In addition, this application found that adopting an iterative approach of exploration and self-improvement makes the model perform better when dealing with more complex problems. For example, the accuracy of the AIME test set increased from 33.3% to 46.7%, and the exploration and self-improvement methods played a role again.
[0292] It is understandable that in the embodiments of the present application, when it comes to data related to problem samples, etc., when the embodiments of the present application are applied to specific products or technologies, user permission or consent is required, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions.
[0293] The following continues to describe the exemplary structure of the training device 455 of the question-and-answer model provided by the embodiments of the present application as a software module. In some embodiments, as Figure 2 shown, the software module stored in the training device 455 of the question-and-answer model in the memory 450 may include: a determination module, configured to determine multiple candidate reasoning paths based on a first question sample carrying an answer label, where the candidate reasoning paths are used to indicate the reasoning steps required to solve the problem corresponding to the first question sample;
[0294] a classification module, configured to classify the multiple candidate reasoning paths to obtain a first reasoning path belonging to a first category and a second reasoning path belonging to a second category, where the reasoning accuracy of the first reasoning path is greater than the reasoning accuracy of the second reasoning path;
[0295] a loss module, configured to determine a first loss value of the question-and-answer model based on the first reasoning path and the answer label, and determine a second loss value of the question-and-answer model based on the first reasoning path and the second reasoning path;
[0296] a training module, configured to train the question-and-answer model based on the first loss value and the second loss value to obtain a target question-and-answer model.
[0297] In some embodiments, the above-mentioned training device of the question-and-answer model further includes: a pre-training module, configured to determine multiple third reasoning paths based on a second question sample through an initial question-and-answer model; determine a fourth reasoning path from the multiple third reasoning paths based on the answer label carried by the second question sample; generate a first answer corresponding to the fourth reasoning path through the initial question-and-answer model based on the fourth reasoning path; train the initial question-and-answer model based on the first answer corresponding to the fourth reasoning path and the answer label carried by the second question sample to obtain the question-and-answer model. The above-mentioned determination module is further configured to determine the multiple candidate reasoning paths based on the first question sample through the question-and-answer model.
[0298] In some embodiments, the above-mentioned pre-training module is further configured to, for each of the third inference paths, generate a first answer corresponding to the third inference path through the initial question-answering model based on the third inference path; determine a first similarity between each of the first answers and the answer label carried by the second question sample, and determine the third inference path corresponding to the first answer with the largest first similarity as the fourth inference path.
[0299] In some embodiments, the above-mentioned pre-training module is further configured to determine a third loss value of the initial question-answering model based on the difference between the first answer corresponding to the fourth inference path and the answer label carried by the second question sample; determine a fourth loss value of the initial question-answering model based on each inference step in the fourth inference path; and train the initial question-answering model based on the third loss value and the fourth loss value to obtain the question-answering model.
[0300] In some embodiments, the above-mentioned pre-training module is further configured to determine a first intermediate loss value of the question-answering model based on the first answer and the answer label carried by the second question sample, and train the initial question-answering model based on the first intermediate loss value to obtain a first question-answering model; generate an i-th answer of the second question sample through the (i - 1)-th question-answering model based on the second question sample; determine an i-th intermediate loss value of the question-answering model based on the i-th answer and the answer label carried by the second question sample, and train the (i - 1)-th question-answering model based on the i-th intermediate loss value to obtain an i-th question-answering model; traverse i until the difference between the i-th intermediate loss value and the (i - 1)-th intermediate loss value is less than a difference threshold, and determine the i-th question-answering model as the question-answering model.
[0301] In some embodiments, the above-mentioned pre-training module is further configured to generate multiple initial inference paths through the question-answering model based on the first question sample, where the path length of the initial inference path is less than the path length of the candidate inference path; for each of the initial inference paths, expand the path of the initial inference path through the question-answering model based on the first question sample to obtain a candidate inference path corresponding to the initial inference path.
[0302] In some embodiments, the above-mentioned classification module is further configured to evaluate the perplexity of each of the candidate inference paths to obtain the perplexity of each of the candidate inference paths, where the magnitude of the perplexity is negatively correlated with the magnitude of the inference accuracy of the candidate inference path; determine the candidate inference path with a perplexity lower than the perplexity threshold as the first inference path, and determine the candidate inference path with a perplexity greater than or equal to the perplexity threshold as the second inference path.
[0303] In some embodiments, the above classification module is further configured to perform the following processing for each of the candidate inference paths: generate a candidate answer corresponding to the candidate inference path through the question-and-answer model, and determine a second similarity between the candidate answer and the answer label; when the second similarity is greater than the similarity threshold, determine the candidate inference path as the first inference path; when the second similarity is less than or equal to the similarity threshold, determine the candidate inference path as the second inference path.
[0304] In some embodiments, the loss module is further configured to, based on the answer label, determine a reference inference path with the greatest answer similarity to the answer label from the multiple first inference paths; generate a reference answer corresponding to the reference inference path through the question-and-answer model based on the reference inference path; determine a first target loss value of the question-and-answer model based on the difference between the reference answer and the answer label; determine a first reference loss value of the question-and-answer model based on each inference step in the reference inference path; perform a weighted sum of the first target loss value and the first reference loss value to obtain the first loss value.
[0305] In some embodiments, the loss module is further configured to, for each of the first inference paths, generate a second answer corresponding to the first inference path through the question-and-answer model based on the first inference path; determine a first target loss value corresponding to the first inference path based on the difference between the second answer and the answer label; determine a first reference loss value corresponding to the first inference path based on each inference step in the first inference path; perform a weighted sum of the first target loss value and the first reference loss value to obtain the first loss value.
[0306] In some embodiments, the loss module is further configured to obtain a first conditional probability of the question model generating each of the first inference paths under the condition of inputting the first question sample, and obtain a second conditional probability of the question model generating each of the second inference paths under the condition of inputting the first question sample; determine the sum of the first conditional probabilities as a third conditional probability, and determine the sum of the second conditional probabilities as a fourth conditional probability; divide the third conditional probability by the fourth conditional probability to obtain the second loss value.
[0307] In some embodiments, the above loss module is further configured to obtain the first conditional probabilities of generating each of the first inference paths respectively by the problem model under the condition of inputting the first problem sample, and obtain the second conditional probabilities of generating each of the second inference paths respectively by the problem model under the condition of inputting the first problem sample; for each of the first conditional probabilities, divide the first conditional probability by each of the second conditional probabilities respectively to obtain first results respectively corresponding to each of the second conditional probabilities, and sum the first results to obtain a second result corresponding to the first conditional probability; sum the second results corresponding to each of the first conditional probabilities to obtain the second loss value.
[0308] In some embodiments, the above loss module is further configured to train the question-and-answer model based on the first loss value and the second loss value to obtain a reference question-and-answer model; through the reference question-and-answer model, generate a third answer for the first inference path based on the first inference path; replace the answer label carried in the first problem sample with the third answer to obtain a supplementary problem sample carrying the third answer; train the reference question-and-answer model based on the supplementary problem sample to obtain the target question-and-answer model.
[0309] Next, the implementation of the answer generation device 555 provided in the embodiments of the present application as an exemplary structure of software modules will be continued. In some embodiments, as Figure 3 shown, the software modules stored in the answer generation device 555 in the memory 550 may include: a generation module, which generates multiple candidate inference paths based on the problem to be solved through the target question-and-answer model, and determines a target inference path for processing the problem from the multiple candidate inference paths through the target question-and-answer model; a solution module, which is configured to answer the problem to be solved based on the target inference path through the target question-and-answer model to obtain the answer to the problem; wherein, the target question-and-answer model is trained by using the training method of the above question-and-answer model.
[0310] The embodiments of the present application provide a computer program product, which includes a computer program or computer-executable instructions, and the computer program or computer-executable instructions are stored in a computer-readable storage medium. The processor of the electronic device reads the computer-executable instructions from the computer-readable storage medium, and the processor executes the computer-executable instructions, so that the electronic device executes the above answer generation method and the training method of the question-and-answer model in the embodiments of the present application.
[0311] An embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions, where the computer-executable instructions, when executed by a processor, cause the processor to execute the answer generation method and the question-and-answer model training method provided by the embodiment of the present application. For example, as Figure 4 shown in the question-and-answer model training method.
[0312] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disc, or CD-ROM; or may be various electronic devices including one or any combination of the above memories.
[0313] In some embodiments, the computer-executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including being deployed as an independent program or being deployed as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0314] As an example, the computer-executable instructions may or may not correspond to a file in the file system, may be stored as part of a file storing other programs or data. For example, stored in one or more scripts in a HyperText Markup Language (HTML) document, stored in a single file dedicated to the program under discussion, or stored in multiple cooperating files (for example, files storing one or more modules, subroutines, or code portions).
[0315] As an example, the computer-executable instructions may be deployed to execute on one electronic device, or on multiple electronic devices located at one location, or on multiple electronic devices distributed at multiple locations and interconnected by a communication network.
[0316] In summary, the embodiment of the present application has the following beneficial effects:
[0317] (1) By determining multiple candidate reasoning paths based on the first question samples with answer tags, where these candidate reasoning paths indicate the reasoning steps required to solve the problem, classifying the multiple candidate reasoning paths to obtain the first reasoning path belonging to the first category and the second reasoning path belonging to the second category, with the reasoning accuracy of the first reasoning path being greater than that of the second reasoning path, enabling the question-and-answer model to distinguish the quality of different candidate reasoning paths, which helps the question-and-answer model to focus on higher-quality reasoning paths during training. Based on the first reasoning path and the answer tag, determine the first loss value of the question-and-answer model, ensuring that the gap between the output of the question-and-answer model and the correct answer is quantified and used as the basis for optimization. At the same time, based on the first reasoning path and the second reasoning path, determine the second loss value of the question-and-answer model, which takes into account the performance of the question-and-answer model on different reasoning paths, helping the question-and-answer model to balance the influence of different paths during training. Based on the first loss value and the second loss value, train the question-and-answer model to obtain the target question-and-answer model. Training the question-and-answer model through the first loss value and the second loss value in different dimensions enables the question-and-answer model to learn how to generate answers more accurately and improve its performance on different reasoning paths, thus effectively improving the problem-solving performance of the question-and-answer model.
[0318] (2) Generate multiple third reasoning paths through the initial question-and-answer model, then determine the fourth reasoning path based on the answer tag and generate the corresponding first answer, and then use the first answer and the answer tag to train the initial question-and-answer model. This process can significantly improve the accuracy and generalization ability of the question-and-answer model. By generating multiple reasoning paths, the initial question-and-answer model increases the diversity and depth of the model's understanding of the question, helping to capture multi-dimensional information of the question. Selecting the fourth reasoning path that best matches the answer tag ensures that the model learns the correct reasoning logic and answer generation strategy during training. The comparison between the generated first answer and the answer tag provides a direct feedback signal for the model. By calculating the loss value and backpropagating, the model can adjust its parameters and optimize its performance in reasoning and answer generation. This iterative training process not only enhances the model's understanding and problem-solving ability for specific question samples but also improves the model's generalization ability and adaptability when facing new questions through continuous learning and optimization.
[0319] (3) The server can effectively determine the fourth inference path that best matches the answer label from multiple third inference paths. First, the server uses the initial Q&A model to generate corresponding first answers based on each third inference path, and then calculates the similarity between these first answers and the answer label carried by the second question sample. By comparing these similarities, the server can identify the first answer that best matches the answer label and determine the corresponding third inference path as the fourth inference path. This not only improves the accuracy of the Q&A model but also enhances the model's in-depth understanding and reasoning ability of questions. The server can continuously optimize the Q&A model to enable it to provide more accurate and expected answers when facing various questions.
[0320] (4) The server can effectively optimize the initial Q&A model and improve its accuracy and reasoning ability in question-solving tasks. First, the server calculates the third loss value based on the difference between the first answer corresponding to the fourth inference path and the answer label carried by the second question sample. This loss value directly reflects the gap between the answer generated by the model and the correct answer. At the same time, the server also analyzes each inference step in the fourth inference path to determine the fourth loss value, which evaluates the accuracy of the model during the inference process. By combining the third loss value and the fourth loss value, the server can comprehensively measure the performance of the model and use an optimization algorithm to adjust the model parameters to minimize the combined loss function. This multi-dimensional loss function design not only ensures a high degree of consistency between the generated answer and the correct answer but also guarantees the rationality and effectiveness of the model's inference process. Therefore, through such a training process, the server can obtain a more intelligent and efficient Q&A model that can provide accurate and expected answers in various application scenarios.
[0321] (5) The server can effectively optimize the initial Q&A model and improve its accuracy and reasoning ability in question-solving tasks. Specifically, the server first calculates the first intermediate loss value based on the difference between the first answer corresponding to the fourth inference path and the answer carried by the second question sample. This loss value directly reflects the gap between the answer generated by the model and the correct answer. Then, the server uses this loss value to train the initial Q&A model, adjust the model's parameters to reduce the loss, and obtain the first Q&A model. Next, the server uses the (i - 1)th Q&A model to generate the ith answer based on the second question sample and calculates the ith intermediate loss value. In this way, the server can continuously optimize the model until the difference between the ith intermediate loss value and the (i - 1)th intermediate loss value is less than the difference threshold, and then determines the ith Q&A model as the final Q&A model. This iterative training process not only ensures a high degree of consistency between the generated answer and the correct answer but also guarantees the rationality and effectiveness of the Q&A model's inference process, thereby improving the performance and reliability of the Q&A model in practical applications.
[0322] (6) Through the Q&A model, based on the first question sample, multiple initial inference paths are generated, and path expansion is performed for each initial inference path to obtain corresponding candidate inference paths. The Q&A model first generates multiple initial inference paths through preliminary processing of the question sample. Although the lengths of these paths are relatively short, they provide a basis for subsequent inferences. Subsequently, for each initial path, the model performs path expansion through further processing and analysis, which not only increases the length of the inference path but also enriches the content of the path, making the inference process more comprehensive and in-depth. This process of path expansion enables the model to better understand and answer questions, improving the accuracy and efficiency of the Q&A system. At the same time, by generating multiple candidate inference paths, the model can analyze questions from multiple perspectives, increasing the diversity and robustness of the inferences, so that the Q&A system can provide more comprehensive and accurate answers when facing complex questions.
[0323] (7) By evaluating the perplexity of each candidate inference path, the server can obtain the uncertainty measure of each path. This measure is negatively correlated with the inference accuracy, that is, the lower the perplexity, the higher the accuracy of the inference path. Based on this, the server sets a perplexity threshold and classifies the paths below this threshold as the first inference paths. These paths have high accuracy and reliability and are suitable for generating the final answer. The paths above or equal to the threshold are classified as the second inference paths. The accuracy of these paths is relatively low and may require further verification or optimization. This not only improves the accuracy and efficiency of the Q&A system but also enhances the robustness of the system, enabling it to better handle complex and changing questions and provide more accurate and satisfactory answers for users.
[0324] (8) The Q&A model generates candidate answers corresponding to the candidate inference paths and determines the second similarity between these candidate answers and the answer tags, enabling the server to quantitatively evaluate the quality of the answers generated by the model. When the second similarity is greater than the similarity threshold, it indicates that the candidate answer is highly consistent with the correct answer semantically, and the inference path of the model is accurate and reliable. Therefore, such paths are determined as the first inference paths, which can ensure that the generated answers have high accuracy and credibility. On the contrary, when the second similarity is less than or equal to the similarity threshold, it indicates that the similarity between the candidate answer and the correct answer is insufficient, and the inference path of the Q&A model may be deviated or incomplete. Therefore, such paths are determined as the second inference paths and need further verification or optimization. This classification method of inference paths based on the similarity threshold not only improves the accuracy and efficiency of the Q&A system but also enhances the robustness, enabling it to better handle complex and changing questions and provide more accurate and satisfactory answers for users.
[0325] (9) The server can effectively evaluate and optimize the performance of the question-and-answer model. First, based on the answer label, the reference inference path with the highest answer similarity to the answer label is determined from multiple first inference paths. This step identifies the most accurate inference path generated by the model by calculating the answer similarity score, such as cosine similarity. Then, through the question-and-answer model, based on the reference inference path, the corresponding reference answer is generated. This step uses the decoding ability of the model to convert the inference path into a natural language answer. Next, based on the difference between the reference answer and the answer label, the first target loss value of the question-and-answer model is determined. The accuracy of the answer generated by the model is quantified through loss functions such as cross-entropy loss. At the same time, based on each inference step in the reference inference path, the first reference loss value of the question-and-answer model is determined. This step evaluates the accuracy of the model when generating inference steps. Finally, the first target loss value and the first reference loss value are weighted and summed to obtain the first loss value. This step comprehensively considers the performance of the model in generating answers and inference steps, providing a comprehensive loss signal for model optimization. Through this rigorous technical derivation process, the server can effectively improve the accuracy and efficiency of the question-and-answer model, providing higher-quality answers for users.
[0326] (10) For each first inference path, the corresponding second answer is generated through the question-and-answer model, using the decoding ability of the model to convert the inference path into a natural language answer. Then, based on the difference between the second answer and the answer label, the first target loss value is determined to quantify the accuracy of the answer generated by the model. At the same time, based on each inference step in the first inference path, the first reference loss value is determined to evaluate the accuracy of the model when generating inference steps. Finally, the first target loss value and the first reference loss value are weighted and summed to obtain the first loss value, comprehensively considering the performance of the model in generating answers and inference steps, providing a comprehensive loss signal for model optimization. The server can effectively improve the accuracy and efficiency of the question-and-answer model, providing higher-quality answers for users.
[0327] (11) The server can effectively evaluate and optimize the performance of the question-and-answer model. Specifically, the server first obtains the first conditional probability of each first inference path generated by the question model under the condition of inputting the first question sample, and obtains the second conditional probability of each second inference path generated by the model. Then, the server determines the third conditional probability as the sum of the first conditional probabilities and the fourth conditional probability as the sum of the second conditional probabilities. Finally, by dividing the third conditional probability by the fourth conditional probability, the second loss value is obtained. This loss value reflects the consistency and accuracy of the model when generating the first inference path and the second inference path, providing an important feedback signal for model optimization. By minimizing the second loss value, the generalization ability and robustness of the model can be improved, enabling it to generate more consistent and accurate answers when facing different questions and inputs.
[0328] (12) First, the server obtains the first conditional probabilities of generating each first inference path under the condition that the question model inputs the first question sample, and obtains the second conditional probabilities of generating each second inference path by the model. Then, for each of the first conditional probabilities, the server divides the first conditional probability by each of the second conditional probabilities to obtain a first result corresponding to each of the second conditional probabilities. Comparing the relationship between the first conditional probability and the second conditional probability reflects the relative uncertainty of the model when generating the first inference path and the second inference path. Next, the server sums up each of the first results to obtain a second result corresponding to the first conditional probability. Combining all the first results forms an overall metric. Finally, by summing up the second results corresponding to each of the first conditional probabilities, the second loss value is obtained. This loss value comprehensively reflects the uncertainty difference of the model when generating the first inference path and the second inference path, providing an important feedback signal for the optimization of the model. By minimizing the second loss value, the generalization ability and robustness of the question-and-answer model can be improved, enabling it to generate more consistent and accurate answers when facing different questions and inputs.
[0329] (13) By training the question-and-answer model based on the first loss value and the second loss value, this application can obtain a reference question-and-answer model, which has high accuracy and generalization ability when answering questions. Then, the third answer generated by the reference question-and-answer model can replace the answer label in the first question sample to obtain a supplementary question sample carrying the third answer. This not only increases the diversity of the training data but also improves the adaptability of the model to different answer forms. Finally, further training the reference question-and-answer model based on the supplementary question sample can obtain a target question-and-answer model, which has higher accuracy and robustness when answering questions. This can not only improve the performance of the target question-and-answer model but also provide an important feedback signal for the optimization of the target question-and-answer model, enabling the training of a more accurate and intelligent question-and-answer model to provide better services and support for users.
[0330] (14) By determining multiple candidate inference paths based on the first question samples carrying answer tags, where these candidate inference paths indicate the inference steps required to solve the problem, classifying the multiple candidate inference paths to obtain a first inference path belonging to the first category and a second inference path belonging to the second category, with the inference accuracy of the first inference path being greater than that of the second inference path, enabling the question-and-answer model to distinguish the quality of different candidate inference paths, which helps the question-and-answer model to focus on higher-quality inference paths during training. Based on the first inference path and the answer tags, determine the first loss value of the question-and-answer model to ensure that the gap between the output of the question-and-answer model and the correct answer is quantified and used as the basis for optimization. At the same time, based on the first inference path and the second inference path, determine the second loss value of the question-and-answer model, which takes into account the performance of the question-and-answer model on different inference paths, helps the question-and-answer model to balance the influence of different paths during training. Based on the first loss value and the second loss value, train the question-and-answer model to obtain the target question-and-answer model. Training the question-and-answer model with the first loss value and the second loss value in different dimensions enables the question-and-answer model to learn how to generate answers more accurately and improve its performance on different inference paths, thus effectively improving the performance of the question-and-answer model in solving problems.
[0331] (15) Through the target question-and-answer model, based on the problem to be solved, generate multiple candidate inference paths, and through the target question-and-answer model, determine the target inference path for processing the problem from the multiple candidate inference paths; through the target question-and-answer model, based on the target inference path, answer the problem to be solved to obtain the answer to the problem. By generating multiple candidate inference paths, the model can comprehensively consider all aspects of the problem, improving the accuracy and comprehensiveness of the answer. By evaluating and selecting the best inference path, the model can ensure that the generated answer is not only correct but also conforms to the context and requirements of the problem. It can improve the robustness of the model, enabling it to still provide effective answers when facing complex and variable problems.
[0332] (16) Significant performance improvements have been achieved in multiple benchmark tests. The following are some key improvements: MATH-OAI: When trained with 3900 distilled examples, the model achieved an accuracy of 90.2%, which is a significant improvement compared to the base model (80.0%). Even with only 1100 examples, the model's accuracy increased to 86.0%, a 7.5% improvement. AIME: With 3900 distilled examples, the model's accuracy was 46.7%, a 251.1% increase compared to the base model (13.3%). With 1100 examples, the model's accuracy was 33.3%, a 153.8% increase. GPQA: For GPQA questions, when trained with 3900 examples, the model's accuracy was 55.1%, a 27% increase compared to the base model (43.4%). Even with 1100 examples, the model had a 10.6% increase and the accuracy reached 48%. In addition, this application found that adopting an iterative approach of exploration and self-improvement enables the model to perform better when dealing with more complex problems. For example, the accuracy of the AIME test set increased from 33.3% to 46.7%, and the exploration and self-improvement method played a role again.
[0333] As described above, the above are only embodiments of this application and are not intended to limit the protection scope of this application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of this application are included in the protection scope of this application.
Claims
1. A training method for a question-answering model, characterized in that, The method includes: Based on a first question sample with an answer label, determining multiple candidate inference paths, where the candidate inference paths are used to indicate the inference steps required to solve the problem corresponding to the first question sample; Classifying the multiple candidate inference paths to obtain a first inference path belonging to a first category and a second inference path belonging to a second category, where the inference accuracy of the first inference path is greater than that of the second inference path; Based on the first inference path and the answer label, determining a first loss value of the question-answering model, and based on the first inference path and the second inference path, determining a second loss value of the question-answering model; Based on the first loss value and the second loss value, training the question-answering model to obtain a target question-answering model.
2. The method according to claim 1, wherein Before determining the multiple candidate inference paths based on the first question sample with an answer label, the method further includes: Through an initial question-answering model, based on a second question sample, determining multiple third inference paths; Based on the answer label carried by the second question sample, determining a fourth inference path from the multiple third inference paths; Through the initial question-answering model, based on the fourth inference path, generating a first answer corresponding to the fourth inference path; Based on the first answer corresponding to the fourth inference path and the answer label carried by the second question sample, training the initial question-answering model to obtain the question-answering model; Determining the multiple candidate inference paths based on the first question sample with an answer label includes: Through the question-answering model, based on the first question sample, determining the multiple candidate inference paths.
3. The method according to claim 2, wherein Determining the fourth inference path from the multiple third inference paths based on the answer label carried by the second question sample includes: For each of the third inference paths, through the initial question-answering model, based on the third inference path, generating a first answer corresponding to the third inference path; Determining a first similarity between each of the first answers and the answer label carried by the second question sample, and determining the third inference path corresponding to the first answer with the maximum first similarity as the fourth inference path.
4. The method according to claim 2, wherein Training the initial question-answering model based on the first answer corresponding to the fourth inference path and the answer label carried by the second question sample to obtain the question-answering model includes: Based on the difference between the first answer corresponding to the fourth inference path and the answer label carried by the second question sample, determining a third loss value of the initial question-answering model; Based on each inference step in the fourth inference path, determining a fourth loss value of the initial question-answering model; Based on the third loss value and the fourth loss value, training the initial question-answering model to obtain the question-answering model.
5. The method according to claim 2, wherein Training the initial question-answering model based on the first answer corresponding to the fourth inference path and the answer label carried by the second question sample to obtain the question-answering model includes: Based on the first answer and the answer label carried by the second question sample, determine the first intermediate loss value of the Q&A model, and based on the first intermediate loss value, train the initial Q&A model to obtain the first Q&A model; Through the (i - 1)-th Q&A model, generate the i-th answer of the second question sample based on the second question sample; Based on the i-th answer and the answer label carried by the second question sample, determine the i-th intermediate loss value of the Q&A model, and based on the i-th intermediate loss value, train the (i - 1)-th Q&A model to obtain the i-th Q&A model; Traverse i until the difference between the i-th intermediate loss value and the (i - 1)-th intermediate loss value is less than the difference threshold, and then determine the i-th Q&A model as the Q&A model.
6. The method according to claim 2, characterized in that The determining of the multiple candidate inference paths based on the first question sample through the Q&A model includes: Through the Q&A model, generate multiple initial inference paths based on the first question sample, where the path length of the initial inference path is less than the path length of the candidate inference path; For each of the initial inference paths, through the Q&A model, expand the path of the initial inference path based on the first question sample to obtain the candidate inference path corresponding to the initial inference path.
7. The method according to claim 1, wherein The classifying of the multiple candidate inference paths to obtain the first inference path belonging to the first category and the second inference path belonging to the second category includes: Perform perplexity evaluation on each of the candidate inference paths to obtain the perplexity of each candidate inference path, where the magnitude of the perplexity is negatively correlated with the magnitude of the inference accuracy of the candidate inference path; Determine the candidate inference paths with perplexity lower than the perplexity threshold as the first inference path, and determine the candidate inference paths with perplexity greater than or equal to the perplexity threshold as the second inference path.
8. The method according to claim 1, characterized in that The classifying of the multiple candidate inference paths to obtain the first inference path belonging to the first category and the second inference path belonging to the second category includes: Perform the following processing for each of the candidate inference paths respectively: Through the Q&A model, generate the candidate answer corresponding to the candidate inference path, and determine the second similarity between the candidate answer and the answer label; When the second similarity is greater than the similarity threshold, determine the candidate inference path as the first inference path; When the second similarity is less than or equal to the similarity threshold, determine the candidate inference path as the second inference path.
9. The method according to claim 1, characterized in that When the number of the first inference paths is multiple, the determining of the first loss value of the Q&A model based on the first inference path and the answer label includes: Based on the answer label, determine the reference inference path with the maximum answer similarity to the answer label from the multiple first inference paths; Through the Q&A model, generate the reference answer corresponding to the reference inference path based on the reference inference path; Based on the difference between the reference answer and the answer label, determine the first target loss value of the Q&A model; Determine a first reference loss value of the question-and-answer model based on each inference step in the reference inference path; Perform a weighted sum of the first target loss value and the first reference loss value to obtain the first loss value.
10. The method according to claim 1, characterized in that, When the number of the first inference paths is multiple, determining the first loss value of the question-and-answer model based on the first inference path and the answer label includes: For each of the first inference paths, through the question-and-answer model, generate a second answer corresponding to the first inference path based on the first inference path; Determine a first target loss value corresponding to the first inference path based on the difference between the second answer and the answer label; Determine a first reference loss value corresponding to the first inference path based on each inference step in the first inference path; Perform a weighted sum of the first target loss value and the first reference loss value to obtain the first loss value.
11. The method according to claim 1, wherein Determining the second loss value of the question-and-answer model based on the first inference path and the second inference path includes: Obtain a first conditional probability of each of the first inference paths generated by the question model under the condition of inputting the first question sample, and obtain a second conditional probability of each of the second inference paths generated by the question model under the condition of inputting the first question sample; Determine the sum of the first conditional probabilities as a third conditional probability, and determine the sum of the second conditional probabilities as a fourth conditional probability; Divide the third conditional probability by the fourth conditional probability to obtain the second loss value.
12. The method according to claim 1, wherein Determining the second loss value of the question-and-answer model based on the first inference path and the second inference path includes: Obtain a first conditional probability of each of the first inference paths generated by the question model under the condition of inputting the first question sample, and obtain a second conditional probability of each of the second inference paths generated by the question model under the condition of inputting the first question sample; For each of the first conditional probabilities, divide the first conditional probability by each of the second conditional probabilities to obtain a first result corresponding to each of the second conditional probabilities, and sum the first results to obtain a second result corresponding to the first conditional probability; Sum the second results corresponding to each of the first conditional probabilities to obtain the second loss value.
13. The method according to claim 1, wherein Training the question-and-answer model based on the first loss value and the second loss value to obtain a target question-and-answer model includes: Train the question-and-answer model based on the first loss value and the second loss value to obtain a reference question-and-answer model; Through the reference question-and-answer model, generate a third answer to the first inference path based on the first inference path; Replace the answer label carried in the first question sample with the third answer to obtain a supplementary question sample carrying the third answer; Train the reference question-and-answer model based on the supplementary question sample to obtain the target question-and-answer model.
14. A training device for a question-and-answer model, characterized in that, The device includes: A determination module, configured to determine multiple candidate inference paths based on a first question sample carrying an answer label, where the candidate inference paths are used to indicate the inference steps that need to be executed to solve the problem corresponding to the first question sample; A classification module, configured to classify the multiple candidate inference paths to obtain a first inference path belonging to a first category and a second inference path belonging to a second category, where the inference accuracy of the first inference path is greater than that of the second inference path; A loss module, configured to determine a first loss value of the question-and-answer model based on the first inference path and the answer label, and determine a second loss value of the question-and-answer model based on the first inference path and the second inference path; A training module, configured to train the question-and-answer model based on the first loss value and the second loss value to obtain a target question-and-answer model.
15. An electronic device, characterized in that, The electronic device includes: A memory, configured to store computer-executable instructions or a computer program; A processor, configured to implement the method according to any one of claims 1 to 13 when executing the computer-executable instructions or the computer program stored in the memory.
16. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that, The computer-executable instructions or the computer program, when executed by the processor, implement the method according to any one of claims 1 to 13.
17. A computer program product, comprising a computer program or computer-executable instructions, characterized in that, The computer program or the computer-executable instructions, when executed by the processor, implement the method according to any one of claims 1 to 13.
Citation Information
Cited By
Question and answer model training method and question and answer task processing method
CN120744074A
Method and device for performing large-model self-distillation absorption depth reasoning by intelligent computing center cloud platform through computing power
CN120745848A
Question and answer large model training method, question and answer method and device, equipment and storage medium
CN120893511A
Medical training data quality inspection method, related equipment and program product
CN121439183A
Prompt word optimization method and device and electronic equipment
CN121660070A