Question and answer method, device and equipment based on large language model and storage medium
By adopting a division of labor and cooperation approach in large language models, using models with fewer parameters to execute thought chain reasoning steps and using models with more parameters to generate answers, the computational cost and response delay problems of large language models in the fields of financial management and medical consulting are solved, and efficient and personalized services are achieved.
Patent Information
- Application Number
- CN202510855839.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-10-10
AI Technical Summary
The existing large language models used in financial management and medical consulting have high computational costs and long response delays, making it difficult to provide efficient and personalized services.
Two large language models are used to work together. The first large language model with fewer parameters performs the task of generating the thought chain reasoning steps, and the second large language model with more parameters performs the task of generating the answer. The accuracy of the thought chain reasoning steps is ensured through knowledge distillation technology.
It reduces computing costs and response delays while ensuring the accuracy and personalization of answers, improving user experience.
Smart Images

Figure CN120763288A_ABST
Abstract
Description
Technical Field
[0001] One or more embodiments of the present application relate to the field of artificial intelligence technology, and in particular, to a question-answering method, apparatus, device, and storage medium based on a large language model. Background Art
[0002] A Large Language Model (LLM) is a natural language processing model based on deep learning technology, boasting powerful language understanding and generation capabilities. LLMs are increasingly being used. For example, intelligent dialogue systems (IDSs) based on LLMs are widely used in numerous fields, including customer service, online education, medical consulting, financial management, and legal aid. In the financial management field, for example, IDSs have become a key tool for improving the quality and efficiency of financial services. By simulating human communication, they provide users with financial-related conversations, Q&A, and query services, such as asset allocation, investment analysis, and risk identification. IDSs are a specific user-facing application of LLMs, designed to understand and answer questions posed by users in natural language and generate concise and clear answers. Specifically, IDSs are based on LLMs, which understand and answer user questions and generate corresponding answers.
[0003] In actual application scenarios, it is usually expected to ensure the generation accuracy of the large language model while effectively reducing the computational cost and response delay of the large language model, so that the large language model can provide services to users accurately and efficiently, improving the user experience when using the large language model. Summary of the Invention
[0004] One or more embodiments of the present application provide the following technical solutions:
[0005] This application provides a question-answering method based on a large language model, the method comprising:
[0006] Get the problems to be solved;
[0007] Inputting the question into a first language model, and generating a thought chain reasoning step corresponding to the question by the first language model;
[0008] The question and the thought chain reasoning steps are input into a second large language model, and the second large language model generates an answer corresponding to the question based on the thought chain reasoning steps; wherein the number of model parameters of the first large language model is less than the number of model parameters of the second large language model.
[0009] The present application also provides a question-answering device based on a large language model, the device comprising:
[0010] Get the module and get the problem to be solved;
[0011] A first generation module inputs the question into a first large language model, and the first large language model generates a thought chain reasoning step corresponding to the question;
[0012] The second generation module inputs the question and the thought chain reasoning steps into the second largest language model, and the second largest language model generates an answer corresponding to the question based on the thought chain reasoning steps; wherein the number of model parameters of the first largest language model is less than the number of model parameters of the second largest language model.
[0013] The present application also provides an electronic device, comprising:
[0014] processor;
[0015] a memory for storing processor-executable instructions;
[0016] The processor implements the steps of any of the above methods by running the executable instructions.
[0017] The present application also provides a computer-readable storage medium having computer instructions stored thereon, which, when executed by a processor, implement the steps of any of the methods described above.
[0018] In the above technical solution, the number of model parameters of the first largest language model used to perform the task of generating thought chain reasoning steps may be less than the number of model parameters of the second largest language model used to perform the question-answering task; for the problem to be solved, the problem can be first input into the first largest language model, and the first largest language model generates the thought chain reasoning steps corresponding to the problem, and then the problem and the thought chain reasoning steps are input into the second largest language model, and the second largest language model generates the answer corresponding to the problem based on the thought chain reasoning steps.
[0019] Using this approach, the reasoning process requiring the output of intermediate reasoning steps can be divided into two nodes: thought chain reasoning process generation and answer generation. The first language model, with a smaller number of model parameters, performs the thought chain reasoning step generation task, while the second language model, with a larger number of model parameters, performs the answer generation task. On the one hand, due to the smaller number of model parameters of the first language model, the computational cost and response latency of the first language model are also relatively small. On the other hand, the first language model can be obtained by performing knowledge distillation on the second language model, thereby ensuring the accuracy of the thought chain reasoning steps generated by the first language model, and thus the accuracy of the answers further generated by the second language model based on these thought chain reasoning steps. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The following is a description of the accompanying drawings required for describing the exemplary embodiments, in which:
[0021] Figure 1 It is a schematic diagram of an intelligent dialogue system shown in an exemplary embodiment of the present application.
[0022] Figure 2 This is a flowchart of a question-answering method based on a large language model, shown as an exemplary embodiment of the present application.
[0023] Figure 3 This is a schematic diagram of a question-answering process based on a large language model, shown as an exemplary embodiment of the present application.
[0024] Figure 4 It is a structural diagram of a device shown in an exemplary embodiment of the present application.
[0025] Figure 5 This is a block diagram of a question-answering device based on a large language model, shown as an exemplary embodiment of the present application. DETAILED DESCRIPTION
[0026] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. When the following description refers to the drawings, identical numbers in different figures represent identical or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with one or more embodiments of the present application. Rather, they are merely examples consistent with certain aspects of one or more embodiments of the present application.
[0027] It should be noted that, in other embodiments, the steps of the corresponding method are not necessarily performed in the order shown and described in this application. In some other embodiments, the method may include more or fewer steps than those described in this application. In addition, a single step described in this application may be broken down into multiple steps for description in other embodiments; and multiple steps described in this application may be combined into a single step for description in other embodiments.
[0028] In real-world applications, users often seek intelligent and efficient methods to help solve their problems. Large language models have garnered widespread attention for their superior natural language understanding and generation capabilities. Trained on large-scale text datasets, these models can understand and process the nuances of human language, providing new insights into various problems.
[0029] For example, in the field of medical consultation, faced with a growing number of patients and a complex and diverse range of disease types, the workload of medical staff continues to increase. Traditional one-on-one manual consultation services are inefficient and unable to meet the needs of large populations. Moreover, for rare or complex diseases, a single doctor may lack sufficient experience and resources to provide optimal diagnosis and treatment recommendations. In addition, the expertise gap between patients and medical providers leads to difficulties in doctor-patient communication. Many patients have difficulty understanding complex medical terminology, which affects their understanding of their health status and their compliance with treatment plans. Therefore, there is a need for a technology that can integrate massive amounts of medical knowledge, achieve fast and accurate query analysis, and provide users with personalized medical advice in an easy-to-understand manner.
[0030] Large language models can not only process and understand large amounts of medical text data, but also help doctors improve their work efficiency and assist in decision-making through natural language interaction. They can also directly provide patients with easy-to-understand medical information, thereby improving the service quality and accessibility of the entire medical consultation field.
[0031] For example, in the field of financial management, with the development of a globalized economy and the increasing complexity of financial markets, investors are increasingly demanding efficient, personalized, and scientific asset management solutions. This is particularly true in the area of fund allocation, where making accurate decisions based on massive amounts of information has become a key issue. Financial market data is vast and complex, including but not limited to macroeconomic indicators, corporate financial reports, market trend analysis, and policy changes. This data comes from a wide range of sources and updates rapidly, making it difficult for traditional data analysis methods to timely and effectively process and extract valuable information to guide investment decisions. Furthermore, investor needs are highly personalized, with clients with varying risk preferences, return expectations, and capital sizes demanding customized investment strategies. However, existing automated investment tools often offer limited options or general solutions, failing to fully account for individual differences and resulting in low customer satisfaction. Furthermore, the uncertainty and volatility of financial markets complicate investment decisions. Accurate forecasts require comprehensive consideration of multiple factors and the ability to quickly adapt to emerging information. Traditional methods are inefficient in handling these dynamic adjustments and struggle to achieve immediate response. Therefore, there is a need for a technology that can integrate multi-source heterogeneous data, automatically learn market patterns, and provide personalized investment advice.
[0032] Large language models can identify potential investment opportunities and risks by understanding and parsing large amounts of text information (such as news reports, research reports, etc.), and dynamically adjust strategies based on real-time market changes, providing investors with more accurate and personalized fund allocation recommendations, thereby improving the scientific nature and accuracy of investment decisions.
[0033] In this application, medical consultation systems, financial management systems and other service systems can be implemented using intelligent dialogue systems (or intelligent question-and-answer systems). Taking the intelligent dialogue system as an example, it provides users with dialogue, question-and-answer, query and other services by simulating human communication methods. The intelligent dialogue system is a specific application form of the large language model for users. It aims to understand and answer questions raised by users in natural language and generate concise and clear answers. Specifically, the intelligent dialogue system is based on the large language model, which understands and answers questions raised by users and generates corresponding answers.
[0034] In actual applications, using an intelligent dialogue system, users can submit questions to a large language model, which will understand and answer the questions raised by the users and generate corresponding answers, allowing users to formulate problem-solving solutions based on the answers corresponding to the questions output by the large language model.
[0035] Large language models are natural language processing models based on deep learning technology, with strong language understanding and generation capabilities. Large language models generally refer to deep learning models trained using large amounts of text data, which can be used to understand the meaning of natural language text or generate natural language text. Large language models can handle a variety of natural language tasks, such as text classification, named entity recognition (NER), question answering, dialogue, etc., and are an important way to artificial intelligence.
[0036] In the field of natural language processing (NLP), large-scale text datasets are often referred to as corpora. Corpora can contain various types of text data, such as literary works, academic papers, legal documents, news reports, daily conversations, emails, online forum posts, etc. By learning the text data in the corpus, large language models can acquire and understand the rules and patterns of natural language, and then effectively process and generate human language.
[0037] Large language models usually adopt the Transformer architecture, that is, large language models are usually deep learning models based on the Transformer architecture. Deep learning models based on the Transformer architecture are a class of neural network models that use the Transformer architecture, which performs well in natural language processing and other fields.
[0038] Transformer is a neural network model for sequence-to-sequence modeling. Transformer does not need to rely on recursive structure, and can parallelize training and inference, speeding up model processing. In deep learning models based on the Transformer architecture, a multi-layer Transformer encoder is usually used to extract features from the input sequence, and a Transformer decoder is used to convert the extracted features into an output sequence. At the same time, such models usually also use self-attention mechanisms to capture long-distance dependencies in input sequences, and use residual connections and normalization methods to speed up training and improve model performance.
[0039] A pretrained model is a large language model pretrained on large amounts of unlabeled text data. Pretrained models are general-purpose models; they are not designed or optimized for specific tasks. To adapt pretrained models to specific application scenarios and task requirements, they require fine-tuning to improve their performance on specific tasks. The large language model that is ultimately put into use is typically a pretrained model that has been further fine-tuned, performing supervised learning on labeled text data. Pretraining and fine-tuning are complementary processes: pretraining enables the model to acquire broad language understanding capabilities, while fine-tuning makes the model more specialized and accurate for specific tasks.
[0040] In other words, the training process of a large language model can be divided into two stages: pre-training and fine-tuning. During the pre-training stage, unsupervised learning (e.g., self-supervised learning) can be used on large-scale, unlabeled text datasets (e.g., online encyclopedias, online articles, books, etc.). Specifically, the model can predict missing parts or the next word based on the context, learn statistical laws such as semantics and syntax, and language structure. Backpropagation and optimization algorithms (e.g., gradient descent) are used to minimize prediction losses, iteratively update model parameters, and gradually improve the model's understanding of language. During the fine-tuning phase, you can select corresponding supervised learning tasks (for example, text classification, named entity recognition, question-answering systems, dialogue systems, etc.) based on the specific application scenarios and task requirements, and prepare task-specific text datasets. You can then use the pre-trained model as the starting point for fine-tuning, and use supervised learning to fine-tune on the task-specific text dataset. Specifically, you can perform the task based on the text dataset, and minimize the loss used to measure the performance of the model in processing specific tasks through backpropagation and optimization algorithms (for example, gradient descent). The model parameters are iteratively updated to gradually improve the model's performance on specific tasks. In practical applications, fine-tuning can flexibly choose supervised learning, unsupervised learning, or semi-supervised learning methods based on the specific application scenarios and the type of available data.
[0041] The language comprehension ability learned by the large language model during the pre-training and fine-tuning stages enables the large language model to perform logical inference, knowledge reasoning, or problem-solving by understanding, analyzing, and integrating text information when faced with complex problems or tasks. This ability is usually referred to as the reasoning ability of the large language model.
[0042] In practical applications, the pre-trained large language model is usually called the base model of the large language model, and the fine-tuned large language model is called the serving model of the large language.
[0043] Large language models are typically guided or stimulated by prompts (often called prompts) to perform specific tasks. A prompt can be an initial text or text fragment provided to the large language model, such as a sentence, a question, or a conversation, intended to guide or stimulate the model to produce the corresponding output. Prompts are key tools for guiding model output and can be very simple or quite complex, including instructions, examples, and descriptions of the desired output format. Prompts explicitly indicate the task expected of the large language model, such as answering a question, simulating a conversation, writing an article, or translating text. Prompts also provide the large language model with necessary background information and context, enabling it to understand the logic, style, theme, or stance that should be followed when generating content. Furthermore, prompts can inspire the large language model to demonstrate its inherent knowledge or specific language abilities, such as explaining complex concepts, citing regulations, or imitating the writing style of a specific author.
[0044] Since large language models are primarily used to understand and generate human language based on text processing, prompts typically appear in the form of text. However, in practical applications, large language models can also accept other forms of input as prompts, such as images, audio, and even video, provided that the large language model is designed or trained to process multimodal data.
[0045] One or more embodiments of the present application provide a technical solution for implementing question answering based on a large language model. In this technical solution, the number of model parameters of the first large language model used to perform the task of generating thought chain reasoning steps may be less than the number of model parameters of the second large language model used to perform the question answering task; for a problem to be solved, the problem can be first input into the first large language model, and the first large language model generates the thought chain reasoning steps corresponding to the problem, and then the problem and the thought chain reasoning steps are input into the second large language model, and the second large language model generates the answer corresponding to the problem based on the thought chain reasoning steps.
[0046] Using this approach, the reasoning process requiring the output of intermediate reasoning steps can be divided into two nodes: thought chain reasoning process generation and answer generation. The first language model, with a smaller number of model parameters, performs the thought chain reasoning step generation task, while the second language model, with a larger number of model parameters, performs the answer generation task. On the one hand, due to the smaller number of model parameters of the first language model, the computational cost and response latency of the first language model are also relatively small. On the other hand, the first language model can be obtained by performing knowledge distillation on the second language model, thereby ensuring the accuracy of the thought chain reasoning steps generated by the first language model, and thus the accuracy of the answers further generated by the second language model based on these thought chain reasoning steps.
[0047] Please refer to Figure 1 , Figure 1 It is a schematic diagram of an intelligent dialogue system shown in an exemplary embodiment of the present application.
[0048] like Figure 1 As shown, the intelligent dialogue system may include a server and at least one client accessing the server via any type of wired or wireless network.
[0049] The above-mentioned server may correspond to a server comprising an independent physical host, or may be a server cluster consisting of multiple independent physical hosts; or, may correspond to a virtual server, cloud server, etc. hosted by a host cluster.
[0050] The above-mentioned client can correspond to terminal devices such as smart phones, tablet computers, laptops, desktop computers, PCs (Personal Computers), PDAs (Personal Digital Assistants), wearable devices (such as smart glasses, smart watches, etc.), smart car devices or game consoles.
[0051] Users can use the intelligent dialogue service provided by the intelligent dialogue system through the client; the client and the server can implement user-oriented intelligent dialogue services through data interaction between each other.
[0052] Specifically, the above-mentioned server can be equipped with a large language model, and the above-mentioned intelligent dialogue system can be based on the large language model. The large language model can perform reasoning based on the question text (which can be called Query or Question) to understand and answer the questions raised by the question text, and generate an answer text corresponding to the question text (which can be called Response or Answer).
[0053] For example, the client can output a corresponding user interface to the user, allowing the user to perform operations such as inputting question text and uploading documents or images to assist in asking questions in the user interface, thereby asking questions to the intelligent dialogue system and using the intelligent dialogue service provided by the intelligent dialogue system. The client can send the question text input by the user to the server, and the server can generate a corresponding answer text based on the question text and output the answer text to the user, that is, return the answer text to the client, and the client can display the answer text to the user through the user interface for the user to view, thereby realizing a user-oriented intelligent dialogue service.
[0054] It should be noted that the question text can be regarded as a special prompt. The question text describes the specific problem that the user wants to solve, and the problem is usually expressed through a carefully designed prompt.
[0055] In practical applications, intelligent dialogue systems rely primarily on the knowledge acquired by their large language models during training by studying static corpora. Due to the limitations of this knowledge, intelligent dialogue systems may experience hallucinations when answering complex or specific questions. Hallucinations occur when the content generated by the large language model appears highly plausible and coherent, sometimes even mimicking human emotions and thinking patterns, creating the illusion of "understanding" the input, even though the content is actually inaccurate or misleading.
[0056] In order to improve the adaptability and response accuracy of intelligent dialogue systems, the RAG (Retrieval-Augmented Generation) approach can be adopted, combining information retrieval and model generation. This allows the intelligent dialogue system to no longer rely solely on the knowledge acquired by the large language model through learning static corpus during the training process when answering questions raised by users. Instead, it can first perform information retrieval in a large document collection based on the question, and then understand and answer the question based on the retrieved relevant documents, and generate the corresponding answer. In other words, the document collection can be combined with the large language model, and during the model generation process, relevant information can be retrieved from the document collection in real time to assist the model in making more accurate and comprehensive answers or decisions. Because the retrieved information and the context of the question are taken into account during the model generation process, it can be ensured that the generated content not only meets actual needs, but is also accurate, reliable, coherent, and natural.
[0057] Specifically, the server can also be equipped with a knowledge base and an information retrieval component. This knowledge base is external to the large language model installed on the server. That is, the data in this knowledge base is not the knowledge acquired by the large language model through learning during training. Instead, it serves as auxiliary information in the large language model's reasoning process, assisting the large language model in generating an answer text corresponding to the question text. During the reasoning process of the large language model, the information retrieval component can perform information retrieval in the knowledge base based on the question text, thereby assisting the large language model in generating an answer text corresponding to the question text through the retrieved relevant information.
[0058] It should be noted that the aforementioned server can also be equipped with other functional components or subsystems such as a prompt generation component. These components or subsystems can work in conjunction with the large language model installed on the server to jointly generate the answer text corresponding to the question text.
[0059] Furthermore, the intelligent dialogue system described above can have only one large language model, which can be used to perform dialogue tasks or question-answering tasks, and can also, depending on actual needs, be used to perform tasks such as generating thought chain reasoning steps, information retrieval tasks, context compression tasks, and relationship extraction tasks. Alternatively, the intelligent dialogue system can have multiple large language models, which, in addition to the large language model used to perform dialogue tasks or question-answering tasks, can also include large language models used to perform thought chain reasoning step generation tasks, large language models used to perform information retrieval tasks, large language models used to perform context compression tasks, and large language models used to perform relationship extraction tasks.
[0060] One or more embodiments of the present application provide a method for implementing a question answering method based on a large language model, which can be applied to Figure 1 The intelligent dialogue system shown (or the specific data processing components in the intelligent dialogue system).
[0061] In the above-mentioned question-answering method based on the large language model, two different large language models can be used, and these two large language models can be respectively referred to as the first large language model and the second large language model. Among them, the number of model parameters of the first large language model can be less than the number of model parameters of the second large language model; that is, the first large language model can be a large language model of small size, and the second large language model can be a large language model of large size. For example, the first large language model can be a Qwen2.5-1.5B-level model (that is, a large language model with 150 million parameters), and the second large language model can be a Qwen2.5-72B-level model (that is, a large language model with 7.2 billion parameters). Since the number of model parameters of the first large language model is small, the computational cost and response delay of the first large language model are also relatively small.
[0062] The first language model can be used to perform the task of generating reasoning steps of a thought chain. That is, the first language model can perform reasoning based on the question as a prompt and generate the reasoning steps of a thought chain corresponding to the question under the guidance of the prompt.
[0063] Chain of Thought is a technology used to improve the performance of AI models in complex reasoning tasks. It mimics the step-by-step thinking process of humans when solving problems, breaking down the problem into multiple intermediate steps and gradually deriving the final answer. Chain of Thought can be introduced into the training and use of large-scale language models, aiming to enhance the model's interpretability and ability to solve complex problems by explicitly expressing intermediate reasoning steps.
[0064] Thought chaining emphasizes breaking down the problem-solving process into a series of sequential, logically related intermediate reasoning steps. This involves identifying the key elements of the problem, listing possible solution paths, and evaluating the feasibility of each path until the final answer is reached. By including the thinking and reasoning behind each step in the model's input or output, the decision-making process becomes transparent, which is critical for understanding the model's decision logic and improving interpretability. Thought chaining aims to make the machine's thinking process more human-like. By mimicking how humans think step by step to solve complex problems, it helps the model make reasonable inferences in the absence of direct training data.
[0065] The specific implementation of the thinking chain can be to directly embed this step-by-step reasoning format in the Prompt of the service model for the large language model, guiding the large language model to generate intermediate reasoning steps while generating the answer corresponding to the question; or, it can also be to construct a question and answer dataset containing intermediate reasoning steps as a thinking chain sample set, and use the thinking chain sample dataset to fine-tune the basic model of the large language model, so that the service model of the large language model obtained by fine-tuning can subsequently perform the thinking chain reasoning step generation task.
[0066] Specifically, thought chain samples can include key content such as problem statements, intermediate reasoning steps, and final answers, which are used to guide and train large language models for step-by-step logical reasoning and problem solving.
[0067] A problem statement is a description of the specific problem or task that needs to be solved. It can be a mathematical problem, logical reasoning, factual inquiry, or any inquiry that requires a series of thinking steps to arrive at an answer.
[0068] The intermediate reasoning steps are the core of the thought chain sample, consisting of a series of logical reasoning, calculations, or analyses that gradually derive the answer from the problem. Each step should be coherent, and the previous step should reasonably lead to the next step until the final answer is reached. For example, when solving math problems, this may involve applying formulas, replacing variables, and simplifying calculations; in logical reasoning tasks, it may include premise analysis and hypothesis verification.
[0069] The final answer is the clear answer to the question after all the intermediate reasoning steps. This answer should be the natural result of the chain of reasoning and closely connected with the previous reasoning steps.
[0070] It should be noted that in order to ensure the accuracy of the thinking chain reasoning steps corresponding to the questions generated by the large language model after training (usually fine-tuning), the thinking chain samples used to train the large language model usually contain three parts: questions, thinking chain reasoning steps and answers. However, in some cases, such as in order to speed up the model training speed, the thinking chain samples used to train the large language model may also only contain two parts: questions and thinking chain reasoning steps. This application does not impose any special restrictions on the specific content of the thinking chain samples used to train the above-mentioned first large language model and the above-mentioned second large language model.
[0071] Thought chain examples can guide the large language model to learn to imitate such reasoning patterns through complete examples of questions, intermediate reasoning steps, and answers. In other words, the large language model can first learn the reasoning pattern shown by the thought chain examples, then imitate this reasoning pattern, reason based on the question, and output the intermediate reasoning steps and inferred answers in this reasoning process.
[0072] The above-mentioned second largest language model can be used to perform dialogue tasks or question-answering tasks. That is, the second largest language model can perform further reasoning based on the above-mentioned question and the thought chain reasoning steps corresponding to the question generated by the above-mentioned first largest language model to generate an answer corresponding to the question. Specifically, a Prompt can be generated based on the question and its corresponding thought chain reasoning steps. For example, the question and its corresponding thought chain reasoning steps can be spliced as a Prompt, and the second largest language model performs further reasoning based on the question and its corresponding thought chain reasoning steps in the Prompt, and generates an answer corresponding to the question under the guidance of the Prompt.
[0073] It should be noted that the first large language model can refer to a service model of the first large language model. In actual application, the first large language model can be pre-trained on a large-scale unlabeled text dataset in an unsupervised learning manner to obtain a base model of the first large language model; then, the thought chain reasoning step generation task can be used as a supervised learning task in fine-tuning, and a text dataset specific to the thought chain reasoning step generation task can be prepared, and then the base model of the first large language model can be used as a starting point for fine-tuning, and the text dataset specific to the thought chain reasoning step generation task can be used for fine-tuning in a supervised learning manner to obtain the service model of the first large language model.
[0074] The text dataset specific to the thought chain reasoning step generation task can include the thought chain sample. In this case, the base model of the first large language model learns the reasoning mode shown in the thought chain sample, and becomes the service model of the first large language model. Subsequently, the service model of the first large language model can imitate the reasoning mode to generate the thought chain reasoning step corresponding to the question.
[0075] In actual application, for a thought chain sample, the question therein can be used as sample content, and the thought chain reasoning step (or the thought chain reasoning step and the answer) therein can be used as sample label. Fine-tuning of the large language model based on the thought chain sample can be supervised training.
[0076] Similarly, the second large language model can refer to a service model of the second large language model. In actual application, the second large language model can be pre-trained on a large-scale unlabeled text dataset in an unsupervised learning manner to obtain a base model of the second large language model; then, the dialogue task or the question and answer task can be used as a supervised learning task in fine-tuning, and a text dataset specific to the dialogue task or the question and answer task can be prepared, and then the base model of the second large language model can be used as a starting point for fine-tuning, and the text dataset specific to the dialogue task or the question and answer task can be used for fine-tuning in a supervised learning manner to obtain the service model of the second large language model.
[0077] In constructing the text dataset specific to the dialogue task or the question and answer task, the collected questions (the questions are training samples) can be labeled, and answers corresponding to the questions can be labeled (the labeled answers are labels of the training samples). In this way, the labeled questions can be used as the text dataset specific to the dialogue task or the question and answer task for supervised training of the second large language model.
[0078] In some embodiments, since the number of model parameters of the second large language model is much larger than the number of model parameters of the first large language model, in order to ensure the model effect of the first large language model as much as possible, thereby ensuring the accuracy of the thought chain reasoning steps corresponding to the question generated by the first large language model, the second large language model (which can be a basic model or a service model that cannot perform the thought chain reasoning step generation task but can perform other tasks) can be fine-tuned using the thought chain samples first, so that the fine-tuned second large language model can be used to perform the thought chain reasoning step generation task, and then the knowledge related to the thought chain reasoning step generation in the second large language model that can be used to perform the thought chain reasoning step generation task is migrated to the first large language model in the manner of model distillation.
[0079] Among them, knowledge distillation (Knowledge Distillation) is a model compression technology, its core idea is to migrate the knowledge of a complex, large and excellent performance "teacher model" (Teacher Model) to a simpler and lighter "student model" (Student Model). It trains the student model to imitate the behavior of the teacher model, which can reduce the computational cost and resource demand of the model while maintaining high model performance, making the student model more suitable for deployment on resource-constrained devices (such as mobile devices or embedded systems).
[0080] Specifically, in the process of migrating the knowledge related to the thought chain reasoning step generation in the second large language model that can be used to perform the thought chain reasoning step generation task to the first large language model in the manner of model distillation, first, on the one hand, the thought chain samples used to train (usually fine-tune) the second large language model that can be used to perform the thought chain reasoning step generation task as the teacher model can be obtained, and on the other hand, for each thought chain sample, the thought chain reasoning steps corresponding to the question in the thought chain sample generated by the second large language model based on the reasoning of the question in the thought chain sample can be obtained; then, for each thought chain sample, the original thought chain reasoning steps in the thought chain sample can be further replaced with the thought chain reasoning steps corresponding to the question in the thought chain sample generated by the second large language model based on the reasoning of the question in the thought chain sample, to update the thought chain sample, and the updated thought chain sample is determined as a distillation sample; finally, the first large language model as the student model can be trained (usually also fine-tuned) based on the distillation sample, so that the distillation loss function corresponding to the first large language model converges.
[0081] Alternatively, for each thought chain sample, the thought chain reasoning steps and answers corresponding to the question generated by the above-mentioned second largest language model based on the question in the thought chain sample can be obtained, and the original thought chain reasoning steps and answers in the thought chain sample can be further replaced with the thought chain reasoning steps and answers corresponding to the question generated by the second largest language model based on the question in the thought chain sample to update the thought chain sample, and the updated thought chain sample can be determined as a distilled sample.
[0082] In some embodiments, in order to speed up the generation of thought chain reasoning steps, the thought chain reasoning steps in the above thought chain samples may be thought chain reasoning steps that have undergone context compression processing.
[0083] Specifically, when the first language model is trained directly based on thought chain samples, the thought chain reasoning steps in each thought chain sample used can be thought chain reasoning steps after context compression.
[0084] When the first language model as a student model is obtained by adopting the knowledge distillation method, the chain of thought reasoning steps in each chain of thought sample used to train the second language model as a teacher model can be the chain of thought reasoning steps after context compression, so that the trained second language model can directly generate compressed chain of thought reasoning steps; or, the chain of thought reasoning steps in each chain of thought sample used to train the second language model as a teacher model can be complete chain of thought reasoning steps, and after generating the complete chain of thought reasoning steps, the second language model can further perform context compression on the generated chain of thought reasoning steps to obtain compressed chain of thought reasoning steps. In this way, the chain of thought reasoning steps in the distilled samples obtained by sample update can be the chain of thought reasoning steps after context compression.
[0085] Context compression is a technique for optimizing the inference efficiency of large language models. It aims to reduce computational costs and response times by reducing redundant information in input prompts, while preserving key semantic content as much as possible to maintain the quality of model output. It is particularly important when processing long text or complex tasks, as the attention mechanism of large models often has input length limitations, and long contexts significantly increase computing resource consumption.
[0086] The core idea of context compression is to identify and retain the part of information that is most relevant to the current task and remove unnecessary or repetitive content. Context compression can be achieved through information filtering, summary generation, structured representation, dynamic clipping, and other methods. Information filtering refers to extracting the most critical information from the input, such as the question itself, relevant facts, or the part of the context that clearly affects the answer. Summary generation refers to using the model itself or other lightweight models to summarize the input content and compress it into a shorter but information-rich form. Structured representation refers to converting the original context into a structured data format (for example: JSON, table, etc.) to reduce redundant descriptions. Dynamic clipping refers to dynamically and selectively retaining part of the content that is closely related to the current reasoning step based on the current query or task requirements.
[0087] In practical applications, the context compression of the thought chain reasoning steps in the above thought chain samples can be completed manually, or tools such as LLMLingua (an open source tool that focuses on text compression in the large language model reasoning process) can be used to complete the context compression of the thought chain reasoning steps in the above thought chain samples. Alternatively, the context compression of the thought chain reasoning steps in the above thought chain samples can be completed by combining manual and tool methods. This application does not impose any special restrictions on this.
[0088] Since the thought chain reasoning steps in the thought chain samples used for training are thought chain reasoning steps after context compression, the above-mentioned first language model that has completed training can generate shorter thought chain reasoning steps, thereby speeding up the speed of the first language model in generating thought chain reasoning steps and reducing the response delay of the first language model.
[0089] In some embodiments, thought chains have a certain templated nature, meaning they typically follow a fixed structure or pattern. For example, they first pose a question, then gradually outline the thought steps required to solve it. These steps often follow a certain logical order and format. Therefore, during the aforementioned knowledge distillation process, the distillation loss function can be optimized to preserve the key reasoning logic.
[0090] In one example, in the above-mentioned knowledge distillation process, higher attention can be paid to the "logical connectives" (for example, "therefore", "because", "so", "hypothesis", "it can be seen from this", etc.) and "conclusion guiding words" (for example, "the result is", "finally concluded" and other reasoning conclusions, as well as mathematical operators such as "+", "-", "×", "÷", "=") in the reasoning steps of the thinking chain, thereby optimizing the distillation loss function and enabling the student model to more accurately inherit the key reasoning capabilities of the teacher model.
[0091] Specifically, the traditional loss function treats all tokens (which can be called word units, specifically words, phrases or sentences) equally, that is, the weights of the sub-loss functions corresponding to each token are equal. However, in this application, since the logical connectives and conclusion guides in the reasoning steps are more important than the detailed descriptions, the tokens containing the logical connectives and the reasoning conclusions can be used as anchor points of the thinking chain (Anchor Point, referring to words, phrases or sentences that play a key role in the reasoning process), and the weight of the sub-loss function corresponding to the token as the anchor point of the thinking chain in the above distillation loss function is made higher than the weights of the sub-loss functions corresponding to other tokens.
[0092] In practical applications, the above distillation loss function can be a KL divergence (Kullback-Leibler Divergence) loss function or a cross entropy loss function. Taking the KL divergence loss function as an example, the formula of the traditional KL divergence loss function can be as follows:
[0093]
[0094] Where i represents Token; P is the probability distribution of the teacher model; Q is the probability distribution of the student model.
[0095] In this application, the traditional KL divergence loss function can be extended to a weighted form:
[0096]
[0097] Where N is the sequence length, that is, the number of tokens in the reasoning step of the thinking chain; ω i is the weight of the ith Token, indicating the importance of the Token in the reasoning step of the thinking chain; i is the probability of the teacher model on the i-th Token; q i is the probability of the student model on the i-th token. At this point, the weight of the token as the anchor point of the thinking chain can be higher than the weight of other tokens.
[0098] In another example, the core of the thought chain is the causal / dependency relationship of the reasoning chain, which cannot be explicitly modeled by traditional loss functions. Therefore, in the above-mentioned knowledge distillation process, it is necessary not only to imitate the probability distribution output by the teacher model, but also to explicitly model and retain the logical relationship structure (such as causality, dependency, etc.), so that the student model can more accurately inherit the reasoning ability of the teacher model.
[0099] Specifically, traditional loss functions (such as KL divergence and cross entropy) calculate the loss equally for each token and fail to capture the semantic relationship or logical dependency between tokens. This may result in the student model imitating the surface text well, but the key reasoning chain is destroyed. However, in this application, a "logical relationship graph" can be extracted from the reasoning steps of the thought chain generated by the teacher model, and a graph structure consistency loss is introduced in the knowledge distillation process, so that the student model can not only imitate the token-level output, but also learn the reasoning structure behind it.
[0100] That is, the above-mentioned distillation loss function may include a graph structure consistency loss function. Among them, the graph structure consistency loss function can specifically be a graph structure consistency loss function calculated for the problems in each thought chain sample based on the entity relationship graph corresponding to the thought chain reasoning step of the problem generated by the above-mentioned second language model, and the entity relationship graph corresponding to the thought chain reasoning step of the problem generated by the above-mentioned first language model. Graph structure consistency generally refers to the similarity or equivalence of the graph data structure, which may involve consistency in multiple aspects such as the number, type and connection method of the nodes (Node) and edges (Edge) of the graph. In practical applications, the graph structure consistency loss can specifically be the cosine distance between graph embeddings, or the graph similarity calculated using the graph kernel function.
[0101] The entity relationship diagram corresponding to the thought chain reasoning steps generated by the above-mentioned second largest language model may include an entity relationship diagram generated based on the semantic relationship between entities obtained by relationship extraction of the thought chain reasoning steps generated by the second largest language model; the entity relationship diagram corresponding to the thought chain reasoning steps generated by the above-mentioned first largest language model may include an entity relationship diagram generated based on the semantic relationship between entities obtained by relationship extraction of the thought chain reasoning steps generated by the first largest language model.
[0102] Among them, relation extraction (RE) aims to automatically identify and extract semantic relationships between entities from unstructured text, and ultimately form structured knowledge.
[0103] The relationship extracted from the thought chain reasoning step can be expressed as a unit group in the form of (Subject, Relation, Object). For example, assuming that the thought chain reasoning step contains the sentence "According to Newton's second law, force is equal to mass multiplied by acceleration", the extracted relationship can be expressed as the following triple: ("Newton's second law", "definition", "force is equal to mass multiplied by acceleration"); assuming that the thought chain reasoning step contains the sentence "Because the weather is fine, we decided to go on a picnic", the extracted relationship can be expressed as the following triple: ("the weather is fine", "leads to", "go on a picnic"). In this case, Subject and Object can be used as nodes, and Relation as an edge to generate an entity relationship graph based on the extracted relationship.
[0104] In practical applications, the above distillation loss function can be the sum of a weighted loss function obtained by optimizing the traditional loss function and a graph structure consistency loss function.
[0105] Please refer to Figure 2 , Figure 2 This is a flowchart of a question answering method based on a large language model, shown in an exemplary embodiment of the present application. The question answering method based on a large language model can be applied to Figure 1 The intelligent dialogue system shown.
[0106] like Figure 2 As shown, the above-mentioned question answering method based on the large language model may include the following steps:
[0107] Step 202: Obtain the problem to be solved.
[0108] In this embodiment, the problem to be solved may be obtained first.
[0109] For example, the question text submitted by the user may be obtained, thereby obtaining the question raised by the user using the question text.
[0110] Step 204: Input the question into the first large language model, and generate the thought chain reasoning steps corresponding to the question by the first large language model.
[0111] In this embodiment, when the above question is obtained, the question can be input into the above first language model, and the first language model can perform reasoning based on the question to generate a thought chain reasoning step corresponding to the question.
[0112] For example, the question text submitted by the user can be input into the above-mentioned first language model as a prompt. Under the guidance of the prompt, the first language model performs reasoning based on the question text, thereby generating descriptive text corresponding to the various intermediate reasoning steps included in the reasoning process based on the question text.
[0113] Step 206: Input the question and the thought chain reasoning steps into the second largest language model, and the second largest language model generates an answer corresponding to the question based on the thought chain reasoning steps; wherein the number of model parameters of the first largest language model is less than the number of model parameters of the second largest language model.
[0114] In this embodiment, when the above-mentioned chain of thought reasoning steps corresponding to the above-mentioned question are obtained, the question and the chain of thought reasoning steps can be further input into the above-mentioned second largest language model, and the second largest language model can perform further reasoning based on the question and the chain of thought reasoning steps to generate an answer corresponding to the question.
[0115] For example, the question text submitted by the user and the descriptive text corresponding to the intermediate reasoning steps included in the reasoning process based on the question text generated by the above-mentioned first language model can be spliced together, and the spliced text can be input into the above-mentioned second language model as a new Prompt. Under the guidance of the Prompt, the second language model can further reason based on the question text in the Prompt and the descriptive text corresponding to the intermediate reasoning steps, so as to generate an answer text corresponding to the question text.
[0116] In some embodiments, the question may be about fund allocation, the chain of reasoning may include an analysis of user preferences and fund categories, and the answer may include fund categories and their allocation ratios. Alternatively, the question may be about medical diagnosis, the chain of reasoning may include the process of reasoning from symptoms to disease, and the answer may be the diagnosis result.
[0117] For example, the question could be: "I am 35 years old, have a stable income, a moderate risk tolerance, and hope to achieve steady growth in my assets. How should I allocate funds?"
[0118] The above thought chain reasoning steps can be as follows:
[0119] “Analyze user basic information and investment goals:
[0120] Age: 35 → Longer investment horizon, but not extremely high risk tolerance;
[0121] Stable income → can make regular fixed investments;
[0122] Medium risk tolerance → Not suitable for high-volatility products, but still want returns higher than deposits;
[0123] Goal: Steady appreciation → Hope that assets can maintain their value and grow moderately.
[0124] Filter fund categories based on user preferences:
[0125] Equity funds: They are more volatile and not suitable for all investments, but can be invested in small amounts to capture growth potential.
[0126] Hybrid funds: moderate risk, suitable as one of the main configurations;
[0127] Bond funds: low risk, stable returns, suitable for conservative investors;
[0128] Money market funds: highly liquid and extremely low risk, suitable for short-term funds or emergency preparedness;
[0129] Index funds / ETFs: low cost, high transparency, and can be used as part of a stock allocation.
[0130] Taking into account the risk diversification principle and the probability of achieving investment goals:
[0131] The need to balance benefits and risks;
[0132] Diversify your investments across different types of funds to avoid the risk of a single asset class.”
[0133] The above answer can be expressed as follows:
[0134] Based on your situation, we recommend the following fund allocation plan:
[0135] Hybrid funds, 40%, have moderate risk, balance returns and stability, and are suitable as core investments;
[0136] Bond funds, 30%, have relatively stable returns and reduce overall portfolio volatility;
[0137] Index funds (such as the CSI 300 ETF), 20%, to obtain long-term capital appreciation and control positions to control risks;
[0138] Money market funds, 10%, for liquidity management and to cope with unexpected expenses.”
[0139] Please refer to Figure 3 , Figure 3 This is a schematic diagram of a question-answering process based on a large language model, shown as an exemplary embodiment of the present application.
[0140] like Figure 3As shown, in the above-mentioned question-answering process based on the large language model, two different large language models can be used, and these two large language models can be respectively referred to as the first large language model and the second large language model. Among them, the number of model parameters of the first large language model can be less than the number of model parameters of the second large language model. The second large language model can be used to perform the task of generating the reasoning steps of the thought chain, so that the knowledge related to the generation of the reasoning steps of the thought chain in the second large language model can be transferred to the first large language model by using the model distillation method.
[0141] In the above situation, you can first get the problem that needs to be solved.
[0142] When the above question is obtained, the question can be input into the above first language model, and the first language model performs reasoning based on the question to generate a thought chain reasoning step corresponding to the question.
[0143] When the above-mentioned thought chain reasoning steps corresponding to the above-mentioned question are obtained, the question and the thought chain reasoning steps can be further input into the above-mentioned second largest language model, and the second largest language model can perform further reasoning based on the question and the thought chain reasoning steps to generate an answer corresponding to the question.
[0144] In the above technical solution, the number of model parameters of the first largest language model used to perform the task of generating thought chain reasoning steps may be less than the number of model parameters of the second largest language model used to perform the question-answering task; for the problem to be solved, the problem can be first input into the first largest language model, and the first largest language model generates the thought chain reasoning steps corresponding to the problem, and then the problem and the thought chain reasoning steps are input into the second largest language model, and the second largest language model generates the answer corresponding to the problem based on the thought chain reasoning steps.
[0145] Using this approach, the reasoning process requiring the output of intermediate reasoning steps can be divided into two nodes: thought chain reasoning process generation and answer generation. The first language model, with a smaller number of model parameters, performs the thought chain reasoning step generation task, while the second language model, with a larger number of model parameters, performs the answer generation task. On the one hand, due to the smaller number of model parameters of the first language model, the computational cost and response latency of the first language model are also relatively small. On the other hand, the first language model can be obtained by performing knowledge distillation on the second language model, thereby ensuring the accuracy of the thought chain reasoning steps generated by the first language model, and thus the accuracy of the answers further generated by the second language model based on these thought chain reasoning steps.
[0146] Corresponding to the aforementioned method embodiments, the present application also provides device embodiments.
[0147] Please refer to Figure 4 , Figure 4 4 is a structural diagram of a device shown in an exemplary embodiment of the present application. At the hardware level, the device includes a processor 402, an internal bus 404, a network interface 406, a memory 408 and a non-volatile memory 410, and of course may also include other required hardware. One or more embodiments of the present application can be implemented based on software, such as the processor 402 reading the corresponding computer program from the non-volatile memory 410 into the memory 408 and then running it. Of course, in addition to software implementation, one or more embodiments of the present application do not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc., that is, the execution subject of the following processing flow is not limited to each logic module, but can also be hardware or logic devices.
[0148] Please refer to Figure 5 , Figure 5 This is a block diagram of a question-answering device based on a large language model, shown as an exemplary embodiment of the present application.
[0149] The above-mentioned question-answering device based on the large language model can be applied to Figure 4 The device shown in the figure is used to implement the technical solution of this application. The device includes:
[0150] Acquisition module 502, acquiring the problem to be solved;
[0151] A first generating module 504 inputs the question into a first large language model, and generates a thought chain reasoning step corresponding to the question by the first large language model;
[0152] The second generation module 506 inputs the question and the thought chain reasoning steps into the second largest language model, and the second largest language model generates an answer corresponding to the question based on the thought chain reasoning steps; wherein the number of model parameters of the first largest language model is less than the number of model parameters of the second largest language model.
[0153] In some embodiments, the apparatus comprises:
[0154] a sample acquisition module, which acquires thought chain samples for training the second language model serving as a teacher model;
[0155] a sample updating module, which obtains a thought chain reasoning step and an answer corresponding to the question generated by the second largest language model through reasoning based on the question in the thought chain sample, updates the thought chain sample based on the generated thought chain reasoning step and the answer, and determines the updated thought chain sample as a distilled sample;
[0156] A distillation module trains the first large language model as a student model based on the distillation samples to converge a distillation loss function corresponding to the first large language model.
[0157] In some embodiments, the thought chain reasoning steps in the thought chain sample are thought chain reasoning steps after context compression.
[0158] In some embodiments, the weight of the sub-loss function corresponding to the word-gram serving as the anchor point of the thought chain in the distillation loss function is higher than the weight of the sub-loss function corresponding to other word-grams in the distillation loss function.
[0159] In some embodiments, the distillation loss function comprises a KL divergence loss function.
[0160] In some embodiments, the distillation loss function includes a graph structure consistency loss calculated based on an entity relationship graph corresponding to the thought chain reasoning step generated by the second largest language model, and an entity relationship graph corresponding to the thought chain reasoning step generated by the first largest language model; the entity relationship graph includes an entity relationship graph generated based on the semantic relationship between entities obtained by relationship extraction of the thought chain reasoning step.
[0161] In some embodiments, the question is related to fund allocation, the thought chain reasoning step is an analysis of user preferences and fund categories, and the answer is the fund category and its allocation ratio; or, the question is related to medical diagnosis, the thought chain reasoning step is the reasoning process from symptoms to disease, and the answer is the diagnosis result.
[0162] For the device embodiments, they basically correspond to the method embodiments, so for relevant details, please refer to the partial description of the method embodiments. The device embodiments described above are merely illustrative, wherein the modules described as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical modules, that is, they may be located in one place or distributed across multiple network modules. Some or all of the modules may be selected according to actual needs to achieve the purpose of the technical solution of this application.
[0163] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer, which may be in the form of a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email transceiver, game console, tablet computer, wearable device, or any combination of these devices.
[0164] In a typical configuration, a computer includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0165] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0166] Computer-readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, disk storage, quantum memory, graphene-based storage media or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0167] It should be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a..." does not preclude the presence of additional identical elements in the process, method, commodity, or apparatus comprising the element.
[0168] The above description is of specific embodiments of the present application. Other embodiments are within the scope of this application. In some cases, the actions or steps described in this application can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0169] The terms used in one or more embodiments of the present application are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of the present application. The singular forms "a," "the," and "the" are also intended to include the plural forms unless the context clearly indicates otherwise. The term "and / or" refers to and includes any or all possible combinations of one or more of the associated listed items.
[0170] The terms "one embodiment," "some embodiments," "example," "specific example," or "one implementation" used in one or more embodiments of the present application mean that the specific features or characteristics described in conjunction with the embodiment are included in at least one embodiment of the present application. The schematic descriptions of these terms do not necessarily refer to the same embodiment. Moreover, the specific features or characteristics described can be combined in a suitable manner in one or more embodiments of the present application. In addition, different embodiments and specific features or characteristics in different embodiments can be combined without conflict.
[0171] It should be understood that although the terms first, second, third, etc. may be used to describe various information in one or more embodiments of the present application, these information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "at the time of", "when the time of" or "in response to determining".
[0172] The above description is merely a preferred embodiment of one or more embodiments of the present application and is not intended to limit one or more embodiments of the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of one or more embodiments of the present application shall be included in the scope of protection of one or more embodiments of the present application.
[0173] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
Claims
1. A question-answering method based on a large language model, the method comprising: Get the problems to be solved; Inputting the question into a first language model, and generating a thought chain reasoning step corresponding to the question by the first language model; The question and the thought chain reasoning steps are input into a second large language model, and the second large language model generates an answer corresponding to the question based on the thought chain reasoning steps; wherein the number of model parameters of the first large language model is less than the number of model parameters of the second large language model.
2. The method according to claim 1, before inputting the question into the first large language model, the method comprising: Obtaining a thought chain sample for training the second largest language model serving as a teacher model; Obtaining a thought chain reasoning step corresponding to the question generated by the second largest language model through reasoning based on the question in the thought chain sample, updating the thought chain sample based on the generated thought chain reasoning step, and determining the updated thought chain sample as a distilled sample; The first large language model as a student model is trained based on the distilled samples so that a distillation loss function corresponding to the first large language model converges.
3. According to the method of claim 2, the thought chain reasoning steps in the thought chain sample are thought chain reasoning steps after context compression.
4. According to the method of claim 2, the weight of the sub-loss function corresponding to the word unit serving as the anchor point of the thought chain in the distillation loss function is higher than the weight of the sub-loss function corresponding to other word units in the distillation loss function.
5. The method according to claim 3, wherein the distillation loss function comprises a KL divergence loss function.
6. According to the method according to claim 2, the distillation loss function includes a graph structure consistency loss function calculated based on the entity relationship graph corresponding to the thought chain reasoning step generated by the second largest language model, and the entity relationship graph corresponding to the thought chain reasoning step generated by the first largest language model; the entity relationship graph includes an entity relationship graph generated based on the semantic relationship between entities obtained by relationship extraction of the thought chain reasoning step.
7. According to the method according to claim 1, the question is a question related to fund allocation, the thought chain reasoning step is an analysis of user preferences and fund categories, and the answer is the fund category and its allocation ratio; or, the question is a question related to medical diagnosis, the thought chain reasoning step is the reasoning process from symptoms to disease, and the answer is the diagnosis result.
8. A question-answering device based on a large language model, comprising: Get the module and get the problem to be solved; A first generation module inputs the question into a first large language model, and the first large language model generates a thought chain reasoning step corresponding to the question; The second generation module inputs the question and the thought chain reasoning steps into the second largest language model, and the second largest language model generates an answer corresponding to the question based on the thought chain reasoning steps; wherein the number of model parameters of the first largest language model is less than the number of model parameters of the second largest language model.
9. An electronic device comprising: processor; a memory for storing processor-executable instructions; The processor implements the method according to any one of claims 1 to 7 by running the executable instructions.
10. A computer-readable storage medium having computer instructions stored thereon, wherein when the instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Training method of clinical auxiliary diagnosis model
CN121747910A