Question and answer method, question and answer large model training method, related equipment and program product

By extracting intermediate hidden state features from question data and generating pattern signals through a large question-answering model, and dynamically adjusting the inference strategy, the problem of low efficiency and waste of resources of large language models in handling complex problems is solved, and efficient and flexible inference mode selection is achieved.

CN121029945APending Publication Date: 2025-11-28IFLYTEK CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511229221.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-29
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

Existing large-scale language models suffer from inefficiency, resource waste, and lack of flexibility when dealing with complex problems, making it difficult to dynamically adjust the reasoning mode according to the complexity of the problem.

Method used

By configuring a large question-answering model, the intermediate hidden state features of the question data are extracted, and pattern signals are generated to represent the appropriate reasoning patterns. The pattern determination module is embedded inside the model to dynamically adjust the reasoning strategy and select the appropriate thought chain pattern for reasoning.

Benefits of technology

It improves the intelligence and adaptability of the question-answering model, ensuring reasoning accuracy while increasing efficiency, optimizing resource utilization, and avoiding redundant reasoning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121029945A_ABST
    Figure CN121029945A_ABST
Patent Text Reader

Abstract

The invention discloses a question and answer method, a question and answer big model training method, related equipment and a program product, a configured question and answer big model can extract middle hidden layer state features of question data, a mode signal representing a reasoning mode adapted to the question data can be generated based on the middle hidden layer state features, and exemplarily, the method is simple and convenient to implement, and the efficiency is high. A short CoT reasoning mode can be generated for a simple problem, a long CoT reasoning mode can be generated for a complex problem, and then the middle hidden layer state features and the generated mode signals can be transmitted backwards for subsequent hidden layer reasoning to generate response information of problem data. Compared with a strategy of a fixed reasoning mode, the method has the advantages that reasoning accuracy can be guaranteed, reasoning efficiency can be improved, and resource utilization can be optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of natural language processing technology, and more specifically, to a question-answering method, a large question-answering model training method, related equipment and software products. Background Technology

[0002] In recent years, large-scale language models have made significant progress in understanding and generating human language. However, for complex problems requiring multi-step logical reasoning, such as mathematical problems, directly generating answers often leads to errors or "illusions." To address this issue, Chain-of-Thought (CoT) technology has been proposed. The core idea of ​​CoT is to guide the model to generate a series of intermediate reasoning steps or thought processes before generating the final answer. This approach significantly improves the performance of models on complex reasoning tasks, including solving mathematical problems.

[0003] However, existing CoT methods typically employ relatively fixed inference patterns. Even for simple problems, the model may generate lengthy and unnecessary inference chains, leading to inefficient inference, wasted computational resources, and increased inference time. Summary of the Invention

[0004] In view of the above problems, this application is proposed to provide a question-answering method, a large-scale question-answering model training method, related equipment and program products, so as to enable the large-scale question-answering model to flexibly select an appropriate reasoning mode for different question data, thereby alleviating the problems of fixed reasoning modes. The specific solution is as follows:

[0005] Firstly, it provides a question-and-answer method, including:

[0006] Obtain the problem data;

[0007] The intermediate hidden state features of the question data are extracted by the configured question-answering big model, and a pattern signal is generated based on the intermediate hidden state features. The pattern signal represents the reasoning mode adapted to the question data. Reasoning is performed based on the intermediate hidden state features and the pattern signal to generate the response information of the question data.

[0008] In one possible design, in another implementation of the first aspect of the embodiments of this application, the process of generating a pattern signal based on the intermediate hidden layer state features includes:

[0009] A pattern recognition matrix is ​​generated based on the intermediate hidden layer state features. The pattern recognition matrix includes weight parameters corresponding to various inference modes.

[0010] A category vector is generated based on the pattern recognition matrix, and the category vector is a probability distribution of various reasoning patterns;

[0011] The pattern recognition matrix is ​​multiplied by the category vector to obtain the pattern signal.

[0012] In one possible design, in another implementation of the first aspect of the embodiments of this application, the probability distribution of each inference mode in the category vector is a one-hot vector.

[0013] In one possible design, in another implementation of the first aspect of the embodiments of this application, the reasoning mode adapted to the problem data includes any of the following:

[0014] Long-chain CoT reasoning mode, short-chain CoT reasoning mode, and no-chain CoT reasoning mode.

[0015] In one possible design, in another implementation of the first aspect of the embodiments of this application, the question-answering big model includes a first hidden layer module, a second hidden layer module, and a pattern determination module embedded between the first hidden layer module and the second hidden layer module;

[0016] The process of generating response information for the question data using a large question-answering model includes:

[0017] The intermediate hidden layer state features of the problem data are extracted through the first hidden layer module;

[0018] The pattern determination module generates the pattern signal based on the intermediate hidden layer state features;

[0019] The second hidden layer module infers and generates response information for the problem data based on the intermediate hidden layer state features and the pattern signal.

[0020] In one possible design, in another implementation of the first aspect of the embodiments of this application, the first hidden layer module and the second hidden layer module form a backbone network, and the embedding position of the pattern determination module in the backbone network is located in the shallow layer of the backbone network.

[0021] In one possible design, in another implementation of the first aspect of this application, the second hidden layer module includes a plurality of hidden layers connected in series. The process of generating response information for the problem data through the second hidden layer module based on the intermediate hidden layer state features and the pattern signal includes:

[0022] For each hidden layer in the second hidden layer module, the output of the previous hidden layer and the combination of the pattern signal are used as input to calculate the output of the current hidden layer, until the last hidden layer generates the response information of the problem data. The first hidden layer in the second hidden layer module takes the combination of the intermediate hidden layer state features and the pattern signal as input.

[0023] In one possible design, in another implementation of the first aspect of the embodiments of this application, the training process of the question-answering large model includes:

[0024] Acquire training data, which includes question samples, correct answer labels corresponding to the question samples, and inference pattern labels that are adapted to the question samples;

[0025] The question sample is fed into the question-answering big model, which extracts the intermediate hidden state features of the question sample and generates a prediction pattern signal based on the intermediate hidden state features. The prediction pattern signal represents the reasoning pattern that the question-answering big model predicts to fit the question sample. Reasoning is performed based on the intermediate hidden state features and the prediction pattern signal to generate the response information of the question sample.

[0026] The pattern recognition loss is calculated based on the predicted pattern signal and the inference pattern label. The generation loss is calculated based on the response information of the question sample and the correct answer label. The total loss is obtained based on the pattern recognition loss and the generation loss. The parameters of the question answering model are updated according to the total loss.

[0027] Secondly, a method for training a large question-answering model is provided, including:

[0028] Acquire training data, which includes question samples, correct answer labels corresponding to the question samples, and inference pattern labels that are adapted to the question samples;

[0029] The question sample is fed into the question-answering big model, which extracts the intermediate hidden state features of the question sample and generates a prediction pattern signal based on the intermediate hidden state features. The prediction pattern signal represents the reasoning pattern that the question-answering big model predicts to fit the question sample. Reasoning is performed based on the intermediate hidden state features and the prediction pattern signal to generate the response information of the question sample.

[0030] The pattern recognition loss is calculated based on the predicted pattern signal and the inference pattern label. The generation loss is calculated based on the response information and the correct answer label. The total loss is obtained based on the pattern recognition loss and the generation loss. The parameters of the question-answering model are updated according to the total loss.

[0031] Thirdly, an electronic device is provided, comprising: a memory and a processor;

[0032] The memory is used to store programs;

[0033] The processor is configured to execute the program to implement the steps of the question-answering method described in any of the first aspects of this application, or to implement the steps of the question-answering large model training method described in the second aspect of this application.

[0034] Fourthly, a readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the steps of the question-answering method described in any of the preceding first aspects of this application, or implements the steps of the question-answering large model training method described in the preceding second aspect of this application.

[0035] Fifthly, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the steps of the question-answering method described in any of the first aspects of this application, or implements the steps of the question-answering large model training method described in the second aspect of this application.

[0036] By employing the aforementioned technical solution, the question-answering big model configured in this application can extract the intermediate hidden layer state features of the question data. Based on these features, it can generate pattern signals representing the inference mode adapted to the question data, such as long thought chain (CoT) inference mode and short CoT mode. The intermediate hidden layer state features and the generated pattern signals can then be passed forward for subsequent hidden layer inference to generate response information for the question data. The question-answering big model of this application can dynamically adjust the inference strategy according to the complexity of the input question data. That is, it generates an adapted inference mode by analyzing the question data. For example, a short CoT inference mode can be generated for simple questions, and a long CoT inference mode can be generated for complex questions. The final result can then be generated based on the pattern signals and intermediate hidden layer state features, improving the intelligence and adaptability of the question-answering big model. Compared to strategies with fixed inference modes, the method of this application can improve inference efficiency and optimize resource utilization while ensuring inference accuracy. Attached Figure Description

[0037] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the scope of this application. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:

[0038] Figure 1 A schematic diagram of an implementation system architecture for the question-answering method and question-answering large model training method provided in the embodiments of this application;

[0039] Figure 2 A schematic diagram illustrating the training process of a large question-answering model provided in this application embodiment;

[0040] Figure 3 This is a schematic flowchart of a question-and-answer method provided in an embodiment of this application;

[0041] Figure 4 This is a schematic diagram of the structure of a large question-and-answer model provided in an embodiment of this application;

[0042] Figure 5 This is a schematic diagram of another question-and-answer large model provided in an embodiment of this application;

[0043] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0044] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0045] It is understood that before using the technical solutions disclosed in the various embodiments of this application, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this application in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0046] CoT (Coding in Reasoning) is a hinting technique that stimulates the problem-solving ability of large, complex models by explicitly generating multi-step reasoning paths. Its essence is a transparent reasoning framework of "problem decomposition → step-by-step derivation → answer generation," and it has become one of the core methods for improving the logical reasoning ability of LLM (Limited Learning Model). CoT includes two reasoning modes: Short CoT (SCoT) and Long CoT (LCoT).

[0047] Short Chain-of-Thought (SCot) is a fast and concise way of thinking that aims to reach a conclusion in the shortest amount of time and with the fewest steps. It is characterized by relatively short and continuous reasoning steps (usually no more than 10–20 steps).

[0048] Long Chain-of-Thought (LCoT) is a thinking method that analyzes problems from multiple dimensions and levels to achieve a comprehensive and in-depth understanding. Its characteristics include long and structured reasoning steps (which may involve branching, iteration, or multi-level planning), and the ability to utilize external tools (such as calculators, code interpreters, search engines, database APIs, etc.) during the reasoning process.

[0049] While existing CoT-based inference methods have improved accuracy, they suffer from the following drawbacks:

[0050] Inefficiency: Existing CoT methods (especially zero-shot and few-shot CoT) tend to generate thought chains of similar or redundant lengths for all problems. For simple problems (such as mathematical reasoning problems like single-digit addition or simple equation solving), generating detailed, long thought chains significantly increases inference time, consumes more computational resources (such as GPU time), and incurs API call costs. This "one-size-fits-all" approach is inefficient.

[0051] Lack of flexibility: Existing models struggle to dynamically adjust the level of detail and number of steps in reasoning based on the actual complexity of the problem. They often can only execute a pre-defined or prompt-guided thought process.

[0052] Performance bottleneck: For particularly simple problems, generating long thought chains may actually introduce unnecessary error risks.

[0053] Therefore, how to flexibly select or generate thought chains of different lengths and complexities according to the actual needs of the problem while ensuring the accuracy of reasoning is a major challenge currently facing large-scale language models in the question-answering field.

[0054] In some possible technical implementations, a cascaded strategy can be used to implement the question-answering process. First, a pre-trained classifier is used to classify the input question into simple or complex problems. If the classification result is determined to be a simple problem, the short-thinking-chain CoT inference mode is activated: that is, the large model is invoked to reason about the input question, and the large model is instructed to use the short-thinking-chain CoT inference mode. If the classification result is determined to be a complex problem, the long-thinking-chain CoT inference mode is activated: that is, the large model is invoked to reason about the input question, and the large model is instructed to use the long-thinking-chain CoT inference mode.

[0055] This approach separates the process of classifying the complexity of the input problem from the reasoning process of the large model. In essence, it still belongs to the large model reasoning according to the specified reasoning pattern. When the problem complexity classification result is inaccurate, the reasoning process of the large model still has the problems listed above.

[0056] To address the issues of inefficiency, lack of flexibility, and failure to adaptively adjust the reasoning process to the complexity of the problem when applying thought chains in existing technologies, this application provides an end-to-end question-answering big model. It embeds a mechanism for adaptively selecting the reasoning mode based on the input question into the forward propagation process of the question-answering big model, enabling the model to dynamically adjust subsequent reasoning strategies based on its initial understanding of the input question, thereby achieving a balance between reasoning efficiency and accuracy.

[0057] This application constructs a large-scale language model that can intelligently determine the complexity of the input problem and accordingly select or generate a thought chain of appropriate length that is just enough to solve the problem—using short thought chains (or even direct solutions without thought chains) for simple problems, and using long thought chains for complex problems, thereby significantly improving reasoning efficiency and optimizing resource utilization while ensuring the accuracy of reasoning.

[0058] This application provides a question-answering method and a training method for a large question-answering model, which can be applied to, for example... Figure 1 The system architecture shown may include a terminal 100 and a server 200. The server 200 may include one or more servers (…). Figure 1 (This example uses a server as an illustration).

[0059] Either terminal 100 or server 200 can be used independently to execute the question-and-answer method provided in this application embodiment. Alternatively, terminal 100 and server 200 can also be used collaboratively to execute the question-and-answer method provided in this application embodiment.

[0060] The training of the large question-answering model can generally be performed on server 200. The trained large question-answering model can be deployed on server 200 or on terminal 100.

[0061] The following description Figure 1 The product form of the mid-terminal 100;

[0062] The terminal 100 in this application embodiment can be a mobile phone, tablet computer, learning machine, teaching large screen, wearable device, vehicle-mounted device, conference terminal, augmented reality (AR) / virtual reality (VR) device, laptop computer, ultra-mobile personal computer (UMPC), netbook, personal digital assistant (PDA), etc., and this application embodiment does not impose any restrictions on it.

[0063] The question-answering method in this application is based on a configured large question-answering model, which can adopt a large language model (LLM) structure to support question-answering task processing. This question-answering method can be applied to various scenarios, such as mathematical reasoning, everyday human-computer dialogue, tutoring, and medical consultation.

[0064] Before introducing the question-answering method of this application, we will first introduce the training process of the large question-answering model.

[0065] This application provides a method for training a large question-answering model. Taking the application of this method to a computer device as an example, the computer device can specifically be... Figure 1 The system consists of server 200 or terminal 100 and server 200. (Refer to...) Figure 2 The training method for this large question-answering model specifically includes the following steps:

[0066] S1, Obtain training data.

[0067] The training data includes question samples, the correct answer labels corresponding to the question samples, and the inference pattern labels that are adapted to the question samples.

[0068] Depending on the applicable scenario of the question-answering model, training data for that scenario can be selected. Taking mathematical reasoning as an example, mathematical questions can be collected as question samples.

[0069] The correct solution label for a problem sample includes the correct solution process. For example, for simple problem samples, the correct solution label may include a short thought process and the final answer, or even omit the thought process and directly output the final answer. For complex problem samples, the correct solution label may include a long thought process and the final answer. Whether the correct solution label for a problem sample includes a thought process, and the length of that thought process, depends on the actual solution process of the problem sample; it can be manually labeled or obtained through other means.

[0070] There can be multiple inference mode labels for problem samples, such as: short thought chain CoT, long thought chain CoT, medium-length thought chain CoT, and no thought chain reasoning mode. Different inference modes indicate that when reasoning about problem samples, the large model can use different inference modes to generate thought chains of different lengths (or even no thought chain), and the final answer.

[0071] The reasoning pattern labels for problem samples can be based on manual annotation or on the length of solution steps, problem type, etc. in the labels of correct answers, and can be determined through heuristic rules.

[0072] Taking the reasoning pattern label as having two types of labels, short thought chain CoT and long thought chain CoT, as an example, an optional heuristic rule is as follows:

[0073] When the number of solution steps in the correct answer label is less than the set threshold T, the reasoning mode label is determined to be short thought chain CoT; when the number of solution steps is not less than the set threshold T, the reasoning mode label is determined to be long thought chain CoT.

[0074] Another alternative heuristic rule is as follows:

[0075] When the problem sample type is the first defined type, the reasoning mode label can be determined as short thought chain (CoT); when the problem sample type is the second defined type, the reasoning mode label can be determined as long thought chain (CoT). The first and second defined types can be manually set based on the actual situation. For example, the first defined type indicates a higher complexity of the problem sample, and the second defined type indicates a lower complexity of the problem sample.

[0076] S2, the question sample is fed into the question-answering big model, the question-answering big model extracts the intermediate hidden layer state features of the question sample, and generates a prediction pattern signal based on the intermediate hidden layer state features. The prediction pattern signal represents the reasoning pattern that the question sample is adapted to by the question-answering big model. Reasoning is performed based on the intermediate hidden layer state features and the prediction pattern signal to generate the response information of the question sample.

[0077] Specifically, the large question-answering model consists of a backbone network composed of several hidden layers. The intermediate hidden layer state features of question samples can be extracted through the first few hidden layers.

[0078] This application improves the structure of the question-answering big model by embedding a pattern determination module within it. This module generates a prediction pattern signal based on the intermediate hidden layer state features. The prediction pattern signal represents the inference pattern adapted to the predicted question sample.

[0079] Furthermore, the intermediate hidden layer state features extracted from the first few hidden layers and the predicted pattern signal obtained by the pattern determination module are passed backward. Through subsequent hidden layers, the intermediate hidden layer state features and predicted pattern information are combined to continue reasoning, and finally, the response information of the question sample is generated. This response information may include the thought chain and the answer, or it may only include the answer without the thought chain.

[0080] For example, when the prediction pattern signal is a no-thought-chain reasoning mode, the generated response information may only contain the answer and not the thought chain. When the prediction pattern signal is a short thought chain or a long thought chain, the generated response information may include a thought chain of the corresponding length and the answer.

[0081] S3 calculates the pattern prediction loss based on the predicted pattern signal and the inference pattern label, calculates the generation loss based on the response information and the correct answer label, obtains the total loss based on the pattern prediction loss and the generation loss, and updates the parameters of the question-answering model according to the total loss.

[0082] This application adopts an end-to-end joint training strategy, in which the model needs to learn the following two capabilities simultaneously during the training process:

[0083] a) Understanding and accurately assessing the complexity of input problem samples to predict correct reasoning patterns;

[0084] b) Given a problem sample and a prediction pattern signal, generate the correct response information.

[0085] To this end, this application designs a joint training loss, which includes both pattern prediction loss and generation loss.

[0086] Among them, the model prediction loss L model Based on the predicted pattern signal and the inference pattern label, the classification loss can be used to measure the difference between the predicted pattern signal and the inference pattern label labeled in the training data.

[0087] Taking the reasoning pattern label as including two reasoning patterns, short thought chain and long thought chain, as an example, the prediction pattern signal is defined as the probability distribution P of each reasoning pattern. pred =[p short p long The reasoning mode is tagged as:

[0088] y model For patterns ∈ {Short, Long} (represented by one-hot vectors, such as [1,0] or [0,1]), the pattern prediction loss L can be calculated using cross-entropy. model One way:

[0089]

[0090] The above formula uses two reasoning modes, short thought chain and long thought chain, as an example, where Short represents a short thought chain and Long represents a long thought chain. Of course, the reasoning mode label can be extended to other reasoning modes by changing the range of values ​​for c in the above formula.

[0091] in, This represents the value of category c in the inference pattern label. This represents the predicted probability of belonging to category c.

[0092] Generation loss L genThe generation loss, calculated based on response information and correct answer labels, can be measured using standard language modeling losses (such as cross-entropy loss) to measure the probability of the model predicting the true target token at each step. It should be noted that the generation loss in this application is calculated given the input question samples and the model's predicted pattern signal S. The generation loss is used to encourage the model to generate the correct token sequence under a specific pattern (predicted pattern signal S):

[0093]

[0094] Where B represents the batch size, and N represents the batch size. i Let be the sequence length of the correct answer labels corresponding to the i-th question sample in the batch. It is the true target of the i-th problem sample at step t (the label of the correct answer at step t). It is the token sequence generated by the model before step t, input i S is the i-th problem sample. i θ is the pattern signal predicted by the model for the i-th problem sample, and θ is the model parameter.

[0095] The calculated pattern prediction loss L model and generation loss L gen The two can then be weighted and summed to obtain the total loss:

[0096] L total =αL model +βL gen

[0097] Here, α and β are weight hyperparameters used to balance the importance of the pattern prediction and generation tasks. These weight parameters can be adjusted experimentally.

[0098] Considering the ultimate goal is to generate correct response information, L model This is set up to facilitate achieving this goal. Therefore, we can set β > α, thereby assigning L... gen Higher weight.

[0099] Furthermore, it can be calculated according to the total loss L total Update the parameters of the large question-answering model.

[0100] When updating the parameters of a large question-answering model, deep learning optimization algorithms can be used, such as the AdamW optimization algorithm, to minimize the total loss L. total Update the parameters of the large question-answering model.

[0101] The question-answering large-scale model training method provided in this embodiment improves the network structure of the question-answering large-scale model. Within the model, it can generate predicted pattern signals based on the features of intermediate hidden layers, and then continue reasoning based on the predicted pattern signals and intermediate hidden layer features to generate response information for question samples. Simultaneously, by jointly training with pattern recognition loss and generation loss, the model can be trained to achieve end-to-end capabilities in question understanding, pattern judgment, and adaptive reasoning. This allows the trained question-answering large-scale model to dynamically adjust its reasoning pattern according to the characteristics of the input question, making the model more intelligent and adaptable. It avoids generating redundant thought chains for simple questions, reducing the possibility of errors and improving the response speed of question answering. Furthermore, it can provide sufficient reasoning depth for complex questions, helping to improve the accuracy of the reasoning results.

[0102] Further integration Figure 2 The network structure of the question-answering model is illustrated in the figure.

[0103] The question-answering model may include a first hidden layer module, a second hidden layer module, and a pattern determination module embedded between the first hidden layer module and the second hidden layer module.

[0104] The first hidden layer module includes several hidden layers, and the second hidden layer module includes several hidden layers. The pattern determination module can be one or more fully connected layers.

[0105] During the training process of the large question-answering model:

[0106] Problem samples are sent to the first hidden layer module, which extracts the intermediate hidden layer state features of the problem samples and passes them to the second hidden layer module and the pattern determination module respectively.

[0107] The pattern determination module generates a predicted pattern signal S based on the intermediate hidden layer state features and then passes the predicted pattern signal S to the second hidden layer module.

[0108] The second hidden layer module infers and generates response information for problem samples based on the state features of the intermediate hidden layer and the prediction pattern signal.

[0109] In one alternative example, the pattern determination module can pass the generated predicted pattern signal S to each hidden layer in the second hidden layer module. That is, each hidden layer in the second hidden layer module uses the output of the previous hidden layer and the predicted pattern signal S as input to calculate the output of the current hidden layer, until the last hidden layer generates the response information for the problem sample. The first hidden layer in the second hidden layer module uses a combination of the intermediate hidden layer state features and the predicted pattern signal S as input.

[0110] By embedding a pattern determination module within the model, the complexity of the problem can be determined using the understanding of the problem by the model's first hidden layer module, and then a predictive pattern signal S can be generated based on the problem complexity. Compared to predicting the inference pattern of a large model through simple rules or independent shallow models, the method of this application can more accurately predict the inference pattern that is suitable for the current input problem.

[0111] In one optional implementation, a backbone network is composed of a first hidden layer module and a second hidden layer module, with a pattern determination module embedded between the first and second hidden layer modules. The embedding position of the pattern determination module within the backbone network can be in a shallow layer. That is, the hidden layers contained in the first hidden layer module belong to the shallow layer of the backbone network, while the hidden layers contained in the second hidden layer module belong to the deep layer of the backbone network.

[0112] In one example, the first hidden layer module contains approximately 10% of the hidden layers in the backbone network. By embedding the pattern determination module in a shallow layer of the backbone network, pattern signal prediction is performed after the first hidden layer module has a preliminary understanding of the input problem's encoding. In other words, the shallow network must consider both the generation task and the inference pattern recognition task adapted to the input problem. Embedding the pattern determination module in a shallow layer of the backbone network saves computational overhead.

[0113] In one possible implementation, the backbone network consisting of the first and second hidden layer modules can adopt a transformer structure. The predicted mode signal generated by the mode determination module is used as an additional conditional vector and fused into the state of each hidden layer in the second hidden layer module. Each hidden layer can use the superposition of the output of the previous hidden layer and the predicted mode signal as its input vector.

[0114] The predicted mode signal S generated by the mode determination module is a two-dimensional vector [s short s long For example, [0,1] represents a long thought chain reasoning pattern, and [1,0] represents a short thought chain reasoning pattern. In a certain feedforward network FFN of the second hidden layer module, the traditional calculation method is:

[0115]

[0116] Where W1 and W2 are weight parameters, and b1 and b2 are bias terms.

[0117] After adopting the scheme of this application and introducing the prediction mode signal S, the calculation method is adjusted as follows:

[0118]

[0119] in, , These are additional weighting parameters.

[0120] Under different prediction mode signals S, the response to input x is different, thus guiding the model to produce different computational paths and generation behaviors.

[0121] The large question-answering model is conditionally influenced by the predicted pattern signal S during the inference process, and the model can exhibit different generative behaviors:

[0122] When S indicates "short CoT", the model's internal attention pattern tends to generate more direct, fewer-step token sequences. For example, it may focus more on extracting key numbers and performing calculations to quickly arrive at the result.

[0123] When S indicates "long CoT", the model generates a more detailed sequence of tokens that includes intermediate steps and explanations. This may involve activating pathways in the model that handle more complex logical relationships.

[0124] In some embodiments of this application, a question-answering method based on a large question-answering model is further described. Taking the application of this method to a computer device as an example, the computer device may specifically be... Figure 1 The system consists of terminal 100 or a combination of terminal 100 and server 200. (Refer to...) Figure 3 The question-and-answer method specifically includes the following steps:

[0125] Step S100: Obtain problem data.

[0126] The question-and-answer method proposed in this application can be applied to different scenarios, such as tutoring scenarios in the education industry and consultation scenarios in the medical industry.

[0127] Depending on the application scenario, the system can retrieve the question data that needs to be answered in that specific scenario. Taking a math tutoring scenario as an example, the retrieved question data could be math problems.

[0128] Step S110: Extract the intermediate hidden state features of the question data through the configured question-answering big model, and generate a pattern signal based on the intermediate hidden state features. The pattern signal represents the reasoning mode adapted to the question data. Reasoning is performed based on the intermediate hidden state features and the pattern signal to generate the response information of the question data.

[0129] The question-answering model can be trained using the training method described in the aforementioned embodiments.

[0130] After receiving the question data P, P is converted into a token sequence, and an initial token embedding vector is obtained through the embedding layer, completing the input encoding process of the question data. The input encoding of the question data is then pre-encoded through the initial hidden layer module of the question-answering model to obtain hidden layer states (intermediate hidden layer state features) containing preliminary contextual information of the question. Further, the question-answering model generates a pattern signal S based on the intermediate hidden layer state features. This pattern signal characterizes the complexity of the question data, and can also be understood as the reasoning pattern adapted to the question data. For example, for simple question data, the adapted reasoning pattern can be a short thought chain CoT reasoning pattern; for complex question data, the adapted reasoning pattern can be a long thought chain CoT reasoning pattern. The intermediate hidden layer state features and the pattern signal are passed to subsequent hidden layers of the model for reasoning, ultimately generating response information for the question data. The response information can include both the thought chain and the answer, or it can only include the answer without the thought chain (e.g., when the pattern signal represents a no-thought chain reasoning pattern, the final generated response information can only contain the answer to the question data, without necessarily including the thought chain of the reasoning process).

[0131] The question-answering model in this application incorporates a pattern recognition mechanism, which can predict and generate applicable reasoning patterns (long and short thought chains, etc.) based on a preliminary understanding of the question. The subsequent reasoning process of the model is regulated by the generated pattern signals, dynamically adjusting the calculation and generation behavior to achieve the output of thought chains of different lengths.

[0132] The question-answering method provided in this application is based on an improved question-answering big model. The question-answering big model can extract intermediate hidden state features from the question data. Based on these features, it can generate pattern signals representing the inference pattern adapted to the question data, such as long thought chain (CoT) inference patterns and short CoT patterns. The intermediate hidden state features and the generated pattern signals can then be passed forward for subsequent hidden layer inference to generate response information for the question data. The question-answering big model of this application can dynamically adjust the inference strategy according to the complexity of the input question data. That is, it generates an adapted inference pattern by analyzing the question data. For example, a short CoT inference pattern can be generated for simple questions, and a long CoT inference pattern can be generated for complex questions. The final result can then be generated based on the pattern signals and intermediate hidden state features, improving the intelligence and adaptability of the question-answering big model. Compared to strategies with fixed inference patterns, the method of this application can improve inference efficiency and optimize resource utilization while ensuring inference accuracy.

[0133] In one possible implementation, the reasoning pattern adapted to the problem data includes, but is not limited to, any of the following:

[0134] Long-thinking-chain CoT reasoning model, short-thinking-chain CoT reasoning model, and CoT reasoning model without thinking chain, etc.

[0135] Examples of different generative behaviors of the model corresponding to different inference modes are as follows:

[0136] When the pattern signal generated by the model indicates a "short CoT", the attention pattern within the model tends to generate more direct, fewer-step token sequences. For example, it may focus more on extracting key numbers and performing calculations to quickly obtain results.

[0137] When the pattern signal indicates a "long CoT", the model generates a more detailed sequence of tokens that includes intermediate steps and explanations. This may involve activating pathways in the model that handle more complex logical relationships.

[0138] When the mode signal indicates "no CoT", the model can output only the final answer without having to output the intermediate steps.

[0139] Of course, the above are just a few examples of possible reasoning patterns. In addition, other types of reasoning patterns can be designed according to the length of the thought chain. This application will not exhaustively list them.

[0140] In some embodiments of this application, the above-described step S110, which generates a mode signal based on the intermediate hidden layer state features, is described.

[0141] In one possible implementation, refer to Figure 4 As shown, the pattern signal S can be directly generated based on the intermediate hidden layer state features through the pattern determination module embedded in the question-answering model. The pattern signal S can be the probability distribution of various inference patterns; for example, the pattern signal S can be represented in one-hot vector form to represent the probability distribution of various inference patterns.

[0142] In another possible implementation, refer to Figure 5 As shown, the implementation process of step S110 may include the following steps:

[0143] S11. Generate a pattern recognition matrix based on the intermediate hidden layer state features. This pattern recognition matrix includes weight parameters corresponding to various inference modes.

[0144] Taking all reasoning patterns, including short thought chain CoT and long thought chain CoT, as an example, the pattern recognition matrix can be [W short W long ], where W short W represents the weight of the short thought chain CoT. long This represents the weight of the long thought chain CoT.

[0145] S12. Generate category vectors based on the pattern recognition matrix. The category vectors are the probability distributions of various reasoning patterns.

[0146] Taking all reasoning patterns, including short thought chain CoT and long thought chain CoT, as an example, based on the pattern recognition matrix generated in the previous step, the category vector can be obtained through the activation function: [s short s long ], where s short s represents the probability of a short thought chain CoT. long This represents the probability of a long thought chain CoT.

[0147] In some possible examples, the probability distribution of each inference pattern in the category vector can be represented by a one-hot encoded vector. For example, [0,1] represents a long thought chain inference pattern, and [1,0] represents a short thought chain inference pattern. By using one-hot encoded vectors to represent the category vector, the discriminativeness of the pattern signal S can be enhanced, making it easier for the large model to infer according to a more explicit inference pattern in subsequent iterations.

[0148] S13. Multiply the pattern recognition matrix by the category vector to obtain the pattern signal S.

[0149] Taking the pattern recognition matrix and category vector from the example above as an example, multiplying them together yields the pattern signal S:

[0150] S=[W short W long ]·[s short s long ] T

[0151] The pattern signal S generated by the pattern determination module can be passed as an additional conditional vector to subsequent hidden layers of the model, guiding the subsequent inference process. Under different pattern signals S, the question-answering model responds differently to the input question, thus guiding the question-answering model to generate different computational paths and generation behaviors.

[0152] Combination Figure 3 , Figure 4 As shown, an optional component structure of the question-answering big data model of this application is presented.

[0153] The question-answering model may include a first hidden layer module, a second hidden layer module, and a pattern determination module embedded between the first hidden layer module and the second hidden layer module.

[0154] The aforementioned step S110, the process of generating response information for question data through a large question-answering model, may include:

[0155] The intermediate hidden layer state features of the problem data are extracted through the first hidden layer module;

[0156] The pattern determination module generates pattern signals based on the intermediate hidden layer state features;

[0157] The second hidden layer module infers and generates response information for the problem data based on the state features and pattern signals of the intermediate hidden layer.

[0158] In one alternative example, the pattern determination module can pass the generated pattern signal to each hidden layer in the second hidden layer module. That is, each hidden layer in the second hidden layer module takes the output of the previous hidden layer and the predicted pattern signal as input to calculate the output of the current hidden layer, until the last hidden layer generates the response information for the problem data. The first hidden layer in the second hidden layer module takes a combination of the intermediate hidden layer state features and the pattern signal as input.

[0159] By passing the pattern signal as a condition vector to each hidden layer in the second hidden layer module, each hidden layer can perform inference calculations under the guidance of the pattern signal, thereby guiding the model to flexibly adjust the inference mode according to the input question.

[0160] This embodiment embeds a pattern determination module within the model, utilizing the understanding of the problem by the model's first hidden layer module to determine the problem's complexity, and then generates a prediction pattern signal S based on the problem's complexity. Compared to predicting the inference pattern of a large model through simple rules or independent shallow models, the method of this application can more accurately predict the inference pattern that is suitable for the current input problem.

[0161] In one optional implementation, a backbone network is composed of a first hidden layer module and a second hidden layer module, with a pattern determination module embedded between the first and second hidden layer modules. The embedding position of the pattern determination module within the backbone network can be in a shallow layer. That is, the hidden layers contained in the first hidden layer module belong to the shallow layer of the backbone network, while the hidden layers contained in the second hidden layer module belong to the deep layer of the backbone network.

[0162] In one example, the first hidden layer module contains approximately 10% of the hidden layers in the backbone network. By embedding the pattern determination module in a shallow layer of the backbone network, pattern signal prediction is performed after the first hidden layer module has a preliminary understanding of the input problem's encoding. In other words, the shallow network must consider both the generation task and the inference pattern recognition task adapted to the input problem. Embedding the pattern determination module in a shallow layer of the backbone network saves computational overhead.

[0163] In one possible implementation, the backbone network consisting of the first and second hidden layer modules can adopt a transformer structure. The mode signal generated by the mode determination module is used as an additional conditional vector and fused into the state of each hidden layer in the second hidden layer module. Each hidden layer can use the superposition of the output of the previous hidden layer and the mode signal as its input vector.

[0164] In summary, unlike existing technologies that use independent modules to predict the reasoning patterns of large models or use fixed reasoning patterns, this application deeply integrates the reasoning pattern selection mechanism based on the input question into the forward propagation process of the question-answering large model. This enables the question-answering large model to dynamically adjust the subsequent reasoning patterns based on the initial understanding of the question, achieving a trade-off between efficiency and accuracy.

[0165] This application introduces a collaborative reasoning mechanism with thought chains of different lengths, which has the following technical advantages:

[0166] Enhance model flexibility and adaptability: The model can dynamically adjust its inference strategy based on the characteristics of the input problem, making it more intelligent and adaptable.

[0167] Improve overall performance: By avoiding generating redundant steps for simple problems, the possibility of errors is reduced; at the same time, providing sufficient reasoning depth for complex problems helps to improve the success rate of solving the problem.

[0168] Improve user experience: Get faster response times for simple questions from users.

[0169] This application also provides an electronic device in its embodiments. (See reference...) Figure 6 The diagram illustrates a structural schematic suitable for implementing the electronic device in the embodiments of this application. The electronic device in the embodiments of this application may include, but is not limited to, fixed terminals such as mobile phones, tablets, learning machines, large teaching screens, wearable devices, etc. Figure 6 The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.

[0170] like Figure 6 As shown, the electronic device may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 1, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 2 or a program loaded from a storage device 8 into a random access memory (RAM) 3, to implement the question-answering method or the question-answering large model training method of the foregoing embodiments of this application. When the electronic device is powered on, the RAM 3 also stores various programs and data required for the operation of the electronic device. The processing unit 1, ROM 2, and RAM 3 are interconnected via a bus 4. An input / output (I / O) interface 5 is also connected to the bus 4.

[0171] Typically, the following devices can be connected to I / O interface 5: input devices 6 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 7 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 8 including, for example, memory cards, hard drives, etc.; and communication devices 9. Communication device 9 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 6 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or have alternatively.

[0172] This application also provides a computer program product including computer-readable instructions, which, when executed on an electronic device, cause the electronic device to implement any of the question-answering methods or question-answering large model training methods provided in this application.

[0173] This application also provides a computer-readable storage medium that carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement any of the question-answering methods or question-answering large model training methods provided in this application.

[0174] It should also be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. In addition, in the device embodiment drawings provided in this application, the connection relationship between modules indicates that they have a communication connection, which can be implemented as one or more communication buses or signal lines.

[0175] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware, or it can be implemented by special-purpose hardware including application-specific integrated circuits, special-purpose CPUs, special-purpose memory, special-purpose components, etc. Generally, any function performed by a computer program can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be diverse, such as analog circuits, digital circuits, or special-purpose circuits. However, for this application, software program implementation is more often the preferred implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a computer floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, training equipment, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0176] In the above embodiments, the implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, in the form of a computer program product.

[0177] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, training device, or data center to another website, computer, training device, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device or data center that integrates one or more available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives (SSDs)).

[0178] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The various embodiments can be combined as needed, and the same or similar parts can be referred to each other.

Claims

1. A question-and-answer method, characterized in that the method include: Obtain the problem data; The intermediate hidden state features of the question data are extracted by the configured question-answering big model, and a pattern signal is generated based on the intermediate hidden state features. The pattern signal represents the reasoning mode adapted to the question data. Reasoning is performed based on the intermediate hidden state features and the pattern signal to generate the response information of the question data.

2. The method according to claim 1, characterized in that, The process of generating pattern signals based on the intermediate hidden layer state features includes: A pattern recognition matrix is ​​generated based on the intermediate hidden layer state features. The pattern recognition matrix includes weight parameters corresponding to various inference modes. A category vector is generated based on the pattern recognition matrix, and the category vector is a probability distribution of various reasoning patterns; The pattern recognition matrix is ​​multiplied by the category vector to obtain the pattern signal.

3. The method according to claim 2, characterized in that, The probability distribution of each reasoning mode in the category vector is a one-hot vector.

4. The method according to claim 1, characterized in that, The reasoning mode adapted to the question data includes any of the following: Long-chain CoT reasoning mode, short-chain CoT reasoning mode, and no-chain CoT reasoning mode.

5. The method according to claim 1, characterized in that, The question-answering model includes a first hidden layer module, a second hidden layer module, and a pattern determination module embedded between the first hidden layer module and the second hidden layer module; The process of generating response information for the question data using a large question-answering model includes: The intermediate hidden layer state features of the problem data are extracted through the first hidden layer module; The pattern determination module generates the pattern signal based on the intermediate hidden layer state features; The second hidden layer module infers and generates response information for the problem data based on the intermediate hidden layer state features and the pattern signal.

6. The method according to claim 5, characterized in that, The first hidden layer module and the second hidden layer module form a backbone network, and the embedding position of the pattern determination module in the backbone network is located in the shallow layer of the backbone network.

7. The method according to claim 5, characterized in that, The second hidden layer module comprises several hidden layers connected in series. The process by which the second hidden layer module infers and generates response information for the problem data based on the intermediate hidden layer state features and the pattern signal includes: For each hidden layer in the second hidden layer module, the output of the previous hidden layer and the combination of the pattern signal are used as input to calculate the output of the current hidden layer, until the last hidden layer generates the response information of the problem data. The first hidden layer in the second hidden layer module takes the combination of the intermediate hidden layer state features and the pattern signal as input.

8. The method according to any one of claims 1-7, characterized in that, The training process of the large question-answering model includes: Acquire training data, which includes question samples, correct answer labels corresponding to the question samples, and inference pattern labels that are adapted to the question samples; The question sample is fed into the question-answering big model, which extracts the intermediate hidden state features of the question sample and generates a prediction pattern signal based on the intermediate hidden state features. The prediction pattern signal represents the reasoning pattern that the question-answering big model predicts to fit the question sample. Reasoning is performed based on the intermediate hidden state features and the prediction pattern signal to generate the response information of the question sample. The pattern recognition loss is calculated based on the predicted pattern signal and the inference pattern label. The generation loss is calculated based on the response information of the question sample and the correct answer label. The total loss is obtained based on the pattern recognition loss and the generation loss. The parameters of the question answering model are updated according to the total loss.

9. A method for training a large question-answering model, characterized in that, include: Acquire training data, which includes question samples, correct answer labels corresponding to the question samples, and inference pattern labels that are adapted to the question samples; The question sample is fed into the question-answering big model, which extracts the intermediate hidden state features of the question sample and generates a prediction pattern signal based on the intermediate hidden state features. The prediction pattern signal represents the reasoning pattern that the question-answering big model predicts to fit the question sample. Reasoning is performed based on the intermediate hidden state features and the prediction pattern signal to generate the response information of the question sample. The pattern recognition loss is calculated based on the predicted pattern signal and the inference pattern label. The generation loss is calculated based on the response information and the correct answer label. The total loss is obtained based on the pattern recognition loss and the generation loss. The parameters of the question-answering model are updated according to the total loss.

10. An electronic device, characterized in that, include: Memory and processor; The memory is used to store programs; The processor is configured to execute the program to implement the steps of the question-answering method as described in any one of claims 1 to 8, or to implement the steps of the question-answering large model training method as described in claim 9.

11. A readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the question-answering method as described in any one of claims 1 to 8, or the steps of the question-answering large model training method as described in claim 9.

12. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the question-answering method as described in any one of claims 1 to 8, or the steps of the question-answering large model training method as described in claim 9.

Citation Information

Cited By

  • Text generation method and device based on large model, training method and device, equipment and medium

    CN121599118A